Multilingual Large Language Models Evaluation Benchmark
-
Updated
Aug 21, 2024 - Python
Multilingual Large Language Models Evaluation Benchmark
GraphParser is a semantic parser which can convert natural language sentences to logical forms and graphs.
Evaluation Dataset for "Bootstrapping Large-Scale Fine-Grained Contextual Advertising Classifier from Wikipedia" (TextGraphs-15 Workshop@NAACL 2021)
for prompts, dataset, and code addressing the task of scientific synthesis
Open evaluation datasets, test cases, and metrics for DSH plugins.
Provider-neutral AI evaluation toolkit for reusable test cases, regression datasets, provider adapters, deterministic and LLM judges, agent trajectory and security evaluation, statistical analysis, quality/cost/latency trade-offs, online experiments, evidence guardrails, and release gates.
To associate your repository with the evaluation-datasets topic, visit your repo's landing page and select "manage topics."