Hi maintainers — I maintain EvalPort, an open, framework-agnostic JSON spec (Apache 2.0) for portable LLM eval test cases and results (TestCase/Grader/EvalSuite/ResultSet/GraderResult, JSON-Schema validated), so eval data isn't locked to one tool's format.
I read rag_experiment_accelerator/artifact/models/query_output.py and rag_experiment_accelerator/evaluation/eval.py before writing this, not just the README:
QueryOutput(question, actual, expected, retrieved_contexts, search_type, search_evals, rerank, cross_encoder_at_k, ...) maps directly onto a TestCase/Result pair: question → TestCase.input, expected → TestCase.expected_output, retrieved_contexts → TestCase.retrieval_context, actual → Result.actual_output.
evaluate_single_prompt() builds a metric_dic keyed by every configured metric_type — and this repo genuinely supports a lot of them: string-similarity (lcsstr, jaccard, levenshtein, fuzzy_score, cosine_ochiai), ROUGE (rouge1/2/L-precision/recall/fmeasure), BERT-based semantic similarity across 6 named models, and LLM-judged (llm_context_precision, llm_answer_relevance, llm_context_recall). Each becomes a GraderResult on the same test case, so a single QueryOutput naturally fans out into several GraderResults under one Result — which is exactly what EvalPort's Result.grader_results: array is for.
search_evals[i]["precision_scores"] (per-k precision, aggregated in evaluate_prompts() into total_precision_scores_by_search_type/MAP@k) is the retrieval-quality half — a natural custom-type grader (params.handler = "rag_experiment_accelerator:precision_at_k") alongside the generation-quality graders above, since EvalPort's well-known types don't have a MAP@k-style retrieval metric built in.
Given how many metric_type values exist already (I count 20+ across plain_metrics, transformer-based, and LLM-based), most would map to custom graders with params.handler = "rag_experiment_accelerator:<metric_type>" rather than forcing them into EvalPort's ~11 well-known types — the spec's own "type openness" rule exists for exactly this.
What I'm proposing: a standalone rag-experiment-accelerator-openeval-adapter package — zero changes to this repo, built only against the public QueryOutput class and metric_dic shape from evaluate_single_prompt() — with to_openeval(query_outputs: list[QueryOutput], metric_dics: list[dict]) -> (EvalSuite, ResultSet). I don't see a process for feature proposals beyond the standard CONTRIBUTING/CLA template, so filing this as an issue first. Happy to build it either as:
- A standalone package in EvalPort's own
adapters/ directory (same pattern as the ~46 other framework adapters already there, e.g. azure-ai-evaluation-openeval-adapter) — zero footprint on this repo, or
- A PR into this repo (e.g. under
rag_experiment_accelerator/evaluation/) if you'd rather it live here.
Real tests would validate against EvalPort's actual JSON Schema, not a mock. Let me know if this is useful, or not — no worries either way.
Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
— Sahi, independent contributor (not affiliated with this project)
Hi maintainers — I maintain EvalPort, an open, framework-agnostic JSON spec (Apache 2.0) for portable LLM eval test cases and results (
TestCase/Grader/EvalSuite/ResultSet/GraderResult, JSON-Schema validated), so eval data isn't locked to one tool's format.I read
rag_experiment_accelerator/artifact/models/query_output.pyandrag_experiment_accelerator/evaluation/eval.pybefore writing this, not just the README:QueryOutput(question, actual, expected, retrieved_contexts, search_type, search_evals, rerank, cross_encoder_at_k, ...)maps directly onto aTestCase/Resultpair:question→TestCase.input,expected→TestCase.expected_output,retrieved_contexts→TestCase.retrieval_context,actual→Result.actual_output.evaluate_single_prompt()builds ametric_dickeyed by every configuredmetric_type— and this repo genuinely supports a lot of them: string-similarity (lcsstr,jaccard,levenshtein,fuzzy_score,cosine_ochiai), ROUGE (rouge1/2/L-precision/recall/fmeasure), BERT-based semantic similarity across 6 named models, and LLM-judged (llm_context_precision,llm_answer_relevance,llm_context_recall). Each becomes aGraderResulton the same test case, so a singleQueryOutputnaturally fans out into severalGraderResults under oneResult— which is exactly what EvalPort'sResult.grader_results: arrayis for.search_evals[i]["precision_scores"](per-k precision, aggregated inevaluate_prompts()intototal_precision_scores_by_search_type/MAP@k) is the retrieval-quality half — a naturalcustom-type grader (params.handler = "rag_experiment_accelerator:precision_at_k") alongside the generation-quality graders above, since EvalPort's well-known types don't have a MAP@k-style retrieval metric built in.Given how many
metric_typevalues exist already (I count 20+ acrossplain_metrics, transformer-based, and LLM-based), most would map tocustomgraders withparams.handler = "rag_experiment_accelerator:<metric_type>"rather than forcing them into EvalPort's ~11 well-known types — the spec's own "type openness" rule exists for exactly this.What I'm proposing: a standalone
rag-experiment-accelerator-openeval-adapterpackage — zero changes to this repo, built only against the publicQueryOutputclass andmetric_dicshape fromevaluate_single_prompt()— withto_openeval(query_outputs: list[QueryOutput], metric_dics: list[dict]) -> (EvalSuite, ResultSet). I don't see a process for feature proposals beyond the standard CONTRIBUTING/CLA template, so filing this as an issue first. Happy to build it either as:adapters/directory (same pattern as the ~46 other framework adapters already there, e.g.azure-ai-evaluation-openeval-adapter) — zero footprint on this repo, orrag_experiment_accelerator/evaluation/) if you'd rather it live here.Real tests would validate against EvalPort's actual JSON Schema, not a mock. Let me know if this is useful, or not — no worries either way.
Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
— Sahi, independent contributor (not affiliated with this project)