Skip to content

Proposal: standalone rag-experiment-accelerator-openeval-adapter (QueryOutput/metric_dic <-> EvalPort Suite/ResultSet) #992

Description

@adhabnr-ux

Hi maintainers — I maintain EvalPort, an open, framework-agnostic JSON spec (Apache 2.0) for portable LLM eval test cases and results (TestCase/Grader/EvalSuite/ResultSet/GraderResult, JSON-Schema validated), so eval data isn't locked to one tool's format.

I read rag_experiment_accelerator/artifact/models/query_output.py and rag_experiment_accelerator/evaluation/eval.py before writing this, not just the README:

  • QueryOutput(question, actual, expected, retrieved_contexts, search_type, search_evals, rerank, cross_encoder_at_k, ...) maps directly onto a TestCase/Result pair: question → TestCase.input, expected → TestCase.expected_output, retrieved_contexts → TestCase.retrieval_context, actual → Result.actual_output.
  • evaluate_single_prompt() builds a metric_dic keyed by every configured metric_type — and this repo genuinely supports a lot of them: string-similarity (lcsstr, jaccard, levenshtein, fuzzy_score, cosine_ochiai), ROUGE (rouge1/2/L-precision/recall/fmeasure), BERT-based semantic similarity across 6 named models, and LLM-judged (llm_context_precision, llm_answer_relevance, llm_context_recall). Each becomes a GraderResult on the same test case, so a single QueryOutput naturally fans out into several GraderResults under one Result — which is exactly what EvalPort's Result.grader_results: array is for.
  • search_evals[i]["precision_scores"] (per-k precision, aggregated in evaluate_prompts() into total_precision_scores_by_search_type/MAP@k) is the retrieval-quality half — a natural custom-type grader (params.handler = "rag_experiment_accelerator:precision_at_k") alongside the generation-quality graders above, since EvalPort's well-known types don't have a MAP@k-style retrieval metric built in.

Given how many metric_type values exist already (I count 20+ across plain_metrics, transformer-based, and LLM-based), most would map to custom graders with params.handler = "rag_experiment_accelerator:<metric_type>" rather than forcing them into EvalPort's ~11 well-known types — the spec's own "type openness" rule exists for exactly this.

What I'm proposing: a standalone rag-experiment-accelerator-openeval-adapter package — zero changes to this repo, built only against the public QueryOutput class and metric_dic shape from evaluate_single_prompt() — with to_openeval(query_outputs: list[QueryOutput], metric_dics: list[dict]) -> (EvalSuite, ResultSet). I don't see a process for feature proposals beyond the standard CONTRIBUTING/CLA template, so filing this as an issue first. Happy to build it either as:

  1. A standalone package in EvalPort's own adapters/ directory (same pattern as the ~46 other framework adapters already there, e.g. azure-ai-evaluation-openeval-adapter) — zero footprint on this repo, or
  2. A PR into this repo (e.g. under rag_experiment_accelerator/evaluation/) if you'd rather it live here.

Real tests would validate against EvalPort's actual JSON Schema, not a mock. Let me know if this is useful, or not — no worries either way.

Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md

— Sahi, independent contributor (not affiliated with this project)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions