Measure whether your model judge agrees with human raters. Chance-corrected agreement statistics, bootstrap confidence intervals, and calibration gates for Swift Testing and CI. Zero dependencies.
-
Updated
Aug 1, 2026 - Swift
Measure whether your model judge agrees with human raters. Chance-corrected agreement statistics, bootstrap confidence intervals, and calibration gates for Swift Testing and CI. Zero dependencies.
Automates the COPUS classroom observation protocol with Google Gemini: codes STEM lecture videos in 2-minute windows, checks agreement with human observers (Cohen's κ / Gwet's AC1), runs a video/audio/transcript ablation study, and generates faculty feedback reports. Streamlit app + CLI.
A self-hosted multi-rater data labeling platform for structured annotation workflows.
A study in "Agent Trajectory Evaluation": judging an AI agent's full Execution Path, Not just its final Answer. A small ReAct Agent generates Trajectories; ~40 are hand-labelled as Ground truth; 'Rule-based and LLM-judge scorers' are measured against 'those labels'. The Deliverable is the agreement analysis, and where Automated Evaluation breaks.
Labeling queue library for managing human labeling workflows
Inter-rater agreement study: 40 legal questions with gold references, an anchored rubric, four LLM judges, and the statistics that say whether they measure the same thing
Design-aware temporal reliability auditing for human annotation studies
Statistical validity checks for human-graded AI evaluations
Do LLM non-determinism and deployment-stack variation alter the conclusions of environmental-health meta-analyses? 36,000 LLM calls across 6 deployment stacks, with a pre-registered dual-human validation of the gold standard.
⚖️ Dual-Judge: 让AI测试结果真正有说服力 | 双LLM交叉验证消除单模型偏见 | 独立于具体Agent的通用评估框架 | Making AI Evaluation Trustworthy
Cross-model LLM-as-judge eval harness: validate AI judges with Fleiss' kappa / Krippendorff's alpha, not accuracy. Ships a real 7-model panel (Claude, GPT, Gemini, Grok, Qwen, DeepSeek, GLM) you can replay in ~30s, no API key. MIT.
Agent skills for decisions under uncertainty, plus the evaluation harness that measures them: pre-registered predictions enforced by git ancestry, blind LLM-as-a-judge relabeling with chance-corrected agreement, and every run published with raw transcripts.
Statistical analysis of inter-rater reliability and quality patterns in LLM evaluation systems using R
Expert review operations for AI-training data: versioned weighted rubrics, exact rubric scoring, and inter-rater agreement (Cohen, Fleiss, Krippendorff) with seeded bootstrap CIs.
Blind pairwise evaluation — seeded blinding, position-bias detection, and inter-rater agreement. The win rate is the number that means least. Zero dependencies; MIT.
Score job descriptions against a 44-domain life-science ontology, calibrated against a clinical expert panel and validated by inter-rater agreement (quadratic weighted kappa) between two independent runs.
Python package implementing measure from my working paper "Kappa-IoU: Inter-Rater Reliability for Spatial Annotation"
How reliable is LLM-as-judge? Inter-model agreement and failure modes across 5 local and frontier models on 200 claim-verification items.
Reproducible LLM-as-a-judge reliability lab: chance-corrected agreement (Cohen's kappa, Krippendorff's alpha) with bootstrap CIs, computed keyless from a committed MT-Bench snapshot and re-derived in CI as a drift gate.
Tool-agnostic inter-coder reliability (Krippendorff alpha, Cohen/Fleiss kappa) and disagreement adjudication for qualitative coding
To associate your repository with the inter-rater-reliability topic, visit your repo's landing page and select "manage topics."