You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Do LLM judges know when they're wrong? Three open-weight judges (Qwen2.5-7B, kev-8b, auto-j-13b) graded against MT-Bench human votes: calibration, position and padding attacks, and what auto-accepting confident verdicts would cost. Interactive site included.
Reproducible LLM-as-a-judge reliability lab: chance-corrected agreement (Cohen's kappa, Krippendorff's alpha) with bootstrap CIs, computed keyless from a committed MT-Bench snapshot and re-derived in CI as a drift gate.
Score an LLM judge against human labels: chance-corrected agreement with cluster-bootstrap intervals, the human-human ceiling it has to be read against, and a written drift protocol. Stdlib only, no API key.