Skip to content
#

inter-rater-reliability

Here are 53 public repositories matching this topic...

Automates the COPUS classroom observation protocol with Google Gemini: codes STEM lecture videos in 2-minute windows, checks agreement with human observers (Cohen's κ / Gwet's AC1), runs a video/audio/transcript ablation study, and generates faculty feedback reports. Streamlit app + CLI.

  • Updated Oct 1, 2026
  • Python

A study in "Agent Trajectory Evaluation": judging an AI agent's full Execution Path, Not just its final Answer. A small ReAct Agent generates Trajectories; ~40 are hand-labelled as Ground truth; 'Rule-based and LLM-judge scorers' are measured against 'those labels'. The Deliverable is the agreement analysis, and where Automated Evaluation breaks.

  • Updated Aug 25, 2026
  • Python

Do LLM non-determinism and deployment-stack variation alter the conclusions of environmental-health meta-analyses? 36,000 LLM calls across 6 deployment stacks, with a pre-registered dual-human validation of the gold standard.

  • Updated Aug 25, 2026
  • Python

Agent skills for decisions under uncertainty, plus the evaluation harness that measures them: pre-registered predictions enforced by git ancestry, blind LLM-as-a-judge relabeling with chance-corrected agreement, and every run published with raw transcripts.

  • Updated Sep 10, 2026
  • Python

Add this topic to your repo

To associate your repository with the inter-rater-reliability topic, visit your repo's landing page and select "manage topics."

Learn more