Temporal Reliability Auditing in Human Annotation Studies
CoderDrift is a design-aware Python workflow for checking whether nominal or ordinal annotation reliability remains stable over annotation order. It keeps descriptive profiles, exploratory change-point candidates, and calibrated primary alerts separate. Primary labels are available only when a dataset and evidence channel match a prospectively locked calibration profile.
CoderDrift does not diagnose fatigue, negligence, bias, or individual coder deterioration. Peer evidence is relative to available overlapping coders, and external references are treated as fallible.
git clone --branch v1.1.0 --depth 1 https://github.com/ecylmz/CoderDrift.git
cd CoderDrift
uv sync --frozen
uv run coderdrift versionThe version command reports CoderDrift 1.1.0.
For the latest development state on main:
git clone https://github.com/ecylmz/CoderDrift.git
cd CoderDrift
uv sync --frozen
uv run coderdrift versionThe development version command should report CoderDrift 1.1.0. The immutable v1.1.0 release is archived at doi:10.5281/zenodo.22165191.
Python 3.12, 3.13, and 3.14 are supported. CoderDrift uses CPU-only open-source dependencies and has no telemetry or hosted-service requirement.
uv run coderdrift validate configs/examples/nominal_descriptive.yml
uv run coderdrift audit configs/examples/nominal_descriptive.ymlAudit bundles contain validated configuration, eligibility decisions, static and temporal summaries, change candidates, simulation-conditional family records, stability regions, warnings, provenance, self-contained HTML, and vector figures.
- Descriptive: static, batch, and rolling profiles; no primary alert.
- Exploratory: PELT candidates with explicit unsupported-condition labels.
- Calibrated primary: requires a compatible immutable profile, eligible length and family, practical effect, support qualification, and applicable family false-alert calibration under its declared null generator.
The built-in profiles are intentionally narrow: MultiPref overall peer
evidence at n >= 250, WebCrowd25K binary-reference evidence at n >= 250,
and WebCrowd25K four-level peer evidence at n >= 500.
CoderDrift does not provide distribution-free FWER control for arbitrary real datasets. A new design must prospectively declare and validate its own null DGP; incompatible or failed designs remain descriptive or exploratory.
- Getting started
- Methodology and interpretation
- Prospective calibration tutorial
- Data acquisition and governance
- Python API
Build the documentation locally with uv run mkdocs build --strict.
The repository includes the 192-cell, 19,200-panel expanded stress benchmark, the accepted locked validation cells, three immutable profiles, an 81-cell frozen-policy robustness benchmark, a MultiPref graph-preserving semi-synthetic validation, and aggregate MultiPref and WebCrowd25K result tables. Record-level real-data audit bundles stay local to respect data terms and worker privacy. Manuscript materials are maintained separately from the public software repository.
Third-party raw datasets are not redistributed. After lawful acquisition, run:
uv sync --frozen
uv run python scripts/reproduce_all.pyWithout third-party raw data, the command verifies frozen scientific inputs, aggregate result hashes, examples, and tests. With both lawful inputs present, it also rebuilds all real-data audits, tables, and figures. Expensive revision experiments are opt-in:
uv run python scripts/reproduce_all.py --rerun-custom-calibration
uv run python scripts/reproduce_all.py --rerun-robustness
uv run python scripts/reproduce_all.py --rerun-semisyntheticThe custom example preserves its failed locked gate and creates no primary profile. Robustness and semi-synthetic results evaluate frozen policies and do not feed profile selection.
The frozen feasibility evidence is documented in PHASE0_REPORT.md; all later
execution decisions and gates are consolidated in PROJECT_REPORT.md.
BSD-3-Clause. Dataset terms remain those of their original providers.
Use the metadata in CITATION.cff. Version 1.1.0 is archived
at doi:10.5281/zenodo.22165191.