Skip to content

Repository files navigation

CoderDrift

DOI CI

Temporal Reliability Auditing in Human Annotation Studies

CoderDrift is a design-aware Python workflow for checking whether nominal or ordinal annotation reliability remains stable over annotation order. It keeps descriptive profiles, exploratory change-point candidates, and calibrated primary alerts separate. Primary labels are available only when a dataset and evidence channel match a prospectively locked calibration profile.

CoderDrift does not diagnose fatigue, negligence, bias, or individual coder deterioration. Peer evidence is relative to available overlapping coders, and external references are treated as fallible.

Installation

Stable release install

git clone --branch v1.1.0 --depth 1 https://github.com/ecylmz/CoderDrift.git
cd CoderDrift
uv sync --frozen
uv run coderdrift version

The version command reports CoderDrift 1.1.0.

Development checkout

For the latest development state on main:

git clone https://github.com/ecylmz/CoderDrift.git
cd CoderDrift
uv sync --frozen
uv run coderdrift version

The development version command should report CoderDrift 1.1.0. The immutable v1.1.0 release is archived at doi:10.5281/zenodo.22165191.

Python 3.12, 3.13, and 3.14 are supported. CoderDrift uses CPU-only open-source dependencies and has no telemetry or hosted-service requirement.

Minimal use

uv run coderdrift validate configs/examples/nominal_descriptive.yml
uv run coderdrift audit configs/examples/nominal_descriptive.yml

Audit bundles contain validated configuration, eligibility decisions, static and temporal summaries, change candidates, simulation-conditional family records, stability regions, warnings, provenance, self-contained HTML, and vector figures.

Evidence modes

  • Descriptive: static, batch, and rolling profiles; no primary alert.
  • Exploratory: PELT candidates with explicit unsupported-condition labels.
  • Calibrated primary: requires a compatible immutable profile, eligible length and family, practical effect, support qualification, and applicable family false-alert calibration under its declared null generator.

The built-in profiles are intentionally narrow: MultiPref overall peer evidence at n >= 250, WebCrowd25K binary-reference evidence at n >= 250, and WebCrowd25K four-level peer evidence at n >= 500.

CoderDrift does not provide distribution-free FWER control for arbitrary real datasets. A new design must prospectively declare and validate its own null DGP; incompatible or failed designs remain descriptive or exploratory.

Documentation

Build the documentation locally with uv run mkdocs build --strict.

Published validation artifacts

The repository includes the 192-cell, 19,200-panel expanded stress benchmark, the accepted locked validation cells, three immutable profiles, an 81-cell frozen-policy robustness benchmark, a MultiPref graph-preserving semi-synthetic validation, and aggregate MultiPref and WebCrowd25K result tables. Record-level real-data audit bundles stay local to respect data terms and worker privacy. Manuscript materials are maintained separately from the public software repository.

Reproduction

Third-party raw datasets are not redistributed. After lawful acquisition, run:

uv sync --frozen
uv run python scripts/reproduce_all.py

Without third-party raw data, the command verifies frozen scientific inputs, aggregate result hashes, examples, and tests. With both lawful inputs present, it also rebuilds all real-data audits, tables, and figures. Expensive revision experiments are opt-in:

uv run python scripts/reproduce_all.py --rerun-custom-calibration
uv run python scripts/reproduce_all.py --rerun-robustness
uv run python scripts/reproduce_all.py --rerun-semisynthetic

The custom example preserves its failed locked gate and creates no primary profile. Robustness and semi-synthetic results evaluate frozen policies and do not feed profile selection.

The frozen feasibility evidence is documented in PHASE0_REPORT.md; all later execution decisions and gates are consolidated in PROJECT_REPORT.md.

License

BSD-3-Clause. Dataset terms remain those of their original providers.

Citation

Use the metadata in CITATION.cff. Version 1.1.0 is archived at doi:10.5281/zenodo.22165191.

About

Design-aware temporal reliability auditing for human annotation studies

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages