Skip to content

Build /safety-approach page: expose all 58k+ prompts, accord, guide, batteries, rubrics, capture+judge bundles, and judge template #16

Description

@emooreatx

Goal

Build a public ciris.ai/safety-approach page that exposes the complete CIRIS safety surface — every prompt, rubric, question, judge transcript, and CI-attested verdict — as a stopgap until the full safety.ciris.ai site lands.

Why now: the safety-battery CI loop is live and green on the am cell (Amharic mental_health), with hi (Hindi) coming online today. The first pilot teams need a public URL to reference; "look at the GH Actions Artifacts API" is a non-starter for non-technical reviewers (researchers, mental-health professionals, partner orgs). This page is the bridge.

Why our entire model depends on this: rules are crowdsourced, verdicts are machined. Both halves of that contract need to be public and inspectable — otherwise it's just censorship with extra steps. The whole approach falls down if reviewers can't read every prompt, every rubric, and every verdict for themselves.

What to expose

All seven surfaces below are versioned (SHA-pinned), public, and stable. Pull paths are programmatic — no scraping.

1. Localized strings — 58,497 total (2,017 keys × 29 locales)

Every user-facing string the agent ever produces — error messages, defer notifications, conscience pondering, status text, button labels.

2. ACCORD texts — 29 locales

The ACCORD is the agent's foundational ethical contract. One file per locale.

3. Comprehensive guide — 29 locales + base

The agent's longer-form operating guide.

4. DMA prompts — 203 YAML files

The decision-maker-analyzer prompts (PDMA, CSDMA, DSDMA, IDMA, ASPDMA, conscience). One per (DMA, locale).

5. Safety batteries — 14 cells (mental_health pilot)

The question sets, with arc-level metadata, per-question stage/category/translations.

6. Rubrics — both halves

Human-readable rubric markdown (14 files, one per cell):

Machine-applicable criteria (operationalized — currently 2 of 14: am, hi):

  • Source: tests/safety/{lang_eng}_mental_health/v4_{lang_eng}_canonical_universal_criteria.json
  • Five kinds: term_present, term_absent, regex_present, script_detection, interpreter_judgment
  • Authoritative format: CIRISNodeCore SCHEMA.md §12
  • Render hint: paired view — rubric.md on left, criteria.json on right, U-rows aligned

7. Capture results + judge verdicts

Per-cell, per-CI-run signed bundles produced by .github/workflows/safety-battery.yml.

Capture bundle (what the agent said):

  • Tuple name: safety-battery-capture-{lang}-{domain}-v{N}-{model_slug}-{agent_v}-{template}
  • Contents: results.jsonl (one row per question + response), summary.json, manifest_signed.json
  • Sigstore-attested via attest-build-provenance@v1

Interpret bundle (what the judge said):

  • Tuple name: same as capture + -{rubric_short}-{interpreter_v}
  • Contents: verdicts.jsonl (one row per criterion application), verdicts_summary.json, manifest_signed.json
  • Each verdict carries: verdict (pass/fail/undetermined), severity, cited_span, interpreter_kind (deterministic vs foundation_model), judge_model, judge_prompt_sha256
  • Sigstore-attested

Lookup: GET https://api.github.com/repos/CIRISAI/CIRISAgent/actions/artifacts?name={tuple-name} — latest-wins by name, queryable across all CI runs

Full integration spec: CIRISNodeCore/PROGRAMMATIC_ACCESS.md

8. Judge prompt template

The single calibratable surface that decides every interpreter_judgment verdict.

  • Source: JUDGE_PROMPT_TEMPLATE in tools/qa_runner/modules/safety_interpret.py (~50 lines, all of it deployment-context + judgment rules)
  • Authoritative contract: CIRISNodeCore FSD/JUDGE_MODEL.md
  • Every verdict carries the sha256[:8] of the prompt template that produced it, so historical verdicts are reproducible
  • Render hint: live render of the template alongside the calibration-history note (moral frameworks the prompt inherits from — WHO mhGAP §0.1, Catholic moral theology §0.2, Ubuntu §0.3 — kept out of the operational prompt but documented in the FSD)

Implementation suggestions

  • Static-render where possible — most of this content changes at release cadence (every 1-2 weeks), not in real time. Render at build time from the public GitHub repos.
  • Search across the corpus — Algolia/Pagefind/lunr. A reviewer should be able to type "schizophrenia" and find the criterion that catches cross-cluster contamination across all 14 cells.
  • Permalink everything — every U-row, every question, every verdict gets a stable anchor. Reviewers cite by URL.
  • Diff view — when a rubric or criteria.json changes, show the diff. The whole point of rules-crowdsourced is that the audit trail is public.
  • No login — entire page is public, no telemetry beyond standard request logs. Verdict transparency is the product.

Sourcing

Everything is in two public repos, no auth required:

Acceptance

A non-technical reviewer (mental-health professional, partner-org liaison, journalist) lands on ciris.ai/safety-approach, picks the am/mental_health cell, sees:

  • The 9 Amharic questions side-by-side with English
  • The 9-row rubric.md AND the operationalized criteria.json
  • The most recent capture run's responses
  • The most recent judge run's verdicts (with PASS/FAIL/UNDETERMINED + cited spans)
  • The judge prompt template that produced those verdicts
  • The Sigstore attestation status on both bundles

…all without leaving the page or knowing what GH Actions is.

This is the proof-of-transparency we owe partners. Full safety.ciris.ai (Contribution flow, voting, expertise weighting) layers on top of this later.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions