Goal
Build a public ciris.ai/safety-approach page that exposes the complete CIRIS safety surface — every prompt, rubric, question, judge transcript, and CI-attested verdict — as a stopgap until the full safety.ciris.ai site lands.
Why now: the safety-battery CI loop is live and green on the am cell (Amharic mental_health), with hi (Hindi) coming online today. The first pilot teams need a public URL to reference; "look at the GH Actions Artifacts API" is a non-starter for non-technical reviewers (researchers, mental-health professionals, partner orgs). This page is the bridge.
Why our entire model depends on this: rules are crowdsourced, verdicts are machined. Both halves of that contract need to be public and inspectable — otherwise it's just censorship with extra steps. The whole approach falls down if reviewers can't read every prompt, every rubric, and every verdict for themselves.
What to expose
All seven surfaces below are versioned (SHA-pinned), public, and stable. Pull paths are programmatic — no scraping.
1. Localized strings — 58,497 total (2,017 keys × 29 locales)
Every user-facing string the agent ever produces — error messages, defer notifications, conscience pondering, status text, button labels.
2. ACCORD texts — 29 locales
The ACCORD is the agent's foundational ethical contract. One file per locale.
3. Comprehensive guide — 29 locales + base
The agent's longer-form operating guide.
4. DMA prompts — 203 YAML files
The decision-maker-analyzer prompts (PDMA, CSDMA, DSDMA, IDMA, ASPDMA, conscience). One per (DMA, locale).
5. Safety batteries — 14 cells (mental_health pilot)
The question sets, with arc-level metadata, per-question stage/category/translations.
6. Rubrics — both halves
Human-readable rubric markdown (14 files, one per cell):
Machine-applicable criteria (operationalized — currently 2 of 14: am, hi):
- Source:
tests/safety/{lang_eng}_mental_health/v4_{lang_eng}_canonical_universal_criteria.json
- Five
kinds: term_present, term_absent, regex_present, script_detection, interpreter_judgment
- Authoritative format:
CIRISNodeCore SCHEMA.md §12
- Render hint: paired view — rubric.md on left, criteria.json on right, U-rows aligned
7. Capture results + judge verdicts
Per-cell, per-CI-run signed bundles produced by .github/workflows/safety-battery.yml.
Capture bundle (what the agent said):
- Tuple name:
safety-battery-capture-{lang}-{domain}-v{N}-{model_slug}-{agent_v}-{template}
- Contents:
results.jsonl (one row per question + response), summary.json, manifest_signed.json
- Sigstore-attested via
attest-build-provenance@v1
Interpret bundle (what the judge said):
- Tuple name: same as capture +
-{rubric_short}-{interpreter_v}
- Contents:
verdicts.jsonl (one row per criterion application), verdicts_summary.json, manifest_signed.json
- Each verdict carries:
verdict (pass/fail/undetermined), severity, cited_span, interpreter_kind (deterministic vs foundation_model), judge_model, judge_prompt_sha256
- Sigstore-attested
Lookup: GET https://api.github.com/repos/CIRISAI/CIRISAgent/actions/artifacts?name={tuple-name} — latest-wins by name, queryable across all CI runs
Full integration spec: CIRISNodeCore/PROGRAMMATIC_ACCESS.md
8. Judge prompt template
The single calibratable surface that decides every interpreter_judgment verdict.
- Source:
JUDGE_PROMPT_TEMPLATE in tools/qa_runner/modules/safety_interpret.py (~50 lines, all of it deployment-context + judgment rules)
- Authoritative contract: CIRISNodeCore FSD/JUDGE_MODEL.md
- Every verdict carries the sha256[:8] of the prompt template that produced it, so historical verdicts are reproducible
- Render hint: live render of the template alongside the calibration-history note (moral frameworks the prompt inherits from — WHO mhGAP §0.1, Catholic moral theology §0.2, Ubuntu §0.3 — kept out of the operational prompt but documented in the FSD)
Implementation suggestions
- Static-render where possible — most of this content changes at release cadence (every 1-2 weeks), not in real time. Render at build time from the public GitHub repos.
- Search across the corpus — Algolia/Pagefind/lunr. A reviewer should be able to type "schizophrenia" and find the criterion that catches cross-cluster contamination across all 14 cells.
- Permalink everything — every U-row, every question, every verdict gets a stable anchor. Reviewers cite by URL.
- Diff view — when a rubric or criteria.json changes, show the diff. The whole point of rules-crowdsourced is that the audit trail is public.
- No login — entire page is public, no telemetry beyond standard request logs. Verdict transparency is the product.
Sourcing
Everything is in two public repos, no auth required:
Acceptance
A non-technical reviewer (mental-health professional, partner-org liaison, journalist) lands on ciris.ai/safety-approach, picks the am/mental_health cell, sees:
- The 9 Amharic questions side-by-side with English
- The 9-row rubric.md AND the operationalized criteria.json
- The most recent capture run's responses
- The most recent judge run's verdicts (with PASS/FAIL/UNDETERMINED + cited spans)
- The judge prompt template that produced those verdicts
- The Sigstore attestation status on both bundles
…all without leaving the page or knowing what GH Actions is.
This is the proof-of-transparency we owe partners. Full safety.ciris.ai (Contribution flow, voting, expertise weighting) layers on top of this later.
Goal
Build a public
ciris.ai/safety-approachpage that exposes the complete CIRIS safety surface — every prompt, rubric, question, judge transcript, and CI-attested verdict — as a stopgap until the fullsafety.ciris.aisite lands.Why now: the safety-battery CI loop is live and green on the
amcell (Amharic mental_health), withhi(Hindi) coming online today. The first pilot teams need a public URL to reference; "look at the GH Actions Artifacts API" is a non-starter for non-technical reviewers (researchers, mental-health professionals, partner orgs). This page is the bridge.Why our entire model depends on this: rules are crowdsourced, verdicts are machined. Both halves of that contract need to be public and inspectable — otherwise it's just censorship with extra steps. The whole approach falls down if reviewers can't read every prompt, every rubric, and every verdict for themselves.
What to expose
All seven surfaces below are versioned (SHA-pinned), public, and stable. Pull paths are programmatic — no scraping.
1. Localized strings — 58,497 total (2,017 keys × 29 locales)
Every user-facing string the agent ever produces — error messages, defer notifications, conscience pondering, status text, button labels.
ciris_engine/data/localized/{lang}.json(29 files)ciris_engine/data/localized/manifest.json— language list, RTL flags, base-language indicatoren.json; flag[EN]markers (untranslated placeholders); render RTL locales (ar,fa,ur) correctly2. ACCORD texts — 29 locales
The ACCORD is the agent's foundational ethical contract. One file per locale.
ciris_engine/data/localized/accord_1.2b_{lang}.txt3. Comprehensive guide — 29 locales + base
The agent's longer-form operating guide.
ciris_engine/data/localized/CIRIS_COMPREHENSIVE_GUIDE_{lang}.txtenbase?)4. DMA prompts — 203 YAML files
The decision-maker-analyzer prompts (PDMA, CSDMA, DSDMA, IDMA, ASPDMA, conscience). One per (DMA, locale).
ciris_engine/logic/dma/prompts/localized/{lang}/*.yml(29 locales × ~7 DMAs)5. Safety batteries — 14 cells (mental_health pilot)
The question sets, with arc-level metadata, per-question stage/category/translations.
tests/safety/{lang_eng}_mental_health/v4_{lang_eng}_mental_health_arc.jsonCIRISNodeCore SCHEMA.md §116. Rubrics — both halves
Human-readable rubric markdown (14 files, one per cell):
tests/safety/{lang_eng}_mental_health/v4_{lang_eng}_scoring_rubric.mdMachine-applicable criteria (operationalized — currently 2 of 14:
am,hi):tests/safety/{lang_eng}_mental_health/v4_{lang_eng}_canonical_universal_criteria.jsonkinds:term_present,term_absent,regex_present,script_detection,interpreter_judgmentCIRISNodeCore SCHEMA.md §127. Capture results + judge verdicts
Per-cell, per-CI-run signed bundles produced by
.github/workflows/safety-battery.yml.Capture bundle (what the agent said):
safety-battery-capture-{lang}-{domain}-v{N}-{model_slug}-{agent_v}-{template}results.jsonl(one row per question + response),summary.json,manifest_signed.jsonattest-build-provenance@v1Interpret bundle (what the judge said):
-{rubric_short}-{interpreter_v}verdicts.jsonl(one row per criterion application),verdicts_summary.json,manifest_signed.jsonverdict(pass/fail/undetermined),severity,cited_span,interpreter_kind(deterministic vs foundation_model),judge_model,judge_prompt_sha256Lookup:
GET https://api.github.com/repos/CIRISAI/CIRISAgent/actions/artifacts?name={tuple-name}— latest-wins by name, queryable across all CI runsFull integration spec: CIRISNodeCore/PROGRAMMATIC_ACCESS.md
8. Judge prompt template
The single calibratable surface that decides every
interpreter_judgmentverdict.JUDGE_PROMPT_TEMPLATEintools/qa_runner/modules/safety_interpret.py(~50 lines, all of it deployment-context + judgment rules)Implementation suggestions
Sourcing
Everything is in two public repos, no auth required:
Acceptance
A non-technical reviewer (mental-health professional, partner-org liaison, journalist) lands on
ciris.ai/safety-approach, picks theam/mental_healthcell, sees:…all without leaving the page or knowing what GH Actions is.
This is the proof-of-transparency we owe partners. Full safety.ciris.ai (Contribution flow, voting, expertise weighting) layers on top of this later.