Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 17 additions & 4 deletions .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,10 +1,23 @@
{
"name": "scriptorium",
"version": "1.0.0",
"description": "Sovereign rigor layer for scientific work — audit, verify, recompute, and grade research with deterministic engines, graduated epistemic status, and an offline-capable reviewer path. Complements generative writing tools; never fabricates a number.",
"author": { "name": "Łukasz Minarowski", "url": "https://orcid.org/0000-0002-2536-3508" },
"version": "1.3.0",
"description": "Sovereign rigor layer for scientific work \u2014 audit, verify, recompute, and grade research with deterministic engines, graduated epistemic status, and an offline-capable reviewer path. Complements generative writing tools; never fabricates a number.",
"author": {
"name": "\u0141ukasz Minarowski",
"url": "https://orcid.org/0000-0002-2536-3508"
},
"repository": "https://github.com/kicrazom/scriptorium",
"homepage": "https://github.com/kicrazom/scriptorium",
"license": "MIT",
"keywords": ["research", "science", "peer-review", "literature", "statistics", "bayesian", "academic-writing", "epistemic", "claude-code-plugin"]
"keywords": [
"research",
"science",
"peer-review",
"literature",
"statistics",
"bayesian",
"academic-writing",
"epistemic",
"claude-code-plugin"
]
}
2 changes: 1 addition & 1 deletion ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ feature count.
- ⏳ Each fixed bug / behavioural failure becomes a regression fixture (ongoing convention).

## v1.0.0 — Stable public release ✅ (released)
- Uniform engine envelope contract (all 9 engines) + JSON schemas + STABILITY.md (SemVer + deprecation policy).
- Uniform engine envelope contract (all 8 engines) + JSON schemas + STABILITY.md (SemVer + deprecation policy).
- Coverage measured honestly (subprocess) + gated; two CI jobs (matrix + R).
- `epistemic_grade` conformed to the envelope contract.

Expand Down
2 changes: 1 addition & 1 deletion STATUS.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@ where R / the package is absent (e.g. on the default CI runner), and run locally
| Bash-capable components routed through the parser (`statistician`, `power-sample-size`) | done ✅ — they call the parser CLI |
| offline read-only agents routed through the parser | not applicable — no `Bash`; they read `profile.md` directly by design |
| `schemas/` (JSON Schema I/O contracts: envelope, finding, power request/response, profile) | implemented ✅, enforced in `tests/test_schemas.py` |
| uniform engine envelope contract (every engine → `ok`/`error` + graded `finding`) | implemented ✅, verified for all 9 engines in `tests/test_envelope_contract.py` (v1.0.0 stable contract) |
| uniform engine envelope contract (every engine → `ok`/`error` + graded `finding`) | implemented ✅, verified for all 8 engines in `tests/test_envelope_contract.py` (v1.0.0 stable contract) |
| per-engine bespoke response schemas (each engine's `data` payload pinned, not just the shared envelope) | implemented ✅, enforced in `tests/test_engine_schemas.py` (representative fixture + negative control per engine; v1.2.0) |

## Behavioral validation (prompt layer)
Expand Down
18 changes: 18 additions & 0 deletions agents/peer-reviewer.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,6 +210,24 @@ Trace the argument: problem → mechanisms → evidence → synthesis → implic

- Citation style consistent and complete (numbered / author-year per the venue); reference
list matches in-text callouts.
- **Citation-manipulation screen (COPE).** Run this on every manuscript and **state the result
either way** — a clean screen is a finding and must be reported as plainly as a hit.
- **The bolt-on fingerprint.** In a Vancouver-numbered list, references are numbered in order
of first appearance. So a citation that *first appears in the Introduction* but carries one
of the *last* numbers in the list was appended after the list was finalised — i.e. inserted
late, typically at a revision round. Trace first-appearance order against the numbering and
report any reference that breaks it.
- **Off-target citations.** Does the cited source actually support the sentence it is attached
to? A source that cannot support its claim (wrong disease, wrong species, wrong method, wrong
field) is the highest-yield signal available offline, and it is checkable from the reference
*title* alone — no network needed.
- Also check: self-citation clusters (especially off-target ones inserted where a topical source
was required); single-publisher or single-group clusters; blocks of references cited exactly
once, together, in a passage they do not bear on; uncited references; duplicate entries under
two numbers; and citations to the target journal (a coercion signal — its *absence* is equally
worth stating).
- **Report, do not accuse.** Describe the pattern and let the editor adjudicate. Coercion may come
from a reviewer or an editor, not the authors — the authors may be the injured party.
- Reproducibility: data- and code-availability statements present; enough methodological
detail to replicate; software / model versions stated.
- Ethics: ethics-committee / IRB approval (with a number), informed consent, and **data-
Expand Down
59 changes: 59 additions & 0 deletions schemas/evidence_selection_neutrality_response.schema.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://github.com/kicrazom/scriptorium/schemas/evidence_selection_neutrality_response.schema.json",
"title": "EvidenceSelectionNeutralityResponse",
"description": "Envelope returned by scripts/core/evidence_selection_neutrality.py: deterministic scan of an AUTHORED INSTRUCTION (prompt, skill body, task spec, reviewer request) for directives that steer the evidence base — a foregone conclusion, cherry-picked sources, statistical steering, coerced citations, concealed AI/COI, suppressed uncertainty, or spin. Distinct from injection_scan: that engine asks whether an instruction came from an unauthorised source; this one asks whether an authorised instruction distorts what counts as evidence. Heuristic — a clean scan is silence, not a certificate of neutrality.",
"type": "object",
"required": ["status"],
"properties": {
"status": {"type": "string", "enum": ["ok", "error"]},
"message": {"type": "string", "description": "Present when status == error."},
"data": {
"type": "object",
"required": ["steering", "verdict", "risk_score", "finding"],
"properties": {
"steering": {
"type": "array",
"items": {
"type": "object",
"required": ["pattern", "category", "severity", "line", "snippet"],
"properties": {
"pattern": {"type": "string"},
"category": {
"type": "string",
"enum": [
"foregone_conclusion",
"source_cherry_picking",
"statistical_steering",
"citation_coercion",
"disclosure_evasion",
"uncertainty_suppression",
"spin_directive"
]
},
"severity": {"type": "string", "enum": ["critical", "warning"]},
"line": {"type": "integer", "minimum": 1},
"snippet": {"type": "string"}
}
}
},
"verdict": {
"type": "string",
"enum": ["pass", "review", "fail"],
"description": "fail when any blocking category is hit (foregone_conclusion, source_cherry_picking, statistical_steering, citation_coercion, disclosure_evasion)."
},
"risk_score": {
"type": "integer",
"minimum": 0,
"maximum": 4,
"description": "0 none | 1-2 reviewable steering | 3 blocking steering | 4 blocking across more than one category."
},
"finding": {"$ref": "finding.schema.json"}
}
}
},
"allOf": [
{"if": {"properties": {"status": {"const": "ok"}}}, "then": {"required": ["data"]}},
{"if": {"properties": {"status": {"const": "error"}}}, "then": {"required": ["message"]}}
]
}
4 changes: 3 additions & 1 deletion schemas/power_response.schema.json
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,14 @@
"properties": {
"n_per_group": {"type": "integer", "minimum": 1},
"n_total": {"type": "integer", "minimum": 1},
"n_group1": {"type": "integer", "minimum": 1},
"n_group2": {"type": "integer", "minimum": 1},
"events_required": {"type": "integer", "minimum": 1},
"method": {"type": "string", "description": "Exact analytic routine used."},
"assumptions": {"type": "object", "description": "alpha, power, alternative, and the effect input echoed back."},
"finding": {"$ref": "finding.schema.json"}
},
"description": "Sample-size designs carry {n_per_group, n_total}; survival carries {events_required}."
"description": "Sample-size designs carry {n_per_group, n_total}; two_sample_t also carries per-arm {n_group1, n_group2} (differ when ratio != 1); survival carries {events_required}."
}
},
"allOf": [
Expand Down
34 changes: 13 additions & 21 deletions scripts/behavioral/judge_behavior.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,30 +35,22 @@ def build_judge_prompt(case, agent_response):
)


_DECODER = json.JSONDecoder()


def _extract_json_object(text):
"""Return the first balanced top-level {...} substring, or None. String-aware so braces
inside quoted values do not unbalance the scan."""
"""Return the first `{...}` substring that decodes as valid JSON, or None.

json.JSONDecoder.raw_decode does the string-aware brace matching for us and, unlike a
plain balance scan, only accepts a `{` that begins genuinely valid JSON.
"""
start = text.find("{")
while start != -1:
depth, in_str, esc = 0, False, False
for i in range(start, len(text)):
ch = text[i]
if in_str:
if esc:
esc = False
elif ch == "\\":
esc = True
elif ch == '"':
in_str = False
elif ch == '"':
in_str = True
elif ch == "{":
depth += 1
elif ch == "}":
depth -= 1
if depth == 0:
return text[start:i + 1]
start = text.find("{", start + 1)
try:
_obj, end = _DECODER.raw_decode(text, start)
return text[start:end]
except json.JSONDecodeError:
start = text.find("{", start + 1)
return None


Expand Down
24 changes: 21 additions & 3 deletions scripts/core/citation_parse.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,16 +16,34 @@
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))
from scripts.lib import json_io, provenance, epistemic # noqa: E402

MARKER = re.compile(r"\[(\d+)\]")
# A numeric marker group: starts with a digit, then only digits, commas, spaces, and
# hyphen/en-dash — so [1], [1,2,5], and [1-3]/[1–3] match, but [see 2] does not.
MARKER = re.compile(r"\[(\d[\d\s,–-]*)\]")
_RANGE = re.compile(r"(\d+)\s*[–-]\s*(\d+)")


def _expand(inner):
"""Expand one marker group's inner text into the set of integers it cites."""
out = set()
for part in inner.split(","):
part = part.strip()
r = _RANGE.fullmatch(part)
if r:
lo, hi = int(r.group(1)), int(r.group(2))
if lo <= hi:
out.update(range(lo, hi + 1))
elif part.isdigit():
out.add(int(part))
return out


def parse(req):
body = req["body_text"]
refs = req["references"]
n_refs = len(refs)

markers = sorted({int(m) for m in MARKER.findall(body)})
cited = set(markers)
cited = set().union(*(_expand(g) for g in MARKER.findall(body))) if MARKER.search(body) else set()
markers = sorted(cited)
orphan = [i for i in range(1, n_refs + 1) if i not in cited]
dangling = [m for m in markers if m < 1 or m > n_refs]

Expand Down
Loading
Loading