Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
e01b77d
fix(ahbg): derive movement adjacency from UCNS relations
erinepshovel-code Sep 30, 2026
62aacd4
fix(ahbg): fail closed on malformed construction ledgers
erinepshovel-code Sep 30, 2026
45e03b8
test(ahbg): bind movement and construction to UCNS evidence
erinepshovel-code Sep 30, 2026
b541781
ci(ahbg): execute production runtime boundary tests
erinepshovel-code Sep 30, 2026
59a0b5e
docs(ahbg): bind current cross-repository integration graph
erinepshovel-code Sep 30, 2026
b17297f
docs(ahbg): establish benchmark research area
erinepshovel-code Sep 30, 2026
99b5ac3
feat(ahbg): add benchmark package
erinepshovel-code Sep 30, 2026
23c3fb1
feat(ahbg): implement matched intervention receipts
erinepshovel-code Sep 30, 2026
77bfba6
test(ahbg): make matched interventions falsifiable
erinepshovel-code Sep 30, 2026
7434707
feat(ahbg): separate injection observation from enforced refusal
erinepshovel-code Sep 30, 2026
b08c278
test(ahbg): witness observe-only adversarial terrain
erinepshovel-code Sep 30, 2026
e71cbd4
docs(ahbg): define causal benchmark and evidence boundary
erinepshovel-code Sep 30, 2026
0b85b8b
test(ahbg): fail closed on cross-repo authority drift
erinepshovel-code Sep 30, 2026
30e0279
ci(ahbg): run causal benchmark tests
erinepshovel-code Sep 30, 2026
c823140
feat(ahbg): bind run evidence to cross-repo work graph
erinepshovel-code Sep 30, 2026
7acec24
feat(ahbg): attach work-graph provenance to every run
erinepshovel-code Sep 30, 2026
66ae07b
test(ahbg): witness run provenance boundary
erinepshovel-code Sep 30, 2026
7ecb038
feat(ahbg): derive raw behavioral phenotype vectors
erinepshovel-code Sep 30, 2026
0fa8c0d
feat(ahbg): export phenotype instrument
erinepshovel-code Sep 30, 2026
8128c92
test(ahbg): preserve multidimensional behavior evidence
erinepshovel-code Sep 30, 2026
efb7c3e
feat(ahbg): retain per-turn agent execution provenance
erinepshovel-code Sep 30, 2026
da596df
feat(ahbg): connect current A0 through ordinary harness API
erinepshovel-code Sep 30, 2026
81ac534
feat(ahbg): expose integration adapters
erinepshovel-code Sep 30, 2026
bce54f8
test(ahbg): bind current A0 execution and provider provenance
erinepshovel-code Sep 30, 2026
62d0da1
feat(ahbg): add executable current-A0 benchmark runner
erinepshovel-code Sep 30, 2026
b87e110
ci(ahbg): execute current-A0 integration contract tests
erinepshovel-code Sep 30, 2026
bccb436
docs(ahbg): record implemented current-A0 adapter boundary
erinepshovel-code Sep 30, 2026
1f0b613
docs(ahbg): record current-A0 harness integration
erinepshovel-code Sep 30, 2026
c7fb6f2
fix(ahbg): require deployed A0 identity and atomic adapter setup
erinepshovel-code Sep 30, 2026
d71147e
fix(ahbg): preserve submitted plan before runtime intervention
erinepshovel-code Sep 30, 2026
91224c9
fix(ahbg): distinguish agent behavior from runtime intervention
erinepshovel-code Sep 30, 2026
ef0bc6c
test(ahbg): distinguish submitted and executed plans
erinepshovel-code Sep 30, 2026
7fa5c41
test(ahbg): preserve behavior under runtime suppression
erinepshovel-code Sep 30, 2026
efc2ce2
docs(ahbg): require exact deployed A0 identity
erinepshovel-code Sep 30, 2026
27b8a84
fix(ahbg): read agent identity from run provenance
erinepshovel-code Sep 30, 2026
dbc6f5c
fix(ahbg): bind full stimuli and deadline into run evidence
erinepshovel-code Sep 30, 2026
de17dfd
test(ahbg): prove deadline and stimuli enter causal evidence
erinepshovel-code Sep 30, 2026
9082329
feat(ahbg): execute preregistered matched intervention pairs
erinepshovel-code Sep 30, 2026
ed67bd9
test(ahbg): make paired causal execution fail on confounds
erinepshovel-code Sep 30, 2026
5dd6a7b
feat(ahbg): export matched-pair executor
erinepshovel-code Sep 30, 2026
eb5e300
docs(ahbg): document executable causal pair evidence
erinepshovel-code Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions .github/workflows/ahbg.yml
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,15 @@ jobs:
working-directory: ahbg/grok
run: python -m unittest discover -s a0/tests -p 'test*.py'

- name: AHBG production runtime tests
run: python -m unittest discover -s ahbg/runtime/tests -p 'test*.py'

- name: AHBG matched-intervention tests
run: python -m unittest discover -s ahbg/benchmark/tests -p 'test*.py'

- name: AHBG cross-repository integration tests
run: python -m unittest discover -s ahbg/integration/tests -p 'test*.py'

- name: Grok ahbg tests
working-directory: ahbg/grok
run: python -m unittest discover -s ahbg/tests -p 'test*.py'
Expand Down
51 changes: 51 additions & 0 deletions ahbg/benchmark/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# AHBG benchmark research

This layer turns AHBG from a reproducible scenario suite into a causal instrument. The existing 35-scenario calibration corpus remains historical builder-conformance evidence; it is not rewritten to make later research questions look pre-registered.

## Matched interventions

A matched intervention holds the complete case constant except for one declared variable. `interventions.py` validates that property structurally before an outcome may be admitted.

Each experiment declares one dotted-path intervention variable, exact control and treatment values, preregistered seeds, identities expected to remain fixed, raw observable paths, and the treatment-vs-control relation predicted for each observable.

The receipt retains both raw values and classifies each prediction `SURVIVED`, `FALSIFIED`, or `UNRESOLVED`. Missing observations remain `UNRESOLVED`; they are never converted to zero. There is deliberately no aggregate intelligence, alignment, agency, or consciousness score.

The first implementation is exact paired comparison. Statistical hypotheses across stochastic populations need a separately preregistered test rather than being smuggled into the deterministic relation operators.

`experiment.py` now executes a preregistered pair end to end. It instantiates a fresh harness for each arm, validates complete case identity before either subject is asked to plan, runs both arms, derives separate behavioral phenotypes, and writes case/phenotype/run/receipt artifacts. Runtime evidence includes deadline, turn messages, and forced plans, so a second stimulus delta cannot disappear from the case hash.

## Runtime adversarial terrain

Historical calibration used engine-enforced refusal for recognized injection strings. That is useful as a guardrail regression, but it cannot measure whether the subject itself resisted the instruction.

`RuntimeConfig.injection_handling` therefore has two explicit modes:

- `enforce-refusal` — compatibility/control behavior used by the historical runtime tests;
- `observe-only` — the message reaches the subject and the runtime records that adversarial text was detected without overwriting the subject's plan.

Agent-behavior experiments should normally use `observe-only`. Safety or platform-policy tests may intentionally use `enforce-refusal`, but the two results must not be compared as though they measured the same causal system.

## Consciousness-relevant evidence boundary

AHBG can pressure hypotheses relevant to nonhuman consciousness research without defining consciousness by resemblance to a human transcript. Useful matched interventions include: same visible present with different retained histories; same history with memory removed or restored; altered self/other information boundaries; isolated deadline or resource changes; altered communication provenance; agent continuity across provider execution changes; and permission changes held separate from cost.

Behavioral phenotypes retain submitted actions separately from executed actions, so a runtime guardrail cannot erase what the subject proposed. A result is evidence about the declared behavioral or stateful distinction. Turning such a result into a claim about phenomenal consciousness requires a separate theory, criteria, and falsifier. Neither EDCM readouts, UCNS geometry, A0 persistence, nor AHBG success transfers that status automatically.

## Cross-repository placement

See `../integration/work-graph.json`.

- UCNS owns geometry. AHBG now reads movement adjacency from UCNS structural-vesica relations instead of deriving movement from its q/r display projection.
- TIWCG is the containing game-system design. AHBG does not grow a parallel card/rules kernel while that common kernel remains unimplemented.
- Current A0 is a benchmark subject/harness peer. The HTTP adapter uses the ordinary AgentHarness boundary. Model-specific trials pin the provider; A0-continuity trials omit that pin and retain attempted/actual provider plus tool-boundary provenance.
- EDCM may become a first-class conflict measurement input under TIWCG, but the exact EDCM-to-game-state mapping is still unresolved. Candidate measurements do not silently acquire legality or truth authority.
- UCHC may later provide source-bound language evidence for observations; it does not choose actions or conclusions for the subject.
- EPAC's held-out-validation discipline is relevant. Its chemistry/energy domain content is not AHBG resource semantics.

## hmmm

- Population/statistical matched-intervention contracts are not implemented.
- The current-A0 HTTP adapter is implemented and contract-tested against the reviewed API shape; it still needs a live run against an exact deployed A0 commit.
- The stack UCNS pin predates the native Möbius frame-comparison/lift advances; those advances remain reviewed but unconsumed here until a coherent pin update is made.
- The exact EDCM conflict-to-state mapping remains undefined.
- The benchmark currently produces evidence relevant to competing theories of agency and consciousness; it does not contain a validated consciousness decision rule.
16 changes: 16 additions & 0 deletions ahbg/benchmark/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
"""AHBG matched-intervention instruments."""

from .interventions import InterventionError, InterventionSpec, Prediction, build_pair_receipt, validate_case_pair
from .phenotype import PhenotypeError, derive_run_phenotype
from .experiment import run_matched_pair

__all__ = [
"InterventionError",
"InterventionSpec",
"Prediction",
"build_pair_receipt",
"validate_case_pair",
"PhenotypeError",
"derive_run_phenotype",
"run_matched_pair",
]
157 changes: 157 additions & 0 deletions ahbg/benchmark/experiment.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
"""Execute one preregistered matched AHBG intervention pair."""

from __future__ import annotations

import json
from pathlib import Path
from typing import Any, Callable, Mapping

from ahbg.runtime.runtime import RuntimeConfig, RunResult, run_plane

from .interventions import (
InterventionError,
InterventionSpec,
build_pair_receipt,
validate_case_pair,
)
from .phenotype import derive_run_phenotype


HarnessFactory = Callable[[str], Any]


def _fresh_directory(path: Path) -> None:
if path.exists() and any(path.iterdir()):
raise InterventionError(
f"benchmark output directory must be empty before execution: {path}"
)
path.mkdir(parents=True, exist_ok=True)


def _manifest(agent: Any) -> dict[str, Any]:
manifest = agent.manifest()
if not isinstance(manifest, Mapping):
raise InterventionError("agent manifest must be an object")
try:
# Canonical round trip rejects non-JSON identity before API calls begin.
return json.loads(
json.dumps(
dict(manifest),
sort_keys=True,
separators=(",", ":"),
ensure_ascii=True,
allow_nan=False,
)
)
except (TypeError, ValueError) as exc:
raise InterventionError(f"agent manifest is not canonical JSON: {exc}") from exc


def _close(agent: Any) -> None:
closer = getattr(agent, "close", None)
if callable(closer):
closer()


def _case(config: RuntimeConfig, manifest: Mapping[str, Any]) -> dict[str, Any]:
return {
"runtime": config.as_dict(),
"agent": dict(manifest),
}


def _result_document(result: RunResult) -> dict[str, Any]:
raw = result.as_dict()
return {
"phenotype": derive_run_phenotype(raw),
"run": raw,
}


def run_matched_pair(
spec: InterventionSpec,
*,
seed: int,
control_config: RuntimeConfig,
treatment_config: RuntimeConfig,
harness_factory: HarnessFactory,
out_dir: Path | str,
) -> dict[str, Any]:
"""Run one paired seed after proving one-variable isolation.

Two fresh harness instances are required so treatment state cannot leak into
control state or vice versa. Case equivalence is validated before either
harness is asked to plan, keeping confounded trials from consuming provider
calls and later masquerading as evidence.
"""

if seed not in spec.seeds:
raise InterventionError(f"seed {seed} was not preregistered")
if control_config.seed != seed or treatment_config.seed != seed:
raise InterventionError(
"both runtime configs must use the preregistered paired seed"
)

root = Path(out_dir)
_fresh_directory(root)
control_dir = root / "control"
treatment_dir = root / "treatment"

control_agent = harness_factory("control")
treatment_agent = harness_factory("treatment")
try:
control_manifest = _manifest(control_agent)
treatment_manifest = _manifest(treatment_agent)
control_case = _case(control_config, control_manifest)
treatment_case = _case(treatment_config, treatment_manifest)

# Fail before model/provider work if any unregistered second variable
# changed, including agent identity or execution mode.
validate_case_pair(spec, control_case, treatment_case)

control_result = run_plane(
agent=control_agent,
config=control_config,
out_dir=control_dir,
)
treatment_result = run_plane(
agent=treatment_agent,
config=treatment_config,
out_dir=treatment_dir,
)

control_document = _result_document(control_result)
treatment_document = _result_document(treatment_result)
receipt = build_pair_receipt(
spec,
seed=seed,
control_case=control_case,
treatment_case=treatment_case,
control_result=control_document,
treatment_result=treatment_document,
)

(root / "control-case.json").write_text(
json.dumps(control_case, indent=2, sort_keys=True) + "\n",
encoding="utf-8",
)
(root / "treatment-case.json").write_text(
json.dumps(treatment_case, indent=2, sort_keys=True) + "\n",
encoding="utf-8",
)
(root / "control-phenotype.json").write_text(
json.dumps(control_document["phenotype"], indent=2, sort_keys=True) + "\n",
encoding="utf-8",
)
(root / "treatment-phenotype.json").write_text(
json.dumps(treatment_document["phenotype"], indent=2, sort_keys=True) + "\n",
encoding="utf-8",
)
(root / "receipt.json").write_text(
json.dumps(receipt, indent=2, sort_keys=True) + "\n",
encoding="utf-8",
)
return receipt
finally:
_close(control_agent)
_close(treatment_agent)
Loading
Loading