feat(testing): CSD flows run on the five-platform matrix; state: now asserts something - #97
Conversation
A CSD reaches `testable` when its flow runs on the matrix (CSD.md §1). Nothing ran one here: flow_spec.py was vendored, but no leg drove a flow. - testing/flows/: the home for flows. Each names its CSD (`csd: CSD-005`); the runner reads that CSD's `shows:` (field ids for `relation`) and `states:` (tags for `state:`). A missing or unparseable CSD, or a flow naming a `proposed:` tag, is a load error before anything boots. - flow_spec.py local delta (VENDORED.md): the `csd:` key, CSD binding at load, and `state:` actually checked. Upstream's body was `pass`. - run_flows.py: pass / fail / refused (by `client:` floor) / cannot-start. Refused stays green; cannot-start is red, since with the floor met it means the flow silently never ran. - run_platform --flows: after a clean smoke walk, in the same app, one sign-in (session_fixture). Flows load before bring-up. All five legs in five-platform-live-qa.yml pass it, and PyYAML is installed per job. - Seeded flow: people.yaml (CSD-005, client >=0.5.224). On a bare node it checks landing with the add card, a search with no match (state: empty), and a cleared search. - testing/test_flows.py (35 tests), wired into build.yml. Every rule was mutation-checked red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
…es what's on screen when a step won't advance The first iOS run stopped at "wizard did not advance past 'you'": the four fields were typed back to back, and iOS needs ~2 s per field for the value to reach the ViewModel's StateFlow (client/CLAUDE.md). A step that still won't advance now reports the tags on screen, so a required field the fixture doesn't fill shows in the failure. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
|
First live run (36199588428): the People flow (CSD-005) passed 3/3 on Linux, Android, macOS and Windows, with a screenshot per step. On iOS it never ran: the session fixture stopped at "wizard did not advance past 'you'". The fixture typed four fields back to back, and iOS needs about 2 s per field for the value to reach the StateFlow (client/CLAUDE.md). Fixed: it now settles between fields, and a step that won't advance reports what's on screen. main is merged in; a fresh full run is dispatched. 🤖 Generated with Claude Code |
…start PR #97 (`feat/csd-flows-on-the-matrix`, open) defines the shape a flow here must have, and it is stricter than main's vendored `flow_spec.py` in three ways. Two were easy to meet; the third decides where these files live. 1. `csd:` is required and BINDING — it resolves to the one FSD/CSD/CSD-NNN-*.md and the runner reads that card's `shows:` and `states:`. 2. Naming a `proposed:` tag in `requires`, `expect` or `do` is a LOAD ERROR, as is `state: X` whose tag the CSD still marks proposed. Run against #97's real binder, 14 of the 15 flows loaded clean and one did not: `csd-090-duty-conferral` asserted `state: empty`, whose tag CSD-090 correctly marks `proposed:duty_source_absent` — the screen renders "this node knows no accord family yet" through `duty_source_error`, so the one state that is not a failure is the one that looks most like one. The step now asserts which rows are absent and says in a comment why it may not claim more. 3. THE RUNNER DOES NOT NAVIGATE, and `cannot-start` is RED on purpose. Sign-in lands on Contacts; every other surface is unreachable. Landing 13 flows that each stop at their first `requires` would turn all five matrix legs red for a reason that has nothing to do with the client. So `csd-005-people.yaml` is DELETED — #97 already ships `testing/flows/people.yaml` for CSD-005, and two flows for one CSD is the drift this repo exists to measure — and the other 14 move to `testing/flows/drafts/`. `flow_spec.discover` globs `p.glob("*.yaml")` and not `rglob`, so a subdirectory is staged and never run; `discover(["testing/flows"])` returns [] here, confirmed against #97's own code. Each draft is complete, loads clean, and promotes by `git mv` alone — the hop is never written into a flow, it is derived from nav_map (CSD.md §2.0). Six of the fourteen wait on navigation and nothing else: CSD-025, 036, 057, 066, 068 and 087. Navigation in the runner unblocks more flows at once than any tag or route in this audit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
…t flows
`test_every_tag_a_seeded_flow_names_exists_in_the_client` builds its literal set
with `re.findall(r'"([a-z][a-z0-9_]+)"', ...)`. There is no `$` in that character
class, so a tag the client composes by interpolation is invisible to it.
Seven such tags are named by four of the staged flows, and all seven are real at
run time: `age_band_adult` (`"age_band_$token"`, SetupScreen.kt:2661),
`trace_consent_yes`/`_no` (`"trace_consent_$token"`, :1095), `radio_cohort_self`
(`"radio_cohort_$value"`, ClaimNodeScreen.kt:383) and
`chk_duty_box_{moderate,review,consent_revocation}` (`"chk_duty_box_$verb"`,
DutyConferralScreen.kt:301). Promoting csd-082, csd-083, csd-085 or csd-090 as
written would redden that test against a client carrying every tag it names.
The fix belongs in the test: keep the part before the first `$` of any literal
that contains one, and accept a tag that starts with such a prefix.
`opt_run_with_ai` is the control — written as a whole literal
(SetupScreen.kt:2567) and seen today.
Recorded in testing/flows/drafts/README.md rather than worked around in the
flows, because a flow bent to fit a test's blind spot stops describing the app.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
…g check sees interpolated tags
Sign-in lands on Contacts, so a flow for any other surface ended
cannot-start. Before step one, run_one now walks nav_map's derived hop
for the flow's first `requires: screen:` on this build (node or agent,
from /state). It waits for each tag, then clicks it.
- A missing hop tag is cannot-start and names the tag and its position.
- A screen with no hop that is not flow-only is cannot-start: "no nav
hop for Screen.X".
- A flow-only screen (screen_atlas.flow_only) is waited for, not walked to.
test_every_tag_a_seeded_flow_names_exists_in_the_client read only whole
literals, so interpolated tags ("age_band_$token", "trace_consent_$token",
"radio_cohort_$value", "chk_duty_box_$verb") looked absent. A tag now
counts if it equals a literal or starts with the pre-`$` prefix of an
interpolated one. Only prefixes with two segments count, so "btn_$x"
cannot vouch for every button.
Every new rule was mutation-checked red.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
…-matrix # Conflicts: # .github/workflows/build.yml
…reviewed, the 0.5.218 wait Where we are: 72 CSDs on main (64 building, 7 sketched, 1 envisioned), the new Network, 0.5.218, communities and key-verification cards, and the cards retired by the dedup rule (001 into 055, 105/106 into 039, 107/108 into 054/020, 034 unused). Merged since #90 in three shapes: cards and fixes, the checks that keep CSDs honest (route gate 121/26 -> 14/23), and the two integration merges #115 and #124. Three PRs in review (#77 conflicting, #97, #122); batch 2 of the gap-closing review is next. The dependency graph stops pretending the circles are a chain: B5, B6 and B7 are partly built ahead of B3 and B4, so the graph shows what still has to come before what. CIRISAgent#1213 is closed and removed; the 0.5.218 cut takes its place as the hexagon four merged cards wait on. Persist#907 and #910 are closed; the households hexagon now names #686/#687/#916. Upstream is regrouped by lane (people/chats/files, consent/data, households/communities, accord/trust root, node ops, agent surfaces, setup/identity) from the open issue lists in the four upstream repos. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…clickable before clicking it On iOS the typed fields reach the ViewModel a beat late, so Next is on screen but disabled, and a disabled control has no click handler: the click was a bare 404 that named no step. Wait (bounded) for can_click, and if it never enables say which step and what was on screen. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…ore Next, and retries once JOIN_FEDERATION advances only when traceConsentAnswered; a Next clicked in the same instant as the answer raced it on Windows and the step reported 'did not advance' with the question still on screen. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…k=false — retry, bounded, then name the step The iOS automation tree omits canClick when false, so the clickable wait passes; the click then 404s with 'No click handler'. Treat that answer as 'not yet' for up to 30 s and, if it stays, say which step and what was on screen instead of a bare 404. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…rge left the second on its own line, run as a command) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
… asks for it On iOS the YOU step keeps Next disabled until the federation-ID label is valid (desktop mints the identity itself); the fixture never filled it, so the walk stopped at 'you'. The bounded wait added before named it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…capability Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
… and the gate driver raises on success:false
iOS answered a failed text input with HTTP 200 {success:false}; the driver
checked only the status, so every field the session fixture typed on iOS
silently took nothing and the wizard's Next stayed disabled. Desktop
already answered 404. Android had the same 200. The driver now treats a
success:false body as a failure whatever the status (test red on the old
driver, green now).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…e typing or clicking it On the iPhone the first-run wizard is taller than the screen: the account fields sit below the age band and the AI choice, and /input refuses a field the person could not see (CIRISClient#33). With #97's 404-on-failure fix the refusal finally surfaced. Scroll down in steps, bounded, on an off-screen refusal; everything else raises as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…led past Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…reports what each scroll answered Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…tep, before calling a wizard field unreachable A phone with the keyboard up has a small viewport; a fixed four steps down turned around before reaching the device-name field. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…rd asks The legs run against a bare node with no LLM; iOS asks the question and the walk stopped at the AI step, whose Next waits for a usable LLM choice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
… of stopping on CircleTab nav_map drops the row hop for a one-card tab because the wide layout opens it directly; phones list it first, so iOS landed on CircleTab. Open the only nav_epistemic_* row; with several, name them instead of guessing. Test added. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
iOS landed on a People list with the community roster and Contacts: the compact layout was still on Neighbours after the circle hop. Pick the row named for the target screen when the list has it; otherwise say which rows were there and ask whether the circle hop applied. Test added. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…-matrix # Conflicts: # client/VENDORING.md
…g main Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…ll marks proposed CSD-005's trust chip shipped on main (people review), so the test that named it went red for the product doing its job. Derive the tag instead (the same fix the two-node branch made). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
CSD flows now run on this repo's five-platform matrix. That's what a CSD needs to reach
testable(CSD.md §1: "floor flips off unreleased; flow runs on the matrix"). Until now the runner (testing/gate/flow_spec.py) was vendored here with no flows and no step that ran it; CIRISAgent's gate ran the only four flows (CSD-001–004).How a flow links to its CSD
testing/flows/*.yaml; each flow namescsd: CSD-0NN, which must match exactly oneFSD/CSD/CSD-0NN-*.md.shows:(field id → tag, sorelationworks on field ids) andstates:(state → tag).check_csd_v3's own pattern, so the checker and the runner can't disagree.proposed:tag inrequires,expectordo(a flow can't assert a tag that doesn't exist); astate:whose tag is missing or proposed; arelationoperand that isn't ashows:field; a flow with nocsd:.A fix in the vendored runner:
state:did nothing, so everystate:assertion passed. It now requires that state's tag on screen and every other state's tag off screen. Recorded inVENDORED.mdas a local delta, with thecsd:key, both to go upstream to CIRISAgent.Runner (
run_platform --flows testing/flows, orrun_flows.pystandalone):client:floor is above the client under test; reported, leg stays green) · cannot-start (floor met but the first precondition never held; red, because a flow that never started silently never ran) · fail (red).reports/<leg>.jsonunderflows, screenshots toshots/flows-<leg>/.five-platform-live-qa.yml.Seeded flow:
people.yaml(CSD-005,client: ">=0.5.224"): landing on Contacts with the add card, a search that matches nothing reachingstate: empty, clearing it bringing the add card back. Every tag is on main, and a test checks each exists in the client source.Checks: 35 new tests (168 in
testing/); every red path planted (11 defects, table in the commit) ·packaging/gates.sh0 · nothing underclient/changed.Live proof: the five-platform run dispatched on this branch; its result is linked in a comment below.
Known risks on a live run: if the app shows Add Federation ID after sign-in, or doesn't land on Contacts within 30 s, the result is cannot-start. The session fixture has never run in CI before. Flows for other surfaces will need
nav_maphops.🤖 Generated with Claude Code
https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK