What happened
ciris-server v0.5.205 (shipped here as 0.3.74) could not open a database that had history. CIRISPersist#840: a one-word edit inside a SQL comment in migration V070 changed its refinery checksum, so any node that applied V070 before 2026-09-09 failed validation inside Engine::with_* before anything bound. It took the canonical down ~7 minutes across 16 crash-loops.
This node is in the affected class. Prod's corpus carries exactly the checksum named upstream:
$ sqlite3 'file:/data/cirisstatus/data/ciris_engine.db?mode=ro' \
"select version,name,checksum from refinery_schema_history where version=70;"
70|ceg_018_at_rest_blob_key_grants|7163486563775091993
That is the same value the upstream fix cites as the canonical's. 0.3.74 on this box would have crash-looped identically. It survived only because 0.3.74 sat undeployed and prod is still on 0.3.70 / persist v42.1.0, below the broken range — luck, not a control.
Why no test caught it, here or upstream
A fresh database applies the current migration and records the current checksum, so there is nothing to mismatch. The failure is structurally invisible to any node starting from an empty data dir. Upstream's 122 integration binaries and both mesh ladders were green on the broken pin.
Ours has the same blind spot. image-runtime-smoke boots the release image on a clean volume and asserts it serves. That is a good test of a new node and it is the one shape of test that cannot observe this bug class. A node whose entire job is to hold a corpus across restarts has no lane that restarts it onto one.
The ask
Add a with-history lane to CI: boot the image on a fresh volume, let it mint an identity and apply migrations, stop it, then boot the new image on that same volume and assert /health still returns 200.
That is the whole test. It costs one extra container start, needs no fixture corpus, and it is the difference between finding a migration break in CI and finding it when prod stops answering. It would have caught CIRISPersist#840 on the 0.3.74 PR.
Worth noting the general shape, because it is not specific to migrations: every upgrade this node performs is an upgrade onto existing state, and every test we run is a test of a node with none. Migration checksums are one way that gap bites; a schema change the reader cannot parse, or a config leaf whose shape moved, would land the same way.
Upstream has pinned every migration checksum so this exact cause cannot recur, and their own commit notes the with-history check is "still a script run by hand rather than a lane." The corresponding lane on our side is this issue.
What happened
ciris-server v0.5.205 (shipped here as 0.3.74) could not open a database that had history. CIRISPersist#840: a one-word edit inside a SQL comment in migration V070 changed its refinery checksum, so any node that applied V070 before 2026-09-09 failed validation inside
Engine::with_*before anything bound. It took the canonical down ~7 minutes across 16 crash-loops.This node is in the affected class. Prod's corpus carries exactly the checksum named upstream:
That is the same value the upstream fix cites as the canonical's. 0.3.74 on this box would have crash-looped identically. It survived only because 0.3.74 sat undeployed and prod is still on 0.3.70 / persist v42.1.0, below the broken range — luck, not a control.
Why no test caught it, here or upstream
A fresh database applies the current migration and records the current checksum, so there is nothing to mismatch. The failure is structurally invisible to any node starting from an empty data dir. Upstream's 122 integration binaries and both mesh ladders were green on the broken pin.
Ours has the same blind spot.
image-runtime-smokeboots the release image on a clean volume and asserts it serves. That is a good test of a new node and it is the one shape of test that cannot observe this bug class. A node whose entire job is to hold a corpus across restarts has no lane that restarts it onto one.The ask
Add a with-history lane to CI: boot the image on a fresh volume, let it mint an identity and apply migrations, stop it, then boot the new image on that same volume and assert
/healthstill returns 200.That is the whole test. It costs one extra container start, needs no fixture corpus, and it is the difference between finding a migration break in CI and finding it when prod stops answering. It would have caught CIRISPersist#840 on the 0.3.74 PR.
Worth noting the general shape, because it is not specific to migrations: every upgrade this node performs is an upgrade onto existing state, and every test we run is a test of a node with none. Migration checksums are one way that gap bites; a schema change the reader cannot parse, or a config leaf whose shape moved, would land the same way.
Upstream has pinned every migration checksum so this exact cause cannot recur, and their own commit notes the with-history check is "still a script run by hand rather than a lane." The corresponding lane on our side is this issue.