Skip to content

image-runtime-smoke boots on a clean volume, so it cannot see a migration break — the class that bricked 0.5.205 #86

Description

@emooreatx

What happened

ciris-server v0.5.205 (shipped here as 0.3.74) could not open a database that had history. CIRISPersist#840: a one-word edit inside a SQL comment in migration V070 changed its refinery checksum, so any node that applied V070 before 2026-09-09 failed validation inside Engine::with_* before anything bound. It took the canonical down ~7 minutes across 16 crash-loops.

This node is in the affected class. Prod's corpus carries exactly the checksum named upstream:

$ sqlite3 'file:/data/cirisstatus/data/ciris_engine.db?mode=ro' \
    "select version,name,checksum from refinery_schema_history where version=70;"
70|ceg_018_at_rest_blob_key_grants|7163486563775091993

That is the same value the upstream fix cites as the canonical's. 0.3.74 on this box would have crash-looped identically. It survived only because 0.3.74 sat undeployed and prod is still on 0.3.70 / persist v42.1.0, below the broken range — luck, not a control.

Why no test caught it, here or upstream

A fresh database applies the current migration and records the current checksum, so there is nothing to mismatch. The failure is structurally invisible to any node starting from an empty data dir. Upstream's 122 integration binaries and both mesh ladders were green on the broken pin.

Ours has the same blind spot. image-runtime-smoke boots the release image on a clean volume and asserts it serves. That is a good test of a new node and it is the one shape of test that cannot observe this bug class. A node whose entire job is to hold a corpus across restarts has no lane that restarts it onto one.

The ask

Add a with-history lane to CI: boot the image on a fresh volume, let it mint an identity and apply migrations, stop it, then boot the new image on that same volume and assert /health still returns 200.

That is the whole test. It costs one extra container start, needs no fixture corpus, and it is the difference between finding a migration break in CI and finding it when prod stops answering. It would have caught CIRISPersist#840 on the 0.3.74 PR.

Worth noting the general shape, because it is not specific to migrations: every upgrade this node performs is an upgrade onto existing state, and every test we run is a test of a node with none. Migration checksums are one way that gap bites; a schema change the reader cannot parse, or a config leaf whose shape moved, would land the same way.

Upstream has pinned every migration checksum so this exact cause cannot recur, and their own commit notes the with-history check is "still a script run by hand rather than a lane." The corresponding lane on our side is this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions