Skip to content

Let deploy recovery reach reservations made by the current stable deploy - #2532

Merged
shiny-code-app[bot] merged 2 commits into
mainfrom
work/recovery-lookup-2531
Sep 27, 2026
Merged

shiny-code-app[bot] merged 2 commits into
mainfrom
work/recovery-lookup-2531

Conversation

@shiny-code-app

Copy link
Copy Markdown
Contributor

Why

RepairShopr (repairshopr-sync / prod) has had every deploy blocked since 2026-09-12 by a fenced reservation from run 34724620086. The owner-approved recovery dry-run (run 36344798857) returned 404 reservation_not_found, so the supported recovery path cannot reach it. The same failure happened on 2026-07-17 (run 29609495343); it was recovered through #2167 / #2187 only because that reservation still used the old key.

Recovery looks a reservation up by its exact Idempotency-Key and the fingerprint of the exact original payload. Since 2026-08-05 the reusable stable deploy builds the key as generic-web-stable-deploy:{product}:{instance}:{artifact_id}:{deploy_reference}:{run_id}:1 and sends deploy.deploy_reference. The recovery action and the workflow's provider-evidence step still rebuilt the older ...:{instance}:{run_id}:{attempt} key without deploy_reference. Neither matched any reservation made after 2026-08-05.

This PR fixes the first finish-line item of #2531 only. Lock expiry, alerting on a held fence, and the ~5-minute trigger remain open there.

What changed

  • .github/actions/generic-web-deploy-recovery-dry-run: builds the current (artifact_scoped) key and sends deploy_reference in the original payload, empty when the request omits it, as the original deploy did. It accepts an optional deploy_reference. Reservations made before 2026-08-05 are still reachable with an explicit deploy_key_format: run_scoped. A new mode: provider-evidence sends the same request to the provider-evidence route.
  • .github/workflows/reusable-generic-web-stable-deploy.yml: both recovery steps now use that action, pinned to 6bac61a in this PR, so the workflow no longer keeps its own shell copy of the key and payload (60 lines removed). The workflow_dispatch recovery path now passes the original deploy_reference.
  • docs/operations.md: documents both key formats and the attempt-1 rule for rerun deploys.

No service code changes. The server already looks up the exact key and fingerprint the client sends.

Safety

Recovery still names exactly one reservation: one full key plus the fingerprint of one full original payload, and the service's exact lookup is unchanged. There is no prefix, wildcard, or format fallback. A wrong format, attempt, or coordinate gets 404 and changes nothing. The key format is chosen explicitly (artifact_scoped by default, run_scoped only on request), and malformed values are rejected before any OIDC or network call. Apply still requires the artifact-backed, digest-bound workflow_run provenance, and provider-evidence mode refuses a digest. What a recovery can release or adopt is still decided server-side by the dry-run classification and the reviewed digest. This PR does not change that.

Test evidence

test_action_reaches_reservation_created_by_current_stable_deploy runs the reusable workflow's real stable-deploy request steps: the key-building shell step, plus launchplane-request with the workflow's own payload and payload-fields. It then requires the recovery action, in dry-run and provider-evidence modes, to send the same Idempotency-Key and an original_deploy identical to the deploy payload. It covers an empty deploy_reference (the RepairShopr case) and a non-empty one. The expected key is never copied into the test.

  • On origin/main's action it fails: generic-web-stable-deploy:example-product:prod:34724620086:1 != generic-web-stable-deploy:example-product:prod:ghcr.io/example/product@sha256:…::34724620086:1, and deploy_reference is rejected as an unsupported field.
  • With this change it passes. The legacy test now pins run_scoped, and a new test rejects an unknown format, run_scoped with deploy_reference, and a padded deploy_reference before any network call.
  • The targeted modules that reference the recovery action, the reusable workflow, or docs/operations.md pass (237 tests). ruff, ruff format and mypy are clean on the changed test. JetBrains inspection (changed files) is GREEN.

Merge note

Land with a merge commit, not squash or rebase, so the pinned action commit 6bac61a stays reachable from main.

After merge

  • No Launchplane service deploy is needed for the dry-run: RepairShopr calls this reusable workflow @main, and the fix is entirely in the workflow and action.
  • Then dispatch a fresh RepairShopr Launchplane Recovery Request with the same inputs as run 36344788668 (original_run_id 34724620086, original_run_attempt 1). It should return a recovery digest instead of reservation_not_found.
  • Apply is separate: RepairShopr's recovery-apply job pins this action at 01198841, which still builds the old key. It needs a product-repo pin bump to a commit containing this fix before a reviewed apply can reach the reservation.

Refs #2531

Recovery rebuilt the pre-2026-08-05 idempotency key and payload, so every
reservation made by the current reusable stable deploy came back as
reservation_not_found. The action now builds the artifact-scoped key and
sends deploy_reference like the original deploy did. The run-scoped key
stays available by explicit deploy_key_format. A provider-evidence mode
lets the reusable workflow reuse this one key builder.

Refs #2531
Pin the reusable workflow's recovery steps to the action that rebuilds
the current stable-deploy key, and replace the workflow's own shell copy
of the provider-evidence key and payload with the action's
provider-evidence mode. The dispatch path now carries the original
deploy_reference. A regression test runs the real stable-deploy request
steps and requires recovery to send the same key and payload.

Refs #2531
@shiny-code-app

Copy link
Copy Markdown
Contributor Author

Review by another model. OpenAI gpt-6-astra, read-only against this branch. The author is Anthropic Opus 5.5. The reviewer found no dry-run-to-apply bypass and no broader reservation matching. It confirmed that required artifact fields, unknown-field rejection, workflow provenance, target checks and digest binding still hold. It raised three medium findings. All three are declined:

  1. Whitespace around coordinates changes the identity. Declined. Stable deploy gets artifact_id and source_git_ref from build outputs, never from typed text, so they carry no padding. Trimming only helps an operator who pastes a value with a stray space. If an original ever had been padded, recovery gets 404 and nothing is released. It cannot pick a different run, because run ID, product, instance and payload are all part of the exact key.
  2. payload-fields JSON-decodes values, so a JSON-looking deploy_reference would fingerprint differently. Declined. The only callers pass no deploy_reference (empty) or an image reference with / and :, and neither parses as JSON. The mismatch would fail closed with 404. The same quirk already applies to source_git_ref on main and is not introduced here.
  3. Pre-2026-08-05 reservations now need an explicit deploy_key_format: run_scoped, which the workflow_dispatch adapter and the product repos' staged requests can't send. Declined for this PR, with the gap noted. The only known pre-August reservation (July, RepairShopr) was recovered in August. Any request for an old reservation now returns 404 and changes nothing. run_scoped is still reachable through the reusable workflow's recovery_request_json input. If a live pre-August fence turns up, wiring the selector into the product request workflows is a small follow-up.

Also noted: deploy_key_format: null is treated as the default format. That is harmless because it still builds one exact key, so I left it alone rather than re-pin the action for it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant