Skip to content

feat(core): a configurable client lease, and a job on its last attempt survives a lost lease - #44

Merged
giraffesyo merged 1 commit into
canaryfrom
feat/lease-ttl
Oct 2, 2026
Merged

giraffesyo merged 1 commit into
canaryfrom
feat/lease-ttl

Conversation

@giraffesyo

Copy link
Copy Markdown
Member

Summary

Two changes to what happens when a client cannot renew its lease, both for programs whose jobs are long calls to another system.

Config.LeaseTTL. The client lease was fixed at 15 seconds, renewed every 5. A database outage longer than that cancels every running job on the client, which is right for short jobs and expensive when a job is forty minutes into a cloud operation. The lease length is now a config field (default unchanged, minimum one second), renewed three times per TTL. It is per client, so each program picks its own; the rescuer only reads expires_at, so clients with different leases can share an install. A longer lease delays the rescue of a crashed client's jobs by the same amount.

A job on its last attempt is no longer cancelled when the lease is lost. Fencing cancels running jobs so that an attempt cannot overlap the one another client starts after the rescue. A job with no attempts left is never run again (the rescuer discards it), so there is nothing to overlap and the cancel only threw its work away. Such a job now runs under the work context, which only a hard stop cancels, instead of the lease generation. Jobs with attempts left behave as before. This is not a setting: I could not find a case where cancelling the last attempt is the better outcome.

One consequence, documented in the plan and the operations guide: if the leader's rescue lands before the job's own result, the job stays dead-lettered as client lost although its work finished. Recording the late result would mean letting a finalize overwrite a rescue, which changes the finalize fence, and is left out of this change.

Testing

  • TestFencedClientLetsALastAttemptFinish: a one-attempt job is running when its client's lease is expired behind its back; the client re-registers, the job's context is not cancelled and it runs once. With the change reverted the test fails (the last attempt was cancelled when the lease was lost).
  • TestLeaseTTLSetsTheLeaseAndItsRenewal and a new validation case for a lease under a second.
  • Full suite with -race against Postgres 18, and golangci-lint, pass. The existing fencing test, which covers a job with attempts left, is unchanged and passes.

No SQL changed, and nothing on the insert, claim or finalize paths, so no hopperbench run.

Checklist

  • make check passes (core module lint and all tests run locally)
  • No cryptography was added (tests run with GODEBUG=fips140=only)
  • New or changed SQL in the claim, finalize, rescue or leader paths has a concurrency or chaos test (no SQL changed; the fencing change has its own test)
  • Behavior changes are reflected in docs/PLAN.md

@giraffesyo
giraffesyo merged commit 23abd36 into canary Oct 2, 2026
9 checks passed
@giraffesyo
giraffesyo deleted the feat/lease-ttl branch October 2, 2026 14:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant