Skip to content

Fix workload broadcast ordering and record dedup keys only after successful emit - #80

Merged
Noah-Tervalon-Nvidia merged 8 commits into
NVIDIA:developfrom
mkalkere:fix/workload-broadcast-reliability
Sep 28, 2026
Merged

Noah-Tervalon-Nvidia merged 8 commits into
NVIDIA:developfrom
mkalkere:fix/workload-broadcast-reliability

Conversation

@mkalkere

@mkalkere mkalkere commented Sep 14, 2026 •

Copy link
Copy Markdown

Changelog title

Workload broadcast reliability: ordered broadcasts, dedup after emit

Changelog body

  • Inter-node workload broadcasts are now serialized in origin order, so a lifecycle remove can no longer overtake its own upsert on a peer and resurrect a ghost workload.
  • Dedup keys are recorded only after the broker emit succeeds; a failed emit's retry is no longer swallowed as a duplicate.

Bumps

  • services: patch
  • nvpair-cluster-manager: none
  • nvpair-engine-manager: none
  • nvpair-errors: none
  • nvpair-job-scheduler: none
  • nvpair-manual-nodes: none
  • nvpair-node-info: none
  • nvpair-node-scanner: none
  • nvpair-node-settings: none
  • nvpair-proxy: none
  • nvpair-tui: none
  • nvpair-ui-broker: none
  • nvpair-workload-manager: patch

Description

Fixes two reliability bugs in services/nvpair-workload-manager that could corrupt the scheduler's view of cluster load:

  1. Inter-node workload broadcasts are now serialized in origin order. Every frame was fanned out in its own goroutine, so a lifecycle remove could overtake its own upsert on a peer and resurrect a ghost workload -- phantom load in the scheduler. A single ordered worker (broadcastCh + broadcastLoop) now emits frames in order; a full queue drops with a warning and the heartbeat/backfill re-syncs.
  2. Dedup keys are recorded only after the broker emit succeeds. Previously the key was recorded before emit, so a failed emit returned 500 and the retry was swallowed as a duplicate -- the event was lost with no repair path. dedupIndex.seen()/add() are now split and the key is added only on success.
sequenceDiagram
    participant Origin as Origin broker
    participant Sender as Origin workload manager
    participant Receiver as Peer workload manager
    participant Peer as Peer broker and workload store

    Note over Origin: Example 1 - messages arrive out of order
    Origin->>Sender: workload:started for workload 7
    Sender->>Sender: Start goroutine A to broadcast started
    Origin->>Sender: workloads:remove for workload 7
    Sender->>Sender: Start goroutine B to broadcast remove
    Note over Sender: Goroutine B finishes first
    Sender->>Receiver: workloads:remove for workload 7
    Receiver->>Peer: workloads:remove for workload 7
    Note over Peer: Workload 7 is absent
    Sender->>Receiver: workload:started for workload 7
    Receiver->>Peer: workloads:upsert for workload 7
    Note over Peer: Workload 7 now appears again

    Note over Origin: Example 2 - a failed emit loses a retry
    Origin->>Sender: workload:started for workload 8
    Sender->>Receiver: broadcast started for workload 8
    Receiver->>Receiver: Record workload 8 dedup key
    Receiver->>Peer: Try to emit workloads:upsert
    Note over Receiver: Write to broker fails
    Receiver-->>Sender: HTTP 500
    Sender->>Receiver: Retry started for workload 8
    Receiver->>Receiver: Key exists, so skip the emit
    Receiver-->>Sender: HTTP 200
    Note over Peer: Workload 8 was never added
Loading

Scope

Changes are confined to services/nvpair-workload-manager: broadcast ordering, per-key deduplication, focused tests, and the component spec. Frozen-node handling remains outside this PR (issue #6). Component versions are declared above; services/versions.json is updated by release automation.

Validation

  • go test -count=1 ./... and go vet ./... passed in services/nvpair-workload-manager on Windows with Go 1.27.0.
  • The broadcast test sends 25 frames through the cluster mTLS path and checks their arrival order.
  • A regression test forces a snapshot/removal interleaving and verifies that the peer ends with the workload removed.
  • Dedup tests cover failed writes and retries, concurrent duplicates, independent keys, and successful-key eviction. Focused dedup tests also passed across 10 repeated runs.

Risk

  • The bounded outbound queue drops new frames with a warning when full. Heartbeat or backfill may repair a missed tracked lifecycle state, but removals are not replayed and a dropped removal may leave stale peer state.
  • A slow peer can delay later frames to healthy peers while the ordered worker finishes the current broadcast round. Delivery of the current frame to other peers remains concurrent.
  • HTTP 200 confirms that the receiver wrote an event to its broker or deduplicated it; it does not confirm that the broker applied the event.

Checklist

  • I have read the Contributing Guidelines.
  • Every commit is signed off (git commit -s), certifying the Developer Certificate of Origin.
  • New or existing tests cover the change.
  • Relevant documentation is updated.
  • I checked the diff, changed filenames, and commit messages for credentials, private data, internal URLs, internal issue identifiers, and generated artifacts.
  • I recorded the validation commands and results above.
  • I declared the component bump in the pair-release-intent:v1 block above (services/versions.json is automation-managed; CI rejects hand edits), and described user-visible changes so they reach the release notes.

@Noah-Tervalon-Nvidia
Noah-Tervalon-Nvidia changed the base branch from main to develop September 21, 2026 21:56
Replace the per-frame goroutine fan-out with a single ordered broadcast
worker draining a bounded queue, so a remove can never overtake the
lifecycle upsert it follows and resurrect a ghost workload on peers.

Signed-off-by: Mallikh Kaula <mallikh@users.noreply.github.com>
Split the dedup check from the record: a failed broker emit answers 500
without recording the key, so the peer retry is emitted instead of being
swallowed as a duplicate. Applies to lifecycle upserts and removals.

Signed-off-by: Mallikh Kaula <mallikh@users.noreply.github.com>
Ordering diagrams for the serialized broadcast worker and the
dedup-after-emit sequence, plus a reading-order entry in the README.

Signed-off-by: Mallikh Kaula <mallikh@users.noreply.github.com>
@mkalkere
mkalkere force-pushed the fix/workload-broadcast-reliability branch from d9e390f to 53b844c Compare September 26, 2026 19:33
@mkalkere

Copy link
Copy Markdown
Author

Rebased onto current develop — the old branch was sitting on a pre-restructure base and couldn't merge anymore. The only rebase fallout was the hand-rolled versions.json bump, which I dropped: versions are automation-managed now, so the bump is declared in the release-intent block at the top of the description instead. Code and tests are unchanged; go test -race passes on the workload-manager module.

@kjlubick

Copy link
Copy Markdown
Collaborator

I'm currently reviewing this change. Overall, this looks like a change we want to address those two bugs (I added a mermaid sequence diagram to the description to help explain them better). I'll upload a new commit with my feedback and then we'll get this merged ASAP

Signed-off-by: Kaylee Lubick <klubick@nvidia.com>
Signed-off-by: Kaylee Lubick <klubick@nvidia.com>
Signed-off-by: Kaylee Lubick <klubick@nvidia.com>

@kjlubick kjlubick left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. I folded your doc changes into spec.md and applied some general style fixes to the tests you provided. I also noted an existing race condition related to the snapshots and fixed that.

Signed-off-by: Kaylee Lubick <klubick@nvidia.com>
Signed-off-by: Kaylee Lubick <klubick@nvidia.com>

@Noah-Tervalon-Nvidia Noah-Tervalon-Nvidia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I really like these fixes. Thanks for updating this! LGTM.

@Noah-Tervalon-Nvidia
Noah-Tervalon-Nvidia merged commit 691372c into NVIDIA:develop Sep 28, 2026
13 checks passed
pair-release-intent Bot added a commit that referenced this pull request Sep 28, 2026
Apply release intent from PR #80.

Applies-PR: #80
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants