Skip to content

Chat: a contact added while reachable_nodes=0 keys a pair room whose bodies read not_granted forever #699

Description

@emooreatx

Measured 2026-09-29 with harness/native (Docker-free two-node mesh under the synthetic trust root, chore/adopt-edge-v33 = 0.5.218 + edge v33.1 / persist v51.1), production routes only. Reproduced 3 of 3 times; the fix-by-ordering passed 2 of 2.

Sequence. Two people, each on their own node. Claim, announce, self-key-record + POST /v1/federation/peering both ways, then POST /v1/contacts {key_id: <peer's owner key>} ~11 s after peering. The reply is 200 with reachable_nodes: 0: the peer's owner→node binding has not yet reached this node at federation scope (the peer's announce was seconds earlier). Both open the pair room (POST /v1/chat); it keys in ~15 s; the peer sends.

What the recipient sees. The message row arrives, the body blob is pulled (blob_swarm::pull … outcome=Stored), and the transcript row is status: live, unopened_reason: not_granted, unopened_detail: "not granted: this viewer holds no key for <sha>". Still so after 5+ minutes; nothing re-issues the grant. The recipient's log shows one key_grant set admitted … axis=Epoch{…, epoch: 0} wraps_offered=1 wraps_written=1 from the sender — the epoch-0 set — and nothing for the epoch the room advanced to when the joiner was added. (Not a rooting effect: rooted_with says no valid root in common in the passing runs too.)

Same flow, one change: re-POST /v1/contacts until reachable_nodes >= 1 (took 4 tries, ~12 s) before opening the room → keyed at 30 s, body on the other node 2 s after send, transcript shows the text. Whole run 33 s.

Ask (either closes it):

  1. POST /v1/contacts refuses by name while the guest resolves to no node (contacts.not_yet_reachable, 409), so a client retries instead of building a room it cannot read; or
  2. the room handshake re-issues the current epoch's key grants when the member's binding lands (the reconciler already sees it land).

CIRISClient's testing/gate/two_node.py waits for the peer's owner key to be known, which is earlier than the binding, so its CSD-091 flow sits in this window (CIRISServer#698 discussion). Repro: python -m harness.native chat --direct --reachable-wait 0 --binary <test-anchor ciris-server> (fails); --reachable-wait 120 (passes).

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions