noise-pq: add Noise_XXhfs_25519+MLKEM768_ChaChaPoly_SHA256 spec (Stage 1 Working Draft) - #716
paschal533 wants to merge 6 commits into
Conversation
- Protocol name: Noise_XXhfs_25519+ML-KEM-768_ChaChaPoly_SHA256 - Protocol ID: /noise-mlkem768-hfs/0.1.0 - KEM: raw ML-KEM-768 (FIPS 203), removing X-Wing composite wrapper - Wire sizes: Msg A 1216 B, Msg B 1200 B (was 1248/1232 with X-Wing) - Add Rust reference impl (royzah/rust-libp2p PR libp2p#1) - Add triangle interop results (JS/Python/Rust -- 3/3 PASS, 2026-06-24)
|
Status update on this draft. Since I opened it in June, the evidence base has changed enough that I think it is worth summarising in one place, because the two questions a spec like this normally has to answer — is it implementable from the text alone? and is it affordable? — now have concrete answers rather than estimates. Four independent implementations
Four different ML-KEM-768 libraries, four codebases written separately against the prose in this document. All six pairwise combinations interoperate
Every pair completed a live handshake over TCP. No implementation required a protocol change to interoperate with any other. That is the main thing I wanted to establish about this text: four people reading it separately produced the same bytes on the wire. Where a pair also exchanged application data after the handshake, that covers both transport keys, since The affordability questionThis is the part that has genuinely changed. The early numbers made hybrid key exchange look expensive — +5.0x over classical in JavaScript, +12.9x in Python. Benchmarking the Nim implementation, which is the only one whose KEM is compiled assembly-optimised C rather than interpreted or WASM, gives a different picture:
Same wire protocol, same 2,288 bytes of overhead, same lattice operations. The cost of quantum resistance in this design is not a property of the protocol — it is a property of whether the runtime ships an optimised ML-KEM-768. With a good one the KEM is under a tenth of handshake time and the hybrid handshake costs 14% more than the classical one. The rest is classical crypto that a deployment was already paying for. Two caveats I would rather state than have someone find. nim-libp2p optimises only the KEM — its X25519 and ChaCha20-Poly1305 come from BearSSL, SHA-256 from a pure-Nim library, Ed25519 from a pure-Nim reference port — so its 2.8 ms classical baseline is not a floor, and 1.14x is the ratio that particular composition pays rather than a general figure. And all interop runs so far were on one machine over loopback, so they establish wire-format agreement, not behaviour under loss or reordering on a real path. What I am asking forI am not asking for a merge. Per the spec lifecycle this needs interest group formation and broader review to move past Stage 1, and I do not think that is something I can do from the PR alone. What would help most:
Happy to split, restructure, or rewrite whatever is useful. All four implementation PRs are gated on this document, so any direction at all is more useful to me than the PR continuing to sit. |
|
Correcting the performance table in my previous comment. The direction of the correction matters for this draft, because the earlier figures overstated the cost of the protocol this document specifies. Corrected measurementsAll four implementations, measured the same week, each holding its own cryptographic backend constant:
Previously I reported 12.9x, 5.0x and 1.14x, with Rust unmeasured. Three faults produced the difference, and every one of them inflated the apparent cost:
Why this bears on the draftThe practical objection to a hybrid handshake is usually cost, and the honest answer is now considerably better than the one I gave in June. Three of four independent implementations put the overhead between 1.1x and 1.5x. The outlier is Python at 10.5x, and that is entirely attributable to Rust is a new data point since my last comment. Its benchmark existed in the repository and had simply never been run. It matters because Rust and Nim use entirely different KEM libraries and land at 1.30x and 1.13x, which is what allows the result to be attributed to the protocol design rather than to any single vendor's optimisation work. The wire overhead of 2,288 bytes, which is the figure this draft specifies, is unchanged and is now derived independently by three implementations. Where the draft standsStill open since 25 June with no review. Four implementations are complete, tested, and interoperable in all six pairwise combinations, and none of them can merge until the protocol exists as a specification. The three questions from my earlier comment stand, and I would still value an answer to any of them:
I am happy to restructure, split, or rewrite whatever would help. Silence is the only outcome I cannot work with. |
…r like The implementation migrated from X-Wing to raw ML-KEM-768 in 8585aba, but NOISE_HFS_SPEC.md and benchmarks/results.md were never updated. Both still described X-Wing, including wire sizes of 1,248/1,232/64. The test vectors, the source and libp2p/specs#716 all say 1,216/1,200/64. The spec now states ML-KEM-768 throughout: encapsulation key 1,184 B, ciphertext 1,088 B, ekem1 1,104 B with its tag, and messages of 1,216/1,200/64 bytes, matching test/fixtures/pqc-test-vectors.json. The benchmark reported +4.9x, which timed noise() against noiseHFS(). Those default to different crypto backends, so that comparison changes the KEM and the whole symmetric/DH backend at once and then attributes the difference to the KEM. paired-passes.mjs measures all four cells: like for like, native backend 1.57x (1.51 to 1.61) like for like, pure JS backend 1.16x (1.13 to 1.18) mismatched backends, as before 3.54x The KEM costs 3.45 to 3.68 ms; the backend choice costs 14.11 ms. In JavaScript the post-quantum primitive is not the expensive part. Adds paired-passes.mjs and its raw output, which were untracked.
py-libp2p PR #1310's post-quantum suite (tests/security/noise/pq) now runs 56 tests at 8e12d013: kem 12, noise state 15, patterns 11, transport 8, vector replay 10.
The excerpted vector description now matches the regenerated js-libp2p-noise fixture (a colon instead of an em dash), and the conformance list uses a comma. No normative change.
ChainSafe/js-libp2p-noise#665, libp2p/py-libp2p#1310 and libp2p/rust-libp2p#6481 are open draft pull requests (gh pr view --json isDraft,state), matching the existing Nim row.
|
Retitled for the protocol name change, and updated the document. Name. The profile is now Identifier. Because of that, the draft moves the libp2p identifier from Interop: all four implementations, both directions, 48/48. TypeScript (ChainSafe/js-libp2p-noise#665), Python (libp2p/py-libp2p#1310), Nim (iftech/nim-libp2p#2811) and Rust (libp2p/rust-libp2p#6481 by @royzah, with our harness from royzah/rust-libp2p#1). A neutral runner tried every ordered listener/dialer pairing, each against itself included, three times each, on 2026-09-17. A run passes only if both sides exit cleanly, each reports the other's actual peer id, and one encrypted greeting goes each way, which covers both Negative controls. The TypeScript implementation rebuilt with only the old name fails all six cross-implementation orderings on message B, while same-implementation pairs pass (control A, re-run with the matrix). A dialer printing a fake peer id completes the handshake and is caught by the identity cross-check (control B, from the earlier run Corrections to the description and my 5 September comment.
Benchmarks. One-session re-run on the renamed suite, 2026-09-17: SUMMARY.md. Each harness measures a different transport, so only the within-language ratio is meaningful, and sampling differs by language (details in the file):
Absolute latencies from this session aren't comparable with earlier sessions. The file lists earlier published figures for reference only and computes no ratio or delta against them. |
|
Closing in favour of #727 by @royzah, which specifies the same suite. One spec is better than two, and #727 came with the rust-libp2p implementation. I've offered the test vectors, the four-implementation interop results and the downgrade note over there. The branch stays available if any of it is wanted: paschal533/specs. |
Our spec PR libp2p/specs#716 was closed on 2026-09-18 in favour of royzah's libp2p/specs#727, which is now the single Working Draft for Noise_XXhfs_25519+MLKEM768_ChaChaPoly_SHA256. The README's work table cites #727 and records #716 as closed, with the closed text still readable on the fork branch it was opened from. Also states /noise-mlkem768-hfs/0.2.0 as what the implementations ship rather than as a spec-endorsed identifier: #727 writes 0.1.0 and lists the identifier string as its first open issue.
libp2p/specs#716, this author's own draft, was closed on 2026-09-18 in favour of royzah's libp2p/specs#727, now the single Working Draft for Noise_XXhfs_25519+MLKEM768_ChaChaPoly_SHA256. NOISE_HFS_SPEC.md cites #727 for the raw ML-KEM-768 pattern and records #716 as closed. Both docs now state /noise-mlkem768-hfs/0.2.0 as what this implementation ships rather than as a spec-endorsed identifier: #727 writes 0.1.0 and lists the identifier string as its first open issue.
libp2p/specs#727 (Stage 1A Working Draft, by royzah) is the single spec for Noise_XXhfs_25519+MLKEM768_ChaChaPoly_SHA256. The separate draft libp2p/specs#716 by this author was closed on 2026-09-18 in its favour; NOISE_HFS_SPEC.md records that and links the closed text. Both docs now state /noise-mlkem768-hfs/0.2.0 as what this implementation ships rather than as a spec-endorsed identifier: iftech#727 writes 0.1.0 and lists the identifier string as its first open issue.
…p#727 libp2p#716 was closed on 2026-09-18 in favour of libp2p#727 by royzah, so one document rather than two carries the proposal for Noise_XXhfs_25519+MLKEM768_ChaChaPoly_SHA256. This text is kept for reference; the material it carries was offered to libp2p#727.
Summary
This PR adds a new protocol spec for
Noise_XXhfs_25519+MLKEM768_ChaChaPoly_SHA256, a post-quantum hybrid extension of the classical Noise XX handshake used in libp2p, under the protocol id/noise-mlkem768-hfs/0.2.0.Update (2026-09-17): The profile was renamed from
Noise_XXhfs_25519+ML-KEM-768_ChaChaPoly_SHA256. Noise (revision 34, §8.2) requires each algorithm name to consist solely of alphanumerics and/, so the hyphenated name was never valid; thanks to @royzah for spotting it. The rename is wire-incompatible: the name is hashed into the handshake hash, so old-name and new-name peers fail when the initiator decrypts Message B. The libp2p identifier therefore moved from/noise-mlkem768-hfs/0.1.0to/noise-mlkem768-hfs/0.2.0. Message sizes are unchanged. See §2.1. #727 by @royzah uses the sameMLKEM768name with the identifier0.1.0.Update (2026-06-24, history): The spec was revised from an earlier X-Wing draft to use raw ML-KEM-768 (FIPS 203). The Noise XXhfs pattern already provides classical security through three independent DH operations (
ee,es,se), so using X-Wing (which bundles an extra X25519 inside the KEM) would be redundant, adding 64 bytes of wire overhead with no security benefit. Raw ML-KEM-768 gives the same hybrid security guarantee with a smaller footprint. The Rust implementation, libp2p/rust-libp2p#6481 by @royzah, also uses raw ML-KEM-768.The spec is filed at Stage 1 (Working Draft) per the libp2p spec lifecycle.
e1andekem1HFS token extensions from the Noise HFS draft, applied to the existing Noise XX patternekem1slotMotivation
NIST finalized ML-KEM (FIPS 203) in August 2024. Harvest-now-decrypt-later attacks make forward secrecy the most urgent deployment target. The XXhfs pattern adds quantum-safe forward secrecy with no changes to the classical authentication layer, giving operators a safe incremental migration path.
Reference Implementations
Four implementations, using four different ML-KEM-768 libraries, have been tested against each other:
@noble/post-quantum; 99 tests, 5 deterministic test vectorskyber-py; 56 testsml-kemcrate (RustCrypto)Interoperability matrix (2026-09-17)
Every ordered (listener, dialer) pairing of the four implementations, including each against itself, was run three times over loopback TCP on one Windows 11 machine: 48 runs, 48 passed.
A run passes only if both sides exit cleanly, each side reports the other's actual peer id, and one encrypted transport message goes each way, which exercises both cipher states from the split. The harnesses start the handshake directly on TCP, without multistream-select, so identifier negotiation is not covered. Results,
versions.txtand all 96 logs: pq-noise-artifacts, run 20260917T134954Z. Two negative controls, the TypeScript implementation rebuilt with only the old name failing all six cross-implementation orderings and a dialer printing a fake peer id caught by the identity cross-check: interop/negative-controls.Earlier results, corrected. An earlier version of this description showed "Triangle interoperability (2026-06-24)" (TypeScript ↔ Python, Rust ↔ TypeScript, Rust ↔ Python, all PASS), named royzah/rust-libp2p PR #1 as the Rust reference implementation, and gave Python 68 tests. Those runs were handshake-only: the runner counted a pass when the dialer exited cleanly and printed a peer id, and no transport frames were exchanged. In the Rust pairs Rust only listened, the listener was our harness from royzah/rust-libp2p#1 rather than part of #6481, and its
Cargo.lockpinned royzah'ssnowfork at a 14 June commit that still used the hyphenated name. The Python dialer was a standalone re-implementation of the handshake. The Rust implementation is #6481; the Python test count is 56, the PQ suite at the branch head. Details: artifacts README.Test plan
encryptAndHashbeforemixKeyinekem1)/noise-mlkem768-hfs/0.2.0here,0.1.0in noise: hybrid post-quantum handshake (X25519 + ML-KEM-768) #727)