feat(cli): cut the menu to 12 flags — opinionated surface, --help-all for the rest (D18) - #110
Conversation
Five places where ArcServe's command line told the user something that was not true, all found by a full inventory of the surface (131 clap args across two binaries, 70 ARC_* env vars read by Rust). 1. `arc` printed "TurboQuant 3.5-bit KV cache compression (lossless, default)" as a fixed banner before any model was loaded. It is false for every MLA model and every head_dim != 128 — nearly all of them, as arc-cli's own module doc already said — and "lossless" has never been measured. The banner now states identity only. 2. Nothing reported the *resolved* state of any subsystem. Added an ArcServe startup summary logged after load, once every auto-fallback has run: resolved KV cache type + block count, prefix-cache state and why it is off, max-seqs. It reads the value back out of the loaded pipeline rather than restating what was requested. 3. `--mcp-port` / `--mcp-config` parsed and did nothing — MCP was wired in the old `mistralrs-server` binary and never carried across to the unified CLI. They now fail loudly instead of yielding a server with no tools. 4. `--format` claimed "Auto-detected if not specified". There is no detection; every consumer does `unwrap_or(ModelFormat::Plain)`, so passing only `-f model.gguf` silently loaded the plain path. 5. `run_ppl.sh --sinkhorn-ab` toggled ARC_FUSED_SINKHORN, which nothing reads — the engine reads ARC_NO_FUSED_SINKHORN and fused is the default. Both arms ran fused-on, so the A/B compared fused against itself. Any "bit-identical" conclusion from the old script is void. Deleted: ARC_TD_MOE_CALIBRATION, the one provably dead env var — parsed, defaulted to 256, threaded through two signatures, bound to `_calibration_set_size` and never read. Removed from both signatures and all call sites; the env var and `--td-moe-calibration` now warn for one release rather than being dropped silently, since the flag is scripted. Hidden (still functional, off `--help`): `--tgt-non-granular-index`, `arc bench --mock` (emits synthetic numbers in a real-looking artifact), `arc validate --o-proj`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uns do not Checked the provenance before retracting anything, and the retraction would have been wrong. The s2 "bit-identical ppl + token-identical 6/6" entry in FACTS.md was measured BEFORE commit 9387e2b (2026-08-13), which flipped the gate from opt-in ARC_FUSED_SINKHORN to ARC_NO_FUSED_SINKHORN and made fused the default — after that measurement and because of it. Two confirmations the toggle was live at s2: the same harness returned a NEGATIVE result at s1 ("ppl drift + 4/6 token divergence"), which a tautological A/B cannot produce, and wave3-H then fixed bit-identity, which s2 re-verified. The breakage is forward-looking: from 2026-08-13 until this fix the script compared fused against fused. No such run is recorded, so no published claim is affected. Narrowed the script comment and PR claim accordingly. Also recorded in FACTS.md, next to "twin-seed ensemble ppl" in Known-unmeasured: ARC_QTIP_ROTATION_SEED has no Rust reader, so the twin-seed ensemble's two bakes would be bit-identical and the ensemble would average a distribution with itself — a guaranteed null that would read as the scientific conclusion "ensembling doesn't help". Wire the seed before spending GPU time. Cleared as non-issues: ARC_QTIP_EXPERT_GREEDY/_VITERBI have Rust readers; ARC_FORCE_GPU_QTIP_QUANTIZE was removed deliberately (12527af); ARC_DISABLE_YARN_STD is an unmerged snippet in a doc, real var is ARC_YARN_ON_STANDARD_LAYERS with inverted polarity. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous pass hid 3 flags out of ~110 — a re-labelling, not a cut. This
is the cut. `serve --help` is now 42 lines and 12 flags, grouped Model /
Hardware / Serving / Diagnostics:
Model -m/--model-id, --isq, --from-uqff, -c/--chat-template
Hardware --cpu, -n/--device-layers
Serving -p/--port, --host, --ui, --max-seqs, --pa-context-len
Diagnostics -l/--log
Everything else is hidden, still parsed, still working, and listed by
`--help-all` on any subcommand (46 flags for serve). Nothing is deleted and no
behavioural default changes in this pass.
The test applied to each flag was "would a competent operator set this, and
would getting it wrong hurt them" — not "could someone conceivably". Flags that
let a user silently make things worse were hidden hardest, and each carries a
doc comment saying why:
--max-seq-len / --max-batch-size device-map planning hints universally
misread as the context limit; --pa-context-len is the real one
--pa-cache-type overriding the resolved KV type disables prefix caching
and turns the safe auto-fallback into a hard error (attribute only —
the doc text is owned by the TurboQuant chain)
--dtype, --arch silently degrade or are silently dropped per model path
--no-kv-cache, --prefix-cache-n 0 look like tuning, read as a mystery
throughput regression
--gqa wrong value yields garbage tokens, no error
--mtp-depth warns and falls back, so it can appear to work
Also fixes the same false claim the runtime banner carried: `arc --help`
still said "Defaults to TurboQuant 3.5-bit KV cache (lossless)", which is
decided per model at load time and cannot be stated in static help.
--help-all is implemented by recursively un-hiding a cloned clap Command
(clap 4.6 mut_args/mut_subcommands), descending to the named subcommand so
`arc bench --help-all` shows bench's flags rather than the root's.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 72 24645 17702 4215 2728 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 74 4600 4597 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 144 14955 12317 810 1828 Shell 26 6312 4307 1564 441 Plain Text 4 3801 0 2479 1322 TOML 33 1488 1293 45 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 197 39646 0 30479 9167 |- BASH 72 1652 1200 331 121 |- C 3 17 17 0 0 |- CUDA 2 84 56 16 12 |- JSON 18 708 708 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 65 2048 1713 77 258 |- TOML 6 207 164 0 43 |- YAML 4 38 33 5 0 (Total) 45409 4679 31021 9709 ───────────────────────────────────────────────────────────────────────────────── Rust 668 318690 275524 15018 28148 |- Markdown 484 26172 471 22580 3121 (Total) 344862 275995 37598 31269 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1295 469159 339700 79625 49834 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
…face # Conflicts: # arc-cli/src/main.rs # mistralrs-cli/src/args/paged_attn.rs
|
Proven causally, not inferred from the step name. On this PR's head Same tree, one word different. That is the whole delta. (The probe was reverted immediately — this branch is untouched at The underlying break is #118 + #115: #118 added a fixture building This PR's own content was verified green before I pushed it — On those conflicts, since a reviewer will want to know what I decided and why: this PR's own comment says the Next step: #131 lands → I re-absorb master here → lanes re-run clean. No action needed from anyone. |
|
Cleared: the Sequence of what happened here, because the stale status is exactly the thing that gets quoted an hour later by someone who wasn't in the thread:
🔑 The lesson I'm recording on the PR rather than just in my own notes: cancelling CI does not merely discard information — it WRITES a false negative into the PR's permanent check record. I had told the coordinator "nothing informative was destroyed", which was true about the runs and false about the record. Now: master is green ( Judge this PR on the new run. If you are reading the old aggregate, it is describing my cancellation, not this code. |
Full inventory of ArcServe's command-line surface: 131 clap args across
mistralrs-cliandarc-cli, 70ARC_*env vars read by Rust, plus the silent default constants. Tiering below; this PR ships the unambiguous half and leaves one default alone deliberately.The five things that were not true
arcbanner claimed a subsystem that usually was not running. It printedTurboQuant 3.5-bit KV cache compression (lossless, default)as a fixed string before any model was loaded. False for every MLA model and everyhead_dim != 128— arc-cli's own module doc already said "in practice almost no model runs TurboQuant today, and none has been measured with it" — and "lossless" has never been measured at all. Banner now states identity only.Nothing reported the resolved state of anything.
serveprinted two lines and named no subsystem. Added an ArcServe summary logged after load, once every auto-fallback has run: resolved KV cache type + block count, prefix-cache state and why it is off, max-seqs. It reads the value back out of the loaded pipeline rather than restating the request — that distinction is the whole point.--mcp-port/--mcp-configparsed and did nothing. MCP is wired in the oldermistralrs-serverbinary and was never carried across to the unified CLI (commands/serve.rsonly readsserver.ui/host/port). A user got a clean startup and no tools. Now a hard error.--formatclaimed "Auto-detected if not specified". There is no detection — every consumer doesformat.unwrap_or(ModelFormat::Plain), so passing only-f model.ggufsilently loaded the plain path.run_ppl.sh --sinkhorn-absilently became a tautology on 2026-08-13. It toggledARC_FUSED_SINKHORN; commit9387e2bc5flipped the gate toARC_NO_FUSED_SINKHORNand made fused the default. From that commit until this PR,env -u ARC_FUSED_SINKHORNdisabled nothing and both arms ran fused-on.Scope, checked rather than assumed: this does not invalidate the s2 "bit-identical ppl + token-identical 6/6" entry in
FACTS.md. That ran before the flip, on the opt-in gate the script drove correctly — and the flip happened because of it. Two independent confirmations the toggle was live at s2: the same harness returned a negative result at s1 (Sinkhorn fused: REJECTED — ppl drift + 4/6 token divergence), which a no-op A/B cannot produce, and wave3-H then fixed bit-identity, which s2 re-verified. Only runs dated between 2026-08-13 and this fix are meaningless, and none is recorded. Scope written intoFACTS.mdnext to the entry.One live landmine found alongside it:
ARC_QTIP_ROTATION_SEEDhas no Rust reader, butensemble_ppl.pydocuments it as the only thing distinguishing bake B from bake A. The twin-seed ensemble would therefore average a distribution with itself — a guaranteed null that reads exactly like "error patterns are correlated, ensembling doesn't help". Twin-seed is in Known-unmeasured, so nothing published rests on it; recorded there as a blocker before any GPU time is spent. Cleared as non-issues:ARC_QTIP_EXPERT_GREEDY/_VITERBIhave Rust readers;ARC_FORCE_GPU_QTIP_QUANTIZEwas removed deliberately (12527af2d);ARC_DISABLE_YARN_STDis an unmerged snippet in a doc (real varARC_YARN_ON_STANDARD_LAYERS, inverted polarity).Tiers — the actual cut
serve --helpis 42 lines, 12 flags, grouped. Everything else is hidden, still parsed, still working, and listed by--help-allon any subcommand.--help-all)Kept — the menu:
-m/--model-id,--isq,--from-uqff,-c/--chat-template--cpu,-n/--device-layers-p/--port,--host,--ui,--max-seqs,--pa-context-len-l/--logHidden — works, documented, off the menu. The test was "would a competent operator set this, and would getting it wrong hurt them", not "could someone conceivably". Each hidden flag carries a doc comment saying why. The ones that let a user silently make things worse:
--max-seq-len/--max-batch-size— device-map planning hints, universally misread as the context limit. Raising--max-seq-lendoes not let conversations grow.--pa-context-lenis the real one, and it stays visible.--pa-cache-type— overriding the resolved KV type disables prefix caching and converts the safe auto-fallback into a hard error. (Attribute only — the doc text is owned by the TurboQuant chain, so this does not collide.)--dtype,--arch— degrade silently, or are silently dropped on the paths most users are on.--no-kv-cache,--prefix-cache-n 0— read as a mystery throughput regression, not a setting.--gqa— wrong value yields garbage tokens with no error.--mtp-depth— warns and falls back on models without an MTP head, so it can appear to work while changing nothing.Deleted — still just 1 (
ARC_TD_MOE_CALIBRATION). Hiding is not deleting: bake scripts and CI keep working untouched.Discovery.
--help-allrecursively un-hides a cloned clapCommand(4.6mut_args/mut_subcommands) and descends to the named subcommand, soarc bench --help-allshows bench's flags, not the root's. Every subcommand's--helpfooter advertises it and points at theArcServe:startup line. Short help says what you may choose; the startup summary says what you got.Two judgement calls worth overriding me on
--lora/--xlorahidden. LoRA serving is a real capability and hiding it costs discovery. I applied the rule literally — a minority serving shape — but this is the cut I'd most expect you to reverse.--format/-fhidden, which removes GGUF from the visible menu entirely. Justified by Arc's own path being safetensors + UQFF/ISQ, but it means a user with a GGUF file has to reach--help-allto discover the loader exists.Deliberately left alone
defaults::PAGED_CACHE_TYPE = TurboQuantstays. I had queried it; the answer is that it is measured — commit4eba13905, 55 tok/s, +46% over baseline on a B200, plus eight CUDA correctness fixes. The "quality is not established" wording in the flag doc is itself one of the false claims being corrected separately. feat(turboquant): CUDA kernels at head_dim 64/128/256/512, Hopper + Blackwell #94's narrowing to head_dim 128 means the default applies only where it has actually run, with explicit--pa-cache-typeopening the wider set. Owner instruction is TurboQuant default in every case; it isbytes' flag and correctly not mine.ARC_V4_XS_PER_SEQ,ARC_MTP_PER_SEQ_KV,ARC_V4_TURBOQUANT,ARC_CUDA_ARCHS,ARC_SEGMENTED_KV,ARC_EP_SIZE,ARC_EP_BALANCE,ARC_SCHED_BUCKETED_DECODE) — they live on unmerged branches owned by other chains. Untouched.PostLoadHooknow has a consumer (arc-engine/src/td_moe_loader.rs:141),CrossPrefixMeteris used inkv_sharing/mod.rs, and the Sage kernels are compiled inmistralrs-quant/build.rs. Nothing to delete there.CPU-only validation (D14):
cargo check --testsgreen on all touched crates,cargo test -p arc-engine141+4 pass, scoped clippy lane exits 0.🤖 Generated with Claude Code