Skip to content

feat(cli): cut the menu to 12 flags — opinionated surface, --help-all for the rest (D18) - #110

Merged
heydryft merged 6 commits into
masterfrom
cli/opinionated-surface
Aug 19, 2026
Merged

heydryft merged 6 commits into
masterfrom
cli/opinionated-surface

Conversation

@heydryft

@heydryft heydryft commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Full inventory of ArcServe's command-line surface: 131 clap args across mistralrs-cli and arc-cli, 70 ARC_* env vars read by Rust, plus the silent default constants. Tiering below; this PR ships the unambiguous half and leaves one default alone deliberately.

The five things that were not true

  1. arc banner claimed a subsystem that usually was not running. It printed TurboQuant 3.5-bit KV cache compression (lossless, default) as a fixed string before any model was loaded. False for every MLA model and every head_dim != 128 — arc-cli's own module doc already said "in practice almost no model runs TurboQuant today, and none has been measured with it" — and "lossless" has never been measured at all. Banner now states identity only.

  2. Nothing reported the resolved state of anything. serve printed two lines and named no subsystem. Added an ArcServe summary logged after load, once every auto-fallback has run: resolved KV cache type + block count, prefix-cache state and why it is off, max-seqs. It reads the value back out of the loaded pipeline rather than restating the request — that distinction is the whole point.

  3. --mcp-port / --mcp-config parsed and did nothing. MCP is wired in the older mistralrs-server binary and was never carried across to the unified CLI (commands/serve.rs only reads server.ui/host/port). A user got a clean startup and no tools. Now a hard error.

  4. --format claimed "Auto-detected if not specified". There is no detection — every consumer does format.unwrap_or(ModelFormat::Plain), so passing only -f model.gguf silently loaded the plain path.

  5. run_ppl.sh --sinkhorn-ab silently became a tautology on 2026-08-13. It toggled ARC_FUSED_SINKHORN; commit 9387e2bc5 flipped the gate to ARC_NO_FUSED_SINKHORN and made fused the default. From that commit until this PR, env -u ARC_FUSED_SINKHORN disabled nothing and both arms ran fused-on.

    Scope, checked rather than assumed: this does not invalidate the s2 "bit-identical ppl + token-identical 6/6" entry in FACTS.md. That ran before the flip, on the opt-in gate the script drove correctly — and the flip happened because of it. Two independent confirmations the toggle was live at s2: the same harness returned a negative result at s1 (Sinkhorn fused: REJECTED — ppl drift + 4/6 token divergence), which a no-op A/B cannot produce, and wave3-H then fixed bit-identity, which s2 re-verified. Only runs dated between 2026-08-13 and this fix are meaningless, and none is recorded. Scope written into FACTS.md next to the entry.

    One live landmine found alongside it: ARC_QTIP_ROTATION_SEED has no Rust reader, but ensemble_ppl.py documents it as the only thing distinguishing bake B from bake A. The twin-seed ensemble would therefore average a distribution with itself — a guaranteed null that reads exactly like "error patterns are correlated, ensembling doesn't help". Twin-seed is in Known-unmeasured, so nothing published rests on it; recorded there as a blocker before any GPU time is spent. Cleared as non-issues: ARC_QTIP_EXPERT_GREEDY/_VITERBI have Rust readers; ARC_FORCE_GPU_QTIP_QUANTIZE was removed deliberately (12527af2d); ARC_DISABLE_YARN_STD is an unmerged snippet in a doc (real var ARC_YARN_ON_STANDARD_LAYERS, inverted polarity).

Tiers — the actual cut

serve --help is 42 lines, 12 flags, grouped. Everything else is hidden, still parsed, still working, and listed by --help-all on any subcommand.

serve run quantize calibrate
visible 12 10 18 15
total (--help-all) 46 42 22 19

Kept — the menu:

  • Model — -m/--model-id, --isq, --from-uqff, -c/--chat-template
  • Hardware — --cpu, -n/--device-layers
  • Serving — -p/--port, --host, --ui, --max-seqs, --pa-context-len
  • Diagnostics — -l/--log

Hidden — works, documented, off the menu. The test was "would a competent operator set this, and would getting it wrong hurt them", not "could someone conceivably". Each hidden flag carries a doc comment saying why. The ones that let a user silently make things worse:

  • --max-seq-len / --max-batch-size — device-map planning hints, universally misread as the context limit. Raising --max-seq-len does not let conversations grow. --pa-context-len is the real one, and it stays visible.
  • --pa-cache-type — overriding the resolved KV type disables prefix caching and converts the safe auto-fallback into a hard error. (Attribute only — the doc text is owned by the TurboQuant chain, so this does not collide.)
  • --dtype, --arch — degrade silently, or are silently dropped on the paths most users are on.
  • --no-kv-cache, --prefix-cache-n 0 — read as a mystery throughput regression, not a setting.
  • --gqa — wrong value yields garbage tokens with no error.
  • --mtp-depth — warns and falls back on models without an MTP head, so it can appear to work while changing nothing.

Deleted — still just 1 (ARC_TD_MOE_CALIBRATION). Hiding is not deleting: bake scripts and CI keep working untouched.

Discovery. --help-all recursively un-hides a cloned clap Command (4.6 mut_args/mut_subcommands) and descends to the named subcommand, so arc bench --help-all shows bench's flags, not the root's. Every subcommand's --help footer advertises it and points at the ArcServe: startup line. Short help says what you may choose; the startup summary says what you got.

Two judgement calls worth overriding me on

  • --lora / --xlora hidden. LoRA serving is a real capability and hiding it costs discovery. I applied the rule literally — a minority serving shape — but this is the cut I'd most expect you to reverse.
  • --format / -f hidden, which removes GGUF from the visible menu entirely. Justified by Arc's own path being safetensors + UQFF/ISQ, but it means a user with a GGUF file has to reach --help-all to discover the loader exists.

Deliberately left alone

  • defaults::PAGED_CACHE_TYPE = TurboQuant stays. I had queried it; the answer is that it is measured — commit 4eba13905, 55 tok/s, +46% over baseline on a B200, plus eight CUDA correctness fixes. The "quality is not established" wording in the flag doc is itself one of the false claims being corrected separately. feat(turboquant): CUDA kernels at head_dim 64/128/256/512, Hopper + Blackwell #94's narrowing to head_dim 128 means the default applies only where it has actually run, with explicit --pa-cache-type opening the wider set. Owner instruction is TurboQuant default in every case; it is bytes' flag and correctly not mine.
  • 8 env vars named in the original brief do not exist on master (ARC_V4_XS_PER_SEQ, ARC_MTP_PER_SEQ_KV, ARC_V4_TURBOQUANT, ARC_CUDA_ARCHS, ARC_SEGMENTED_KV, ARC_EP_SIZE, ARC_EP_BALANCE, ARC_SCHED_BUCKETED_DECODE) — they live on unmerged branches owned by other chains. Untouched.
  • The BACKLOG "wired-but-dead" list is stale: PostLoadHook now has a consumer (arc-engine/src/td_moe_loader.rs:141), CrossPrefixMeter is used in kv_sharing/mod.rs, and the Sage kernels are compiled in mistralrs-quant/build.rs. Nothing to delete there.

CPU-only validation (D14): cargo check --tests green on all touched crates, cargo test -p arc-engine 141+4 pass, scoped clippy lane exits 0.

🤖 Generated with Claude Code

heydryft and others added 3 commits August 17, 2026 20:18
Five places where ArcServe's command line told the user something that
was not true, all found by a full inventory of the surface (131 clap
args across two binaries, 70 ARC_* env vars read by Rust).

1. `arc` printed "TurboQuant 3.5-bit KV cache compression (lossless,
   default)" as a fixed banner before any model was loaded. It is false
   for every MLA model and every head_dim != 128 — nearly all of them,
   as arc-cli's own module doc already said — and "lossless" has never
   been measured. The banner now states identity only.

2. Nothing reported the *resolved* state of any subsystem. Added an
   ArcServe startup summary logged after load, once every auto-fallback
   has run: resolved KV cache type + block count, prefix-cache state and
   why it is off, max-seqs. It reads the value back out of the loaded
   pipeline rather than restating what was requested.

3. `--mcp-port` / `--mcp-config` parsed and did nothing — MCP was wired
   in the old `mistralrs-server` binary and never carried across to the
   unified CLI. They now fail loudly instead of yielding a server with
   no tools.

4. `--format` claimed "Auto-detected if not specified". There is no
   detection; every consumer does `unwrap_or(ModelFormat::Plain)`, so
   passing only `-f model.gguf` silently loaded the plain path.

5. `run_ppl.sh --sinkhorn-ab` toggled ARC_FUSED_SINKHORN, which nothing
   reads — the engine reads ARC_NO_FUSED_SINKHORN and fused is the
   default. Both arms ran fused-on, so the A/B compared fused against
   itself. Any "bit-identical" conclusion from the old script is void.

Deleted: ARC_TD_MOE_CALIBRATION, the one provably dead env var — parsed,
defaulted to 256, threaded through two signatures, bound to
`_calibration_set_size` and never read. Removed from both signatures and
all call sites; the env var and `--td-moe-calibration` now warn for one
release rather than being dropped silently, since the flag is scripted.

Hidden (still functional, off `--help`): `--tgt-non-granular-index`,
`arc bench --mock` (emits synthetic numbers in a real-looking artifact),
`arc validate --o-proj`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uns do not

Checked the provenance before retracting anything, and the retraction would
have been wrong. The s2 "bit-identical ppl + token-identical 6/6" entry in
FACTS.md was measured BEFORE commit 9387e2b (2026-08-13), which flipped the
gate from opt-in ARC_FUSED_SINKHORN to ARC_NO_FUSED_SINKHORN and made fused the
default — after that measurement and because of it. Two confirmations the
toggle was live at s2: the same harness returned a NEGATIVE result at s1 ("ppl
drift + 4/6 token divergence"), which a tautological A/B cannot produce, and
wave3-H then fixed bit-identity, which s2 re-verified.

The breakage is forward-looking: from 2026-08-13 until this fix the script
compared fused against fused. No such run is recorded, so no published claim is
affected. Narrowed the script comment and PR claim accordingly.

Also recorded in FACTS.md, next to "twin-seed ensemble ppl" in
Known-unmeasured: ARC_QTIP_ROTATION_SEED has no Rust reader, so the twin-seed
ensemble's two bakes would be bit-identical and the ensemble would average a
distribution with itself — a guaranteed null that would read as the scientific
conclusion "ensembling doesn't help". Wire the seed before spending GPU time.

Cleared as non-issues: ARC_QTIP_EXPERT_GREEDY/_VITERBI have Rust readers;
ARC_FORCE_GPU_QTIP_QUANTIZE was removed deliberately (12527af);
ARC_DISABLE_YARN_STD is an unmerged snippet in a doc, real var is
ARC_YARN_ON_STANDARD_LAYERS with inverted polarity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous pass hid 3 flags out of ~110 — a re-labelling, not a cut. This
is the cut. `serve --help` is now 42 lines and 12 flags, grouped Model /
Hardware / Serving / Diagnostics:

  Model      -m/--model-id, --isq, --from-uqff, -c/--chat-template
  Hardware   --cpu, -n/--device-layers
  Serving    -p/--port, --host, --ui, --max-seqs, --pa-context-len
  Diagnostics -l/--log

Everything else is hidden, still parsed, still working, and listed by
`--help-all` on any subcommand (46 flags for serve). Nothing is deleted and no
behavioural default changes in this pass.

The test applied to each flag was "would a competent operator set this, and
would getting it wrong hurt them" — not "could someone conceivably". Flags that
let a user silently make things worse were hidden hardest, and each carries a
doc comment saying why:

  --max-seq-len / --max-batch-size  device-map planning hints universally
      misread as the context limit; --pa-context-len is the real one
  --pa-cache-type   overriding the resolved KV type disables prefix caching
      and turns the safe auto-fallback into a hard error (attribute only —
      the doc text is owned by the TurboQuant chain)
  --dtype, --arch   silently degrade or are silently dropped per model path
  --no-kv-cache, --prefix-cache-n 0   look like tuning, read as a mystery
      throughput regression
  --gqa             wrong value yields garbage tokens, no error
  --mtp-depth       warns and falls back, so it can appear to work

Also fixes the same false claim the runtime banner carried: `arc --help`
still said "Defaults to TurboQuant 3.5-bit KV cache (lossless)", which is
decided per model at load time and cannot be stated in static help.

--help-all is implemented by recursively un-hiding a cloned clap Command
(clap 4.6 mut_args/mut_subcommands), descending to the named subcommand so
`arc bench --help-all` shows bench's flags rather than the root's.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@heydryft heydryft changed the title feat(cli): make the surface state what it actually does — five false claims, one dead env var (D18) feat(cli): cut the menu to 12 flags — opinionated surface, --help-all for the rest (D18) Aug 17, 2026
@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     72        24645        17702         4215         2728
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     74         4600         4597            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  144        14955        12317          810         1828
 Shell                    26         6312         4307         1564          441
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1488         1293           45          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                197        39646            0        30479         9167
 |- BASH                  72         1652         1200          331          121
 |- C                      3           17           17            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  18          708          708            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  65         2048         1713           77          258
 |- TOML                   6          207          164            0           43
 |- YAML                   4           38           33            5            0
 (Total)                            45409         4679        31021         9709
─────────────────────────────────────────────────────────────────────────────────
 Rust                    668       318690       275524        15018        28148
 |- Markdown             484        26172          471        22580         3121
 (Total)                           344862       275995        37598        31269
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1295       469159       339700        79625        49834
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

@heydryft

Copy link
Copy Markdown
Contributor Author

⚠️ This PR's red lane is INHERITED FROM MASTER, not this PR. Do not triage it as broken.

cargo check (cuda, workspace) failed at the step "Type-check CUDA-gated tests (no codegen, no run)", exit 101. I absorbed master into this branch while master's test target was red — my sequencing error, not a defect here.

Proven causally, not inferred from the step name. On this PR's head 67cd73171:

cargo check -p mistralrs-core --tests            -> exit 101
  mistralrs-core/src/paged_attention/scheduler.rs:653:13
  error[E0308]: expected `Mutex<SequenceGroup>`, found a different `Mutex<SequenceGroup>`

... then cherry-pick #131's one-word fix on top ...

cargo check -p mistralrs-core --tests            -> exit 0, 0 errors

Same tree, one word different. That is the whole delta. (The probe was reverted immediately — this branch is untouched at 67cd73171.)

The underlying break is #118 + #115: #118 added a fixture building Arc<tokio::sync::Mutex<SequenceGroup>>, #115 then moved Sequence::new_waiting onto std::sync::Mutex. Textually disjoint, no merge conflict, both a genuine 17/17 — GitHub tests each PR against the base as it stood, so neither run contained the other. #131 fixes it.

This PR's own content was verified green before I pushed it — cargo check -p arc-cli -p mistralrs-cli --tests exit 0, unpiped, after resolving the two conflicts with #111.

On those conflicts, since a reviewer will want to know what I decided and why: this PR's own comment says the --pa-cache-type doc text is "owned by the TurboQuant chain", so I kept master's (#111's) text verbatim and re-applied only this PR's hide = true attribute. For the startup banner I kept this PR's deletion: the banner printed a TurboQuant claim before any model was loaded, and this PR replaces it with the post-load ArcServe: kv-cache=…, prefix-cache=…, max-seqs=… line, which reports what was actually resolved rather than what was hoped. #111's substantive correction — unset means TurboQuant, and it is not "lossless" — survives verbatim on the flag's own --help text.

Next step: #131 lands → I re-absorb master here → lanes re-run clean. No action needed from anyone.

@heydryft

Copy link
Copy Markdown
Contributor Author

Cleared: the CI complete: FAILURE above is a CANCELLATION ARTIFACT of mine, not a verdict on this PR. Fresh lanes are running now — read those, not the stale aggregate.

Sequence of what happened here, because the stale status is exactly the thing that gets quoted an hour later by someone who wasn't in the thread:

  1. Master's test target broke (fix(ArcSched): stop bucketing paged decode by sequence length (8.7x on spread traffic) #118 + perf(engine): kill the SequenceGroup spin-lock and the await-under-lock response dispatch — and the measurement that refutes their premise [MEASURED] #115, a semantic conflict between two textually disjoint PRs that each had a genuine 17/17 — GitHub tests each PR against the base as it stood, so neither run could see the other). Fixed by fix(build): master's test target is red — #118's fixture builds a tokio Mutex where #115 now requires std's #131, now merged.
  2. I absorbed master into this branch while master was red, so this PR inherited that failure. Proven causally rather than inferred: on this branch's then-head, cargo check -p mistralrs-core --tests exited 101; cherry-picking fix(build): master's test target is red — #118's fixture builds a tokio Mutex where #115 now requires std's #131's one-word commit on the same tree took it to exit 0, 0 errors.
  3. To free runners for fix(build): master's test target is red — #118's fixture builds a tokio Mutex where #115 now requires std's #131 — which was queued behind ~20 runs re-deriving the break it fixes — I cancelled this PR's in-flight runs. That was my call and it had a consequence I did not anticipate: ci-complete uses if: always() over needs: every lane, so a CANCELLED dependency makes the aggregate FAILURE. The red you see is that, not a defect.

🔑 The lesson I'm recording on the PR rather than just in my own notes: cancelling CI does not merely discard information — it WRITES a false negative into the PR's permanent check record. I had told the coordinator "nothing informative was destroyed", which was true about the runs and false about the record.

Now: master is green (cca5e5c6e, verified unpiped — cargo check --workspace --tests exit 0, cargo test -p mistralrs-core --lib 537 passed / 0 failed), this branch has absorbed it, and full lanes are re-running.

Judge this PR on the new run. If you are reading the old aggregate, it is describing my cancellation, not this code.

@heydryft
heydryft merged commit a44ae20 into master Aug 19, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant