Add Qwen3.5 FP4 B200 AgentX MTP - #2420
Conversation
Add the initial SGLang NEXTN AgentX recipe with golden synthetic acceptance and 256k traces.\n\n中文:新增 Qwen3.5 FP4 B200 AgentX MTP 初始配置,使用 SGLang NEXTN、黄金合成接受长度和 256k 轨迹数据集。
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Append the Qwen3.5 FP4 B200 AgentX MTP benchmark entry.\n\n中文:追加 Qwen3.5 FP4 B200 AgentX MTP 基准测试触发记录。
Probe the allocated DGXC node for the complete Qwen3.5 NVFP4 checkpoint on /raid and fall back to shared Lustre when it is absent.\n\n中文:在分配到的 DGXC 节点上探测 /raid 中完整的 Qwen3.5 NVFP4 权重;若未暂存,则回退到共享 Lustre。
Add strict two-pool HiCache accounting, request cache reporting, session-sticky attention-DP routing, and AgentX request/graph limits for targeted B200 tuning.\n\n中文:为 B200 定向调优加入严格的双池 HiCache 容量核算、请求级缓存统计、会话粘性的注意力数据并行路由,以及 AgentX 请求数与 CUDA Graph 上限。
Move the configuration into the AgentX section, document the measured cache model, and add session-sticky DEP8 no-offload probes for the throughput frontier.\n\n中文:将配置移入 AgentX 区域,记录实测缓存模型,并为吞吐量前沿加入具备会话粘性路由的 DEP8 无卸载探测点。
中文:添加 B200 TP2 c16 HiCache 对照探针,并严格使用按 GPU 数量计算的 DRAM 上限。
中文:将 B200 TP2 HiCache 验证点设为独立的 c20,避免一次性重复启动无卸载配置。
Add the isolated TP4 c96 no-offload point after c64 remained healthy, so the HBM working-set cliff can be measured before adding matched HiCache coverage. 中文:在 c64 保持健康后新增独立的 TP4 c96 无卸载测试点,以便在添加匹配的 HiCache 覆盖前测量 HBM 工作集拐点。
Remove the DEP8 search arm after its c128 probe fragmented prefix reuse, completed no profiled requests in three minutes, and delivered less than half of TP4 c64 aggregate prefill throughput. 中文:DEP8 c128 测试出现前缀复用碎片化,三分钟内没有完成任何正式请求,聚合预填充吞吐量不足 TP4 c64 的一半,因此移除该搜索分支。
Account for SGLang target KV, Mamba, and native NEXTN draft host pools as H * 31/15 per rank, with page-alignment reserve, so projected use stays within runner-generated DRAM. 中文:按每个 rank 的 H * 31/15 计入 SGLang 目标 KV、Mamba 与原生 NEXTN 草稿主机缓存,并预留页对齐空间,确保预计用量不超过 runner 生成的 DRAM 上限。
Add matched no-offload and DRAM HiCache c72 points immediately after the measured c64 KV-pressure onset. 中文:在实测 c64 KV 压力拐点之后,添加匹配的 c72 无卸载与 DRAM HiCache 对照测试点。
中文:将最新 main 合并到 B200 提交分支,并将本 PR 的性能变更记录重新追加到文件末尾。
Simplify the B200 launcher script to TP-only serving after the measured DEP arm was removed from the search space. 中文:实测 DEP 分支已从搜索空间移除,因此将 B200 启动脚本简化为仅使用 TP 的服务路径。
Replace the dominated c24-c64 no-offload tail with dense c12-c22 points and matched HiCache coverage beginning at c16. 中文:移除性能受压的 c24-c64 无卸载尾部,改为密集采样 c12-c22,并从 c16 开始提供匹配的 HiCache 覆盖。
Densify TP4 around the measured c64-c72 transition and extend the HiCache curve beyond the HBM cliff.\n\n中文:细化 B200 Qwen TP4 在实测 c64-c72 缓存拐点附近的并发点,并将 HiCache 曲线扩展到 HBM 拐点之后。
Remove no-offload c22 after live profiling placed it past the TP2 cache cliff, while retaining the HiCache arm and c20 boundary.\n\n中文:实测表明无卸载 c22 已越过 TP2 缓存拐点,因此移除该点,同时保留 HiCache 配置与 c20 边界点。
中文:在缓存抖动前收紧 B200 TP4 并发范围。
中文:B200 TP2 在并发 14 后切换至 HiCache。
中文:裁剪性能受支配的 B200 TP4 并发 66 配置。
中文:使用 SGLang 默认单 tokenizer 启动路径,避免长时间 HiCache 初始化期间的多 tokenizer 共享内存竞态。
中文:移除已被 c64 严格支配的 B200 c68 无卸载点,并让 HiCache 从缓存悬崖边界开始接管。
Use six tokenizer workers for TP4 long-context replay while retaining SGLang default single-worker startup for TP2 HiCache. 中文:TP4 长上下文回放使用 6 个 tokenizer worker;TP2 HiCache 保持 SGLang 默认单 worker 启动路径,避免共享内存初始化竞态。
Keep the measured Pareto region and validate every node-local checkpoint shard before selecting NVMe weights.\n\n中文:保留实测帕累托区间,并在选择 NVMe 权重前校验节点本地检查点的全部分片。
Pin the AIPerf trace idle-gap branch and set the Qwen B200 replay cap to 300 seconds.
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30564461716 |
Signed-off-by: Cam Quilici <cjquilici@gmail.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30565108476 |
Signed-off-by: Cam Quilici <cjquilici@gmail.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30569689996 |
Update AIPerf after removing runtime trace idle enforcement. Signed-off-by: Cam Quilici <cjquilici@gmail.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30571066450 |
Signed-off-by: Cam Quilici <cjquilici@gmail.com>
Signed-off-by: Cam Quilici <cjquilici@gmail.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30577529970 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30584321246 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30590080837 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30597334596 |
Update the Qwen3.5 B200 AgentX submission to the validated AIPerf revision containing profiling handoff anchoring, cache-bust continuity, spawn/join timing, and runtime idle caps.\n\n中文:将 Qwen3.5 B200 AgentX 提交固定到已验证的 AIPerf 版本,包含 profiling 交接锚定、cache-bust 连续性、spawn/join 时序以及运行时空闲上限修复。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30666626028 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30666626028 |
|
/stage-results 30666626028 |
|
@cquil11 staged run 30666626028: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-31~r30666626028 This run remains available across future @cquil11 已将运行 30666626028 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-31~r30666626028 后续的 |
中文:将 main 的最新更改合并到 Qwen B200 AgentX 分支。
Pin Qwen B200 AgentX to the validated AIPerf scheduling revision and run GSM8K instead of SWE-bench for the full-hour sweep. 中文:将 Qwen B200 AgentX 固定到已验证的 AIPerf 调度版本,并在一小时完整扫描中使用 GSM8K 替代 SWE-bench。
Summary
Add Qwen3.5-397B-A17B NVFP4 AgentX coverage on B200 with SGLang native NEXTN MTP, golden synthetic acceptance, prefix caching, the 256k trace dataset, and a 300-second per-trace idle-gap cap.