Conversation
将 DSV4 GB300 AgentX 扫描适配到共享变体配方。
- Bump the aggregate and disaggregated GB300 DeepSeek-V4-Pro Dynamo+SGLang recipes and master entries to lmsysorg/sglang:nightly-dev-20261001-b37e6f79. - Drop enable-w4a4-mxfp4-megamoe from the disaggregated recipes. - Replace the deprecated `dp-size N` + `enable-dp-attention` spelling with `attn-dp-size N`. - Remove SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT; its only reader, the SWA radix cache, was removed upstream. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks for the contribution!
中文感谢你的贡献!
|
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline pr-link finding, I also checked the two other tuning asymmetries in the new disagg variants: override_1p1d_c240 and override_2p1d_c480 leave cuda-graph-max-bs-decode (256) below prefill max-running-requests (512), and only override_1p1d_c64's prefill sets hicache-mem-layout: page_first_direct. Both mirror the pre-existing pattern from the superseded override_1p1d_c480 variant rather than being new regressions, so they read as tuning choices, not correctness bugs — still worth a human's eye alongside the pr-link issue.
Extended reasoning...
The diff refreshes DSV4 GB300 SGLang agentic recipes/master-config to a new nightly image and adds a perf-changelog entry; no security-sensitive surface is touched (benchmark YAML config only). A CONFIRMED inline finding shows the new changelog entry's pr-link still points to PR #3218, which the PR description itself says this PR supersedes, so a human should fix that before merge. I independently verified two other previously-flagged-and-ruled-out asymmetries (cuda-graph-max-bs-decode vs max-running-requests, and hicache-mem-layout only being set on one of three new variants) and found they match pre-existing patterns from the variant being replaced, not new regressions.
| - "Relevant ATOM changes in the range: V4 decode reuses the sparse prefill ASM (ROCm/ATOM#2271), greedy sampler picks via aiter.topk_select (ROCm/ATOM#2244), and FP8 block scales declared scale_fmt ue8m0 are now stored as E8M0 on gfx950 by default (ROCm/ATOM#2419; previously FP32 unless ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE=1). The upstream DeepSeek-V4 recipes are unchanged across the range, and every recipe flag and choice (all2all-backend rccl, dp-load-balance least_tokens, moe-backend standard) remains valid." | ||
| - "No data-type or precision change to the DeepSeek-V4-Pro-0813 DSpark draft: no online quantization is configured, so it keeps its checkpoint precision; the E8M0 scale storage represents the checkpoint's power-of-two block scales exactly. kv-cache-dtype and index-cache-dtype touch cache storage only." | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3605 | ||
|
|
||
| - config-keys: | ||
| - dsv4-fp4-gb300-dynamo-sglang-agentic-agg | ||
| - dsv4-fp4-gb300-dynamo-sglang-agentic-disagg | ||
| scenario-type: | ||
| - agentic-coding | ||
| description: | ||
| - "Refresh DeepSeek-V4-Pro-0813 GB300 Dynamo+SGLang AgentX on nightly-dev-20261001-b37e6f79 and Dynamo 1.5.0.dev20260914. Keep TP8 c1 and TP4 c4 aggregate points, and DEP8 prefill / DEP16 decode points at c64, c240, c480, c960, c1440 and c1920." | ||
| - "Replace the long-TTFT 1P1D c480 point with 1P1D c240 and 2P1D c480, retain DSpark K=6 on both serving stages, and use the current SGLang CUDA graph flags. Adapt the recipes and registry to the shared variant layout under inferencex-e2e/." | ||
| - "Drop enable-w4a4-mxfp4-megamoe from the disaggregated recipes, replace the deprecated dp-size plus enable-dp-attention spelling with attn-dp-size, and remove SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT, which the removed SWA radix cache no longer reads." | ||
| - "将 DeepSeek-V4-Pro-0813 GB300 Dynamo+SGLang AgentX 更新至 nightly-dev-20261001-b37e6f79 和 Dynamo 1.5.0.dev20260914。聚合配置保留 TP8 c1 与 TP4 c4;DEP8 预填充、DEP16 解码配置覆盖 c64、c240、c480、c960、c1440 和 c1920。" | ||
| - "将首 token 延迟较长的 1P1D c480 替换为 1P1D c240 与 2P1D c480,在预填充和解码阶段均使用 DSpark K=6,并采用当前的 SGLang CUDA graph 参数。配方和注册项适配 inferencex-e2e/ 下的共享变体布局。" | ||
| - "分离式配方移除 enable-w4a4-mxfp4-megamoe;将已弃用的 dp-size 加 enable-dp-attention 写法改为 attn-dp-size;并移除已随 SWA radix cache 删除而失效的 SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT。" | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3218 |
There was a problem hiding this comment.
🟡 (optional) Maintainers tracing this benchmark change via the changelog land on the wrong, already-superseded pull request. The new entry's pr-link at line 9184 is https://github.com/SemiAnalysisAI/InferenceX/pull/3218, but per the PR description this change ships in a different PR that explicitly "Supersedes #3218" (opened 2026-10-01, title mentions nightly-dev-20261001). configuration-procedures.md's changelog procedure requires pr-link to be "the real URL" of the PR making the change, filled in "immediately after creating the PR" - here it still cites the stale, superseded PR instead of the one actually introducing these recipe/master-config edits. Fix: set pr-link to the URL of the PR that actually lands this diff, not #3218.
Why this was flagged
The new perf-changelog.yaml entry (appended at line 9171-9184) documents the recipe/master-config changes in this diff, but its pr-link field (line 9184) points to pull/3218, the PR the author's own description says is superseded by this one. configuration-procedures.md ('Append the changelog safely', step 3) requires pr-link to hold the real URL of the PR introducing the change, filled in once the PR exists; this PR was already opened (2026-10-01T08:18:00Z per the description) when this entry was written. A maintainer or auditor following this link later to see the actual commit/review history for this image bump lands on a different, closed/superseded PR instead of the one that merged these exact YAML changes, breaking the traceability the changelog exists to provide. No validator check in this diff catches it since the file is syntactically a valid append.
Verification: nit. The appended changelog entry at inferencex-e2e/perf-changelog.yaml:9184 sets pr-link to pull/3218, but the PR author's own description states "Supersedes #3218" (this PR is a new one opened 2026-10-01, title "refresh DSV4 GB300 SGLang on nightly-dev-20261001"). So #3218 is the earlier, superseded PR, not the PR actually introducing this change. This violates configuration-procedures.md (step 3): "Replace it with the real URL immediately after creating the PR."
… recipes AgentX power measurement on single-node runs requires TP, PP_SIZE and PCP_SIZE in the benchmark environment; the aggregate recipe only set TP, so srt_agentic.sh exited before sending requests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36835579692 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36835579692 |
[by Claude Code]
Supersedes #3218. Carries over its DeepSeek-V4-Pro-0813 GB300 Dynamo+SGLang AgentX refresh (rebased onto current
main), plus:lmsysorg/sglang:nightly-dev-20261001-b37e6f79(Dynamo1.5.0.dev20260914unchanged).enable-w4a4-mxfp4-megamoe: truefrom the disaggregated prefill and decode recipes.--dp-size N --enable-dp-attentionis deprecated in favor of--attn-dp-size Nand resolves to the same layout.SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERTis removed because upstream deleted the SWA radix cache that read it.Topologies and concurrency points are unchanged:
Validation:
main.srtctl dry-runsucceeds for all eight variant selectors.attn-dp-sizeequals TP on every disaggregated role.dp-size,enable-dp-attention, or W4A4 flag remains.ServerArgs.AI model disclosure
Claude Code (Claude Opus 5.5) rebased the earlier change, applied the image, flag, and recipe updates, and ran the validation. No delegated AI agents were used.
中文
取代 #3218。保留其 GB300 上 DeepSeek-V4-Pro-0813 Dynamo+SGLang AgentX 的配置更新(已 rebase 至当前
main),并新增:lmsysorg/sglang:nightly-dev-20261001-b37e6f79(Dynamo1.5.0.dev20260914不变)。enable-w4a4-mxfp4-megamoe: true。--dp-size N --enable-dp-attention改为--attn-dp-size N(解析后的并行布局相同);移除SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT,因读取它的 SWA radix cache 已在上游删除。拓扑与并发点不变:聚合 TP8 c1、TP4 c4;DEP8 预填充 / DEP16 解码 1P1D c64、1P1D c240、2P1D c480、2P1D c960、3P1D c1440、4P1D c1920。
验证:变更日志检查通过;八个变体选择器的
srtctl dry-run均成功;展开后的变体中每个分离式角色的attn-dp-size均等于 TP,且不再含dp-size、enable-dp-attention或 W4A4 参数;所有配方参数在新版ServerArgs中均存在;MegaMoE 每 rank token 预算仍满足新版检查(预填充 65536/8=8192 ≤ 9216;解码最多 512×7=3584 ≤ 4096)。AI 模型披露
由 Claude Code(Claude Opus 5.5)完成 rebase、镜像、参数与配方更新及验证。未使用委派的 AI agent。
🤖 Generated with Claude Code