Skip to content

feat(agentx): refresh DSV4 GB300 SGLang on nightly-dev-20261001 / 基于 nightly-dev-20261001 更新 DSV4 GB300 SGLang 配置 - #3630

Open
nvpohanh wants to merge 4 commits into
mainfrom
dsv4-gb300-dynamo-sgl-agentx-nightly1001
Open

nvpohanh wants to merge 4 commits into
mainfrom
dsv4-gb300-dynamo-sgl-agentx-nightly1001

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

[by Claude Code]

Supersedes #3218. Carries over its DeepSeek-V4-Pro-0813 GB300 Dynamo+SGLang AgentX refresh (rebased onto current main), plus:

  • Moves both recipes and master entries to SGLang lmsysorg/sglang:nightly-dev-20261001-b37e6f79 (Dynamo 1.5.0.dev20260914 unchanged).
  • Drops enable-w4a4-mxfp4-megamoe: true from the disaggregated prefill and decode recipes.
  • Updates flags for the new nightly. --dp-size N --enable-dp-attention is deprecated in favor of --attn-dp-size N and resolves to the same layout. SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT is removed because upstream deleted the SWA radix cache that read it.

Topologies and concurrency points are unchanged:

  • Aggregated: TP8 at c1 and TP4 at c4.
  • Disaggregated, DEP8 prefill with DEP16 decode: 1P1D at c64 and c240, 2P1D at c480 and c960, 3P1D at c1440, and 4P1D at c1920.

Validation:

  • The changelog validator passes against current main.
  • srtctl dry-run succeeds for all eight variant selectors.
  • Expanding the resolved variants confirms these points:
    • attn-dp-size equals TP on every disaggregated role.
    • No dp-size, enable-dp-attention, or W4A4 flag remains.
  • Every recipe flag exists in the new nightly's ServerArgs.
  • The MegaMoE per-rank token budgets still satisfy the new nightly's checks:
    • Prefill: 65536 / 8 = 8192 tokens per rank, within the 9216 limit.
    • Decode: at most 512 × 7 = 3584 tokens per rank, within the 4096 limit.

AI model disclosure

Claude Code (Claude Opus 5.5) rebased the earlier change, applied the image, flag, and recipe updates, and ran the validation. No delegated AI agents were used.

中文

取代 #3218。保留其 GB300 上 DeepSeek-V4-Pro-0813 Dynamo+SGLang AgentX 的配置更新(已 rebase 至当前 main),并新增:

  • 两个配方及主配置均更新为 SGLang lmsysorg/sglang:nightly-dev-20261001-b37e6f79(Dynamo 1.5.0.dev20260914 不变)。
  • 分离式预填充与解码配方移除 enable-w4a4-mxfp4-megamoe: true。
  • 适配新版 nightly 的参数:已弃用的 --dp-size N --enable-dp-attention 改为 --attn-dp-size N(解析后的并行布局相同);移除 SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT,因读取它的 SWA radix cache 已在上游删除。

拓扑与并发点不变:聚合 TP8 c1、TP4 c4;DEP8 预填充 / DEP16 解码 1P1D c64、1P1D c240、2P1D c480、2P1D c960、3P1D c1440、4P1D c1920。

验证:变更日志检查通过;八个变体选择器的 srtctl dry-run 均成功;展开后的变体中每个分离式角色的 attn-dp-size 均等于 TP,且不再含 dp-size、enable-dp-attention 或 W4A4 参数;所有配方参数在新版 ServerArgs 中均存在;MegaMoE 每 rank token 预算仍满足新版检查(预填充 65536/8=8192 ≤ 9216;解码最多 512×7=3584 ≤ 4096)。

AI 模型披露

由 Claude Code(Claude Opus 5.5)完成 rebase、镜像、参数与配方更新及验证。未使用委派的 AI agent。

🤖 Generated with Claude Code

nvpohanh and others added 2 commits October 1, 2026 01:12
将 DSV4 GB300 AgentX 扫描适配到共享变体配方。
- Bump the aggregate and disaggregated GB300 DeepSeek-V4-Pro Dynamo+SGLang
  recipes and master entries to lmsysorg/sglang:nightly-dev-20261001-b37e6f79.
- Drop enable-w4a4-mxfp4-megamoe from the disaggregated recipes.
- Replace the deprecated `dp-size N` + `enable-dp-attention` spelling with
  `attn-dp-size N`.
- Remove SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT; its only reader, the SWA radix
  cache, was removed upstream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline pr-link finding, I also checked the two other tuning asymmetries in the new disagg variants: override_1p1d_c240 and override_2p1d_c480 leave cuda-graph-max-bs-decode (256) below prefill max-running-requests (512), and only override_1p1d_c64's prefill sets hicache-mem-layout: page_first_direct. Both mirror the pre-existing pattern from the superseded override_1p1d_c480 variant rather than being new regressions, so they read as tuning choices, not correctness bugs — still worth a human's eye alongside the pr-link issue.

Extended reasoning...

The diff refreshes DSV4 GB300 SGLang agentic recipes/master-config to a new nightly image and adds a perf-changelog entry; no security-sensitive surface is touched (benchmark YAML config only). A CONFIRMED inline finding shows the new changelog entry's pr-link still points to PR #3218, which the PR description itself says this PR supersedes, so a human should fix that before merge. I independently verified two other previously-flagged-and-ruled-out asymmetries (cuda-graph-max-bs-decode vs max-running-requests, and hicache-mem-layout only being set on one of three new variants) and found they match pre-existing patterns from the variant being replaced, not new regressions.

Comment thread inferencex-e2e/perf-changelog.yaml Outdated
Comment on lines +9168 to +9184
- "Relevant ATOM changes in the range: V4 decode reuses the sparse prefill ASM (ROCm/ATOM#2271), greedy sampler picks via aiter.topk_select (ROCm/ATOM#2244), and FP8 block scales declared scale_fmt ue8m0 are now stored as E8M0 on gfx950 by default (ROCm/ATOM#2419; previously FP32 unless ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE=1). The upstream DeepSeek-V4 recipes are unchanged across the range, and every recipe flag and choice (all2all-backend rccl, dp-load-balance least_tokens, moe-backend standard) remains valid."
- "No data-type or precision change to the DeepSeek-V4-Pro-0813 DSpark draft: no online quantization is configured, so it keeps its checkpoint precision; the E8M0 scale storage represents the checkpoint's power-of-two block scales exactly. kv-cache-dtype and index-cache-dtype touch cache storage only."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3605

- config-keys:
- dsv4-fp4-gb300-dynamo-sglang-agentic-agg
- dsv4-fp4-gb300-dynamo-sglang-agentic-disagg
scenario-type:
- agentic-coding
description:
- "Refresh DeepSeek-V4-Pro-0813 GB300 Dynamo+SGLang AgentX on nightly-dev-20261001-b37e6f79 and Dynamo 1.5.0.dev20260914. Keep TP8 c1 and TP4 c4 aggregate points, and DEP8 prefill / DEP16 decode points at c64, c240, c480, c960, c1440 and c1920."
- "Replace the long-TTFT 1P1D c480 point with 1P1D c240 and 2P1D c480, retain DSpark K=6 on both serving stages, and use the current SGLang CUDA graph flags. Adapt the recipes and registry to the shared variant layout under inferencex-e2e/."
- "Drop enable-w4a4-mxfp4-megamoe from the disaggregated recipes, replace the deprecated dp-size plus enable-dp-attention spelling with attn-dp-size, and remove SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT, which the removed SWA radix cache no longer reads."
- "将 DeepSeek-V4-Pro-0813 GB300 Dynamo+SGLang AgentX 更新至 nightly-dev-20261001-b37e6f79 和 Dynamo 1.5.0.dev20260914。聚合配置保留 TP8 c1 与 TP4 c4;DEP8 预填充、DEP16 解码配置覆盖 c64、c240、c480、c960、c1440 和 c1920。"
- "将首 token 延迟较长的 1P1D c480 替换为 1P1D c240 与 2P1D c480,在预填充和解码阶段均使用 DSpark K=6,并采用当前的 SGLang CUDA graph 参数。配方和注册项适配 inferencex-e2e/ 下的共享变体布局。"
- "分离式配方移除 enable-w4a4-mxfp4-megamoe;将已弃用的 dp-size 加 enable-dp-attention 写法改为 attn-dp-size;并移除已随 SWA radix cache 删除而失效的 SGLANG_OPT_SWA_SPLIT_LEAF_ON_INSERT。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3218

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Maintainers tracing this benchmark change via the changelog land on the wrong, already-superseded pull request. The new entry's pr-link at line 9184 is https://github.com/SemiAnalysisAI/InferenceX/pull/3218, but per the PR description this change ships in a different PR that explicitly "Supersedes #3218" (opened 2026-10-01, title mentions nightly-dev-20261001). configuration-procedures.md's changelog procedure requires pr-link to be "the real URL" of the PR making the change, filled in "immediately after creating the PR" - here it still cites the stale, superseded PR instead of the one actually introducing these recipe/master-config edits. Fix: set pr-link to the URL of the PR that actually lands this diff, not #3218.

Why this was flagged

The new perf-changelog.yaml entry (appended at line 9171-9184) documents the recipe/master-config changes in this diff, but its pr-link field (line 9184) points to pull/3218, the PR the author's own description says is superseded by this one. configuration-procedures.md ('Append the changelog safely', step 3) requires pr-link to hold the real URL of the PR introducing the change, filled in once the PR exists; this PR was already opened (2026-10-01T08:18:00Z per the description) when this entry was written. A maintainer or auditor following this link later to see the actual commit/review history for this image bump lands on a different, closed/superseded PR instead of the one that merged these exact YAML changes, breaking the traceability the changelog exists to provide. No validator check in this diff catches it since the file is syntactically a valid append.

Verification: nit. The appended changelog entry at inferencex-e2e/perf-changelog.yaml:9184 sets pr-link to pull/3218, but the PR author's own description states "Supersedes #3218" (this PR is a new one opened 2026-10-01, title "refresh DSV4 GB300 SGLang on nightly-dev-20261001"). So #3218 is the earlier, superseded PR, not the PR actually introducing this change. This violates configuration-procedures.md (step 3): "Replace it with the real URL immediately after creating the PR."

… recipes

AgentX power measurement on single-node runs requires TP, PP_SIZE and
PCP_SIZE in the benchmark environment; the aggregate recipe only set TP,
so srt_agentic.sh exited before sending requests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant