Skip to content

Add DSV4 GB200 Dynamo+SGLang AgentX configs without W4A4 MegaMoE / 添加不启用 W4A4 MegaMoE 的 DSV4 GB200 Dynamo+SGLang AgentX 配置 - #3629

Open
nvpohanh wants to merge 5 commits into
mainfrom
dsv4-fp4-gb200-dyn-sgl-no-megamoe
Open

nvpohanh wants to merge 5 commits into
mainfrom
dsv4-fp4-gb200-dyn-sgl-no-megamoe

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

[by Claude Code]

Add GB200 Dynamo+SGLang AgentX configurations for deepseek-ai/DeepSeek-V4-Pro-0813 using its bundled DSpark head and HiCache, without W4A4 MXFP4 MegaMoE.

This is a copy of #3182, rebased onto current main, with one change: enable-w4a4-mxfp4-megamoe: true is removed from every prefill and decode worker in the six disaggregated recipes. The MegaMoE all-to-all backend (moe-a2a-backend: megamoe), the FP4 indexer and all other recipe settings are unchanged. The two TP8 aggregate recipes never set the flag and are unchanged.

The sweep covers TP8 aggregate concurrency 1 and 4; 1P1D DEP8/DEP16 concurrency 64 and 128; 1P1D DEP16/DEP32 concurrency 256; and 2P1D DEP16/DEP32 concurrency 768, 1024, and 1280.

AI model disclosure

  • Claude Opus 5.5 (Claude Code): rebased the branch onto main, resolved the changelog conflict as an append-only tail, removed the W4A4 MegaMoE flag, validated the YAML and the generated matrix for both config keys, and prepared this PR.
中文

为 deepseek-ai/DeepSeek-V4-Pro-0813 添加 GB200 Dynamo+SGLang AgentX 配置,使用其内置的 DSpark 草稿头和 HiCache,且不启用 W4A4 MXFP4 MegaMoE。

本 PR 复制自 #3182,已 rebase 到当前 main,仅有一处改动:从六个分离式(disagg)recipe 的所有 prefill 和 decode worker 中移除 enable-w4a4-mxfp4-megamoe: true。MegaMoE all-to-all 后端(moe-a2a-backend: megamoe)、FP4 indexer 及其他 recipe 设置均保持不变。两个 TP8 聚合 recipe 本来就未设置该选项,保持不变。

本次 sweep 覆盖 TP8 聚合模式并发 1 和 4;1P1D DEP8/DEP16 并发 64 和 128;1P1D DEP16/DEP32 并发 256;以及 2P1D DEP16/DEP32 并发 768、1024 和 1280。

AI 模型披露

  • Claude Opus 5.5 (Claude Code):将分支 rebase 到 main,以仅追加末尾的方式解决 changelog 冲突,移除 W4A4 MegaMoE 选项,验证 YAML 和两个配置键生成的 matrix,并准备此 PR。

🤖 Generated with Claude Code

nvpohanh and others added 4 commits October 1, 2026 00:56
Remove enable-w4a4-mxfp4-megamoe from every prefill and decode worker in
the six GB200 Dynamo+SGLang AgentX disaggregated recipes. The MegaMoE
all-to-all backend and FP4 indexer settings are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked the PR's core claim — that enable-w4a4-mxfp4-megamoe was removed from all prefill/decode workers in the six disaggregated recipes while moe-a2a-backend: megamoe stays — and confirmed no enable-w4a4-mxfp4-megamoe key remains anywhere in the new recipe files, and container.image matches the master-config image field consistently across all eight new scenarios.

Extended reasoning...

The diff adds 7 new srt-slurm recipe YAMLs plus matching master-config entries, a runners.yaml model-path mapping, a lanes.py policy-table entry, and an append-only changelog addition for a new DeepSeek-V4-Pro-0813 GB200 FP4 AgentX workload; no auth/crypto/permission code is touched, only benchmark config and launch policy tables. A confirmed finding already flags a real routing problem in runners.yaml (new model lacks a numa1-local alias, so it is not visible for node-local staging as the changelog implies), which is sufficient reason a human should look before merge.

deepseek-r1-0528: {root: lustre-models, dir: deepseek-r1-0528}
deepseek-r1-0528-fp4-v2: {root: lustre-models, dir: deepseek-r1-0528-fp4-v2}
DeepSeek-V4-Pro: {root: lustre-models, dir: DeepSeek-V4-Pro}
DeepSeek-V4-Pro-0813: {root: lustre-models, dir: DeepSeek-V4-Pro-0813}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) The changelog promises node-local checkpoint staging for the new gb200-nv dsv4-pro-0813 AgentX workload, but the config actually routes every run through shared Lustre. runners.yaml:611 adds DeepSeek-V4-Pro-0813: {root: lustre-models, ...} with no @ numa1 alias, and lustre-models has no visibility: node-local (runners.yaml:630), so it defaults to shared (infx/clusters/slurm.py:59). The sibling DeepSeek-V4-Pro@ numa1 entry (runners.yaml:619) is how same-cluster non-overridden dsv4 jobs normally get node-local picks via models.py's min(copies, key=lambda name: not node_local(name)). Fix: add a staged DeepSeek-V4-Pro-0813@ numa1 node-local entry, or correct perf-changelog.yaml:9179/9193's claim.

Why this was flagged

Trigger: any of the 8 new gb200-nv dsv4-pro-0813 dynamo-sglang AgentX recipes launches (agg or disagg, all added in this diff) and models.checkpoint() in infx/launch/drivers/srt/models.py resolves MODEL=deepseek-ai/DeepSeek-V4-Pro-0813 with no OVERRIDES entry for gb200-nv, so it falls to the basename scan over cluster.models.entries. Only one copy exists, DeepSeek-V4-Pro-0813 on lustre-models (shared, runners.yaml:611/630), unlike DeepSeek-V4-Pro@ numa1 (runners.yaml:619) which lets the default node-local-preferring logic pick local storage for the sibling model. So every node in every one of these jobs (up to 32 GPUs at concurrency 1280) reads the checkpoint over the shared Lustre mount at job start, contradicting perf-changelog.yaml:9179 and :9193's explicit claim of 'node-local DeepSeek-V4-Pro checkpoint staging', which this PR itself introduces as a factual record of the change.

Verification: The mechanism is real and reachable: basename matches only runners.yaml:611 (DeepSeek-V4-Pro-0813 on lustre-models), lustre-models (runners.yaml:630) sets no visibility, and Volume.visibility defaults to "shared" (clusters/base.py:22), so node_local is False. All 8 new recipes stage from shared Lustre, while perf-changelog.yaml claims "node-local" staging, and there is no DeepSeek-V4-Pro-0813@ numa1 node-local copy.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant