Skip to content

[NVIDIA][AgentX] GB300 DeepSeek-V4.1-Flash Dynamo+SGLang aggregated and 1PxD recipes / [NVIDIA][AgentX] GB300 DeepSeek-V4.1-Flash Dynamo+SGLang 聚合与 1PxD 配方 - #3598

Open
nvpohanh wants to merge 7 commits into
mainfrom
pohanh/dsv41flash-gb300-agentx-curve
Open

nvpohanh wants to merge 7 commits into
mainfrom
pohanh/dsv41flash-gb300-agentx-curve

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Add DeepSeek-V4.1-Flash FP4 AgentX recipes for GB300 with Dynamo + SGLang:

  • three aggregated points: TP4/EP1 c1 and TP2/EP2 c64 with prefill-decode interval 16 or 8;
  • six Mooncake disaggregated points: 1P1D c8/c16/c64/c96, 1P2D c64, and 1P4D c64;
  • registry and performance-changelog entries for both curves.

The recipes run on the default gb300-nv Slurm partition, account and shared caches, like the other GB300 DeepSeek-V4.1-Flash configs. No launcher or cluster-schema changes.

The stack is digest-pinned lmsysorg/sglang:nightly-dev-20260928-81f27fb3 with ai-dynamo 1.6.0.dev20260928, Dynamo KV routing with session affinity, DSpark block size 5, and Mooncake KV transfer for disaggregated points.

Validation

  • parsed and schema-validated configs/runners.yaml (unchanged from main);
  • generated all nine exact-key and filtered full-sweep matrix entries with node counts 1/1/1/2/2/2/2/2/3;
  • passed infx.workflows.validate_perf_changelog against the latest main;
  • passed the 25-test cluster-config suite and the repository-required Ruff check/format pass;
  • verified both recipe files still match the exercised configurations byte-for-byte.

AI model disclosure

  • GPT-5 (Codex) ported the configuration set onto the latest main, resolved the launcher-refactor conflict, prepared this description, and monitored CI. Claude Code (Claude Opus 5.5) monitored the sweep and removed a GB300 Slurm route that CI runners could not use, so the recipes run on the default route. Final configuration selection and submission are by Po-Han Huang.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml
中文

说明

新增 GB300 上 DeepSeek-V4.1-Flash FP4 的 Dynamo + SGLang AgentX 配方:

  • 三个聚合点:TP4/EP1 c1,以及 prefill-decode interval 为 16 或 8 的 TP2/EP2 c64;
  • 六个 Mooncake 分离式点:1P1D c8/c16/c64/c96、1P2D c64 和 1P4D c64;
  • 两条曲线对应的注册表与性能变更日志条目。

配方使用默认的 gb300-nv Slurm 分区、账号和共享缓存,与其他 GB300 DeepSeek-V4.1-Flash 配置一致;不修改启动器或集群 schema。

软件栈固定为摘要锁定的 lmsysorg/sglang:nightly-dev-20260928-81f27fb3、ai-dynamo 1.6.0.dev20260928、带会话亲和性的 Dynamo KV 路由、DSpark block size 5;分离式点使用 Mooncake KV 传输。

验证

  • 已解析并通过 schema 校验 configs/runners.yaml(与 main 相同);
  • 精确 key 与过滤后的 full-sweep 均生成 9 条矩阵记录,节点数为 1/1/1/2/2/2/2/2/3;
  • infx.workflows.validate_perf_changelog 在最新 main 上通过;
  • 25 项集群配置测试以及仓库要求的 Ruff 检查/格式化均通过;
  • 两个配方文件仍与已执行验证的配置逐字节一致。

AI 模型披露

  • GPT-5(Codex)将配置集迁移到最新 main,解决启动器重构冲突,准备本说明并监控 CI。Claude Code(Claude Opus 5.5)监控 sweep,并移除了 CI runner 无法使用的 GB300 Slurm 路由,使配方使用默认路由。最终配置选择和提交由 Po-Han Huang 完成。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it changes SLURM launcher routing (new partition/account/shared-root branch) for multi-node GB300 jobs, a human look would still be worthwhile to confirm the operational details before merge.

What was reviewed:

  • Launcher branch in launch_gb300-nv.sh: check_env_vars now covers both SLURM_PARTITION and SLURM_ACCOUNT, and the new values are set via if/else rather than :- fallback.
  • GB300_SHARED_ROOT is "/data/home/slurm-shared/gharunners" on the new restricted route vs "sa-shared" used everywhere else in the repo — plausible if the restricted partition has its own mount, but not verifiable from source alone.
  • perf-changelog.yaml entry is appended at the file's tail as required; its pr-link is a "pull/0" placeholder rather than the documented pre-PR "TBD" (docs/configuration-procedures.md:636).
  • The new agg/disagg recipe files are paired with matching master-config entries as required for multi-node srt-slurm additions.
Extended reasoning...

The diff adds a GB300 multi-node AgentX benchmark route (two srt-slurm recipe files, master-config wiring, a changelog entry, and new conditional branching in the shared launch_gb300-nv.sh script that selects a different SLURM partition/account and a new shared-storage root for this model/framework combination). It touches no auth, crypto, or data-exposure surface; the risk is purely operational (job routing and shared-cache paths for a NVIDIA benchmark CI runner). The bug hunter reported no findings, but the launcher change introduces production-affecting branching logic and an unverified shared-root path divergence from repo convention, plus a changelog pr-link that deviates from the documented pre-merge placeholder convention, which together are enough that a human familiar with the cluster layout should confirm before merge.

This review covers commit 74b1996, which is no longer the latest commit on this pull request; later commits are not covered by it.

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

Port the nine exercised aggregated and disaggregated AgentX points to the pluggable launcher. Add a named restricted GB300 Slurm route for the alternate partition and shared storage.\n\n将九个已验证的聚合与分离式 AgentX 点迁移到可插拔启动器,并为备用分区和共享存储添加命名的 GB300 restricted Slurm 路由。
@nvpohanh
nvpohanh force-pushed the pohanh/dsv41flash-gb300-agentx-curve branch from 6ff9e47 to 7ababb3 Compare September 30, 2026 00:52
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

@nvpohanh nvpohanh changed the title [NVIDIA][AgentX] GB300 DeepSeek-V4.1-Flash Dynamo+SGLang aggregated and 1PxD recipes [NVIDIA][AgentX] GB300 DeepSeek-V4.1-Flash Dynamo+SGLang aggregated and 1PxD recipes / [NVIDIA][AgentX] GB300 DeepSeek-V4.1-Flash Dynamo+SGLang 聚合与 1PxD 配方 Sep 30, 2026
nvpohanh and others added 4 commits September 29, 2026 18:55
Cover every vertex of the measured Dynamo+SGLang Pareto curve in the two AgentX recipe files: add the two-GPU TP2/EP2 c1 arm, the 1P2D c48 and spread c16 cells, and the HiCache prefill tier at 1P1D c160 and 2P1D c256; move the multi-decode cells (1P2D c48/c64, 1P4D c64, 2P1D c256) to the 1 s session-affinity TTL they were measured and GSM8K-gated with; refresh the measured numbers in the file headers.

在两个 AgentX 配方文件中覆盖已测得的 Dynamo+SGLang Pareto 曲线的每个顶点:新增双 GPU 的 TP2/EP2 c1 配置、1P2D c48 与跨节点 c16 单元,以及 1P1D c160 与 2P1D c256 的 HiCache prefill 层;将多 decode 单元(1P2D c48/c64、1P4D c64、2P1D c256)改为其实际测量与 GSM8K 验证所用的 1 秒会话亲和 TTL;同时更新文件头中的测量数据。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Trim the aggregated and disaggregated AgentX recipes and the nvidia-master sweep to the eight measured Pareto vertices: agg TP4/EP1 c1 and TP2/EP2 c1; disagg 1P2D spread c16, 1P2D c48, 1P2D c64, 1P1D c96, 1P1D HiCache c160 and 2P1D HiCache c256. Drop the dominated aggregated c64 (prefill-decode-interval 8/16), 1P1D c8/c16/c64 and 1P4D c64 arms, add master entries for the new overrides, and update the perf-changelog description to the shipped set.

将聚合与分离式 AgentX 配方以及 nvidia-master sweep 精简为实测 Pareto 曲线的八个顶点:聚合 TP4/EP1 c1 与 TP2/EP2 c1;分离式 1P2D 跨节点 c16、1P2D c48、1P2D c64、1P1D c96、1P1D HiCache c160 与 2P1D HiCache c256。移除被支配的聚合 c64(prefill-decode-interval 8/16)、1P1D c8/c16/c64 与 1P4D c64 配置,为新增 override 添加 master 条目,并将 perf-changelog 描述更新为最终交付集合。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Re-append the perf-changelog entry after main's newer entries so the changelog stays append-only.

将 upstream/main 合并进 pohanh/dsv41flash-gb300-agentx-curve,并把 perf-changelog 条目重新追加到 main 的新条目之后,保持只增不删。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…m route

Drop the named `restricted` GB300 Slurm route (batch_2 partition, restricted
account, alternate shared caches). The CI runners are not associated with that
account, so the canary failed at image import with "Invalid account or
account/partition combination specified". The recipes now use the default
gb300-nv partition, account and shared caches, like the other GB300
DeepSeek-V4.1-Flash configs. Revert the route plumbing in the Slurm settings,
SRT launcher, tests and docs, and drop the matching changelog bullet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nvpohanh and others added 2 commits October 1, 2026 02:30
Resolve the perf-changelog.yaml conflict by keeping main and appending this PR's entry at the end.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egated recipe

The aggregated AgentX power path now requires TP, PP_SIZE and PCP_SIZE before
replay. Set PP_SIZE=1 and PCP_SIZE=1 in the base benchmark env, as the Qwen3.5
GB300 aggregated recipe does, so the replay no longer exits on missing inputs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants