Skip to content

Add Mooncake B300 Dynamo+SGLang AgentX configs for DeepSeek V4 / 添加基于 Mooncake 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 - #3190

Open
nvpohanh wants to merge 5 commits into
mainfrom
dsv4-fp4-b300-dynamo-sglang-agentic
Open

nvpohanh wants to merge 5 commits into
mainfrom
dsv4-fp4-b300-dynamo-sglang-agentic

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

[by Codex]

Adds DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX recipes:

  • Aggregate: TP8 c1 and TP4 c4.
  • Disaggregated with Mooncake KV transfer: 1P1D DEP8 c64/c240 and 2P1D DEP8 c480.
  • Installs only the public mooncake-transfer-engine-efa-cuda13==0.3.13.post1 wheel during disaggregated container setup; no SGLang source patch is applied.
  • Relies on the DSXE-managed fabric injection without explicit /opt/amazon/efa or /opt/amazon/ofi-nccl mounts.
  • Uses the current Python launcher and cluster registry, including persistent AgentX and Hugging Face caches.
  • Removes explicit PROMETHEUS_MULTIPROC_DIR settings and custom Prometheus-directory creation.
  • Keeps prefill and decode on the same DSpark K=6 method.

The c64 recipe follows the existing B200 DSV4 c64 precedent while retaining the B300 c240 recipe's platform settings. Its decode max-running-requests is 128, twice the client concurrency.

Local validation passed: YAML parsing, shell syntax, exact-key matrix generation, append-only perf-changelog validation, focused launcher/changelog tests, topology/speculation checks, and pinned srt-slurm recipe migration verification for all five recipes.

AI model disclosure

  • GPT-5 (Codex, including one delegated agent): prepared the original submission and the current Mooncake rebase, configuration update, validation, PR refresh, and CI review.
  • GPT-6 (Codex): diagnosed and prepared preceding runtime repairs.
中文

新增 DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX 配方:

  • 聚合式:TP8 c1 与 TP4 c4。
  • 使用 Mooncake KV 传输的分离式:1P1D DEP8 c64/c240 与 2P1D DEP8 c480。
  • 分离式容器初始化时仅安装公开的 mooncake-transfer-engine-efa-cuda13==0.3.13.post1 wheel,不修改 SGLang 源码。
  • 依赖 DSXE 管理的网络栈注入,不显式挂载 /opt/amazon/efa 或 /opt/amazon/ofi-nccl。
  • 使用当前 Python 启动器和集群注册表,并保留持久化 AgentX 与 Hugging Face 缓存。
  • 移除显式 PROMETHEUS_MULTIPROC_DIR 设置及自定义 Prometheus 目录创建。
  • Prefill 与 decode 保持使用相同的 DSpark K=6 方法。

c64 配方沿用现有 B200 DSV4 c64 先例,同时保留 B300 c240 配方的平台设置。其 decode max-running-requests 为 128,即客户端并发的两倍。

本地验证全部通过:YAML 解析、Shell 语法、精确配置键矩阵生成、仅追加 perf-changelog 验证、启动器/变更日志专项测试、拓扑/投机解码检查,以及全部五个配方针对固定版本 srt-slurm 的迁移验证。

AI 模型披露

  • GPT-5(Codex,包括一个委派代理):准备最初提交与本次 Mooncake rebase、配置更新、验证、PR 刷新及 CI 审查。
  • GPT-6(Codex):诊断并准备前序运行时修复。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

nvpohanh added a commit that referenced this pull request Sep 16, 2026
将 B300 AgentX 性能变更记录关联到上游 PR #3190。
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the two inline findings, I checked the decode-role env var SGLANG_OPT_DEEPGEMM_MEGA_MOE (line 163 of the 1p1d disagg recipe), which lacks the _USE_ infix used on the prefill side and in perf-changelog.yaml — it matches the identical pattern already present in the pre-existing sibling recipe gb300-fp4/agentx/disagg-1p1d-dep8-dep16-c480-mtp-kvoffload.yaml, so it's an existing convention being copied forward here, not a new bug.

Extended reasoning...

This run reported two confirmed findings (the AGENTS.md-prohibited set -euo pipefail in the new Mooncake/EFA install script, and the missing AgentX HF/dataset-mmap cache mounts in the new multi-node srtslurm block) that will post as inline comments, so a human review is warranted regardless. I independently spot-checked one additional candidate — the SGLANG_OPT_DEEPGEMM_MEGA_MOE vs SGLANG_OPT_USE_DEEPGEMM_MEGA_MOE naming discrepancy between the prefill and decode roles in the 1p1d disagg recipe — by grepping sibling recipes under benchmarks/multi_node/srt-slurm-recipes/dsv4/sglang/gb300-fp4/agentx/, and found the exact same asymmetric naming already present there, confirming it's a copied existing pattern rather than something newly introduced by this PR. I did not find independent grounds to override the two confirmed findings or to add further net-new issues beyond what's already queued as inline comments.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Comment thread runners/launch_b300-dsxe.sh Outdated
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch 2 times, most recently from 059fc7f to daee924 Compare September 22, 2026 15:33
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch from 8e64e2c to 0231870 Compare September 22, 2026 19:15
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch from 8e52842 to 9a27748 Compare September 23, 2026 14:20
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch from 9a27748 to 8a36b64 Compare September 28, 2026 10:13
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch 5 times, most recently from ec75560 to ec0e63d Compare September 28, 2026 22:58
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch from ec0e63d to 1e8abd9 Compare September 29, 2026 01:26
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch from 1e8abd9 to 18c53d0 Compare September 29, 2026 08:35
@hshrivastava-droid

hshrivastava-droid commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

/use 36543651714

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch from 18c53d0 to 75c5b3f Compare September 30, 2026 02:27
@nvpohanh nvpohanh changed the title Add B300 Dynamo+SGLang AgentX configs for DeepSeek V4 / 添加 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 Add Mooncake B300 Dynamo+SGLang AgentX configs for DeepSeek V4 / 添加基于 Mooncake 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 Sep 30, 2026
@nvpohanh

Copy link
Copy Markdown
Collaborator Author

[by Codex]

Rebased onto current main at 44d6f8ad9 and migrated the runner-side change to the Python launcher configuration. The deleted Bash launcher and static srt-slurm cluster profile are not restored; the required persistent AgentX cache is declared in configs/runners.yaml. The branch is conflict-free at 75c5b3fe2.

中文

已 rebase 到当前 main(44d6f8ad9),并将 runner 侧改动迁移到 Python 启动器配置。未恢复已删除的 Bash 启动器和静态 srt-slurm 集群配置;所需的持久化 AgentX 缓存在 configs/runners.yaml 中声明。当前分支在 75c5b3fe2 上无冲突。

nvpohanh and others added 4 commits September 30, 2026 09:11
新增基于 Mooncake 的 B300 DeepSeek V4 AgentX 配方,并适配 Python 启动器。
移除配方中与集群级每 GPU CPU 配置冲突的任务级 CPU 参数。
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gentX

The pluggable launcher resolves DeepSeek-V4-Pro-0813 on b300-dsxe to the
shared /data/models copy for every non-vLLM framework, including multi-node
Dynamo+SGLang, which previously read the node-local /scratch/models copy.
With several points loading concurrently from shared storage, the TP8 c1 agg
worker spent ~2.9h in weight loading and lost its etcd lease; the TP4 c4 agg
worker never became healthy within the 4h window on the previous head.

Pin every point of both configs to /scratch/models/DeepSeek-V4-Pro-0813 via
the MODEL_PATH additional-setting, restoring the storage path under which
this recipe set last passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-b300-dynamo-sglang-agentic branch from 0d69ea6 to e7f6480 Compare September 30, 2026 16:11
…oint

A point-level MODEL_PATH is opaque to the launcher, so srtctl's model
preflight (srt-slurm v2.39.1) ran on the runner host and rejected the
node-local /scratch/models path. Instead, add dynamo-sglang to the b300-dsxe
override that already routes vLLM to DeepSeek-V4-Pro-0813@scratch: the
checkpoint is then known to be node-local, model_paths maps the recipe
alias to /scratch/models/DeepSeek-V4-Pro-0813, and preflight is skipped as
it is for vLLM. Single-node SGLang keeps the shared /data/models copy.

Drop the MODEL_PATH additional-settings added in the previous commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

4 participants