[NV] Move GB300 DeepSeek-V4-Pro AgentX disagg to the Mooncake external linker / 将 GB300 DeepSeek-V4-Pro AgentX 分离式配置切换到 Mooncake external linker - #3187
Conversation
8d51100 to
50dce57
Compare
50dce57 to
d77ecff
Compare
| - "Prefill 改为按 2048 token 切分单个请求并延迟中间 chunk 的 KV 传输,mem-fraction-static 从 0.85 提升到 0.92,镜像升级到 lmsysorg/sglang:nightly-dev-cu13-20260908-20ca564b。" | ||
| - "在相同 golden acceptance length 下相比原配置实测:1P1D、2P1D、4P1D 的单卡总吞吐分别提升 13.2%、18.4%、17.5%,平均 TTFT 分别下降 21.6%、30.7%、37.3%,三个点的稳态前缀命中率均为 96.0%-96.2%。" | ||
| - "临时方案:服务端改动仍在 sgl-project/sglang#39694 评审中,因此 runner 将该分支克隆到 srt-slurm 的 configs 目录,服务进程通过 PYTHONPATH 导入,同时继续使用容器内编译好的 kernel。配套 setup 脚本在该源码树缺失或版本不符时直接让任务失败,避免静默回退到容器自带的 SGLang。待改动进入固定镜像后应移除克隆、setup 脚本和 recipe 中的 PYTHONPATH。" | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/PLACEHOLDER |
There was a problem hiding this comment.
🔴 New changelog entry's pr-link uses .../pull/PLACEHOLDER, which the changelog validator rejects for PR runs, so CI will fail and block this PR from merging. validate_added_pr_link in infx/workflows/validate_perf_changelog.py only accepts the exact pull/{pr_number} link or the literal placeholders "XXX" / "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX" (PR_LINK_PLACEHOLDERS, lines 24-27); PLACEHOLDER matches neither. Fix: replace the pr-link value with XXX (or the real PR number once known) so validate_added_pr_link accepts it.
Extended reasoning...
docs/configuration-procedures.md says the model+hardware playbook permits pr-link: TBD before the PR exists, but the actual validator infx/workflows/validate_perf_changelog.py never allows 'TBD' or 'PLACEHOLDER' — it only allows 'XXX' or the exact pull/ URL (PR_LINK_PLACEHOLDERS set, lines 24-27, checked in validate_added_pr_link lines 144-159). This diff's new entry at perf-changelog.yaml:7589 sets pr-link to 'https://github.com/SemiAnalysisAI/InferenceX/pull/PLACEHOLDER'. When validate_perf_changelog.py runs in CI with pr_number set (a PR run), it computes expected='https://github.com/SemiAnalysisAI/InferenceX/pull/{pr_number}', sees link not in PR_LINK_PLACEHOLDERS and link != expected, and raises ChangelogValidationError('new PR entry must use ... or an XXX placeholder; found ...'). This blocks the PR's changelog check from passing until a maintainer or the author manually edits the line to 'XXX' or the true PR number.
Verification: normal. The new changelog entry at perf-changelog.yaml:7589 sets pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/PLACEHOLDER. The validator infx/workflows/validate_perf_changelog.py rejects this in every mode: (a) on PR runs (pr_number set), validate_added_pr_link (lines 144-160) requires the link be in PR_LINK_PLACEHOLDERS = {"XXX",… | normal. The new changelog entry sets…
| @@ -6664,7 +9063,7 @@ dsv4-fp4-gb300-dynamo-sglang-agentic-disagg: | |||
| - spec-decoding: draft_model | |||
| conc-list: [480] | |||
| kv-offloading: dram | |||
| kv-offload-backend: { name: hicache } | |||
| kv-offload-backend: { name: mooncake } | |||
There was a problem hiding this comment.
🟡 (optional) Operators reading nvidia-master.yaml lose the Mooncake component version that CONFIGS.md requires, making sweep provenance for this backend unrecoverable if the pinned Mooncake build ever needs to be identified after the fact. All four new entries set kv-offload-backend: { name: mooncake } with no version, but configs/CONFIGS.md:136-137 explicitly names Mooncake as an example of an independently-versioned backend that should supply version (unlike framework-native vLLM/HiCache). The Pydantic validator (KVOffloadBackendMetadata in infx/matrix/validation.py:127-132) doesn't enforce this, so nothing blocks the omission. …
Extended reasoning...
…Fix: add a non-image version (e.g. the Mooncake release/commit) for the mooncake kv-offload-backend entries, covering all four sites (configs/nvidia-master.yaml:9066,9084,9102,9120), or update CONFIGS.md if this backend is now treated as framework-native.
configs/CONFIGS.md:133-138 says kv-offload-backend requires non-empty name, version optional only for framework-native backends (vLLM built-in, SGLang HiCache), and explicitly lists Mooncake as needing version since it's independently versioned. The new entries at lines 9066, 9084, 9102, 9120 all set { name: mooncake } with no version key. infx/matrix/validation.py's KVOffloadBackendMetadata makes version fully Optional with no name-based branching (confirmed by reading lines 127-143, and by utils/matrix_logic/test_validation.py tests that never require version for any specific backend name). So CI will not fail, but the documented provenance convention is silently violated: nobody can tell from the master config which Mooncake build/commit these Pareto points were validated against, unlike the equivalent LMCache examples in…
Verification: nit. The deviation is real but the consequence is mild/overstated. CONFIGS.md:135-138 states version is optional only for framework-native backends and explicitly directs "Supply version for independently versioned backends such as LMCache or Mooncake." All four new search-space entries in configs/nvidia-master.yaml (lines 9066, 9084, 9102, 9120) set kv-offload-backend: { name: mooncake }…
| SGLANG_MOONCAKE_OPT_URL="https://github.com/weireweire/sglang.git" | ||
| SGLANG_MOONCAKE_OPT_PIN="d2cf19e69fe7f9d69a7504012c3617fe6669acbd" | ||
| git init configs/sglang-mooncake-opt | ||
| git -C configs/sglang-mooncake-opt fetch --depth 1 \ | ||
| "$SGLANG_MOONCAKE_OPT_URL" "$SGLANG_MOONCAKE_OPT_PIN" | ||
| git -C configs/sglang-mooncake-opt checkout FETCH_HEAD | ||
| cp "$GITHUB_WORKSPACE/benchmarks/multi_node/srt-slurm-recipes/configs/dsv4-gb300-sglang-mooncake-opt.sh" \ | ||
| configs/dsv4-gb300-sglang-mooncake-opt.sh |
There was a problem hiding this comment.
🟡 (optional) Every gb300 dsv4 dynamo-sglang agentic launch now clones an external, unmerged fork, not just the 4 Mooncake recipes. The elif at runners/launch_gb300-nv.sh:302 copies the whole sglang/deepseek-v4/agentic directory, which also contains agg-gb300-tp4-mtp-lowlatency.yaml and agg-gb300-tp8-mtp-lowlatency.yaml (neither uses hicache/mooncake/setup_script). The new git init/fetch/checkout block at lines 323-330 runs unconditionally in that branch, and set -exo pipefail (line 5) aborts the whole launch if the fetch fails. Fix: gate the clone on the specific recipe(s) that set setup_script: dsv4-gb300-sglang-mooncake-opt.sh (e.g. check MODEL_SUFFIX/recipe filename or grep the copied recipe for that key) so unrelated gb300 dsv4 launches don't depend on github.com/weireweire/sglang.git availability.
Extended reasoning...
runners/launch_gb300-nv.sh:302 elif matches IS_AGENTIC=1, FRAMEWORK=dynamo-sglang, MODEL_PREFIX=dsv4 for ANY recipe under sglang/deepseek-v4/agentic, including the two lowlatency aggregated recipes that keep no hicache/mooncake settings at all. Lines 325-328 unconditionally run git init, git fetch --depth 1 against https://github.com/weireweire/sglang.git pinned to commit d2cf19e69fe7f9d69a7504012c3617fe6669acbd, and git checkout FETCH_HEAD. Script header has set -exo pipefail (line 5), so if that fetch fails — fork renamed/deleted, GH rate limit, network blip, or the author force-pushes over the pinned commit — the whole launch script exits nonzero and the job never gets submitted, even for the lowlatency recipes that never read configs/sglang-mooncake-opt or PYTHONPATH. Before this diff, launching agg-gb300-tp4-mtp-lowlatency.yaml or agg-gb300-tp8-mtp-lowlatency.yaml never touched any third-party GitHub fork; after merging, both silently gain a hard dependency on an unrelated contributor's personal repo staying reachable and unchanged at that SHA.
Verification: Severity: nit. The mechanism is real and reachable; the launch abort is conditional on the external clone failing. runners/launch_gb300-nv.sh:302 elif matches on coarse env vars (IS_AGENTIC==1, FRAMEWORK==dynamo-sglang, MODEL_PREFIX==dsv4), not per-recipe, so it handles EVERY dsv4/dynamo-sglang/agentic launch. Lines 325-328 run git init configs/sglang-mooncake-opt, `git fetch --depth 1… | nit.…
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36809646116 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36809646116 |
同步 main,使用字节保留流程恢复 PR #3187 的 perf-changelog 追加项,并记录 CUDA graph 参数修复。
同步 main,使用字节保留流程恢复 PR #3187 的 perf-changelog 追加项,并记录 CUDA graph 参数修复。
ab576df to
9fcb8a6
Compare
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
7565d48 to
6317a45
Compare
|
rabased, glad we have utilized the override feature in srt-slurm. We may also able to utilize zip_override in the future |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
…linker Replace hierarchical cache with the Mooncake unified-cache external linker on the image's own SGLang and bump the image to nightly-dev-cu13-20260916-c9a8fba9. Decode nodes lend host DRAM to the store through four 180GB standalone mooncake_client daemons each, started by the recipe's setup script and registered with the HTTP metadata server; the job fails unless every daemon serves and resolves. Prefill runs 16k-token chunks per DP rank with MegaMoE sized to match and mem-fraction-static 0.80. The benchmark client moves to the dedicated etcd/nats node, where it no longer OOM-kills prefill_0, and Tachometer is off for this ladder. The gb300-nv srt lane gains role_env, which sets this cluster's RDMA devices on the recipe's prefill and decode roles in place of the removed bash launcher block.
6317a45 to
e8afb89
Compare
Dynamo 1.5.0.dev20260914 falls back when ServerArgs.get_model_config is missing (ai-dynamo/dynamo#14234), matching the B200 recipes on the same image, so the sitecustomize shim and its PYTHONPATH entry go away.
…er reads Decode ranks no longer mount a store segment, so MOONCAKE_GLOBAL_SEGMENT_SIZE and MOONCAKE_STANDALONE_STORAGE had no effect there; the daemons take their size from MC_SIDECAR_SEGMENT_SIZE.
…he launcher Set MOONCAKE_DEVICE on the prefill and decode roles directly, as the GB300 vLLM Mooncake recipes already do, and drop the srt lane role_env hook.
…er runs The image ships mooncake-transfer-engine-cuda13 0.3.13; CONFIGS.md asks independently versioned offload backends to state their version.
Replace the recipe setup script with four mooncake-store services on the decode nodes (180gb each, ports 8800-8803). srt-slurm injects the master and HTTP metadata server, starts them before the workers, and gates on their readiness ports, so the custom daemon script goes away.
Decode workers no longer act as Mooncake store clients; the decode stores carry their own MOONCAKE_DEVICE/PROTOCOL, and P/D KV transfer does not read them.
…80 point, and give decode stores 600 s to start
…0 disagg point and drop the DEP32 point
…x-mooncake-opt # Conflicts: # inferencex-e2e/perf-changelog.yaml
…the AgentX power topology check
Summary / 概要
Move the GB300 DeepSeek-V4-Pro Dynamo-SGLang AgentX disaggregated Pareto ladder (1P1D c120 and c480, 2P1D c960, 3P1D c1440, 4P1D c2400) from hierarchical cache to the Mooncake unified-cache external linker, using the image's own SGLang.
将 GB300 DeepSeek-V4-Pro Dynamo-SGLang AgentX 分离式 Pareto 配置(1P1D c120 与 c480、2P1D c960、3P1D c1440、4P1D c2400)从 hierarchical cache 切换到 Mooncake unified-cache external linker,直接使用镜像自带的 SGLang。
Serve the host KV tier through the unified-cache external linker on Mooncake, with group semantics enabled. Prefill keeps a 140GB segment per rank.
Lend decode-node host DRAM to the store through srt-slurm
mooncake-storeservices: four 180GB standalone stores per decode node (ports 8800-8803). srt-slurm points them at the managed master and its HTTP metadata server, starts them before the workers, and gates on their readiness ports with a 600-second timeout (a cold container start on a decode node exceeded srtctl's 120-second default in CI). SGLang itself is unmodified and no custom script is involved.Prefill runs 16k-token chunks per DP rank (
chunked-prefill-size131072,SGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK17408) withmem-fraction-static0.80, which the DSV4 indexer needs at that chunk size.Run the benchmark client on the dedicated etcd/nats node (
benchmark.placement.node: dedicated). On the default head node it shares prefill_0's host and its ~300GB footprint OOM-kills that worker.Disable Tachometer for this ladder; its per-node exporters slowed decode steps by a few percent.
Bump the image to
lmsysorg/sglang:nightly-dev-cu13-20260916-c9a8fba9and record the offload backend as mooncake in the master config.Bump the Dynamo wheel to
1.5.0.dev20260914(same as the B200 recipes on this image); its SGLang worker falls back whenServerArgs.get_model_configis missing, so no compatibility shim is needed.主机侧 KV 层改由 Mooncake 上的 unified-cache external linker 提供,并启用 group semantics;prefill 每个 rank 保留 140GB segment。
decode 节点的主机 DRAM 通过 srt-slurm 的
mooncake-store服务借给 store:每个 decode 节点 4 个 180GB 的独立 store(端口 8800-8803)。srt-slurm 负责将其指向托管的 master 及其 HTTP metadata server、在 worker 之前启动,并以就绪端口作为启动检查,超时 600 秒(CI 中曾有 decode 节点冷启动容器超过 srtctl 默认的 120 秒)。SGLang 本身不做任何修改,也不再需要自定义脚本。prefill 每个 DP rank 按 16k token 切分(
chunked-prefill-size131072,SGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK17408),mem-fraction-static降至 0.80,这是该切分大小下 DSV4 indexer 所需。压测客户端运行在独立的 etcd/nats 节点上(
benchmark.placement.node: dedicated)。默认的 head 节点与 prefill_0 共用主机,客户端约 300GB 的内存占用会导致该 worker 被 OOM kill。该阶梯关闭 Tachometer,其逐节点 exporter 会使 decode 单步变慢数个百分点。
镜像升级到
lmsysorg/sglang:nightly-dev-cu13-20260916-c9a8fba9,master config 中的 offload backend 记为 mooncake。Dynamo wheel 升级到
1.5.0.dev20260914(与同镜像的 B200 配方一致),其 SGLang worker 在缺少ServerArgs.get_model_config时会自动兼容,因此不再需要兼容补丁。RDMA devices / RDMA 设备
The recipe sets
MOONCAKE_DEVICE=mlx5_0,mlx5_1,mlx5_2,mlx5_3on the prefill role and on each decode store service, the processes that act as Mooncake store clients, as the GB300 vLLM Mooncake recipes already name their devices. Decode workers only use Mooncake for P/D KV transfer, which does not read it. Without an explicit list the store client falls back to NVLink for same-domain peers, whose address lookup fails and aborts every worker during the store warmup put. No launcher change is needed.配方在 prefill 角色以及每个 decode store 服务上(即作为 Mooncake store 客户端的进程)设置
MOONCAKE_DEVICE=mlx5_0,mlx5_1,mlx5_2,mlx5_3,与 GB300 vLLM Mooncake 配方直接写明设备的做法一致。decode worker 只用 Mooncake 做 P/D 间的 KV 传输,不读取该变量。未显式指定时,store 客户端会对同 NVLink 域的对端退回 NVLink,地址查找失败并在 warmup put 阶段让所有 worker 退出。无需修改启动器。Validation / 验证
The 4P1D c1920 point was run end to end on GB300 through the same srtctl path CI uses, with the same topology, workload, 3600-second measurement, AgentX warmup, and golden synthetic acceptance length as the current configuration. It completed the full measurement with zero request errors, every decode node's four stores were serving, and steady-state prefix cache hit rate was 96%.
The 4P1D topology at c2400 completed the same end-to-end run with zero request errors: 135,102 total tok/s/GPU (vs 126,755 at c1920), 25.1 ms ITL; prefill still had headroom at c1920.
The 1P1D point at c120 completed the same end-to-end run: 26,681 total tok/s/GPU at 6.73 ms P90 TPOT (149 tok/s/user). The aggregate TP4 c8 point it replaces is 12,415 tok/s/GPU at 6.85 ms P90 TPOT (146 tok/s/user) on the official dashboard, so the aggregate TP4 point is removed from dsv4-fp4-gb300-dynamo-sglang-agentic-agg.
All disaggregated points resolve and dry-run with the CI srtctl; the perf-changelog gate passes.
The 1P1D, 2P1D and 3P1D points take the same configuration delta and are covered by this PR's sweep.
4P1D c1920 已在 GB300 上通过与 CI 相同的 srtctl 路径完整跑通,拓扑、负载、3600 秒测量、AgentX warmup 及 golden synthetic acceptance length 均与当前配置一致;完整跑满测量且请求错误为零,每个 decode 节点的 4 个 store 均已就绪,稳态前缀命中率为 96%。
4P1D 拓扑在并发 2400 下同样完整跑通且请求错误为零:135,102 total tok/s/GPU(c1920 为 126,755),ITL 25.1 ms;c1920 时 prefill 仍有余量。
1P1D 并发 120 同样完整跑通:26,681 total tok/s/GPU,P90 TPOT 6.73 ms(149 tok/s/user)。被其替代的聚合 TP4 c8 在官方 dashboard 上为 12,415 tok/s/GPU、P90 TPOT 6.85 ms(146 tok/s/user),因此从 dsv4-fp4-gb300-dynamo-sglang-agentic-agg 中移除聚合 TP4 点。
所有分离式点均可由 CI srtctl 解析并通过 dry-run,perf-changelog 检查通过。
1P1D、2P1D、3P1D 采用相同的配置改动,由本 PR 的 sweep 覆盖。