Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
将 B300 AgentX 性能变更记录关联到上游 PR #3190。
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36742392724 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36742392724 |
There was a problem hiding this comment.
Beyond the two inline findings, I checked the decode-role env var SGLANG_OPT_DEEPGEMM_MEGA_MOE (line 163 of the 1p1d disagg recipe), which lacks the _USE_ infix used on the prefill side and in perf-changelog.yaml — it matches the identical pattern already present in the pre-existing sibling recipe gb300-fp4/agentx/disagg-1p1d-dep8-dep16-c480-mtp-kvoffload.yaml, so it's an existing convention being copied forward here, not a new bug.
Extended reasoning...
This run reported two confirmed findings (the AGENTS.md-prohibited set -euo pipefail in the new Mooncake/EFA install script, and the missing AgentX HF/dataset-mmap cache mounts in the new multi-node srtslurm block) that will post as inline comments, so a human review is warranted regardless. I independently spot-checked one additional candidate — the SGLANG_OPT_DEEPGEMM_MEGA_MOE vs SGLANG_OPT_USE_DEEPGEMM_MEGA_MOE naming discrepancy between the prefill and decode roles in the 1p1d disagg recipe — by grepping sibling recipes under benchmarks/multi_node/srt-slurm-recipes/dsv4/sglang/gb300-fp4/agentx/, and found the exact same asymmetric naming already present there, confirming it's a copied existing pattern rather than something newly introduced by this PR. I did not find independent grounds to override the two confirmed findings or to add further net-new issues beyond what's already queued as inline comments.
Findings marked 🟡 are optional suggestions and need no follow-up push.
059fc7f to
daee924
Compare
8e64e2c to
0231870
Compare
8e52842 to
9a27748
Compare
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
9a27748 to
8a36b64
Compare
ec75560 to
ec0e63d
Compare
ec0e63d to
1e8abd9
Compare
1e8abd9 to
18c53d0
Compare
|
|
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
18c53d0 to
75c5b3f
Compare
|
[by Codex] Rebased onto current 中文已 rebase 到当前 |
069ba7d to
1f0889b
Compare
新增基于 Mooncake 的 B300 DeepSeek V4 AgentX 配方,并适配 Python 启动器。
移除配方中与集群级每 GPU CPU 配置冲突的任务级 CPU 参数。
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gentX The pluggable launcher resolves DeepSeek-V4-Pro-0813 on b300-dsxe to the shared /data/models copy for every non-vLLM framework, including multi-node Dynamo+SGLang, which previously read the node-local /scratch/models copy. With several points loading concurrently from shared storage, the TP8 c1 agg worker spent ~2.9h in weight loading and lost its etcd lease; the TP4 c4 agg worker never became healthy within the 4h window on the previous head. Pin every point of both configs to /scratch/models/DeepSeek-V4-Pro-0813 via the MODEL_PATH additional-setting, restoring the storage path under which this recipe set last passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0d69ea6 to
e7f6480
Compare
…oint A point-level MODEL_PATH is opaque to the launcher, so srtctl's model preflight (srt-slurm v2.39.1) ran on the runner host and rejected the node-local /scratch/models path. Instead, add dynamo-sglang to the b300-dsxe override that already routes vLLM to DeepSeek-V4-Pro-0813@scratch: the checkpoint is then known to be node-local, model_paths maps the recipe alias to /scratch/models/DeepSeek-V4-Pro-0813, and preflight is skipped as it is for vLLM. Single-node SGLang keeps the shared /data/models copy. Drop the MODEL_PATH additional-settings added in the previous commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
[by Codex]
Adds DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX recipes:
mooncake-transfer-engine-efa-cuda13==0.3.13.post1wheel during disaggregated container setup; no SGLang source patch is applied./opt/amazon/efaor/opt/amazon/ofi-ncclmounts.PROMETHEUS_MULTIPROC_DIRsettings and custom Prometheus-directory creation.The c64 recipe follows the existing B200 DSV4 c64 precedent while retaining the B300 c240 recipe's platform settings. Its decode
max-running-requestsis 128, twice the client concurrency.Local validation passed: YAML parsing, shell syntax, exact-key matrix generation, append-only perf-changelog validation, focused launcher/changelog tests, topology/speculation checks, and pinned srt-slurm recipe migration verification for all five recipes.
AI model disclosure
中文
新增 DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX 配方:
mooncake-transfer-engine-efa-cuda13==0.3.13.post1wheel,不修改 SGLang 源码。/opt/amazon/efa或/opt/amazon/ofi-nccl。PROMETHEUS_MULTIPROC_DIR设置及自定义 Prometheus 目录创建。c64 配方沿用现有 B200 DSV4 c64 先例,同时保留 B300 c240 配方的平台设置。其 decode
max-running-requests为 128,即客户端并发的两倍。本地验证全部通过:YAML 解析、Shell 语法、精确配置键矩阵生成、仅追加 perf-changelog 验证、启动器/变更日志专项测试、拓扑/投机解码检查,以及全部五个配方针对固定版本 srt-slurm 的迁移验证。
AI 模型披露