[Kimi B300] repair RDMA startup and Mooncake recovery / 修复 RDMA 启动及加载恢复 - #3088
edwingao28 wants to merge 7 commits into
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
154c576 to
99a4698
Compare
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34817066729 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34818543261 |
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36852970338 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36852970338 |
| add_device(tmp_path, "rdmap86s0", driver="efa", layer="Unknown") | ||
| result = select(tmp_path) | ||
| assert result.returncode != 0 | ||
| assert result.stdout == "" |
There was a problem hiding this comment.
🟡 (optional) The new RDMA-rail and Mooncake-patch tests give no real CI safety net: neither this file nor the new cases appended to runners/test_kimik3_b300.py are referenced by any GitHub workflow. test-changelog-gate.yml explicitly lists runners/test_slurm_utils.py in both its paths: trigger and its pytest invocation, but these two new files are wired into neither it nor any other workflow. A future regression in select_mooncake_rdma_device (benchmarks/benchmark_lib.sh) or in patch_kimik3_mooncake_recovery.py will merge cleanly through CI and only surface at B300 DSXE benchmark runtime. Fix: add both files to test-changelog-gate.yml's paths: list and its pytest command (or a runners/test_*.py glob), matching the existing runners/test_slurm_utils.py wiring.
Extended reasoning...
Repo-wide grep for "test_mooncake_rdma_device" and "test_kimik3_b300" across every .github/workflows/*.yml returns zero hits. test-changelog-gate.yml's paths trigger (lines ~7-30) lists benchmarks/benchmark_lib.sh and runners/test_slurm_utils.py, and its pytest step (around line 123) runs an explicit file list ending in runners/test_slurm_utils.py -- the two new files added by this PR are absent from both. test-process-result.yml and other workflows only run unrelated test files. So these tests execute only when a developer runs pytest manually, exactly as the PR description's '11 local checks' implies. If a later PR breaks select_mooncake_rdma_device's driver/state/link_layer matching or breaks a patch context string in patch_kimik3_mooncake_recovery.py, CI shows green, and the break is only discovered when a Kimi-K3 B300 Mooncake benchmark run fails to find an RDMA rail or the launch script exits on 'unexpected patch context'.
Verification: nit. The factual claim is verified: neither new test file is wired into CI. A repo-wide grep for "test_mooncake_rdma_device" and "test_kimik3_b300" across .github/workflows/*.yml returns zero hits. test-changelog-gate.yml lists runners/test_slurm_utils.py in its paths trigger (line 33) and its pytest step (line 123), but not the two files this PR touches. test-process-result.yml runs a fixed…
合并 main 并把本 PR 的 changelog 条目重新追加到末尾,历史条目保持不变。 # Conflicts: # perf-changelog.yaml
| ["bash", "-ec", 'source "$1"; select_mooncake_rdma_device "$2"; ' | ||
| 'printf "%s %s\n" "$MOONCAKE_RAIL" "$MC_GID_INDEX"', |
|
InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest |
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Discover active Mellanox RDMA rails by mlx5_core driver on DSXE, load the host mlx5 provider, backport Mooncake load-failure recovery, and raise the master lease to 60s so transferred keys survive the observed get tail. 修复 Kimi B300 上 Mooncake 的 RDMA 启动与加载恢复:按驱动选择活动网卡、 加载主机 mlx5 provider、回移植加载失败恢复,并将 master lease 提高到 60 秒。 Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
kimik3-b300-mooncake.sh sources benchmark_lib.sh with --validation-only, but select_mooncake_rdma_device lived below that early return, so the setup script failed with exit 127 and cancelled the fail-fast canary. 将 select_mooncake_rdma_device 移到 --validation-only 门禁之上,并让测试 以相同方式 source,避免再次出现 command not found。 Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
53f504d to
2238563
Compare
Tip fail-fast 36800796192 killed agentic c40 after ~279s of engine silence under the default 300s VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS cap (shm_broadcast waits, then EngineDead / ProfileAborted). Restore the supported 1800s knob used by sibling Kimi Mooncake recipes; keep the 60s Mooncake master lease. 将 B300 Mooncake AgentX 的 execute_model 超时恢复为 1800 秒,保留 60 秒 master 租约。 Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
Append-only perf-changelog entry for restoring VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 after tip fail-fast 36800796192 c40 EngineDead / ProfileAborted. 为恢复 execute_model 1800 秒超时追加 changelog。 Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
Tip c48 already had device_name=ibp198s0f0 after setup (Wrote line is pre-patch); crash was GPU memcpy / DCP multimem under load. Match proven GB300 DCP8 compact_group_io + max_load_batch_keys=2 and MC_MAX_MR_SIZE, and refuse empty device_name or missing host mlx5 provider. 将 B300 Mooncake DCP8 IO 与已验证的 GB300 设置对齐,并在空 rail / 缺少主机 mlx5 provider 时明确失败。c48 的 Wrote device_name='' 是 patch 前转储;真实失败是高负载下的 GPU memcpy / DCP multimem。 Co-authored-by: Cursor <cursoragent@cursor.com>
c8 on tip ee33eeb had a real rail (Patched device_name=ibp198s0f0) but compact_group_io stormed TRANSFER_FAIL on ~25MiB puts then hung at 0 tok/s. Revert compact_group_io, keep max_load_batch_keys=2, and set VLLM_USE_DIRECT_DCP_KV_GATHER=0 after the prior multimem/memcpy signature. 撤销 B300 Mooncake 的 compact_group_io,关闭 direct KV gather。c8 已有真实 rail,但 compact_group_io 对约 25MiB put 大量 TRANSFER_FAIL 后挂起。 Co-authored-by: Cursor <cursoragent@cursor.com>
Tip c2 (15820fd / run 36852970338) had a real ibp198s0f0 rail, then register_buffer failed (-600) on the ~39GiB cuMem KV region and stormed AddressNotRegistered TRANSFER_FAIL until sample_tokens hang. 中文:tip c2 已有真实 rail,但 cuMem KV 的 register_buffer 失败(-600) 导致 AddressNotRegistered TRANSFER_FAIL,最终 sample_tokens 挂起;关闭 enable-cumem-allocator。 Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
Description
Restore B300 Mooncake startup with active Mellanox discovery and the host mlx5 provider, and backport request-local recomputation after failed KV loads.
Also raise
VLLM_EXECUTE_MODEL_TIMEOUT_SECONDSto 1800. Concurrency 48 and 56 died withEngineDeadErrorafter 293 s and 299 s of engine silence: a Mooncake load blocks insideexecute_model, and vLLM caps that RPC at 300 s by default.VLLM_RPC_TIMEOUT, already raised here, does not cover that path. Concurrency 70 recovered from a 240 s stall with ten times the failed KV loads, so the stall tail is normal for this workload and the cap left under a minute of headroom.Testing: a new test asserts the timeout across every vLLM recipe that configures a
MooncakeStoreConnector; it fails on the old recipe and passes on the new one. Changelog and matrix plan (11 benchmark + 11 eval jobs on one image) validate. Sweep running.Note: the earlier concurrency 4 and 24 cancellations were queue contention, not a hang — neither ever started the benchmark.
batch_1has 16 usable nodes of 18.中文
修复 B300 的活动 Mellanox 网卡识别与主机 mlx5 驱动加载,并回移 KV 加载失败后的请求级本地重算逻辑。
同时将
VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS调高到 1800。并发 48 与 56 在引擎静默 293 秒和 299 秒后以EngineDeadError退出:引擎等待 Mooncake 加载时阻塞在execute_model内部,而 vLLM 对该 RPC 的默认上限是 300 秒;配方此前已调高的VLLM_RPC_TIMEOUT覆盖不到这条路径。并发 70 在 KV 加载失败次数高出十倍的情况下仍从 240 秒停顿中恢复,说明该停顿尾部对这一 workload 属正常范围,300 秒只留下不到一分钟余量。测试: 新增测试对所有配置了
MooncakeStoreConnector的 vLLM 配方断言该超时值,在旧配方上失败、在新配方上通过。changelog 与矩阵规划(11 个 benchmark + 11 个 eval 作业,同一镜像)校验通过;sweep 运行中。说明: 此前并发 4 与 24 的取消是排队争用而非卡死 —— 两者都没有真正启动 benchmark。
batch_1的 18 个节点中仅 16 个可用。Related Issue
Follow-up to #3040 (failed run).
跟进 #3040 的启动失败。
Type of Change
Checklist
perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/reuse-sweep-runon this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.