Skip to content

[Kimi B300] repair RDMA startup and Mooncake recovery / 修复 RDMA 启动及加载恢复 - #3088

Open
edwingao28 wants to merge 7 commits into
mainfrom
fix/kimik3-b300-mooncake-transport
Open

edwingao28 wants to merge 7 commits into
mainfrom
fix/kimik3-b300-mooncake-transport

Conversation

@edwingao28

@edwingao28 edwingao28 commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Restore B300 Mooncake startup with active Mellanox discovery and the host mlx5 provider, and backport request-local recomputation after failed KV loads.

Also raise VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS to 1800. Concurrency 48 and 56 died with EngineDeadError after 293 s and 299 s of engine silence: a Mooncake load blocks inside execute_model, and vLLM caps that RPC at 300 s by default. VLLM_RPC_TIMEOUT, already raised here, does not cover that path. Concurrency 70 recovered from a 240 s stall with ten times the failed KV loads, so the stall tail is normal for this workload and the cap left under a minute of headroom.

Testing: a new test asserts the timeout across every vLLM recipe that configures a MooncakeStoreConnector; it fails on the old recipe and passes on the new one. Changelog and matrix plan (11 benchmark + 11 eval jobs on one image) validate. Sweep running.

Note: the earlier concurrency 4 and 24 cancellations were queue contention, not a hang — neither ever started the benchmark. batch_1 has 16 usable nodes of 18.

中文

修复 B300 的活动 Mellanox 网卡识别与主机 mlx5 驱动加载,并回移 KV 加载失败后的请求级本地重算逻辑。

同时将 VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS 调高到 1800。并发 48 与 56 在引擎静默 293 秒和 299 秒后以 EngineDeadError 退出:引擎等待 Mooncake 加载时阻塞在 execute_model 内部,而 vLLM 对该 RPC 的默认上限是 300 秒;配方此前已调高的 VLLM_RPC_TIMEOUT 覆盖不到这条路径。并发 70 在 KV 加载失败次数高出十倍的情况下仍从 240 秒停顿中恢复,说明该停顿尾部对这一 workload 属正常范围,300 秒只留下不到一分钟余量。

测试: 新增测试对所有配置了 MooncakeStoreConnector 的 vLLM 配方断言该超时值,在旧配方上失败、在新配方上通过。changelog 与矩阵规划(11 个 benchmark + 11 个 eval 作业,同一镜像)校验通过;sweep 运行中。

说明: 此前并发 4 与 24 的取消是排队争用而非卡死 —— 两者都没有真正启动 benchmark。batch_1 的 18 个节点中仅 16 个可用。

Related Issue

Follow-up to #3040 (failed run).

跟进 #3040 的启动失败。

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /reuse-sweep-run on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have。

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

Comment thread inferencex-e2e/runners/test_mooncake_rdma_device.py Outdated
@edwingao28
edwingao28 force-pushed the fix/kimik3-b300-mooncake-transport branch from 154c576 to 99a4698 Compare September 14, 2026 07:15
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@edwingao28 edwingao28 changed the title [Kimi B300] discover Mooncake RDMA devices by driver / 按驱动识别网卡 [Kimi B300] repair RDMA startup and Mooncake recovery / 修复 RDMA 启动及加载恢复 Sep 14, 2026
@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

@edwingao28
edwingao28 marked this pull request as ready for review September 18, 2026 22:33
@edwingao28
edwingao28 requested a review from a team September 18, 2026 22:33

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

add_device(tmp_path, "rdmap86s0", driver="efa", layer="Unknown")
result = select(tmp_path)
assert result.returncode != 0
assert result.stdout == ""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) The new RDMA-rail and Mooncake-patch tests give no real CI safety net: neither this file nor the new cases appended to runners/test_kimik3_b300.py are referenced by any GitHub workflow. test-changelog-gate.yml explicitly lists runners/test_slurm_utils.py in both its paths: trigger and its pytest invocation, but these two new files are wired into neither it nor any other workflow. A future regression in select_mooncake_rdma_device (benchmarks/benchmark_lib.sh) or in patch_kimik3_mooncake_recovery.py will merge cleanly through CI and only surface at B300 DSXE benchmark runtime. Fix: add both files to test-changelog-gate.yml's paths: list and its pytest command (or a runners/test_*.py glob), matching the existing runners/test_slurm_utils.py wiring.

Extended reasoning...

Repo-wide grep for "test_mooncake_rdma_device" and "test_kimik3_b300" across every .github/workflows/*.yml returns zero hits. test-changelog-gate.yml's paths trigger (lines ~7-30) lists benchmarks/benchmark_lib.sh and runners/test_slurm_utils.py, and its pytest step (around line 123) runs an explicit file list ending in runners/test_slurm_utils.py -- the two new files added by this PR are absent from both. test-process-result.yml and other workflows only run unrelated test files. So these tests execute only when a developer runs pytest manually, exactly as the PR description's '11 local checks' implies. If a later PR breaks select_mooncake_rdma_device's driver/state/link_layer matching or breaks a patch context string in patch_kimik3_mooncake_recovery.py, CI shows green, and the break is only discovered when a Kimi-K3 B300 Mooncake benchmark run fails to find an RDMA rail or the launch script exits on 'unexpected patch context'.

Verification: nit. The factual claim is verified: neither new test file is wired into CI. A repo-wide grep for "test_mooncake_rdma_device" and "test_kimik3_b300" across .github/workflows/*.yml returns zero hits. test-changelog-gate.yml lists runners/test_slurm_utils.py in its paths trigger (line 33) and its pytest step (line 123), but not the two files this PR touches. test-process-result.yml runs a fixed…

edwingao28 added a commit that referenced this pull request Sep 19, 2026
合并 main 并把本 PR 的 changelog 条目重新追加到末尾,历史条目保持不变。

# Conflicts:
#	perf-changelog.yaml
@edwingao28 edwingao28 added priority Preempt other runs on this sweep's runners; restore them at the end (org members only) ci-patchwork-waived full-sweep-enabled and removed full-sweep-fail-fast labels Sep 19, 2026
Comment thread runners/test_mooncake_rdma_device.py Outdated
Comment on lines +23 to +24
["bash", "-ec", 'source "$1"; select_mooncake_rdma_device "$2"; '
'printf "%s %s\n" "$MOONCAKE_RAIL" "$MC_GID_INDEX"',
@functionstackx

Copy link
Copy Markdown
Collaborator

InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest main into this PR: we have migrated single-node AgentX onto native srt-slurm (#3428), so AgentX configs are now declarative YAML recipes, not per-config 1000+ line bash slop scripts. Please also delete the old benchmarks/single_node/** scripts (see this recipe for the new format).

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

edwingao28 and others added 2 commits October 1, 2026 01:21
Discover active Mellanox RDMA rails by mlx5_core driver on DSXE, load the
host mlx5 provider, backport Mooncake load-failure recovery, and raise the
master lease to 60s so transferred keys survive the observed get tail.

修复 Kimi B300 上 Mooncake 的 RDMA 启动与加载恢复:按驱动选择活动网卡、
加载主机 mlx5 provider、回移植加载失败恢复,并将 master lease 提高到 60 秒。

Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
kimik3-b300-mooncake.sh sources benchmark_lib.sh with --validation-only,
but select_mooncake_rdma_device lived below that early return, so the
setup script failed with exit 127 and cancelled the fail-fast canary.

将 select_mooncake_rdma_device 移到 --validation-only 门禁之上,并让测试
以相同方式 source,避免再次出现 command not found。

Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
@cursor
cursor Bot force-pushed the fix/kimik3-b300-mooncake-transport branch from 53f504d to 2238563 Compare October 1, 2026 01:22
edwingao28 and others added 5 commits September 30, 2026 22:34
Tip fail-fast 36800796192 killed agentic c40 after ~279s of engine
silence under the default 300s VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS
cap (shm_broadcast waits, then EngineDead / ProfileAborted). Restore
the supported 1800s knob used by sibling Kimi Mooncake recipes; keep
the 60s Mooncake master lease.

将 B300 Mooncake AgentX 的 execute_model 超时恢复为 1800 秒,保留 60
秒 master 租约。

Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
Append-only perf-changelog entry for restoring
VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 after tip fail-fast
36800796192 c40 EngineDead / ProfileAborted.

为恢复 execute_model 1800 秒超时追加 changelog。

Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
Tip c48 already had device_name=ibp198s0f0 after setup (Wrote line is
pre-patch); crash was GPU memcpy / DCP multimem under load. Match proven
GB300 DCP8 compact_group_io + max_load_batch_keys=2 and MC_MAX_MR_SIZE,
and refuse empty device_name or missing host mlx5 provider.

将 B300 Mooncake DCP8 IO 与已验证的 GB300 设置对齐,并在空 rail /
缺少主机 mlx5 provider 时明确失败。c48 的 Wrote device_name='' 是
patch 前转储;真实失败是高负载下的 GPU memcpy / DCP multimem。

Co-authored-by: Cursor <cursoragent@cursor.com>
c8 on tip ee33eeb had a real rail (Patched device_name=ibp198s0f0) but
compact_group_io stormed TRANSFER_FAIL on ~25MiB puts then hung at 0 tok/s.
Revert compact_group_io, keep max_load_batch_keys=2, and set
VLLM_USE_DIRECT_DCP_KV_GATHER=0 after the prior multimem/memcpy signature.

撤销 B300 Mooncake 的 compact_group_io,关闭 direct KV gather。c8 已有真实
rail,但 compact_group_io 对约 25MiB put 大量 TRANSFER_FAIL 后挂起。

Co-authored-by: Cursor <cursoragent@cursor.com>
Tip c2 (15820fd / run 36852970338) had a real ibp198s0f0 rail, then
register_buffer failed (-600) on the ~39GiB cuMem KV region and stormed
AddressNotRegistered TRANSFER_FAIL until sample_tokens hang.

中文:tip c2 已有真实 rail,但 cuMem KV 的 register_buffer 失败(-600)
导致 AddressNotRegistered TRANSFER_FAIL,最终 sample_tokens 挂起;关闭
enable-cumem-allocator。

Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-patchwork-waived full-sweep-fail-fast priority Preempt other runs on this sweep's runners; restore them at the end (org members only) skip_queue

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants