Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
df92415
fix: repair Kimi B300 Mooncake RDMA startup and recovery
edwingao28 Oct 1, 2026
2238563
fix: expose Mooncake RDMA helper under --validation-only
edwingao28 Oct 1, 2026
f5de7a5
fix: restore B300 Mooncake execute_model timeout to 1800s
edwingao28 Oct 1, 2026
76c772e
docs: changelog for B300 Mooncake execute_model timeout restore
edwingao28 Oct 1, 2026
ee33eeb
fix: align B300 Mooncake DCP8 IO with GB300 and fail empty rail
edwingao28 Oct 1, 2026
15820fd
fix: drop B300 Mooncake compact_group_io; disable direct KV gather
edwingao28 Oct 1, 2026
1d9c54d
fix: drop B300 Mooncake cumem after register_buffer -600
edwingao28 Oct 1, 2026
e55883e
fix: drop B300 Mooncake MC_MAX_MR_SIZE after register -600
edwingao28 Oct 1, 2026
97f83d1
chore: touch perf-changelog to retrigger fail-fast sweep
edwingao28 Oct 1, 2026
860c1cc
merge: sync main into B300 Mooncake transport branch
edwingao28 Oct 1, 2026
e51c58f
fix(kimik3-b300): cut Mooncake max_load_batch_keys to 1
edwingao28 Oct 1, 2026
670fb9f
fix(kimik3-b300): disable Mooncake load_async on DSXE
edwingao28 Oct 1, 2026
ea88d65
merge: sync main into B300 Mooncake transport branch
edwingao28 Oct 1, 2026
d7176e6
fix(kimik3-b300): restore Mooncake load_async=true (connector assert)
edwingao28 Oct 1, 2026
1f837c4
fix(kimik3-b300): disable direct DCP Q gather on DSXE
edwingao28 Oct 1, 2026
59bfac1
fix(kimik3-b300): restore DCP Q gather; disable A2A
edwingao28 Oct 2, 2026
21dac22
fix(changelog): restore dsv41flash-fp4-mi355x entry from main
edwingao28 Oct 2, 2026
b70e426
merge: origin/main; keep restored mi355x changelog entry
edwingao28 Oct 2, 2026
bed9f1c
fix(kimik3-b300): lower high-conc gpu-memory-utilization to 0.85
edwingao28 Oct 2, 2026
68cdcc5
fix(kimik3-b300): re-enable direct DCP KV gather
edwingao28 Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions inferencex-e2e/benchmarks/benchmark_lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -135,6 +135,35 @@ run_amd_multinode_after_preflight() {
--kill-on-bad-exit=1 --signal=TERM@30 --unbuffered "$@"
}

# Setup scripts (for example kimik3-b300-mooncake.sh) source this library with
# --validation-only, so the Mooncake rail helper must stay above that gate.
select_mooncake_rdma_device() {
local sysfs_root="${1:-/sys/class/infiniband}"
local device
MOONCAKE_RAIL=""
for device in "$sysfs_root"/*; do
# DSXE has both EFA and Mellanox adapters. The latter may be renamed
# ibp*, so identify the driver rather than assuming an mlx5_* name.
[[ "$(readlink "$device/device/driver" 2>/dev/null)" == */mlx5_core ]] || continue
grep -qx '4: ACTIVE' "$device/ports/1/state" 2>/dev/null || continue
case "$(cat "$device/ports/1/link_layer" 2>/dev/null)" in
InfiniBand) MC_GID_INDEX=0 ;;
Ethernet) MC_GID_INDEX=3 ;;
*) continue ;;
esac
MOONCAKE_RAIL="${device##*/}"
if [[ -z "$MOONCAKE_RAIL" || "$MOONCAKE_RAIL" == "*" ]]; then
echo "Error: resolved an empty Mooncake RDMA rail name from $device" >&2
return 1
fi
export MOONCAKE_RAIL
export MC_GID_INDEX
return 0
done
echo "Error: no active Mellanox RDMA rail; Mooncake cannot initialise" >&2
return 1
}

# Launchers may load only input validation, without benchmark initialization.
if [[ "${1-}" == "--validation-only" ]]; then
return 0
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,12 @@
#!/usr/bin/env bash
# Pin the worker's Mooncake client and point its store at one active RDMA rail.
set -euo pipefail
# Pin the worker's Mooncake client, backport load-failure recovery, and point
# the store at one active Mellanox RDMA rail (driver-selected, including ibp*).
set -eo pipefail

ws=/infmax-workspace
# Temporary upstream #55297 backport; docs/waiver/3088.md is pending review.
python3 "$ws/runners/patch_kimik3_mooncake_recovery.py"

pip_install=(python3 -m pip install)
if python3 -m pip install --help 2>/dev/null | grep -q -- --break-system-packages; then
pip_install+=(--break-system-packages)
Expand All @@ -10,33 +16,42 @@ fi
python3 -c "from mooncake.store import MooncakeDistributedStore" >/dev/null

# Rail-isolated nodes: two RNICs cannot reach each other even within a node, so
# every rank uses one rail. mlx5_0 is down on some nodes, and topology discovery
# then finds no HCA, so take the first active rail at runtime. DSXE nodes name
# their rails rdmap*.
rail=""
for device in mlx5_0 mlx5_1 mlx5_2 mlx5_3 mlx5_4 mlx5_5 mlx5_8 mlx5_9 \
mlx5_10 mlx5_11 mlx5_16 mlx5_17 mlx5_20 mlx5_21 mlx5_22 mlx5_23 \
$(ls /sys/class/infiniband 2>/dev/null | grep '^rdmap' | sort -V); do
if grep -q ACTIVE "/sys/class/infiniband/$device/ports/1/state" 2>/dev/null; then
rail="$device"
break
fi
done
# every rank uses one Mellanox rail. Identify by driver (mlx5_core) rather than
# assuming mlx5_* names — DSXE may rename them ibp*. EFA rails are skipped.
# shellcheck source=/dev/null
source "$ws/benchmarks/benchmark_lib.sh" --validation-only
select_mooncake_rdma_device
rail="$MOONCAKE_RAIL"
if [[ -z "$rail" ]]; then
echo "Error: no active RDMA rail on $(hostname); Mooncake cannot initialise" >&2
for state in /sys/class/infiniband/*/ports/*/state; do
echo "$state: $(cat "$state" 2>&1)" >&2
done
echo "Error: select_mooncake_rdma_device returned an empty rail" >&2
exit 1
fi
echo "Mooncake rail: $rail (MC_GID_INDEX=$MC_GID_INDEX)"

# The enroot EFA hook binds the host libibverbs over the image's, and
# libibverbs only loads providers built against its own private ABI, so the
# image's mlx5 provider never loads and the Mellanox rail vanishes.
# runners.yaml mounts the host library directory at /host-usr-lib; libibverbs
# appends its own -rdmavNN suffix to an absolute RDMAV_DRIVERS entry.
if [[ ! -d /host-usr-lib/libibverbs ]]; then
echo "Error: /host-usr-lib/libibverbs missing; host mlx5 provider required" >&2
exit 1
fi
export RDMAV_DRIVERS=/host-usr-lib/libibverbs/libmlx5

config="${MOONCAKE_CONFIG_PATH:-/logs/mooncake_store_config.json}"
python3 - "$config" "$rail" <<'PY'
import json, sys
path, rail = sys.argv[1:]
if not rail.strip():
raise SystemExit("Error: refusing to write empty Mooncake device_name")
with open(path) as handle:
config = json.load(handle)
config["device_name"] = rail
with open(path, "w") as handle:
json.dump(config, handle, indent=2)
written = json.load(open(path))
if not str(written.get("device_name") or "").strip():
raise SystemExit(f"Error: mooncake store config still has empty device_name: {written!r}")
print(f"Patched mooncake_store_config device_name={written['device_name']!r}")
PY
echo "Mooncake rail: $rail"
Original file line number Diff line number Diff line change
Expand Up @@ -26,17 +26,20 @@ base:
interval_seconds: 10
max_attempts: 360
# Embedded Mooncake: each TP rank contributes TOTAL_CPU_DRAM_GB / 8 GB. The
# setup script pins the client to the master's version and fills in the
# node's active RDMA rail. DSXE rails are InfiniBand without a netdev, so the
# transfer engine picks its own GID (a RoCE v2 index 3 does not exist).
# setup script pins the client, backports load-failure recovery, selects one
# active Mellanox rail by driver (ibp* or mlx5_*), sets MC_GID_INDEX from the
# link layer, requires the host mlx5 provider via RDMAV_DRIVERS at
# /host-usr-lib, and refuses an empty device_name. Master lease is 60s.
# device_name starts empty; kimik3-b300-mooncake.sh patches it before engines
# start (srt-slurm's "Wrote mooncake_store_config" line is the pre-patch dump).
setup_script: kimik3-b300-mooncake.sh
services:
- name: mooncake-master
type: mooncake-master
preamble: >-
python3 -m pip install --break-system-packages --quiet --no-cache-dir --no-deps
--force-reinstall mooncake-transfer-engine-cuda13==0.3.11.post1
args: ["--eviction_high_watermark_ratio=0.95", "--eviction_ratio=0.10"]
args: ["--eviction_high_watermark_ratio=0.95", "--eviction_ratio=0.10", "--default_kv_lease_ttl=60000"]
options:
store_config:
mode: embedded
Expand Down Expand Up @@ -65,25 +68,49 @@ base:
load-format: fastsafetensors
moe-backend: auto
no-enable-flashinfer-autotune: true
enable-cumem-allocator: true
# Keep cuMem/VMM off on this DSXE Mooncake path (LMCache-style caution for
# VMM buffers). The register_buffer -600 / AddressNotRegistered storm on
# tip 15820fd2e/1d9c54d was driven by MC_MAX_MR_SIZE=4GiB (removed below),
# not by cumem alone — pre-MC_MAX_MR tips registered with cumem on.
enable-prefix-caching: true
prefix-match-unit: 128
kv-cache-dtype: fp8
stream-interval: 10
attention-backend: TOKENSPEED_MLA
attention-config: '{"mla_prefill_backend":"TRTLLM_RAGGED","use_prefill_query_quantization":true}'
disable-uvicorn-access-log: true
kv-transfer-config: '{"kv_connector":"MooncakeStoreConnector","kv_role":"kv_both","kv_load_failure_policy":"recompute","kv_connector_extra_config":{"load_async":true,"lookup_async":true,"enable_offload":false}}'
# Do NOT enable compact_group_io on DSXE single-rail Mooncake: tip ee33eeba
# c8 (run 36838397438) enabled it and immediately stormed TRANSFER_FAIL on
# compact-group-io-v1 ~25MiB puts (7177 fails), then hung at 0 tok/s until
# the 1800s sample_tokens cap. Keep max_load_batch_keys=1 after tips
# 860c1ccf c48 / e51c58f5 c32 DCP PYNCCL _ALLGATHER_BASE hangs under
# ~98-100% GPU KV. Tip ea88d652 / sweep 36935627690 canary c1 crashed
# immediately with AssertionError "load_async must be True for better
# performance" in mooncake store worker get_finished when load_async was
# false — restore load_async=true (required by this vLLM Mooncake path).
kv-transfer-config: '{"kv_connector":"MooncakeStoreConnector","kv_role":"kv_both","kv_load_failure_policy":"recompute","kv_connector_extra_config":{"load_async":true,"lookup_async":true,"max_load_batch_keys":1,"enable_offload":false}}'
env:
VLLM_ALLREDUCE_USE_FLASHINFER: '1'
VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION: '1'
VLLM_USE_V2_MODEL_RUNNER: '1'
# These default to auto; name them so the measured DCP a2a path runs.
VLLM_USE_DIRECT_DCP_A2A: '1'
# Tip 1f837c46 eval-only c8 hung ~8.5m after dcp:0 then failed EP
# PyNccl ncclCommInitRank (remote process exited) with Q_GATHER=0;
# restore Q gather. Keep A2A off (dcp allreduce stays PYNCCL). load_async
# must stay true.
VLLM_USE_DIRECT_DCP_A2A: '0'
VLLM_USE_DIRECT_DCP_Q_GATHER: '1'
# Tip bed9f1ce c40 (Slurm 6765): with KV_GATHER=0 the hang is PyNCCL
# kv_gather _ALLGATHER_BASE (dcp.py:1413; last started work: -1) after
# ~42m serve / GPU KV ~86-100%, Mooncake failed_keys=0. A2A=0 and util
# 0.85 did not clear it. Re-enable direct KV gather (GB300 DCP8 /
# B200 Mooncake stock); prior multimem/memcpy with KV_GATHER=1 was under
# util 0.92 before the 0.85 headroom cut.
VLLM_USE_DIRECT_DCP_KV_GATHER: '1'
VLLM_ENGINE_READY_TIMEOUT_S: '3600'
VLLM_RPC_TIMEOUT: '600000'
# Mooncake loads can block inside execute_model past the 300s default;
# c40 on run 36800796192 went silent ~279s then EngineDead at the cap.
VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS: '1800'
VLLM_PREFIX_CACHE_RETENTION_INTERVAL: '0'
VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS: '0'
PYTHONNOUSERSITE: '1'
Expand All @@ -95,6 +122,9 @@ base:
MC_STORE_MEMCPY: '1'
MC_ENABLE_DEST_DEVICE_AFFINITY: '1'
MC_SLICE_SIZE: '1048576'
# Do NOT set MC_MAX_MR_SIZE on DSXE: tip ee33eeba+ (4GiB) made every rank
# register_buffer fail (-600) on the ~40GiB KV region, then AddressNotRegistered
# TRANSFER_FAIL. Pre-MC_MAX_MR tips (c40/c48) registered cleanly without it.
MC_WORKERS_PER_CTX: '4'
WITH_NVIDIA_PEERMEM: '0'
VLLM_MOONCAKE_LOAD_RECV_THREADS: '4'
Expand All @@ -108,8 +138,12 @@ base:
KV_OFFLOADING: dram
TOTAL_CPU_DRAM_GB: '2249'

# One variant per point. Admission is 2x CONC; CONC 56 and 70 keep more memory
# headroom. Graphs capture (1 + drafts) x 1..min(2x CONC, 128) tokens, then the
# One variant per point. Admission is 2x CONC. CONC 24+ use gpu-memory-utilization
# 0.85: with VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0, tip b70e4260a c48
# (Slurm 6737) over-allocated GPU KV to 45.2 GiB (vLLM suggested 36.78 GiB once
# CUDA graphs are counted) and OOMed in flashinfer FP4 MoE prepare_moe (~2.89 GiB
# needed, ~2.3 GiB free) ~4m after Application startup. Keep 0.92 on speculative
# c1–c16. Graphs capture (1 + drafts) x 1..min(2x CONC, 128) tokens, then the
# larger powers of two to 8192. Throughput runs switch DSpark to synthetic
# rejection at the golden acceptance length; points above CONC 16 do not draft
# and keep the matrix's mtp label.
Expand Down Expand Up @@ -178,7 +212,7 @@ override_c24:
agg:
args:
max-num-seqs: 48
gpu-memory-utilization: 0.92
gpu-memory-utilization: 0.85
compilation-config: '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,64,128,256,512,1024,2048,4096,8192]}'
benchmark:
env:
Expand All @@ -190,7 +224,7 @@ override_c32:
agg:
args:
max-num-seqs: 64
gpu-memory-utilization: 0.92
gpu-memory-utilization: 0.85
compilation-config: '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,128,256,512,1024,2048,4096,8192]}'
benchmark:
env:
Expand All @@ -202,7 +236,7 @@ override_c40:
agg:
args:
max-num-seqs: 80
gpu-memory-utilization: 0.92
gpu-memory-utilization: 0.85
compilation-config: '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,128,256,512,1024,2048,4096,8192]}'
benchmark:
env:
Expand All @@ -214,7 +248,7 @@ override_c48:
agg:
args:
max-num-seqs: 96
gpu-memory-utilization: 0.92
gpu-memory-utilization: 0.85
compilation-config: '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,128,256,512,1024,2048,4096,8192]}'
benchmark:
env:
Expand All @@ -226,7 +260,7 @@ override_c56:
agg:
args:
max-num-seqs: 112
gpu-memory-utilization: 0.9
gpu-memory-utilization: 0.85
compilation-config: '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,128,256,512,1024,2048,4096,8192]}'
benchmark:
env:
Expand All @@ -238,7 +272,7 @@ override_c70:
agg:
args:
max-num-seqs: 140
gpu-memory-utilization: 0.9
gpu-memory-utilization: 0.85
compilation-config: '{"cudagraph_mode":"FULL_AND_PIECEWISE","cudagraph_capture_sizes":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,256,512,1024,2048,4096,8192]}'
benchmark:
env:
Expand Down
5 changes: 5 additions & 0 deletions inferencex-e2e/configs/runners.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -594,6 +594,11 @@ clusters:
single-node-models: staged
container-aliases: [dynamo-trtllm, dynamo-sglang, dynamo-vllm]
nginx-aliases: [nginx-sqsh]
# Host libibverbs ABI for Mooncake RDMA: enroot EFA hook overlays the
# image provider, so mount the host library tree and load mlx5 via
# RDMAV_DRIVERS in kimik3-b300-mooncake.sh.
mounts:
/usr/lib/x86_64-linux-gnu: /host-usr-lib
gb200-nv:
gpus-per-node: 4
available-cpu-dram-mib: 860_160
Expand Down
41 changes: 35 additions & 6 deletions inferencex-e2e/docs/configuration-procedures.md
Original file line number Diff line number Diff line change
Expand Up @@ -267,7 +267,7 @@ All ten AgentX throughput points use DSpark K6 (target verification length 7)
and the committed golden AL 3.77. C1/2/4/8/16 use TP8/EP1;
C48/64/96/128/256 use TP8/DPA8/EP8 with native RCCL. Each point runs for
3600 seconds. The C256 full GSM8K eval omits forced acceptance. Keep the
pinned `rocm/atom-dev:nightly_202609291501` image and GPU-only KV. C1 through C16 use
pinned `rocm/atom-dev:nightly_202609161445` image and GPU-only KV. C1 through C16 use
BF16 KV, while C48 and above retain FP8 KV; all points use the FP4 index cache,
8192-token checkpoints and DEP dense FULL graph ladder. Fixed q7 graphs are
captured in each new server; confirm target and DSpark draft capture in
Expand All @@ -280,12 +280,8 @@ record model/source identity and requested settings. Successful startup,
graph capture and requests require runtime log evidence.

The pinned image is the official ATOM nightly
`rocm/atom-dev:nightly_202609291501` (ATOM `0.1.7.dev46+g74fd942b0`, ROCm 7.2.4),
which includes the merged
`rocm/atom-dev:nightly_202609161445`, which includes the merged
[ROCm/ATOM#2233](https://github.com/ROCm/ATOM/pull/2233) inference-mode fix.
From this image ATOM stores the checkpoint's `ue8m0` FP8 block scales as E8M0
on gfx950 by default ([ROCm/ATOM#2419](https://github.com/ROCm/ATOM/pull/2419));
the powers-of-two scales are represented exactly.
The recipe does not patch AITER source at runtime; TP communication
fusion, DSpark K6 and graph capture use the implementation shipped in the image.

Expand Down Expand Up @@ -333,6 +329,39 @@ The H200 DSpark recipe uses the same minimum capture size and preserves the same

B300 uses the same minimum capture size at c1/c2/c4. Its c1 CI comparison reduced request ITL P90/P99 from 38.74/41.42 ms to 2.62/3.45 ms; c2/c4 require CI confirmation.

Kimi-K3 on B300 selects one active Mellanox adapter by its sysfs driver, including
DSXE `ibp*` names; EFA devices are excluded from this RDMA recipe. The embedded
Mooncake ranks share that adapter. InfiniBand uses GID index 0 and RoCE retains
index 3. If no compatible active adapter exists, or if the host mlx5 provider mount
is missing, startup fails before serving. The recipe YAML may leave `device_name`
empty; `kimik3-b300-mooncake.sh` patches a real rail into the store config and
refuses to continue with an empty name (srt-slurm's earlier "Wrote
mooncake_store_config" line is the pre-patch dump). On DSXE the container's
libibverbs comes from the host through the enroot EFA hook, so
`configs/runners.yaml` mounts the host library directory at `/host-usr-lib` and
the setup script requires `RDMAV_DRIVERS` to load its mlx5 provider. Keep
`max_load_batch_keys: 1` and `load_async: true` (tips 860c1ccf c48 /
e51c58f5 c32 hung in DCP PYNCCL `_ALLGATHER_BASE` under ~98–100% GPU KV with
async loads and clean Mooncake metrics; `last started work: -1`. Tip ea88d652
canary c1 crashed with Mooncake `AssertionError: load_async must be True for
better performance` when `load_async` was set false, so restore the required
stock true), but do not enable `compact_group_io` on this DSXE single-rail path
(it storm-failed ~25 MiB compact-group puts at c8). Direct DCP A2A is off (`VLLM_USE_DIRECT_DCP_A2A=0`); keep
`VLLM_USE_DIRECT_DCP_Q_GATHER=1` and `VLLM_USE_DIRECT_DCP_KV_GATHER=1` (tip
1f837c46 eval-only c8 hung ~8.5m after `dcp:0` then failed EP
`ncclCommInitRank` with Q gather off; tip bed9f1ce c40 hung in PyNCCL
`kv_gather` `_ALLGATHER_BASE` with `last started work: -1` when KV gather was
off — A2A=0 and util 0.85 did not clear it). Do not set `MC_MAX_MR_SIZE` here:
with 4GiB every rank hit `register_buffer failed ... -600` on the ~40 GiB KV
region and stormed `AddressNotRegistered` TRANSFER_FAIL (c2/c32); pre-`MC_MAX_MR`
tips registered cleanly. Keep `enable-cumem-allocator` off on this path. Keep
`gpu-memory-utilization` at 0.85 for CONC 24+ (tip b70e4260a c48 reached
Application startup with KV 45.2 GiB at 0.92 under
`VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0`, then OOMed in flashinfer FP4 MoE
`prepare_moe` allocating ~2.89 GiB with ~2.3 GiB free; vLLM suggested ~36.78 GiB
KV once CUDA graphs are counted).


The AgentX-only `dsv41flash-fp4-<sku>-vllm-agentic-dspark` recipes use the per-SKU
`image` pinned in [`nvidia-master.yaml`](../configs/nvidia-master.yaml) (originally `vllm/vllm-openai:deepseekv41-flash-0909`, which B300 still uses) at TP4 on Blackwell SKUs with native five-token DSpark,
probabilistic drafting. Throughput uses the [committed golden AL](../infx/golden_al_distribution/dsv41flash_dspark.yaml) of 3.51 for thinking on and five draft tokens, with synthetic rejection sampling and adaptive verification disabled. Accuracy evals retain real block rejection and adaptive verification. `--engram-config '{"cpu_offload":true}'`
Expand Down
Loading
Loading