diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4/agentx/agg-variants.yaml b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4/agentx/agg-variants.yaml new file mode 100644 index 0000000000..e855c7db4c --- /dev/null +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4/agentx/agg-variants.yaml @@ -0,0 +1,168 @@ +# AgentX dsv41flash sglang gb300-fp4 recipes (Dynamo frontend + SGLang, AGGREGATED +# topology): shared settings in base, one override per benchmark point. Select one +# with CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/agg-variants.yaml:override_. +# +# One SGLang worker serves prefill and decode with the checkpoint's bundled DSpark +# draft (block size 5); the Dynamo frontend routes with the KV-aware router and +# session affinity. The aggregated arms are the two lowest-latency points of the +# curve; everything from c16 to c256 comes from the disaggregated recipes in +# disagg-variants.yaml. Measured on GB300 NVL72 (tokens/s/GPU @ P90 interactivity, +# P90 TTFT; both GSM8K gates passed): +# override_tp4_c1 : 6,215 @ 385.0 tok/s/user, 1.2 s (4 GPUs; GSM8K 0.9735) +# override_tp2_c1 : 12,501 @ 377.1 tok/s/user, 0.9 s (2 GPUs; GSM8K 0.9735) +# The aggregated c64 cells (prefill-decode-interval 8: 159,510 @ 62.7; interval 16: +# 153,542 @ 83.3) were measured as well but are dominated by the disaggregated +# HiCache cells and are not shipped. The two c1 arms are the highest-interactivity +# points of the curve (above the +# published MI355X ATOM curve's last vertex at 336.7 tok/s/user); at c1 the AgentX +# client replays the trace's idle gaps, so a single lane is client-bound and a second +# user on the same worker only lowers the P90 (TP2 c2: 12,538 @ 306.8). The harness +# applies the golden acceptance length (3.51 at K=5) itself. + +schema: 2 + +base: + name: dsv41flash-fp4-gb300-dynamo-sglang-agentx-agg + model: + path: hf:deepseek-ai/DeepSeek-V4.1-Flash + container: lmsysorg/sglang:nightly-dev-20260928-81f27fb3@sha256:d9e4917808cfaa4b3be033a0c85a3a73d71eb93c17accbfa2ac3e96f551f3a33 + precision: fp4 + identity: + model: + repo: deepseek-ai/DeepSeek-V4.1-Flash + dynamo: + install: true + source: + wheel: "1.6.0.dev20260928" + slurm: + time_limit: "4:00:00" + # Cold weight loads from the shared HF cache plus graph capture take 15-20 min. + health_check: + max_attempts: 1440 + interval_seconds: 10 + resources: + gpu_type: gb300 + gpus_per_node: 4 + # Dynamo 1.6 uses the TCP request plane; etcd is the only discovery service needed. + services: + - name: etcd + type: etcd + placement: + node: infra + frontend: + type: dynamo + enable_multiple_frontends: false + env: + PIP_BREAK_SYSTEM_PACKAGES: "1" + DYN_NATS_REQUEST_TIMEOUT_SECS: "1800" + args: + router-mode: kv + router-session-affinity-ttl-secs: "3600" + active-decode-blocks-threshold: "None" + active-prefill-tokens-threshold: "None" + active-prefill-tokens-threshold-frac: "None" + engine: sglang + roles: + agg: + nodes: 1 + workers: 1 + gpus: 2 + env: + PIP_BREAK_SYSTEM_PACKAGES: "1" + PYTHONNOUSERSITE: "1" + PYTHONUNBUFFERED: "1" + HF_HUB_CACHE: /hf_hub_cache + # Outlast AIPerf's pooled connections past Uvicorn's keep-alive. + SGLANG_TIMEOUT_KEEP_ALIVE: "900" + # AgentX measures thinking on, the regime of the golden acceptance curve. + SGLANG_DEFAULT_THINKING: "1" + SGLANG_DSV41_REASONING_EFFORT: high + SGLANG_DSPARK_OPT_MARKOV_W2_BF16: "True" + # TP2 keeps the row-sharded Engram tables in host DRAM. + SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE: "1" + SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT: per_rank + args: + served-model-name: deepseek-ai/DeepSeek-V4.1-Flash + trust-remote-code: true + tensor-parallel-size: 2 + expert-parallel-size: 2 + mem-fraction-static: 0.8 + # Keep active decode requests progressing while long prefixes are queued. + prefill-decode-interval: 16 + # Admission never exceeds the captured decode graph batch. + cuda-graph-max-bs-decode: 64 + # DSpark is the checkpoint's bundled draft; the block size is its only knob. + speculative-algorithm: DSPARK + speculative-dspark-block-size: 5 + reasoning-parser: auto + tool-call-parser: auto + # Draft passes under long-context load outlast the 1800 s default watchdog. + watchdog-timeout: 3600 + enable-metrics: true + weight-loader-prefetch-checkpoints: true + weight-loader-drop-cache-after-load: false + model-loader-extra-config: '{"enable_multithread_load":true}' + sbatch_directives: + mem: "0" + exclusive: "" + srun_options: + container-remap-root: "" + benchmark: + type: custom + command: bash /infmax-workspace/benchmarks/srt_agentic.sh + env: + INFMAX_CONTAINER_WORKSPACE: /infmax-workspace + RESULT_DIR: /logs/agentic + PORT: "8000" + IS_MULTINODE: "false" + MODEL: deepseek-ai/DeepSeek-V4.1-Flash + AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID: "true" + AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING: "0" + AIPERF_REQUIRED_SERVER_METRIC_PREFIX: "sglang:" + AIPERF_DATASET_MMAP_CACHE_DIR: /aiperf_mmap_cache + HF_HUB_CACHE: /hf_hub_cache + # Warmup can leave bytes unacknowledged past AIPerf's 30 s default. + AIPERF_HTTP_TCP_USER_TIMEOUT: "900000" + TP: "2" + PP_SIZE: "1" + PCP_SIZE: "1" + +# Lowest-latency arm: one pure TP4 (EP1) worker on all four GPUs at c1, the standalone +# recipe's override_tp4_c1 behind the Dynamo frontend: static ragged verify, Engram +# tables in HBM (host-table path off), admission 2x CONC, 4096-token prefill chunks. +override_tp4_c1: + name: agg-gb300-tp4ep1-c1 + roles: + agg: + gpus: 4 + args: + tensor-parallel-size: 4 + expert-parallel-size: 1 + chunked-prefill-size: 4096 + max-running-requests: 2 + env: + SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE: "0" + SGLANG_RAGGED_VERIFY_MODE: static + benchmark: + env: + CONC: "1" + TP: "4" + AGENTIC_WARMUP_GRACE_PERIOD: "1800" + +# Two-GPU c1 arm: the base TP2/EP2 worker with the Engram tables in host DRAM, static +# ragged verify, 4096-token prefill chunks and admission 2x CONC: 12,501 tokens/s/GPU +# @ 377.1 tok/s/user, P90 TTFT 0.90 s - twice the tokens/s/GPU of the TP4 c1 arm at a +# 2% lower P90 interactivity. +override_tp2_c1: + name: agg-gb300-tp2ep2-c1 + roles: + agg: + args: + chunked-prefill-size: 4096 + max-running-requests: 2 + env: + SGLANG_RAGGED_VERIFY_MODE: static + benchmark: + env: + CONC: "1" + AGENTIC_WARMUP_GRACE_PERIOD: "1800" diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml new file mode 100644 index 0000000000..dbd3e5d12a --- /dev/null +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml @@ -0,0 +1,299 @@ +# AgentX dsv41flash sglang gb300-fp4 recipes (Dynamo frontend + SGLang, DISAGGREGATED +# topology with Mooncake KV transfer): shared settings in base, one override per +# benchmark point. Select one with +# CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml:override_. +# +# Prefill and decode are separate TP2/EP2 SGLang workers (two GPUs each) behind the +# Dynamo KV router; with the DSpark draft in PD-disaggregated mode both sides must +# use the same TP, so capacity is added with more TP2 workers: one prefill worker +# feeds one or two decode workers (1P1D, 1P2D), which sets the sessions per decode +# worker and with it the interactivity band, and two prefill workers feed one decode +# worker (2P1D) at the throughput end. Measured on GB300 NVL72 (tokens/s/GPU @ P90 +# interactivity, P90 TTFT; every override's GSM8K gate passed): +# override_1p2d_spread_c16 : 19,280 @ 327.4 tok/s/user, 0.7 s (6 GPUs, 3 nodes) +# override_1p2d_c48 : 59,740 @ 244.3 tok/s/user, 1.4 s (6 GPUs, 2 nodes) +# override_1p2d_c64 : 80,583 @ 207.6 tok/s/user, 2.0 s (6 GPUs, 2 nodes) +# override_1p1d_c96 : 144,448 @ 132.2 tok/s/user, 4.1 s (4 GPUs, 2 nodes) +# override_1p1d_hicache_c160 : 172,192 @ 98.8 tok/s/user, 21.7 s (4 GPUs, 2 nodes) +# override_2p1d_hicache_c256 : 174,562 @ 68.1 tok/s/user, 18.9 s (6 GPUs, 3 nodes) +# Against the published MI355X ATOM curve at matched P90 interactivity: 1P2D spread +# c16 1.73x, c48 2.11x, c64 2.21x; 1P1D c96 2.72x, HiCache c160 2.40x; 2P1D HiCache +# c256 1.94x tokens/s/GPU. Other measured cells (1P1D c8/c16/c64, 1P2D c24/c32, 1P4D +# c64) lie on or under the line through these vertices and are not shipped. The +# multi-decode cells run the Dynamo frontend with a 1 s session-affinity TTL (per-turn +# load-aware decode pick); the single-decode cells keep the 3600 s default they were +# measured with. The harness applies the golden acceptance length (3.51 at K=5) itself. + +schema: 2 + +base: + name: dsv41flash-fp4-gb300-dynamo-sglang-agentx-disagg + model: + path: hf:deepseek-ai/DeepSeek-V4.1-Flash + container: lmsysorg/sglang:nightly-dev-20260928-81f27fb3@sha256:d9e4917808cfaa4b3be033a0c85a3a73d71eb93c17accbfa2ac3e96f551f3a33 + precision: fp4 + identity: + model: + repo: deepseek-ai/DeepSeek-V4.1-Flash + dynamo: + install: true + source: + wheel: "1.6.0.dev20260928" + slurm: + # c96 AgentX runs take ~2.5 h including the 60 min profiling phase and the gate. + time_limit: "6:00:00" + health_check: + max_attempts: 1440 + interval_seconds: 10 + resources: + gpu_type: gb300 + gpus_per_node: 4 + # Dynamo 1.6 uses the TCP request plane; etcd is the only discovery service needed + # and runs with the frontend on the prefill node (two nodes total for 1P1D). + services: + - name: etcd + type: etcd + placement: + node: infra + frontend: + type: dynamo + enable_multiple_frontends: false + env: + PIP_BREAK_SYSTEM_PACKAGES: "1" + DYN_NATS_REQUEST_TIMEOUT_SECS: "1800" + args: + router-mode: kv + # Session-affinity TTL as measured for the single-decode 1P1D cells; the + # multi-decode overrides below set it to 1 s (per-turn load-aware decode pick). + router-session-affinity-ttl-secs: "3600" + active-decode-blocks-threshold: "None" + active-prefill-tokens-threshold: "None" + active-prefill-tokens-threshold-frac: "None" + engine: sglang + roles: + prefill: + nodes: 1 + workers: 1 + gpus: 2 + env: + PIP_BREAK_SYSTEM_PACKAGES: "1" + PYTHONNOUSERSITE: "1" + PYTHONUNBUFFERED: "1" + HF_HUB_CACHE: /hf_hub_cache + SGLANG_TIMEOUT_KEEP_ALIVE: "900" + SGLANG_DEFAULT_THINKING: "1" + SGLANG_DSV41_REASONING_EFFORT: high + SGLANG_DSPARK_OPT_MARKOV_W2_BF16: "True" + SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE: "1" + SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT: per_rank + SGLANG_DISAGGREGATION_WAITING_TIMEOUT: "900" + SGLANG_ENABLE_PREFILL_WAR_READ_DONE: "1" + SGLANG_MOONCAKE_CUSTOM_MEM_POOL: "True" + NCCL_MNNVL_ENABLE: "1" + NCCL_CUMEM_ENABLE: "1" + MC_FORCE_MNNVL: "1" + args: + served-model-name: deepseek-ai/DeepSeek-V4.1-Flash + trust-remote-code: true + tensor-parallel-size: 2 + expert-parallel-size: 2 + disaggregation-transfer-backend: mooncake + # Long-context first turns (P90 ISL ~250k tokens): one 16k-token chunk per step. + chunked-prefill-size: 16384 + max-prefill-tokens: 16384 + max-running-requests: 64 + swa-prefix-tails: 4096 + mem-fraction-static: 0.85 + speculative-algorithm: DSPARK + speculative-dspark-block-size: 5 + reasoning-parser: auto + tool-call-parser: auto + watchdog-timeout: 3600 + enable-metrics: true + weight-loader-prefetch-checkpoints: true + weight-loader-drop-cache-after-load: false + model-loader-extra-config: '{"enable_multithread_load":true}' + # KV events for the Dynamo KV router; srtctl allocates the ZMQ ports per worker. + kv_events: true + decode: + nodes: 1 + workers: 1 + gpus: 2 + env: + PIP_BREAK_SYSTEM_PACKAGES: "1" + PYTHONNOUSERSITE: "1" + PYTHONUNBUFFERED: "1" + HF_HUB_CACHE: /hf_hub_cache + SGLANG_TIMEOUT_KEEP_ALIVE: "900" + SGLANG_DEFAULT_THINKING: "1" + SGLANG_DSV41_REASONING_EFFORT: high + SGLANG_DSPARK_OPT_MARKOV_W2_BF16: "True" + SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE: "1" + SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT: per_rank + SGLANG_DISAGGREGATION_WAITING_TIMEOUT: "900" + SGLANG_MOONCAKE_CUSTOM_MEM_POOL: "True" + NCCL_MNNVL_ENABLE: "1" + NCCL_CUMEM_ENABLE: "1" + MC_FORCE_MNNVL: "1" + args: + served-model-name: deepseek-ai/DeepSeek-V4.1-Flash + trust-remote-code: true + tensor-parallel-size: 2 + expert-parallel-size: 2 + disaggregation-transfer-backend: mooncake + max-running-requests: 128 + cuda-graph-max-bs-decode: 128 + mem-fraction-static: 0.85 + speculative-algorithm: DSPARK + speculative-dspark-block-size: 5 + reasoning-parser: auto + tool-call-parser: auto + watchdog-timeout: 3600 + enable-metrics: true + weight-loader-prefetch-checkpoints: true + weight-loader-drop-cache-after-load: false + model-loader-extra-config: '{"enable_multithread_load":true}' + kv_events: true + sbatch_directives: + mem: "0" + exclusive: "" + srun_options: + container-remap-root: "" + benchmark: + type: custom + command: bash /infmax-workspace/benchmarks/srt_agentic.sh + env: + INFMAX_CONTAINER_WORKSPACE: /infmax-workspace + RESULT_DIR: /logs/agentic + PORT: "8000" + IS_MULTINODE: "true" + MODEL: deepseek-ai/DeepSeek-V4.1-Flash + AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID: "true" + AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING: "0" + AIPERF_REQUIRED_SERVER_METRIC_PREFIX: "sglang:" + AIPERF_DATASET_MMAP_CACHE_DIR: /aiperf_mmap_cache + HF_HUB_CACHE: /hf_hub_cache + AIPERF_HTTP_TCP_USER_TIMEOUT: "900000" + AGENTIC_WARMUP_GRACE_PERIOD: "3600" + +# 1 prefill TP2/EP2 + 1 decode TP2/EP2 (4 GPUs, 2 nodes). The decode worker alone +# sets the interactivity (96 sessions give 132 tok/s/user); the single prefill worker +# holds P90 TTFT under 4.1 s up to c96 and saturates between c96 and c128, where the +# HiCache override below takes over. 144,448 tokens/s/GPU @ 132.2 tok/s/user with a +# 4.1 s P90 TTFT. +override_1p1d_c96: + name: disagg-gb300-1p1d-tp2ep2-c96 + benchmark: + env: + CONC: "96" + +# 1 prefill TP2/EP2 + 2 decode TP2/EP2 (6 GPUs). Two decode workers halve the sessions +# per decode worker at a given concurrency, which sets the interactivity band. The +# frontend's session-affinity TTL is 1 s, so every turn takes a load-aware decode pick +# instead of pinning the session to one decode worker for an hour: +4% P90 +# interactivity at 8 sessions per decode worker, neutral at 12 and more (c24 31,465 @ +# 299.0 vs 31,339 @ 298.1; c32 41,910 @ 276.1 vs 41,768 @ 277.3 at TTL 3600). +# +# Same-node layout (both decode workers on one node, 2 nodes total): +# c48: 59,740 tokens/s/GPU @ 244.3 tok/s/user, P90 TTFT 1.36 s +# c64: 80,583 tokens/s/GPU @ 207.6 tok/s/user, P90 TTFT 2.01 s (GSM8K 0.9742) +override_1p2d_c48: + name: disagg-gb300-1p2d-tp2ep2-c48 + frontend: + args: + router-session-affinity-ttl-secs: "1" + roles: + decode: + workers: 2 + args: + max-running-requests: 64 + cuda-graph-max-bs-decode: 64 + benchmark: + env: + CONC: "48" + +override_1p2d_c64: + name: disagg-gb300-1p2d-tp2ep2-c64 + frontend: + args: + router-session-affinity-ttl-secs: "1" + roles: + decode: + workers: 2 + args: + max-running-requests: 64 + cuda-graph-max-bs-decode: 64 + benchmark: + env: + CONC: "64" + +# Spread layout (one decode worker per node, 3 nodes total), 8 sessions per decode +# worker: 19,280 tokens/s/GPU @ 327.4 tok/s/user, P90 TTFT 0.67 s (GSM8K 0.9735) - +# the highest-interactivity disaggregated point (the same cell on one node at TTL +# 3600 gave 19,227 @ 308.9). +override_1p2d_spread_c16: + name: disagg-gb300-1p2d-spread-tp2ep2-c16 + frontend: + args: + router-session-affinity-ttl-secs: "1" + roles: + decode: + nodes: 2 + workers: 2 + args: + max-running-requests: 64 + cuda-graph-max-bs-decode: 64 + benchmark: + env: + CONC: "16" + +# Throughput end. The prefill worker keeps its KV prefix cache in a host-memory tier +# (SGLang HiCache: ratio 1 = one host copy of the device pool, write-back, direct I/O), +# which removes the LRU evictions that pushed P90 TTFT past the 25 s ceiling beyond c96 +# (1P1D c128 without it: 25.4 s). 1P1D c160: 172,192 tokens/s/GPU @ 98.8 tok/s/user, +# P90 TTFT 21.7 s (GSM8K 0.9727). c192 breaches the ceiling (29.2 s): the single +# prefill worker is saturated, so the next step is a second prefill worker (below). +override_1p1d_hicache_c160: + name: disagg-gb300-1p1d-hicache-tp2ep2-c160 + roles: + prefill: + args: + enable-hierarchical-cache: true + hicache-ratio: 1 + hicache-write-policy: write_back + hicache-io-backend: direct + decode: + args: + max-running-requests: 192 + cuda-graph-max-bs-decode: 192 + benchmark: + env: + CONC: "160" + +# Two HiCache prefill workers (one per node) feeding one decode worker (6 GPUs, 3 nodes), +# 128 sessions per prefill worker, decode admission and graph batch 256, session-affinity +# TTL 1 s: 174,562 tokens/s/GPU @ 68.1 tok/s/user, P90 TTFT 18.9 s (GSM8K 0.9712) - the +# highest tokens/s/GPU of the campaign (the same cell without HiCache: 151,671 @ 70.8 +# with a 36 s P90 TTFT). +override_2p1d_hicache_c256: + name: disagg-gb300-2p1d-hicache-tp2ep2-c256 + frontend: + args: + router-session-affinity-ttl-secs: "1" + roles: + prefill: + nodes: 2 + workers: 2 + args: + max-running-requests: 128 + enable-hierarchical-cache: true + hicache-ratio: 1 + hicache-write-policy: write_back + hicache-io-backend: direct + decode: + args: + max-running-requests: 256 + cuda-graph-max-bs-decode: 256 + benchmark: + env: + CONC: "256" diff --git a/inferencex-e2e/configs/nvidia-master.yaml b/inferencex-e2e/configs/nvidia-master.yaml index 63dcf8fff8..8cfd67b436 100644 --- a/inferencex-e2e/configs/nvidia-master.yaml +++ b/inferencex-e2e/configs/nvidia-master.yaml @@ -8927,3 +8927,136 @@ dsv41flash-fp4-gb300-sglang-agentic-dspark: - { tp: 2, ep: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4-mtp/agentic.yaml } - { tp: 4, ep: 1, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4-mtp/agentic.yaml } - { tp: 4, ep: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4-mtp/agentic.yaml } +dsv41flash-fp4-gb300-dynamo-sglang-agentic-agg: + image: lmsysorg/sglang:nightly-dev-20260928-81f27fb3@sha256:d9e4917808cfaa4b3be033a0c85a3a73d71eb93c17accbfa2ac3e96f551f3a33 + model: deepseek-ai/DeepSeek-V4.1-Flash + model-prefix: dsv41flash + runner: cluster:gb300-nv + precision: fp4 + framework: dynamo-sglang + router: { name: dynamo-router, version: "1.6.0.dev20260928" } + multinode: true + disagg: false + scenarios: + agentic-coding: + - dram-utilization: 0.80 + search-space: + - spec-decoding: mtp + conc-list: [1] + num-nodes: 1 + worker: + num-worker: 1 + tp: 4 + ep: 1 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/agg-variants.yaml:override_tp4_c1" + - spec-decoding: mtp + conc-list: [1] + num-nodes: 1 + worker: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/agg-variants.yaml:override_tp2_c1" +dsv41flash-fp4-gb300-dynamo-sglang-agentic-disagg: + image: lmsysorg/sglang:nightly-dev-20260928-81f27fb3@sha256:d9e4917808cfaa4b3be033a0c85a3a73d71eb93c17accbfa2ac3e96f551f3a33 + model: deepseek-ai/DeepSeek-V4.1-Flash + model-prefix: dsv41flash + runner: cluster:gb300-nv + precision: fp4 + framework: dynamo-sglang + router: { name: dynamo-router, version: "1.6.0.dev20260928" } + kv-p2p-transfer: mooncake + multinode: true + disagg: true + scenarios: + agentic-coding: + - dram-utilization: 0.80 + search-space: + - spec-decoding: mtp + conc-list: [16] + prefill: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml:override_1p2d_spread_c16" + decode: + num-worker: 2 + tp: 2 + ep: 2 + dp-attn: false + - spec-decoding: mtp + conc-list: [48] + prefill: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml:override_1p2d_c48" + decode: + num-worker: 2 + tp: 2 + ep: 2 + dp-attn: false + - spec-decoding: mtp + conc-list: [64] + prefill: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml:override_1p2d_c64" + decode: + num-worker: 2 + tp: 2 + ep: 2 + dp-attn: false + - spec-decoding: mtp + conc-list: [96] + prefill: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml:override_1p1d_c96" + decode: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + - spec-decoding: mtp + conc-list: [160] + prefill: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml:override_1p1d_hicache_c160" + decode: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false + - spec-decoding: mtp + conc-list: [256] + prefill: + num-worker: 2 + tp: 2 + ep: 2 + dp-attn: false + additional-settings: + - "CONFIG_FILE=recipes/dsv41flash/sglang/gb300-fp4/agentx/disagg-variants.yaml:override_2p1d_hicache_c256" + decode: + num-worker: 1 + tp: 2 + ep: 2 + dp-attn: false diff --git a/inferencex-e2e/perf-changelog.yaml b/inferencex-e2e/perf-changelog.yaml index 9b9c48a0ef..df423ce255 100644 --- a/inferencex-e2e/perf-changelog.yaml +++ b/inferencex-e2e/perf-changelog.yaml @@ -9168,3 +9168,19 @@ - "Relevant ATOM changes in the range: V4 decode reuses the sparse prefill ASM (ROCm/ATOM#2271), greedy sampler picks via aiter.topk_select (ROCm/ATOM#2244), and FP8 block scales declared scale_fmt ue8m0 are now stored as E8M0 on gfx950 by default (ROCm/ATOM#2419; previously FP32 unless ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE=1). The upstream DeepSeek-V4 recipes are unchanged across the range, and every recipe flag and choice (all2all-backend rccl, dp-load-balance least_tokens, moe-backend standard) remains valid." - "No data-type or precision change to the DeepSeek-V4-Pro-0813 DSpark draft: no online quantization is configured, so it keeps its checkpoint precision; the E8M0 scale storage represents the checkpoint's power-of-two block scales exactly. kv-cache-dtype and index-cache-dtype touch cache storage only." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3605 + +- config-keys: + - dsv41flash-fp4-gb300-dynamo-sglang-agentic-agg + - dsv41flash-fp4-gb300-dynamo-sglang-agentic-disagg + scenario-type: + - agentic-coding + description: + - "Add DeepSeek-V4.1-Flash FP4 AgentX on GB300 with the Dynamo frontend (KV router, session affinity) and SGLang nightly-dev-20260928-81f27fb3 / ai-dynamo 1.6.0.dev20260928, DSpark block size 5." + - "Use the mounted shared Hugging Face cache for every SGLang server role." + - "Aggregated arms: the two lowest-latency points of the curve, a pure TP4 (EP1) c1 point (the standalone recipe's TP4 c1 settings, static ragged verify and Engram in HBM, behind Dynamo; 6,215 tok/s/GPU @ 385.0 tok/s/user P90) and a TP2/EP2 c1 point with static ragged verify and Engram in host DRAM (12,501 @ 377.1, P90 TTFT 0.9 s); both passed their GSM8K gates." + - "Disaggregated curve with Mooncake KV transfer and TP2/EP2 workers: 1P2D with a 1 s Dynamo session-affinity TTL at c16 (one decode worker per node; 19,280 tok/s/GPU @ 327.4 tok/s/user P90, P90 TTFT 0.7 s), c48 (59,740 @ 244.3, 1.4 s) and c64 (80,583 @ 207.6, 2.0 s); 1P1D at c96 (144,448 @ 132.2, 4.1 s) and, with the SGLang HiCache host-memory prefix-cache tier on the prefill worker, at c160 (172,192 @ 98.8, 21.7 s); 2P1D HiCache at c256 (174,562 @ 68.1, 18.9 s). Every point passed its GSM8K gate; at matched P90 interactivity the curve is 1.7-2.7x the published MI355X ATOM tokens/s/GPU." + - "新增 GB300 上 DeepSeek-V4.1-Flash FP4 AgentX 的 Dynamo 前端(KV 路由、会话亲和)+ SGLang nightly-dev-20260928-81f27fb3 / ai-dynamo 1.6.0.dev20260928 配方,DSpark block size 5。" + - "所有 SGLang server role 使用已挂载的共享 Hugging Face cache。" + - "聚合分支:曲线上延迟最低的两个点——纯 TP4(EP1)c1(沿用单机配方的 TP4 c1 设置、static ragged verify 与 HBM 内 Engram,置于 Dynamo 之后;6,215 tok/s/GPU @ 385.0 tok/s/user P90)与采用 static ragged verify、Engram 置于主机内存的 TP2/EP2 c1(12,501 @ 377.1,P90 TTFT 0.9 s);两者均通过 GSM8K 验证。" + - "分离式曲线(Mooncake KV 传输,TP2/EP2 worker):1P2D 配合 1 秒 Dynamo 会话亲和 TTL 的 c16(每节点一个 decode worker;19,280 tok/s/GPU @ 327.4 tok/s/user P90,P90 TTFT 0.7 s)、c48(59,740 @ 244.3,1.4 s)与 c64(80,583 @ 207.6,2.0 s);1P1D 的 c96(144,448 @ 132.2,4.1 s)以及在 prefill worker 上启用 SGLang HiCache 主机内存前缀缓存层后的 c160(172,192 @ 98.8,21.7 s);2P1D HiCache 的 c256(174,562 @ 68.1,18.9 s)。所有点均通过 GSM8K 验证;在相同 P90 交互性下,该曲线为已发布 MI355X ATOM tokens/s/GPU 的 1.7-2.7 倍。" + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3598