Skip to content

[AMD][Qwen3.5] Extend the MI355X FP8 AgentX TP4 HiCache band to concurrency 80 - #3611

Open
yichiche wants to merge 3 commits into
mainfrom
amd/qwen35-fp8-tp4-hicache-48-80
Open

yichiche wants to merge 3 commits into
mainfrom
amd/qwen35-fp8-tp4-hicache-48-80

Conversation

@yichiche

@yichiche yichiche commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Performance: https://inferencex.semianalysis.com/inference?unofficialRun=36743632401
Accuracy: https://inferencex.semianalysis.com/evaluation?unofficialRun=36743632401

Summary

  • Extend the TP4 HiCache band on qwen3.5-fp8-mi355x-sglang-agentic-mtp from concurrency 40 up to 80, adding 48, 56, 64, 72, 80. The arm goes from 14 to 19 points.
  • Size the HiCache host pool in tiers by concurrency: 16-28 keep ratio-only sizing, 32-56 pin hicache-size: 253, 64-80 pin hicache-size: 400.
  • Bump this arm's image to lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260929 (Docker Hub tag HTTP 200). No other arm is touched.

Details

Search space

TP KV conc-list points
4 GPU-resident 1, 4, 8, 12, 14 5
4 HiCache (dram) 16, 18, 20, 22, 24, 28, 32, 36, 40, 48, 56, 64, 72, 80 14

The GPU-resident band is unchanged. The original grid was copied from the TP4 band of qwen3.5-fp8-b200-sglang-agentic-mtp, which stops at 32 because a B200 node has far less HBM per GPU. MI355X at 288 GB per GPU has the device KV headroom to keep feeding the host-DRAM tier past that, so these five points sample the throughput end the tier makes available.

The new points reuse the existing HiCache override shape exactly. Admission continues to follow 2x CONC with the decode graph batch capped at 128, so concurrency 72 and 80 clamp cuda-graph-max-bs-decode to 128 while max-running-requests reaches 144 and 160.

HiCache host pool

The host tier is now sized in tiers rather than by ratio alone. hicache-size overrides hicache-ratio where it is set.

conc host pool rationale
16, 18, 20, 22, 24, 28 ratio only (1.5x device KV) unchanged from what merged in #3602
32, 36, 40, 48, 56 hicache-size: 253 the value already proven on the MI355X MXFP4 AgentX arm
64, 72, 80 hicache-size: 400 the largest points need the deepest host tier to hold prefix across the 256k-capped corpus

Of the points that merged in #3602, only concurrency 32, 36 and 40 change behaviour, so only their numbers are expected to move. Concurrency 16 through 28 are byte-for-byte as merged.

Verification

  • The image tag returns HTTP 200 on Docker Hub under lmsysorg/sglang-rocm.
  • All 19 matrix points map 1:1 onto the 19 recipe overrides, with no duplicate (tp, gpus, CONC, KV_OFFLOADING) signature and no unused override.
  • enable-hierarchical-cache is present on exactly the 14 points declaring KV_OFFLOADING: dram; hicache-size appears only on the 8 points at concurrency 32 and above, and on no GPU-resident override.
  • Arm metadata still matches the recipe base for image and model path, which is what validate_recipe compares at launch.
  • perf-changelog.yaml is append-only, ends with a trailing newline, and contains no tabs, CR characters or non-ASCII bytes.

AI model disclosure

Claude Opus 5 (1M context), run through Claude Code, prepared the recipe change, the changelog entry and this PR text.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@yichiche
yichiche force-pushed the amd/qwen35-fp8-tp4-hicache-48-80 branch from e7edae8 to 9fac19e Compare September 30, 2026 16:11
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked whether pairing hicache-ratio: 1.5 with hicache-size: 400 on the same override is unprecedented in this repo — other single-node/multi-node recipes (e.g. b200-fp8-mtp/agentic.yaml) that pin hicache-size do so alone, without a hicache-ratio, but nothing in the codebase indicates SGLang rejects or misbehaves when both are set together, so I couldn't confirm this as a bug beyond the pattern mismatch already flagged as ruled out.

Extended reasoning...

This is a config-only change to an AMD MI355X SGLang recipe (concurrency band extension, a fixed hicache-size pin, and an image tag bump) plus a matching master-config and changelog update, touching no auth/crypto/permission code. A CONFIRMED finding (stale header comment now contradicting the newly-pinned hicache-size) is already queued as an inline comment, and two additional candidates about the hicache-ratio+hicache-size combination and DRAM budget sizing were investigated but could not be confirmed as bugs from the repo alone.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🟡 inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/qwen3.5/sglang/mi355x-fp8-mtp/agentic.yaml — nit: The file-level comment at lines 72-75 still says HiCache sizing is ratio-only ("the ratio keeps the host pool proportional to TP without pinning a byte count"), but this diff pins hicache-size: 400 on every HiCache override (lines 151-152 etc.), which the PR description itself says overrides hicache-ratio. A future editor reading this header comment will get stale guidance that no longer matches the recipe. Fix: update the comment to state HiCache now pins a fixed 400 GB per-rank host pool via hicache-size, with hicache-ratio left in place but superseded.

    Why this was flagged

    The header comment at agentic.yaml:72-75 describes the pre-change sizing model (ratio-only, no byte pinning). The diff adds hicache-size: 400 to all 14 HiCache override blocks (9 existing plus 5 new), which the changelog entry at perf-changelog.yaml:9159 confirms overrides hicache-ratio. On the base branch the comment was accurate; after this merge it misdescribes the current sizing mechanism for every HiCache point in the file. A maintainer adding a new HiCache point later, or another AMD arm's recipe author copying this pattern, could rely on the stale comment and omit hicache-size, reintroducing ratio-only sizing inconsistent with the rest of the band.

    Verification: nit: The file-level comment at agentic.yaml:72-75 reads "HiCache holds 1.5x the device KV pool and skips non-reusable blocks; the ratio keeps the host pool proportional to TP without pinning a byte count." The diff leaves this comment untouched but adds hicache-size: 400 to all 14 HiCache override blocks (line 152 for c16, and identically for c18/c20/c22/c24/c28/c32/c36/c40 plus the new c48/c56/c64/c72/c80).

@yichiche

yichiche commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

@1am9trash

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 36743632401

@1am9trash 1am9trash left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this. Link: https://inferencex.semianalysis.com/inference?unofficialRun=36743632401
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this. Link: https://inferencex.semianalysis.com/evaluation?unofficialRun=36743632401
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Recipe:
  • Draft precision:
    • The MTP head embedded in the served Qwen/Qwen3.5-397B-A17B-FP8 checkpoint. There is no separate draft checkpoint and the recipe sets no speculative-draft-model-path.
    • Stored precision: block-FP8. weight_scale_inv is present on the mtp.* expert weights; mtp.fc and the MTP gates are listed in modules_to_not_convert and stay at the checkpoint's unquantized dtype.
    • Default handling in the pinned image: lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260927 loads the head natively. The recipe applies no draft path, no quantization override and no dtype override, so nothing is re-quantized or up-converted at load. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not set in the recipe and does not appear in the run logs.
    • Effective serving precision: as stored — block-FP8 for the MTP experts, checkpoint dtype for mtp.fc and the gates. The draft shares the target's FP8 KV cache (kv-cache-dtype: fp8_e4m3) and has no separate KV pool.

Signed: @1am9trash

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

✅✅✅ Verdict: PASS ✅✅✅

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — @1am9trash is a listed owner of inferencex-e2e/configs/amd-master.yaml; the recipe and perf-changelog.yaml fall under the catch-all, which any recognized CODEOWNER satisfies.

✅ Check 1 (Sweep + evals on in-PR commit): PASS — In-PR commit 4a673e7 has all 19 executed agentic / jobs (c1–c80) and the agentic eval / job at success in run 36743632401. This is an AgentX-only PR, so the single-node */ and eval / jobs are correctly skipped. The pinned head 7bf5154 only merges main and leaves the PR diff unchanged.

✅ Check 2 (Evals pass): PASS — GSM8K em_strict is 0.9765 ± 0.0042 (n=1319, infrastructure_success: true) at TP4 HiCache c80, from eval_results_all/agg_eval_all.json in run 36743632401. The eval ran on lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260929, the same image as the PR config.

✅ Check 3 (Recipe linked/merged/complete): PASS — The linked SGLang cookbook Qwen3.5 page is published. Its "397B FP8 on MI355X" HiCache recipe matches all major args: --tp 4, --attention-backend aiter, --enable-aiter-allreduce-fusion, --kv-cache-dtype fp8_e4m3, page-size 16, HiCache write_through_selective/kernel/page_first, and the AITER/FlyDSL/Mamba-bf16 env vars. MTP is EAGLE 3/1/4 and the parsers are qwen3/qwen3_coder. Informational only (InferenceX tuning): hicache-size 253/400, max-running-requests, cuda-graph-max-bs-decode, chunked-prefill, scheduler-recv-interval, and the image tag (the cookbook lists 20260927, the PR uses 20260929).

✅ Check 4 (Reuse command): PASS — 1am9trash (COLLABORATOR) posted /reuse-sweep-run 36743632401 as a whole-line comment.

✅ Check 5 (Latest checklist template): PASS — Every item in the current PR_REVIEW_CHECKLIST.md template is present and checked, including the draft-precision, golden-AL, MODELS.md and Pareto items.

✅ Check 6 (Upstream images / engine-first): PASS — qwen3.5-fp8-mi355x-sglang-agentic-mtp uses framework: sglang with the upstream lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260929 image on MI355X. Engine-first ordering does not apply because this is an SGLang entry.

✅ Check 7 (No deprecated models/scenarios): PASS — On 2026-10-01, MODELS.md lists qwen3.5 agentic coding as active for fp8/fp4. Only bf16, 1k1k and 1k8k are deprecated for this model.

✅ Check 8 (No arch-reducing hacks / metrics publication): PASS — There are no --hf-overrides, JSON model overrides or model-file edits. publish_events_and_metrics is a TRT-LLM setting that SGLang does not support. The effective launch command (c80 job log) has --enable-metrics --enable-cache-report.

✅ Check 9 (Spec-decode via chat template): PASS — The recipe base sets AIPERF_APPLY_CHAT_TEMPLATE: 'true' for every MTP point.

✅ Check 10 (No engine patches): PASS — No .patch files, git apply, inline source rewrites, monkey-patching or engine wheel installs. The pinned image runs as shipped.

✅ Check 11 (Agentic spec-decode golden AL): PASS — The harness (infx/srt_slurm/synthetic_acceptance.py) injected SGLANG_SIMULATE_ACC_LEN=3.39, match-expected and real-draft-token into the agg role (c80 job log). This equals the qwen3.5_mtp.yaml golden AL for thinking_on at speculative-num-steps 3.

➖ Check 12 (Append-only): N/A — The new perf-changelog.yaml entry does not set append-only: true.

✅ Check 13 (Draft runs as shipped): PASS — The draft is the embedded MTP head of Qwen/Qwen3.5-397B-A17B-FP8. It is stored as block-FP8, with mtp.fc and mtp.layers.0.mlp.{gate,shared_expert_gate} in modules_to_not_convert (verified from the HF config.json). The recipe sets no draft path, draft quantization or dtype override. The FP8 KV cache is shared with the target, which is allowed. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is absent from the effective server env. Minor: the sign-off cites image 20260927, but the pinned image is 20260929; no submission-side draft setting exists in either case.

✅ Check 14 (Pareto coverage): PASS — Curve: qwen3.5 / agentic-coding / MI355X SGLang FP8 / sglang-rocm 20260929, from results_bmk/agg_bmk.json in run 36743632401 (source SHA 4a673e7). Metric is tput_per_gpu vs E2EL. P90: 17 frontier points out of 19 measured. P75: 19 out of 19. Both were computed with infx.workflows.pareto_coverage, with 0 invalid points.

Assessed commit: 7bf5154960fb04e99629f155824d2b4921f88734.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants