Skip to content

Enable MiniMax-M3 H100 FP8 indexer cache / 启用 MiniMax-M3 H100 FP8 索引缓存 - #3620

Open
RohitNagraj wants to merge 6 commits into
mainfrom
codex/minimaxm3-h100-fp8-indexer
Open

RohitNagraj wants to merge 6 commits into
mainfrom
codex/minimaxm3-h100-fp8-indexer

Conversation

@RohitNagraj

@RohitNagraj RohitNagraj commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Enable the FP8 MSA indexer cache for minimaxm3-fp8-h100-vllm-agentic-mtp using the official vLLM image nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89.

  • Set attention-config to {"indexer_kv_dtype":"fp8"}.
  • Set cpu-offload-gb to 5, gpu-memory-utilization to 0.93, and VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS to 0.
  • Route the config to cluster:h100-dsxe, dispatch through the registered h100-dgxc-new runner label, and use the staged MiniMax-M3 checkpoint.
  • Set the Mooncake store segments to 134GB for the concurrency 6 and 8 variants, aligned with the selected host-memory budget, and use TCP transport for their single-node DRAM stores.
  • Append the corresponding changelog entries.

Validation: local exact-key and affected-family matrix generation, all seven benchmark and seven eval-only native SRT recipe bindings, srtctl dry-run for the resident and both DRAM variants, YAML and runner inventory parsing, changelog validation, and a one-node DSXE Slurm allocation with model and Enroot access. GPU sweep and eval validation remain pending.

AI model disclosure

Prepared with GPT-6 through Codex. The exact model/version identifier was not exposed by the runtime and could not be verified.

中文

使用官方 vLLM 镜像 nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89,为 minimaxm3-fp8-h100-vllm-agentic-mtp 启用 FP8 MSA 索引缓存。

  • 将 attention-config 设为 {"indexer_kv_dtype":"fp8"}。
  • 将 cpu-offload-gb 设为 5、gpu-memory-utilization 设为 0.93,并将 VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS 设为 0。
  • 将配置切换到 cluster:h100-dsxe,通过已注册的 h100-dgxc-new 运行器标签调度,并读取已暂存的 MiniMax-M3 模型权重。
  • 根据所选的主机内存预算,将并发数为 6 和 8 的变体中的 Mooncake 存储段设为 134GB,并将其单节点 DRAM 存储切换为 TCP 传输。
  • 在变更日志末尾追加对应条目。

验证:本地精确配置键及相关配置组的矩阵生成、全部七个基准测试和七个仅评测变体的原生 SRT 配方绑定、常驻显存及两个 DRAM 变体的 srtctl dry-run、YAML 与运行器清单解析、变更日志校验,以及在 DSXE 上分配单节点并确认模型和 Enroot 可用。GPU 全量测试与准确性评测尚未完成。

AI 模型披露

通过 Codex 使用 GPT-6 完成准备工作。运行环境未提供精确的模型及版本标识,因此无法核实。

启用 MiniMax-M3 H100 FP8 索引缓存,更新 vLLM 镜像并调整内存与 Mooncake 设置。
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

补充 MiniMax-M3 H100 变更日志条目的 PR 链接。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good — straightforward config change. Reviewed the vLLM image tag bump (kept in sync between the recipe and nvidia-master.yaml), the new attention-config/env flags, the gpu-memory-utilization and cpu-offload-gb adjustments, and the recomputed Mooncake global_segment_size (140GB) — the arithmetic in the updated comment matches the new 5 GiB offload value. The changelog entry is appended at the tail, consistent with the append-only convention.

Extended reasoning...

Three-file config-only diff (recipe YAML, master-config image entry, changelog append) for the MiniMax-M3 H100 vLLM agentic recipe: bumps the vLLM image tag, adds an attention-config flag enabling FP8 indexer cache, tunes gpu-memory-utilization/cpu-offload-gb/env vars, and recomputes the Mooncake global_segment_size with a comment showing consistent math. No security-sensitive surface (auth, crypto, permissions) is touched, and the image tag stays synced across the recipe and master-config as required by convention. The bug-hunting system reported no findings and the one candidate issue it examined (CUDA-graph memory pre-reservation interaction) was ruled out, so this small mechanical config bump is safe to approve without further human review.

This review covers commit 1f06681, which is no longer the latest commit on this pull request; later commits are not covered by it.

将 MiniMax-M3 H100 AgentX 配置切换到 CoreWeave,并按该集群的主机内存预算调整 Mooncake 存储段。
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

将 MiniMax-M3 H100 扫描任务切换到 DSXE 运行器,并添加对应的 Slurm 集群配置。
让 DSXE 上的 MiniMax-M3 单节点任务读取已暂存的模型权重。
将 MiniMax-M3 的 Mooncake DRAM 存储切换为 TCP 传输,避免在单节点上注册大型 RDMA 内存段。

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant