Enable MiniMax-M3 H100 FP8 indexer cache / 启用 MiniMax-M3 H100 FP8 索引缓存 - #3620
RohitNagraj wants to merge 6 commits into
Conversation
启用 MiniMax-M3 H100 FP8 索引缓存,更新 vLLM 镜像并调整内存与 Mooncake 设置。
|
Thanks for the contribution!
中文感谢你的贡献!
|
补充 MiniMax-M3 H100 变更日志条目的 PR 链接。
There was a problem hiding this comment.
Looks good — straightforward config change. Reviewed the vLLM image tag bump (kept in sync between the recipe and nvidia-master.yaml), the new attention-config/env flags, the gpu-memory-utilization and cpu-offload-gb adjustments, and the recomputed Mooncake global_segment_size (140GB) — the arithmetic in the updated comment matches the new 5 GiB offload value. The changelog entry is appended at the tail, consistent with the append-only convention.
Extended reasoning...
Three-file config-only diff (recipe YAML, master-config image entry, changelog append) for the MiniMax-M3 H100 vLLM agentic recipe: bumps the vLLM image tag, adds an attention-config flag enabling FP8 indexer cache, tunes gpu-memory-utilization/cpu-offload-gb/env vars, and recomputes the Mooncake global_segment_size with a comment showing consistent math. No security-sensitive surface (auth, crypto, permissions) is touched, and the image tag stays synced across the recipe and master-config as required by convention. The bug-hunting system reported no findings and the one candidate issue it examined (CUDA-graph memory pre-reservation interaction) was ruled out, so this small mechanical config bump is safe to approve without further human review.
This review covers commit 1f06681, which is no longer the latest commit on this pull request; later commits are not covered by it.
将 MiniMax-M3 H100 AgentX 配置切换到 CoreWeave,并按该集群的主机内存预算调整 Mooncake 存储段。
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36828042798 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36828042798 |
将 MiniMax-M3 H100 扫描任务切换到 DSXE 运行器,并添加对应的 Slurm 集群配置。
让 DSXE 上的 MiniMax-M3 单节点任务读取已暂存的模型权重。
将 MiniMax-M3 的 Mooncake DRAM 存储切换为 TCP 传输,避免在单节点上注册大型 RDMA 内存段。
Enable the FP8 MSA indexer cache for
minimaxm3-fp8-h100-vllm-agentic-mtpusing the official vLLM imagenightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89.attention-configto{"indexer_kv_dtype":"fp8"}.cpu-offload-gbto5,gpu-memory-utilizationto0.93, andVLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHSto0.cluster:h100-dsxe, dispatch through the registeredh100-dgxc-newrunner label, and use the staged MiniMax-M3 checkpoint.134GBfor the concurrency 6 and 8 variants, aligned with the selected host-memory budget, and use TCP transport for their single-node DRAM stores.Validation: local exact-key and affected-family matrix generation, all seven benchmark and seven eval-only native SRT recipe bindings,
srtctl dry-runfor the resident and both DRAM variants, YAML and runner inventory parsing, changelog validation, and a one-node DSXE Slurm allocation with model and Enroot access. GPU sweep and eval validation remain pending.AI model disclosure
Prepared with GPT-6 through Codex. The exact model/version identifier was not exposed by the runtime and could not be verified.
中文
使用官方 vLLM 镜像
nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89,为minimaxm3-fp8-h100-vllm-agentic-mtp启用 FP8 MSA 索引缓存。attention-config设为{"indexer_kv_dtype":"fp8"}。cpu-offload-gb设为5、gpu-memory-utilization设为0.93,并将VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS设为0。cluster:h100-dsxe,通过已注册的h100-dgxc-new运行器标签调度,并读取已暂存的 MiniMax-M3 模型权重。134GB,并将其单节点 DRAM 存储切换为 TCP 传输。验证:本地精确配置键及相关配置组的矩阵生成、全部七个基准测试和七个仅评测变体的原生 SRT 配方绑定、常驻显存及两个 DRAM 变体的
srtctl dry-run、YAML 与运行器清单解析、变更日志校验,以及在 DSXE 上分配单节点并确认模型和 Enroot 可用。GPU 全量测试与准确性评测尚未完成。AI 模型披露
通过 Codex 使用 GPT-6 完成准备工作。运行环境未提供精确的模型及版本标识,因此无法核实。