[PowerX] MiniMax-M3 B200 AgentX measured-power coverage / 补齐 MiniMax-M3 B200 AgentX 功耗覆盖 - #3628
edwingao28 wants to merge 2 commits into
Conversation
Add require-power on MiniMax-M3 B200 vLLM and TRT-LLM AgentX scenario arms, and wire agentic matrix/workflow support so the flag takes effect. 为 MiniMax-M3 B200 AgentX(vLLM + TRT-LLM)启用 require-power,并接通 AgentX 矩阵与 sweep 工作流,使该标记生效。 Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
|
Thanks for the contribution!
中文感谢你的贡献!
|
将 MiniMax-M3 B200 PowerX changelog 的 pr-link 更新为 #3628。 Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
|
Claude finished @edwingao28's task in 59s —— View job PR ReviewLGTM - no blocking issues found I couldn't run pytest in this job because the command needed approval. This review is based on reading the code. |
| if is_multinode: | ||
| entry[Fields.DISAGG.value] = disagg | ||
| entry[Fields.SCENARIO_TYPE.value] = "agentic-coding" | ||
| require_power = scenario.get( |
There was a problem hiding this comment.
📈 Line Count Report
- Total Lines: 1684
- Base Lines: 1679
- Change: +5 lines
There was a problem hiding this comment.
LGTM — straightforward config/plumbing change that follows existing patterns.
What was reviewed: the require-power flag threading through run-sweep.yml → generate.py's _agentic_entries → validation.py's Pydantic models, and the new nvidia-master.yaml entries setting the flag on three MiniMax-M3 B200 AgentX scenarios. Checked whether the missing multi-node safeguard noted elsewhere applies here — both edited configs are multinode: false, so that gap isn't exercised by this PR's actual config changes. The perf-changelog entry is appended at the tail as required.
Extended reasoning...
The diff touches five files: a GitHub Actions workflow (adds a computed require-power boolean passed to reusable job workflows), a YAML config (sets require-power: true on three single-node AgentX scenarios), generate.py (stamps the flag onto generated agentic entries), validation.py (adds three new Pydantic fields using the existing kebab-case-alias-to-snake_case pattern), and an append-only changelog. No auth, crypto, or data-exposure surface is touched; this is CI/config plumbing. Decided defer rather than approve because inline findings exist (missing safeguard for a hypothetical future multi-node require-power misuse, and the canary not enforcing require-power) that a human should weigh, even though I verified this PR's own config edits are single-node only and don't trigger that gap.
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36831965498 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36831965498 |
Description
Require validated GPU power on the active MiniMax-M3 B200 AgentX configs (issue #3592 gap: AgentX 11/26 strict-v2 power rows). No prior PowerX PR for this pair.
Config
minimaxm3-fp4-b200-vllm-agentic-mtp—require-power: trueon bothagentic-codingarms (15 points)minimaxm3-fp4-b200-trtllm-agentic-mtp—require-power: trueon the single arm (11 points)Plumbing (required for AgentX
require-powerto take effect on current main)require-poweronAgenticCodingConfig/ agentic matrix entriesrequire-powerintosweep-agentic/sweep-multi-node-agenticinrun-sweep.ymlSingle-node AgentX already collects GPU power via
ENABLE_AGENTX_POWER+ shared monitor; no recipe telemetry blocks added.Testing
uv run --locked python -m infx.matrix.generate test-config ... --config-keys minimaxm3-fp4-b200-vllm-agentic-mtp minimaxm3-fp4-b200-trtllm-agentic-mtp --scenario-type agentic-coding --no-evals→ 26 rows, allrequire-power: true(15 vLLM + 11 TRT)uv run --locked python -m infx.workflows.validate_perf_changelog --changelog-file perf-changelog.yaml --base-ref origin/main --head-ref HEAD→ passedDo not add
full-sweep-fail-fast/priority/skip_queuein this PR open step (operators label when ready).中文
为 MiniMax-M3 B200 现网 AgentX 配置强制校验 GPU 功耗(#3592 缺口:AgentX 11/26)。vLLM 两个 scenario arm、TRT-LLM 一个 arm 均设置
require-power: true;拓扑/并发/镜像不变。当前 main 上 AgentX 需接通 schema/矩阵/sweep 才能使该标记生效;单节点已通过ENABLE_AGENTX_POWER采集功耗,未新增 recipe telemetry。本地矩阵生成 26/26 带require-power,changelog 校验通过。请勿在本 PR 打开时添加 fail-fast/priority 标签。AI model disclosure
Related Issue
Related to #3592.
Type of Change
require-powermatrix/workflow wiring needed for the flag to take effectChecklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.