Skip to content

[PowerX] MiniMax-M3 B200 AgentX measured-power coverage / 补齐 MiniMax-M3 B200 AgentX 功耗覆盖 - #3628

Open
edwingao28 wants to merge 2 commits into
mainfrom
cursor/minimax-m3-b200-powerx-2039
Open

edwingao28 wants to merge 2 commits into
mainfrom
cursor/minimax-m3-b200-powerx-2039

Conversation

@edwingao28

Copy link
Copy Markdown
Collaborator

Description

Require validated GPU power on the active MiniMax-M3 B200 AgentX configs (issue #3592 gap: AgentX 11/26 strict-v2 power rows). No prior PowerX PR for this pair.

Config

  • minimaxm3-fp4-b200-vllm-agentic-mtp — require-power: true on both agentic-coding arms (15 points)
  • minimaxm3-fp4-b200-trtllm-agentic-mtp — require-power: true on the single arm (11 points)
  • Topology, concurrency, images, and draft settings unchanged

Plumbing (required for AgentX require-power to take effect on current main)

  • Accept require-power on AgenticCodingConfig / agentic matrix entries
  • Forward the flag from scenario arms into generated AgentX matrix rows
  • Pass require-power into sweep-agentic / sweep-multi-node-agentic in run-sweep.yml

Single-node AgentX already collects GPU power via ENABLE_AGENTX_POWER + shared monitor; no recipe telemetry blocks added.

Testing

  • uv run --locked python -m infx.matrix.generate test-config ... --config-keys minimaxm3-fp4-b200-vllm-agentic-mtp minimaxm3-fp4-b200-trtllm-agentic-mtp --scenario-type agentic-coding --no-evals → 26 rows, all require-power: true (15 vLLM + 11 TRT)
  • uv run --locked python -m infx.workflows.validate_perf_changelog --changelog-file perf-changelog.yaml --base-ref origin/main --head-ref HEAD → passed

Do not add full-sweep-fail-fast / priority / skip_queue in this PR open step (operators label when ready).

中文

为 MiniMax-M3 B200 现网 AgentX 配置强制校验 GPU 功耗(#3592 缺口:AgentX 11/26)。vLLM 两个 scenario arm、TRT-LLM 一个 arm 均设置 require-power: true;拓扑/并发/镜像不变。当前 main 上 AgentX 需接通 schema/矩阵/sweep 才能使该标记生效;单节点已通过 ENABLE_AGENTX_POWER 采集功耗,未新增 recipe telemetry。本地矩阵生成 26/26 带 require-power,changelog 校验通过。请勿在本 PR 打开时添加 fail-fast/priority 标签。

AI model disclosure

  • Model/version: Composer (exact runtime model identifier could not be verified)
  • Role: Implementation, local validation, and PR preparation

Related Issue

Related to #3592.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe): AgentX require-power matrix/workflow wiring needed for the flag to take effect

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.
Open in Web Open in Cursor 

Add require-power on MiniMax-M3 B200 vLLM and TRT-LLM AgentX scenario
arms, and wire agentic matrix/workflow support so the flag takes effect.

为 MiniMax-M3 B200 AgentX(vLLM + TRT-LLM)启用 require-power,并接通
AgentX 矩阵与 sweep 工作流,使该标记生效。

Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

将 MiniMax-M3 B200 PowerX changelog 的 pr-link 更新为 #3628。

Co-authored-by: Wenyao Gao <edwingao28@users.noreply.github.com>
@edwingao28
edwingao28 marked this pull request as ready for review October 1, 2026 07:43
@edwingao28
edwingao28 requested a review from a team October 1, 2026 07:43
@claude

claude Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @edwingao28's task in 59s —— View job


PR Review

LGTM - no blocking issues found

I couldn't run pytest in this job because the command needed approval. This review is based on reading the code.

if is_multinode:
entry[Fields.DISAGG.value] = disagg
entry[Fields.SCENARIO_TYPE.value] = "agentic-coding"
require_power = scenario.get(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📈 Line Count Report

  • Total Lines: 1684
  • Base Lines: 1679
  • Change: +5 lines

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — straightforward config/plumbing change that follows existing patterns.

What was reviewed: the require-power flag threading through run-sweep.yml → generate.py's _agentic_entries → validation.py's Pydantic models, and the new nvidia-master.yaml entries setting the flag on three MiniMax-M3 B200 AgentX scenarios. Checked whether the missing multi-node safeguard noted elsewhere applies here — both edited configs are multinode: false, so that gap isn't exercised by this PR's actual config changes. The perf-changelog entry is appended at the tail as required.

Extended reasoning...

The diff touches five files: a GitHub Actions workflow (adds a computed require-power boolean passed to reusable job workflows), a YAML config (sets require-power: true on three single-node AgentX scenarios), generate.py (stamps the flag onto generated agentic entries), validation.py (adds three new Pydantic fields using the existing kebab-case-alias-to-snake_case pattern), and an append-only changelog. No auth, crypto, or data-exposure surface is touched; this is CI/config plumbing. Decided defer rather than approve because inline findings exist (missing safeguard for a hypothetical future multi-node require-power misuse, and the canary not enforcing require-power) that a human should weigh, even though I verified this PR's own config edits are single-node only and don't trigger that gap.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant