Skip to content

[PowerX] centralize DCGM profiles and backfill power coverage / 集中配置 DCGM profiles 并补齐功耗覆盖 #3592

Description

@edwingao28

Is your feature request related to a problem? Please describe.

Power collection settings are repeated across recipes, and published measured-power coverage remains uneven.

Describe the solution you'd like

Centralize optional DCGM profiles (inherit, stock, noprof) while preserving defaults, recipe overrides and independent CPU/AMD collectors. Validate a representative model per materially distinct runtime environment; reuse compatible evidence. Keep required PR sweeps/evals for affected rollout scopes.

flowchart LR
    A["Recipe / run configuration"] --> B["Shared collection configuration"]
    B --> C["GPU: DCGM profile / AMD collector"]
    B --> D["CPU power collector"]
    C --> E["Validated power and energy"]
    D --> E
    E --> F["Model × hardware qualification"]
    F --> G["Production and backfill"]
Loading

Collection & sampling

  • Shared DCGM profile selector, native command generation, configuration tests and provenance — PR pending.
  • Producer sampling diagnostics — srt-slurm #23.
  • Consumer gap reporting and measurement-window validation — #3353.
  • Legacy ACPI CPU power discovery — upstream #524.
  • Grace CPU power collection — #3296.
  • AMD monitor window and worker lifecycle — #3055.
  • Required NVIDIA 8k1k power collection — #3027.

Model coverage & backfill

Each checkbox identifies one model × hardware follow-up; attach its PR when scoped. Recover qualified artifacts first, then run the required sweep/evals for confirmed measurement gaps. Include all active engine, precision, topology and concurrency variants. DeepSeek-R1-0528 is excluded from backfill because it is planned for deprecation.

API screen, 2026-09-29: numbers are strict-v2 power rows / returned rows for that workload. Partial counts can include legacy variants; compare active configs before scheduling. Reuse candidates have strict power on every returned row, but still require matrix, eval and raw-telemetry acceptance. This screen alone does not mark delivery complete.

Kimi-K3

  • H200 — AgentX 10/35 · #3550
  • B300 — AgentX 0/11
  • GB300 — AgentX 5/11
  • MI355X — AgentX 10/16

Reuse candidates: B200 (AgentX 7/7), GB200 (AgentX 13/13).

GLM-5.3

  • MI355X — AgentX 0/1

GLM-5.2

  • H200 — AgentX 0/5
  • B200 — AgentX 10/15
  • GB200 — AgentX 0/6
  • GB300 — AgentX 0/13

Reuse candidates: B300 (AgentX 16/16), MI325X (AgentX 7/7), MI355X (AgentX 28/28).

DeepSeek-V4-Pro

  • B200 — AgentX 24/30
  • GB200 — AgentX 0/7 · #3358
  • GB300 — AgentX 1/27
  • MI355X — AgentX 31/42

Reuse candidates: H200 (AgentX 5/5), B300 (AgentX 29/29).

DeepSeek-V4.1-Flash

  • B200 — AgentX 26/28
  • GB200 — AgentX 16/32
  • MI300X — AgentX 2/12
  • MI325X — AgentX 7/18
  • MI355X — AgentX 26/33

Reuse candidates: H100 (AgentX 18/18), H200 (AgentX 32/32), B300 (AgentX 31/31), GB300 (AgentX 32/32).

MiniMax-M3

  • B200 — AgentX 11/26
  • B300 — AgentX 11/19
  • GB200 — AgentX 0/13
  • GB300 — AgentX 0/6
  • MI325X — AgentX 0/9
  • MI355X — AgentX 30/31

Reuse candidates: H100 (AgentX 7/7), H200 (AgentX 7/7), MI300X (AgentX 6/6).

Qwen3.8-Flash-Next

Reuse candidates: H200 (AgentX 5/5), B200 (AgentX 4/4), B300 (AgentX 5/5).

Qwen3.5-397B-A17B

  • H200 — AgentX 8/8; 8k1k 0/11 · 8k1k #3172
  • B200 — AgentX 39/71; 8k1k 40/68 · 8k1k #3378
  • B300 — AgentX 0/71; 8k1k 12/58 · 8k1k #3379
  • GB200 — AgentX 7/11; 8k1k 21/23
  • GB300 — AgentX 2/11; 8k1k 9/53
  • MI300X — AgentX 0/10; 8k1k 0/10
  • MI325X — AgentX 0/32; 8k1k 0/20
  • MI355X — AgentX 16/16; 8k1k 33/104 · FP8 8k1k #3380

Reuse candidates: H100 (8k1k 13/13; AgentX 9/9).

GLM-5.1 / TileRT

  • B200 — 8k1k 1/1; 1k1k 0/1

Production

  • Require validated AgentX power for publication — #3165.
  • Persist telemetry, recover retained artifacts and verify DB → API → page delivery — InferenceX-app #1167.

Describe alternatives you've considered

Per-recipe exporter command overrides.

Additional context

Parent: #2681. Coverage snapshot: NVIDIA, AMD: 60 model × hardware pairs, 70 workload cells in the source snapshot; 51 pairs and 61 workload cells in scope after excluding R1. API screen uses latest benchmark results, retaining exact raw model keys.

GLM-5.1 B200 AgentX is documented and has a published observation but lacks an active master entry; reconcile its support status separately. Exclude retired and other out-of-master observations from new sweep targets.

中文

集中配置可选 DCGM profiles(inherit、stock、noprof),保留默认行为、recipe 覆盖及独立 CPU/AMD 采集。按有实质差异的运行环境选择代表模型验证,复用兼容证据;实际启用范围仍保留正式 sweep/eval 要求。

上方按 model 排列本次覆盖范围。DeepSeek-R1-0528 计划弃用,已移出 backfill 范围,不安排补测。每个 checkbox 对应一个 model × hardware 后续交付,确定范围后补 PR。优先恢复合格 artifacts,确认测量缺口后再补正式实验。

数字是 2026-09-29 API 中各 workload 的 strict-v2 功耗行数/返回行数,部分差异可能来自旧配置,需对照 active configs。Reuse candidates 表示返回行均有有效功耗,仍需核对完整矩阵、eval 与原始遥测。当前筛查不代表正式验收完成。

采集、sampling、CPU/AMD 后端和发布 PR 共用上方清单。生产验收覆盖 artifacts → DB → API → 页面。GLM-5.1 B200 的 AgentX 文档、发布数据和 active config 状态需要单独对齐;已退役及其他未进入当前主配置的历史数据不自动触发新 sweep。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions