Skip to content

[Research] Matched DSpark GPU benchmarks / [Research] 等条件 DSpark GPU 基准测试 - #3607

Closed
Oseltamivir wants to merge 4 commits into
mainfrom
bench/dsv41flash-matched-gpu
Closed

Oseltamivir wants to merge 4 commits into
mainfrom
bench/dsv41flash-matched-gpu

Conversation

@Oseltamivir

@Oseltamivir Oseltamivir commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Research-only DeepSeek-V4.1-Flash fixed-sequence GPU comparisons. The first matrix covers H200, B200 and B300 with eight GPUs, batch 1, 8192 input tokens, 256 output tokens and seven DSpark draft tokens. Further 32-GPU/128K points remain in progress. This draft is not a merge request or a claim of measured performance.

The manual workflow accepts an explicit mean committed-token count per verification step and exact-length controls. Comparison acceptance is rejected for AgentX and evaluations; the measured golden curves and default behavior are unchanged. No engine patches are introduced. B200/B300 use the native default MXFP4/MXFP8 path. The research-only H200 recipe explicitly selects native FP8 activations for both target and DSpark with user authorization; it is not a contribution or official benchmark submission and does not satisfy the default draft-as-shipped contribution rule.

Validation

  • 94 focused SRT acceptance, single-node and benchmark-client tests passed.
  • All three matrix entries and recipe variants validate locally.
  • Ruff check/format passed; workflow security audit reported no findings.
  • Initial runtime probes failed: H200 rejects the selected DeepGEMM FP4 architecture; B200/B300 fail a packed-scale shape assertion.
  • B200/B300 now use the pinned native flashinfer_mxfp4 backend, whose default is MXFP4 weights with MXFP8 activations. Matrix/recipe binding and append-only changelog validation pass.
  • B200 rerun is green: 10/10 requests, each exactly 8192 input / 256 output tokens; mean TPOT 1.04788 ms, measured output throughput 576.719 tokens/s across 8 GPUs. Power validation passed. This is a short concurrency-1 probe; configured synthetic AL is 5.7, with decode-log windows 5.60–5.80.
  • B300 rerun is still initializing. No H200 or 32-GPU/128K results claimed.
  • The user explicitly authorized H200 target/draft W4A8 for this research-only comparison. H200 rerun is dispatched on 032cc4f568f3f38ca881f28550d557340897ddc8. Default wo_a conversion and released weights are retained.
  • B200/B300 config, tokenizer and weight-index SHA256 hashes agree; all 48 staged shard download identifiers match revision dba1be0a40aa45a94ad051997016db3960a90277.

AI model disclosure

The exact underlying AI model/version was not exposed by the runtime and could not be verified. One assistant prepared the code, configuration, documentation and tests. No delegated agents or other contributing models were used.

中文

目的

本草稿用于 DeepSeek-V4.1-Flash 固定序列 GPU 研究对比。首批配置覆盖 H200、B200、B300,使用八张 GPU、batch 1、8192 输入 token、256 输出 token 和七个 DSpark 草稿 token。32 GPU / 128K 配置仍在进行中。本草稿不请求合并,也不声称已获得实测性能。

手动工作流支持显式指定每次验证步骤平均提交的 token 数,以及精确序列长度。AgentX 和评测拒绝对比接受长度覆盖;实测黄金曲线及默认行为保持不变。未引入引擎补丁。B200/B300 使用原生默认 MXFP4/MXFP8 路径;H200 研究专用配置经用户明确授权,对目标模型和 DSpark 均启用原生 FP8 激活。这不是贡献或正式基准提交,也不满足默认草稿保持发布精度的贡献规则。

验证

94 项 SRT 接受长度、单节点与基准客户端测试通过;三个矩阵条目及 recipe 均通过本地验证;Ruff 检查、格式检查和工作流安全审计通过。首轮运行失败:H200 的 DeepGEMM FP4 路径不支持该架构,B200/B300 触发打包缩放因子形状断言。B200/B300 已改用固定镜像自带的 flashinfer_mxfp4,其默认精度为 MXFP4 权重和 MXFP8 激活;矩阵、recipe 绑定与仅追加 changelog 校验均通过。重新运行见上方链接。B200 已全绿:10/10 个请求,每个请求均为 8192 输入 / 256 输出 token;平均 TPOT 1.04788 ms,八张 GPU 的实测输出吞吐为 576.719 tokens/s,功耗验证通过。这是短时 concurrency-1 探测,合成 AL 设为 5.7,解码日志窗口为 5.60–5.80。B300 仍在初始化,尚无 H200 或 32 GPU / 128K 结果。用户已明确授权 H200 的目标模型和 DSpark 使用 W4A8,重新运行已提交,链接见上文;保留发布权重和默认 wo_a 转换。B200/B300 的配置、tokenizer 和权重索引 SHA256 一致,全部 48 个分片下载标识与 dba1be0a40aa45a94ad051997016db3960a90277 版本匹配。

AI 模型披露

运行环境未暴露底层 AI 模型的精确名称或版本,因此无法核实。由单个助手完成代码、配置、文档和测试,未使用委派代理或其他参与模型。

Add eight-GPU 8k256 comparison points with seven draft tokens and caller-supplied mean acceptance. Reject comparison overrides for AgentX and evals, preserve golden defaults, and forward exact-length controls through CI.

中文:新增八 GPU、8k256、七个草稿 token 的对比配置,由调用方指定平均接受长度;拒绝覆盖 AgentX 黄金曲线及评测,并在 CI 中传递精确长度控制。
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

Use the released numeric high-effort header and account for its tokens in fixed-sequence requests. Link the research changelog entries to PR 3607.

中文:使用 V4.1 发布版本的数字化 high-effort 前缀,并在固定序列请求中计入其 token;将研究日志关联至 PR 3607。
B200/B300 对比配置改用原生 W4A8 MoE 后端,避开 DeepGEMM 缩放布局断言,并记录 H200 草稿精度限制。
按用户明确授权,研究对比中的 H200 目标模型和 DSpark 使用原生 W4A8 路径;保留发布权重和 wo_a 默认转换,不作为正式基准贡献。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant