[Research] Matched DSpark GPU benchmarks / [Research] 等条件 DSpark GPU 基准测试 - #3607
Closed
Oseltamivir wants to merge 4 commits into
Closed
Oseltamivir wants to merge 4 commits into
Oseltamivir wants to merge 4 commits into
Conversation
Add eight-GPU 8k256 comparison points with seven draft tokens and caller-supplied mean acceptance. Reject comparison overrides for AgentX and evals, preserve golden defaults, and forward exact-length controls through CI. 中文:新增八 GPU、8k256、七个草稿 token 的对比配置,由调用方指定平均接受长度;拒绝覆盖 AgentX 黄金曲线及评测,并在 CI 中传递精确长度控制。
Contributor
|
Thanks for the contribution!
中文感谢你的贡献!
|
Use the released numeric high-effort header and account for its tokens in fixed-sequence requests. Link the research changelog entries to PR 3607. 中文:使用 V4.1 发布版本的数字化 high-effort 前缀,并在固定序列请求中计入其 token;将研究日志关联至 PR 3607。
B200/B300 对比配置改用原生 W4A8 MoE 后端,避开 DeepGEMM 缩放布局断言,并记录 H200 草稿精度限制。
按用户明确授权,研究对比中的 H200 目标模型和 DSpark 使用原生 W4A8 路径;保留发布权重和 wo_a 默认转换,不作为正式基准贡献。
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Research-only DeepSeek-V4.1-Flash fixed-sequence GPU comparisons. The first matrix covers H200, B200 and B300 with eight GPUs, batch 1, 8192 input tokens, 256 output tokens and seven DSpark draft tokens. Further 32-GPU/128K points remain in progress. This draft is not a merge request or a claim of measured performance.
The manual workflow accepts an explicit mean committed-token count per verification step and exact-length controls. Comparison acceptance is rejected for AgentX and evaluations; the measured golden curves and default behavior are unchanged. No engine patches are introduced. B200/B300 use the native default MXFP4/MXFP8 path. The research-only H200 recipe explicitly selects native FP8 activations for both target and DSpark with user authorization; it is not a contribution or official benchmark submission and does not satisfy the default draft-as-shipped contribution rule.
Validation
flashinfer_mxfp4backend, whose default is MXFP4 weights with MXFP8 activations. Matrix/recipe binding and append-only changelog validation pass.032cc4f568f3f38ca881f28550d557340897ddc8. Defaultwo_aconversion and released weights are retained.dba1be0a40aa45a94ad051997016db3960a90277.AI model disclosure
The exact underlying AI model/version was not exposed by the runtime and could not be verified. One assistant prepared the code, configuration, documentation and tests. No delegated agents or other contributing models were used.
中文
目的
本草稿用于 DeepSeek-V4.1-Flash 固定序列 GPU 研究对比。首批配置覆盖 H200、B200、B300,使用八张 GPU、batch 1、8192 输入 token、256 输出 token 和七个 DSpark 草稿 token。32 GPU / 128K 配置仍在进行中。本草稿不请求合并,也不声称已获得实测性能。
手动工作流支持显式指定每次验证步骤平均提交的 token 数,以及精确序列长度。AgentX 和评测拒绝对比接受长度覆盖;实测黄金曲线及默认行为保持不变。未引入引擎补丁。B200/B300 使用原生默认 MXFP4/MXFP8 路径;H200 研究专用配置经用户明确授权,对目标模型和 DSpark 均启用原生 FP8 激活。这不是贡献或正式基准提交,也不满足默认草稿保持发布精度的贡献规则。
验证
94 项 SRT 接受长度、单节点与基准客户端测试通过;三个矩阵条目及 recipe 均通过本地验证;Ruff 检查、格式检查和工作流安全审计通过。首轮运行失败:H200 的 DeepGEMM FP4 路径不支持该架构,B200/B300 触发打包缩放因子形状断言。B200/B300 已改用固定镜像自带的
flashinfer_mxfp4,其默认精度为 MXFP4 权重和 MXFP8 激活;矩阵、recipe 绑定与仅追加 changelog 校验均通过。重新运行见上方链接。B200 已全绿:10/10 个请求,每个请求均为 8192 输入 / 256 输出 token;平均 TPOT 1.04788 ms,八张 GPU 的实测输出吞吐为 576.719 tokens/s,功耗验证通过。这是短时 concurrency-1 探测,合成 AL 设为 5.7,解码日志窗口为 5.60–5.80。B300 仍在初始化,尚无 H200 或 32 GPU / 128K 结果。用户已明确授权 H200 的目标模型和 DSpark 使用 W4A8,重新运行已提交,链接见上文;保留发布权重和默认 wo_a 转换。B200/B300 的配置、tokenizer 和权重索引 SHA256 一致,全部 48 个分片下载标识与dba1be0a40aa45a94ad051997016db3960a90277版本匹配。AI 模型披露
运行环境未暴露底层 AI 模型的精确名称或版本,因此无法核实。由单个助手完成代码、配置、文档和测试,未使用委派代理或其他参与模型。