Skip to content

refactor(bench): replace benchmark_lib.sh with the infx.bench package - #3610

Draft
adibarra wants to merge 2 commits into
mainfrom
refactor/remove-benchmark-lib
Draft

adibarra wants to merge 2 commits into
mainfrom
refactor/remove-benchmark-lib

Conversation

@adibarra

Copy link
Copy Markdown
Collaborator

Summary

Deletes benchmarks/benchmark_lib.sh and moves its behavior into infx/bench/, a container-side package (stdlib only, Python 3.10) behind one entrypoint: python3 -m infx.bench {wait,monitor,fixed-seq,agentic,eval}.

  • Shared plumbing lives once: proc.py (exit statuses, calls, tee, signal relay/defer, PYTHONPATH), server.py (readiness), env.py (inputs, BenchError reported as one ERROR: line).
  • Registries as dicts (commands, eval frameworks, vendor providers, trace loaders), matching infx/launch.
  • srt entry scripts are thin shims; recipe command paths are unchanged.
  • benchmarks/check_env.sh keeps check_env_vars for workflows, hooks and legacy lanes.
  • Legacy amd_utils/TileRT lanes call the new commands.

Removed

  • SWE-bench Lite and Modal (unused by any config).
  • Unreachable llm-d lane and amd_utils/server_sglang.sh.
  • ~15 dead lib helpers and source-time side effects.
  • AgentX single-concurrency check moved from the workflow into LaunchRequest.

Behavior changes

  • PROFILE trace relay dropped (unreachable on srt; profile.yml relay steps are now dead).
  • Fixes: wrong AgentX corpus pre-download, invalid gb300 meta_env.json with batched EVAL_CONC, vendor failure writers on Python 3.10, SPEED-Bench aborting on the first dead server.
  • PYTHONSAFEPATH no longer leaks into children (would break amd-smi).

Verification

  • Full suite passes (2132), ruff clean, all commands import under Python 3.10, bash -n on every changed script.
  • Stubbed end-to-end smokes of every shim/command.
  • Not yet run on hardware. Needed before merge: AMD fixed-seq point (amd-smi sampling), single-node and multi-node AgentX points, one eval-only run.

Delete benchmarks/benchmark_lib.sh and move its behavior into the
container-side Python package infx/bench (stdlib only, Python 3.10):
wait, monitor, fixed-seq, agentic and eval commands behind one
`python3 -m infx.bench` entrypoint, with shared process, readiness and
input helpers. srt entry scripts become thin shims; recipe paths are
unchanged. Also removes SWE-bench Lite/Modal, the unreachable llm-d lane
and amd_utils/server_sglang.sh, and moves AgentX single-concurrency
validation into LaunchRequest.

中文:用 infx.bench 包替换 benchmark_lib.sh。删除 benchmark_lib.sh,将其功能迁移到容器内运行的 Python 包 infx/bench(仅依赖标准库,兼容 Python 3.10),通过统一的 `python3 -m infx.bench` 入口提供 wait、monitor、fixed-seq、agentic 与 eval 命令。srt 入口脚本改为轻量封装,recipe 路径保持不变。同时移除 SWE-bench Lite/Modal、不可达的 llm-d 流水线与 amd_utils/server_sglang.sh,并将 AgentX 单并发校验移至 LaunchRequest。
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

Main removed the legacy AMD/TileRT launch lanes, so the code only they
used goes too: server_watch, agentic --server-pid, the monitor command
and --require-vendor, fixed-seq point mode, and the native power
collector scripts. Ports main's lib changes: the multi-node sweep honors
CLIENT_BACKEND, BENCHMARK_SERVED_MODEL_NAME and NUM_PROMPTS, and the
lm-eval native context reads config.json before transformers. New recipe
setup scripts source check_env.sh.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant