feat(bench): benchmark models served locally through vllm-on-tap - #668
Open
using-system wants to merge 11 commits into
Open
using-system wants to merge 11 commits into
using-system wants to merge 11 commits into
Conversation
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…80 GB Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n table Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n A100 80 GB Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…default Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oint the example at a shipped preset Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
.llms-benchmark/README.md:## Resultssplit into Remote serving (the existing tables, unchanged) and Local serving (Preset linked to its YAML, GPU, no Cost / $/confirmed, Scoring renormalized to 3/7, 2/7, 1/7, 1/7)./launch-llms-benchmark local <vot environment> <preset> [effort]: serves the preset once with vllm-on-tap, hands opencode the endpoint per launch throughOPENCODE_CONFIG(OpenRouter configuration and key untouched), runs the mission once, destroys the unit at the end..vot/environments/(azure-sweden, local-metal) and the benchmark's presets,.vot/presets/odd-qwen3-6-35b-a3b.yamland.vot/presets/odd-qwen3-8-27b.yaml.odd-qwen3-6-35b-a3bfeat: bootstrap oddyssey as an apm package for observability-driven development #1 (Scoring 18.5, 3/9, 21m38s),odd-qwen3-8-27bfeat: grafana proxy routing, observe-local-run agent, simplified readme #2 (Scoring 9.0, 7/11, 150m02s).Closes #667
Row
odd-qwen3-6-35b-a3bvot/qwen3.6-35b-a3b, no--variant: Effortdefault.Qwen/Qwen3.6-35B-A3B-FP8(MoE, 3B active), FP8 weights through Marlin, MTP speculative decoding (3 tokens), 262144 context,qwen3_codertool parser,qwen3reasoning parser, thinking on; vLLMv0.30.0. Measured before the run: 222 tok/s decode on a short prompt, 55 tok/s at 82k context, 7884 tok/s prefill.Rulings (3 / 9):
create_default_contextis 59% of the MCP's CPU self time; the catalog client opens anhttpx.Clientper call.GET /products/{sku}and 50 send spans;search_productsfetches each hit's detail (DETAIL_FANOUT=25). The report blames the agent for the fan-out, the tool does it.gen_ai.client.operation.durationis absent; onlygen_ai_client_token_usageis emitted.Row
odd-qwen3-8-27bvot/qwen3.8-27b, no--variant(the served model exposes none): Effortdefault.Qwen/Qwen3.8-27B-FP8, FP8 weights through Marlin, MTP speculative decoding (3 tokens), 262144 context,qwen3_codertool parser,qwen3reasoning parser; vLLMv0.30.0.gpu-a100profile (about 2.7 GPU-hours); run 13:43:55 → 16:13:57 UTC.—). 56 turns, median 65.7 s.Rulings (7 / 11):
8a. did not hold up: the attribute is present on the spans cited.
8b. held up: the attribute is emitted and is a retired name.
8c. held up: the series is absent from the metric listing.
Protocol changes
Local serving in
.claude/commands/launch-llms-benchmark.md: the arguments form, preflight (environment, preset flags, GPU cell), serve once, provider file withpermissiondenying a nestedopencode, output limit 32768, native maximum context, log collection from the unit's creation, one run, destroy at the end, local tables and Scoring.Non-conclusive presets tried and removed:
gemma4-12b-qat-agent: 0/5, no telemetry queried.odd-gpt-oss-120b: a nested opencode run on another model, then a stop after the preflight.odd-gpt-oss-20b: explained what it would do instead of acting, even athigheffort.odd-devstral-small-2-24b: does not start on vLLMv0.30.0(its Mistral3 and Pixtral classes import a symbol the image's transformers 5.17 no longer has, vllm#58755, fixed after the release); the image stays pinned.odd-nemotron-3-5-30b-a3b(BF16, hybrid Mamba-2 and MoE, MTP): the fastest measured (108 tok/s at 76k, 71 tok/s at 172k), but it never called the package's MCP tools or skills: it probed the backends withcurland drove its own k6 command; stopped after 9 minutes.odd-ornith-1-5-35b-a3b(FP8, MTP): with thinking it followed the harness but reached 40 minutes before writing its report (the vLLM traces show decode at 37 tok/s above 100k context, 780 s of decode over 34 requests); without thinking it drove in 50 s, then lost itself in the observability container's internals and queried a window a year off; stopped.odd-k2-horizon-36b-a4b(FP8): 256k does not fit beside the weights (full attention, BF16 KV), 128k needs--gpu-memory-utilization 0.96, the chat template needs--chat-template-content-format string; it then decoded 27 tok/s short and 22.5 tok/s at 73k, too slow to end within 40 minutes - not run.Qwen/Qwen-AgentWorld-35B-A3B(a world model simulating environments, not an agent) andQwen/Qwen3.8-Flash-Next(about 180 GB in FP8).odd-gemma4-26b-a4b(BF16, MTP drafter, 256k): 156 tok/s short, 45 tok/s at 82k. Without thinking: 9m56s, 0/1, no report persisted (its queries missed the services and it declared the telemetry absent). With thinking (enable_thinkingper request): 20m07s, 1/3, report persisted.odd-gemma4-31b(QAT w4a16, MTP drafter): 256k needs 70.8 GiB of BF16 KV (46.6 available); FP8 KV is refused on an A100; withint8_per_token_headKV it decoded 7.9 tok/s at 82k (prefill 601 tok/s) while the runs' prompts reach 200k - not run, as it could not end within 40 minutes.Review
A separate reviewer sub-agent checked the branch against main: two blocking findings (an environment exporting traces into the observed store; an example naming a removed preset), both fixed in 237e182; the re-check returned no blocking finding. The
odd-qwen3-6-35b-a3brow was added afterwards without a review round, at the maintainer's request.🤖 Generated with Claude Code