Skip to content

feat(bench): benchmark models served locally through vllm-on-tap - #668

Open
using-system wants to merge 11 commits into
mainfrom
feat/update-benchmark
Open

using-system wants to merge 11 commits into
mainfrom
feat/update-benchmark

Conversation

@using-system

@using-system using-system commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • .llms-benchmark/README.md: ## Results split into Remote serving (the existing tables, unchanged) and Local serving (Preset linked to its YAML, GPU, no Cost / $/confirmed, Scoring renormalized to 3/7, 2/7, 1/7, 1/7).
  • /launch-llms-benchmark local <vot environment> <preset> [effort]: serves the preset once with vllm-on-tap, hands opencode the endpoint per launch through OPENCODE_CONFIG (OpenRouter configuration and key untouched), runs the mission once, destroys the unit at the end.
  • .vot/environments/ (azure-sweden, local-metal) and the benchmark's presets, .vot/presets/odd-qwen3-6-35b-a3b.yaml and .vot/presets/odd-qwen3-8-27b.yaml.
  • Local rows on an A100 80 GB: odd-qwen3-6-35b-a3b feat: bootstrap oddyssey as an apm package for observability-driven development #1 (Scoring 18.5, 3/9, 21m38s), odd-qwen3-8-27b feat: grafana proxy routing, observe-local-run agent, simplified readme #2 (Scoring 9.0, 7/11, 150m02s).

Closes #667

Row odd-qwen3-6-35b-a3b

  • CLI: opencode 1.18.32 (official install script binary), model vot/qwen3.6-35b-a3b, no --variant: Effort default.
  • Preset: Qwen/Qwen3.6-35B-A3B-FP8 (MoE, 3B active), FP8 weights through Marlin, MTP speculative decoding (3 tokens), 262144 context, qwen3_coder tool parser, qwen3 reasoning parser, thinking on; vLLM v0.30.0. Measured before the run: 222 tok/s decode on a short prompt, 55 tok/s at 82k context, 7884 tok/s prefill.
  • Run 20:19:42 → 20:41:20 UTC, ended on its own after persisting its report; no compaction.
  • Tokens: input 9,018,163, output 25,908 + reasoning 10,870; 76 turns, median 7.5 s. 4/4 signals queried.

Rulings (3 / 9):

  1. held up: create_default_context is 59% of the MCP's CPU self time; the catalog client opens an httpx.Client per call.
  2. held up: the cited trace carries 26 GET /products/{sku} and 50 send spans; search_products fetches each hit's detail (DETAIL_FANOUT=25). The report blames the agent for the fan-out, the tool does it.
  3. did not hold up: the 4923 span calls are the API's server spans, not an overcount against rooted traces.
  4. did not hold up: the 22 out-of-stock rejections are the API's correct behavior, as the report itself says.
  5. did not hold up: the p99 above the worst trace is bucket interpolation over 8 calls.
  6. held up: gen_ai.client.operation.duration is absent; only gen_ai_client_token_usage is emitted.
  7. did not hold up: the Python profiler collects CPU only; the memory types listed belong to another service.
  8. did not hold up: no HTTP client spans on the API is expected, as the report itself says.
  9. did not hold up: a trace-level histogram outside PromQL is not a defect.

Row odd-qwen3-8-27b

  • CLI: opencode 1.18.32 (official install script binary), model vot/qwen3.8-27b, no --variant (the served model exposes none): Effort default.
  • Preset: Qwen/Qwen3.8-27B-FP8, FP8 weights through Marlin, MTP speculative decoding (3 tokens), 262144 context, qwen3_coder tool parser, qwen3 reasoning parser; vLLM v0.30.0.
  • Unit served 13:32 → 16:14 UTC on the Container Apps gpu-a100 profile (about 2.7 GPU-hours); run 13:43:55 → 16:13:57 UTC.
  • Tokens: input 5,000,171, output 42,711 + reasoning 139,401; vLLM reports no cached share through the API (Cache —). 56 turns, median 65.7 s.
  • The run was stopped at 16:13:57, after its report was persisted and committed, when the root session started a post-report code reading that could no longer change the report.
  • One context compaction at 15:04 (the prompt plus the reserved 32k output exceeded 262144); the run re-measured its figures before writing the report.
  • 4/4 signals queried; source files read only after the measurements that led to them; no traffic of its own outside the stored scenario; the report carries a replayable measurement protocol.

Rulings (7 / 11):

  1. held up: the cited span counts and the exemplar match the store; the code does it.
  2. held up: the profile frame's self share matches to the hundredth; the code does it.
  3. held up: the cited quantiles and profile shares match the store; the code does it.
  4. held up: the counter delta and the span calls match; the increment sits after the early return.
  5. did not hold up: the routes that ran before the drive are not instrumented, so the absence is no traffic, not an export lag.
  6. did not hold up: the lines exist; they fall seconds after the window the run queried.
  7. did not hold up: the arrival count is run-dependent (8 or 9); nothing observed is at fault.
    8a. did not hold up: the attribute is present on the spans cited.
    8b. held up: the attribute is emitted and is a retired name.
    8c. held up: the series is absent from the metric listing.
  8. held up: the label carries one constant value; the code writes it.

Protocol changes

Local serving in .claude/commands/launch-llms-benchmark.md: the arguments form, preflight (environment, preset flags, GPU cell), serve once, provider file with permission denying a nested opencode, output limit 32768, native maximum context, log collection from the unit's creation, one run, destroy at the end, local tables and Scoring.

Non-conclusive presets tried and removed:

  • gemma4-12b-qat-agent: 0/5, no telemetry queried.
  • odd-gpt-oss-120b: a nested opencode run on another model, then a stop after the preflight.
  • odd-gpt-oss-20b: explained what it would do instead of acting, even at high effort.
  • odd-devstral-small-2-24b: does not start on vLLM v0.30.0 (its Mistral3 and Pixtral classes import a symbol the image's transformers 5.17 no longer has, vllm#58755, fixed after the release); the image stays pinned.
  • odd-nemotron-3-5-30b-a3b (BF16, hybrid Mamba-2 and MoE, MTP): the fastest measured (108 tok/s at 76k, 71 tok/s at 172k), but it never called the package's MCP tools or skills: it probed the backends with curl and drove its own k6 command; stopped after 9 minutes.
  • odd-ornith-1-5-35b-a3b (FP8, MTP): with thinking it followed the harness but reached 40 minutes before writing its report (the vLLM traces show decode at 37 tok/s above 100k context, 780 s of decode over 34 requests); without thinking it drove in 50 s, then lost itself in the observability container's internals and queried a window a year off; stopped.
  • odd-k2-horizon-36b-a4b (FP8): 256k does not fit beside the weights (full attention, BF16 KV), 128k needs --gpu-memory-utilization 0.96, the chat template needs --chat-template-content-format string; it then decoded 27 tok/s short and 22.5 tok/s at 73k, too slow to end within 40 minutes - not run.
  • Skipped without a serve: Qwen/Qwen-AgentWorld-35B-A3B (a world model simulating environments, not an agent) and Qwen/Qwen3.8-Flash-Next (about 180 GB in FP8).
  • odd-gemma4-26b-a4b (BF16, MTP drafter, 256k): 156 tok/s short, 45 tok/s at 82k. Without thinking: 9m56s, 0/1, no report persisted (its queries missed the services and it declared the telemetry absent). With thinking (enable_thinking per request): 20m07s, 1/3, report persisted.
  • odd-gemma4-31b (QAT w4a16, MTP drafter): 256k needs 70.8 GiB of BF16 KV (46.6 available); FP8 KV is refused on an A100; with int8_per_token_head KV it decoded 7.9 tok/s at 82k (prefill 601 tok/s) while the runs' prompts reach 200k - not run, as it could not end within 40 minutes.

Review

A separate reviewer sub-agent checked the branch against main: two blocking findings (an environment exporting traces into the observed store; an example naming a removed preset), both fixed in 237e182; the re-check returned no blocking finding. The odd-qwen3-6-35b-a3b row was added afterwards without a review round, at the maintainer's request.

🤖 Generated with Claude Code

using-system and others added 11 commits September 27, 2026 10:44
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…80 GB

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n table

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n A100 80 GB

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…default

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oint the example at a shipped preset

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(bench): benchmark models served locally through vllm-on-tap

1 participant