Skip to content

chore: sync models with dashboard API - #12

Open
amaan-ai20 wants to merge 3 commits into
mainfrom
model-sync/2026-09-21-2136
Open

amaan-ai20 wants to merge 3 commits into
mainfrom
model-sync/2026-09-21-2136

Conversation

@amaan-ai20

@amaan-ai20 amaan-ai20 commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Automated model-catalog sync. The dashboard API is the source of truth; every value below was taken from it. Merging publishes v3.11.0 (a minor bump — every sync is) to npm.

Corrected

Model Field Was Now
deepseek-v4.1-flash input / output per 1M $0.30 / $1.20 $0.14 / $0.57
qwen3-30b-a3b-fp8 input / output per 1M $0.05 / $0.30 $0.10 / $0.45
llama-3.1-8b-instruct-fast input / output per 1M $0.02 / $0.05 $0.15 / $0.28

ZGPU_FALLBACK is unchanged — glm-5.2 still holds both the highest input and the highest output rate in the catalog, so the falls back to a default ZeroGPU rate equality test still stands.

The values the sample call test hard-codes llama-3.1-8b-instruct-fast's rate in its expected cost, so that assertion was moved to the new rate alongside the table.

Added

Text Generation — full chat surfaces

All four carry a sample_responses_body, so they are routed responses. That route can be widened later if the platform confirms a Chat Completions path for any of them.

Model Context Input / output per 1M
glm-5.3-flash 1,048,576 $0.10 / $0.35
gpt-5.6-luna 272,000 $0.20 / $1.20
gpt-4.1-mini 1,047,576 $0.40 / $1.60
gpt-5.4-nano 400,000 $0.20 / $1.25

Each gets a ZGPU_PRICING + test CATALOG entry, a CHAT_MODELS entry, a row in both chat Models tables, and a mention in the DOCUMENTATION §1 reasoning-and-tool-use bullet.

parameters is null for all four, so their notes cells open with a capability phrase from pricing.description instead of a parameter count. glm-5.3-flash has an empty use_cases, so its cell is sourced from pricing.description alone.

Audio — pricing only

These two appeared in the API after this branch was first cut and were added in a follow-up commit. Neither is Text Generation, so per the sync rules they get a ZGPU_PRICING + test CATALOG entry and nothing else — no chat --model entry, no docs table row, no new per-task command.

Model Task Params Input / output per 1M
whisper-tiny Speech-to-Text 39M $0 / $0
chatterbox-nano Text-to-Speech 110M $0 / $0

Removed

deepseek-v4-flash-0731 — the API no longer returns it. Dropped from ZGPU_PRICING, the test CATALOG and the never overstates savings model list, CHAT_MODELS and the comment above it, its README row / -m example / routing sentence, its DOCUMENTATION row, §1 bullet, routing paragraph and §5 exceptions sentence, and the ADDING_COMMANDS.md Chat-Completions-only list. No command named it, so none was deleted. The Chat Completions routing sentences still name qwen3-30b-a3b-fp8 and glm-5.2, so nothing cascaded to an empty list.

Rebased onto 3.10.1

main moved while this was open (#13, #14). The 3.10.1 patch release collided with this branch's version bump — both sides only moved the version, so the resolution takes main's 3.10.1 and re-runs npm run bump:minor on it, which lands on the same 3.11.0 this branch already targeted. Merged in as a merge commit; nothing else conflicted.

Notes

Ordering choices, where the skill left slack: the four new Text Generation models are inserted in API (dashboard) order, in the same position in every surface — after deepseek-v4.1-flash, which is the last responses entry and where the removed model sat. The two audio models sit above the embedding block, under their own comment, since the existing out: 0 comment there is specifically about embeddings.

Worth a maintainer's eye: whisper-tiny and chatterbox-nano are priced but not reachable from any current command — the endpoint commands (responses, chat_completions, moderations, embeddings) are all text endpoints. Pricing them is what the sync rules call for; wiring an audio endpoint is a separate change.

Verification

  • audit-models.py --strict — 0 findings (24 models)
  • npm run lint, npm run build, npm test (56 tests) pass on the merged tree
  • zerogpu chat --help lists exactly the API's 11 Text Generation models
  • grep confirms deepseek-v4-flash-0731 appears nowhere in src/, tests/, docs/ or README.md

🤖 Generated with Claude Code

https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u

- deepseek-v4.1-flash: $0.30/$1.20 -> $0.14/$0.57 (ZGPU_PRICING, test CATALOG)
- qwen3-30b-a3b-fp8: $0.05/$0.30 -> $0.10/$0.45 (ZGPU_PRICING, test CATALOG)
- llama-3.1-8b-instruct-fast: $0.02/$0.05 -> $0.15/$0.28 (ZGPU_PRICING, test
  CATALOG, and the sample-call assertion that hard-codes the rate)
- add glm-5.3-flash, gpt-5.6-luna, gpt-4.1-mini, gpt-5.4-nano: pricing,
  chat --model (Responses), README + DOCUMENTATION rows, DOCUMENTATION §1 bullet
- remove deepseek-v4-flash-0731: pricing, test CATALOG and model list,
  CHAT_MODELS, README row + example + routing sentence, DOCUMENTATION row,
  §1 bullet, routing paragraph and §5 exceptions, ADDING_COMMANDS list
- version 3.10.0 -> 3.11.0

ZGPU_FALLBACK is unchanged: glm-5.2 still holds both the highest input and the
highest output rate in the catalog.

Source: https://api-dashboard.zerogpu.ai/api/models

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u
@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 54336eb5-d9ba-4a36-bf38-4d6c44bf9825


Comment @coderabbitai help to get the list of available commands.

Resolves the package.json / package-lock.json version conflict with the
3.10.1 patch release (#14). Both sides only moved the version, so the
resolution is main's 3.10.1 re-bumped with `npm run bump:minor` — the same
3.11.0 this branch already targeted, now cut from the current base.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u
The dashboard API added two models after this branch was cut:

- whisper-tiny (Speech-to-Text, 39M, $0/$0)
- chatterbox-nano (Text-to-Speech, 110M, $0/$0)

Neither is Text Generation, so per the model-sync rules they need a
ZGPU_PRICING and test CATALOG entry only — no chat --model entry, no docs
table row, and no new per-task command.

Source: https://api-dashboard.zerogpu.ai/api/models

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants