chore: sync models with dashboard API - #12
Open
amaan-ai20 wants to merge 3 commits into
Open
amaan-ai20 wants to merge 3 commits into
amaan-ai20 wants to merge 3 commits into
Conversation
- deepseek-v4.1-flash: $0.30/$1.20 -> $0.14/$0.57 (ZGPU_PRICING, test CATALOG) - qwen3-30b-a3b-fp8: $0.05/$0.30 -> $0.10/$0.45 (ZGPU_PRICING, test CATALOG) - llama-3.1-8b-instruct-fast: $0.02/$0.05 -> $0.15/$0.28 (ZGPU_PRICING, test CATALOG, and the sample-call assertion that hard-codes the rate) - add glm-5.3-flash, gpt-5.6-luna, gpt-4.1-mini, gpt-5.4-nano: pricing, chat --model (Responses), README + DOCUMENTATION rows, DOCUMENTATION §1 bullet - remove deepseek-v4-flash-0731: pricing, test CATALOG and model list, CHAT_MODELS, README row + example + routing sentence, DOCUMENTATION row, §1 bullet, routing paragraph and §5 exceptions, ADDING_COMMANDS list - version 3.10.0 -> 3.11.0 ZGPU_FALLBACK is unchanged: glm-5.2 still holds both the highest input and the highest output rate in the catalog. Source: https://api-dashboard.zerogpu.ai/api/models Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: Comment |
Resolves the package.json / package-lock.json version conflict with the 3.10.1 patch release (#14). Both sides only moved the version, so the resolution is main's 3.10.1 re-bumped with `npm run bump:minor` — the same 3.11.0 this branch already targeted, now cut from the current base. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u
The dashboard API added two models after this branch was cut: - whisper-tiny (Speech-to-Text, 39M, $0/$0) - chatterbox-nano (Text-to-Speech, 110M, $0/$0) Neither is Text Generation, so per the model-sync rules they need a ZGPU_PRICING and test CATALOG entry only — no chat --model entry, no docs table row, and no new per-task command. Source: https://api-dashboard.zerogpu.ai/api/models Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Automated model-catalog sync. The dashboard API is the source of truth; every value below was taken from it. Merging publishes v3.11.0 (a minor bump — every sync is) to npm.
Corrected
deepseek-v4.1-flashqwen3-30b-a3b-fp8llama-3.1-8b-instruct-fastZGPU_FALLBACKis unchanged —glm-5.2still holds both the highest input and the highest output rate in the catalog, so thefalls back to a default ZeroGPU rateequality test still stands.The
values the sample calltest hard-codesllama-3.1-8b-instruct-fast's rate in its expected cost, so that assertion was moved to the new rate alongside the table.Added
Text Generation — full chat surfaces
All four carry a
sample_responses_body, so they are routedresponses. That route can be widened later if the platform confirms a Chat Completions path for any of them.glm-5.3-flashgpt-5.6-lunagpt-4.1-minigpt-5.4-nanoEach gets a
ZGPU_PRICING+ testCATALOGentry, aCHAT_MODELSentry, a row in bothchatModels tables, and a mention in the DOCUMENTATION §1 reasoning-and-tool-use bullet.parametersisnullfor all four, so their notes cells open with a capability phrase frompricing.descriptioninstead of a parameter count.glm-5.3-flashhas an emptyuse_cases, so its cell is sourced frompricing.descriptionalone.Audio — pricing only
These two appeared in the API after this branch was first cut and were added in a follow-up commit. Neither is
Text Generation, so per the sync rules they get aZGPU_PRICING+ testCATALOGentry and nothing else — nochat --modelentry, no docs table row, no new per-task command.whisper-tinychatterbox-nanoRemoved
deepseek-v4-flash-0731— the API no longer returns it. Dropped fromZGPU_PRICING, the testCATALOGand thenever overstates savingsmodel list,CHAT_MODELSand the comment above it, its README row /-mexample / routing sentence, its DOCUMENTATION row, §1 bullet, routing paragraph and §5 exceptions sentence, and theADDING_COMMANDS.mdChat-Completions-only list. No command named it, so none was deleted. The Chat Completions routing sentences still nameqwen3-30b-a3b-fp8andglm-5.2, so nothing cascaded to an empty list.Rebased onto 3.10.1
mainmoved while this was open (#13, #14). The 3.10.1 patch release collided with this branch's version bump — both sides only moved the version, so the resolution takes main's 3.10.1 and re-runsnpm run bump:minoron it, which lands on the same 3.11.0 this branch already targeted. Merged in as a merge commit; nothing else conflicted.Notes
Ordering choices, where the skill left slack: the four new Text Generation models are inserted in API (dashboard) order, in the same position in every surface — after
deepseek-v4.1-flash, which is the lastresponsesentry and where the removed model sat. The two audio models sit above the embedding block, under their own comment, since the existingout: 0comment there is specifically about embeddings.Worth a maintainer's eye:
whisper-tinyandchatterbox-nanoare priced but not reachable from any current command — the endpoint commands (responses,chat_completions,moderations,embeddings) are all text endpoints. Pricing them is what the sync rules call for; wiring an audio endpoint is a separate change.Verification
audit-models.py --strict— 0 findings (24 models)npm run lint,npm run build,npm test(56 tests) pass on the merged treezerogpu chat --helplists exactly the API's 11 Text Generation modelsgrepconfirmsdeepseek-v4-flash-0731appears nowhere insrc/,tests/,docs/orREADME.md🤖 Generated with Claude Code
https://claude.ai/code/session_014jFCNTvLBuLipdPRdTbg6u