fix(free): replace the delisted nemotron-3-nano-30b, and document the live free lineup - #73
Merged
Merged
Conversation
… live free lineup blockrun delisted nvidia/nemotron-3-nano-30b on 2026-09-08 (NVIDIA deprovisioned it for the account). The gateway redirects the id to nano-omni. Code - router_adapter.FREE_TIERS: nano-30b was SIMPLE's primary and a fallback in MEDIUM, COMPLEX and REASONING. The adapter already dropped it at runtime, since it is not in the catalog, so SIMPLE had silently been opening on lightning. Every slot now goes to nemotron-3-nano-omni-30b-a3b-reasoning: the gateway's own redirect target, the same family and size, and the model those slots were actually served by. The 2026-08-31 reason for excluding it (it answered as nano-30b) is gone with nano-30b. - router_core: re-synced to router-core e38958b (BlockRunAI/router-core#3). eco SIMPLE fallback[0] goes nano-30b -> nano-omni; nano-30b is dropped from model_capabilities; nano-omni gets the supportsVision:false override. The decision fixture is a verbatim copy of upstream's regenerated snapshot, and the parity test passes on it. - New test: FREE_TIERS may only name ids from the $0 chat set /v1/models published on 2026-09-29, and never nano-30b. The existing membership check compared the table against a hand-kept list and could not catch this. Docs - README "Available free models", the `free` routing-profile row, the hero line, both code examples, the NVIDIA section and the FAQ count now name the six live free models instead of the 08-12 lineup (step-3.7-flash, mistral-nemotron, nemotron-nano v2, gpt-oss). gpt-oss-120b/20b were retired 2026-09-03, so their "direct calls still work" rows and privacy note are gone. Counts use the models.free marker. - client.py streaming docstring: example ids moved off the retired deepseek-v4-flash / llama-4-maverick. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
VickyXAI
added a commit
that referenced
this pull request
Sep 30, 2026
… live free lineup (#73) blockrun delisted nvidia/nemotron-3-nano-30b on 2026-09-08 (NVIDIA deprovisioned it for the account). The gateway redirects the id to nano-omni. Code - router_adapter.FREE_TIERS: nano-30b was SIMPLE's primary and a fallback in MEDIUM, COMPLEX and REASONING. The adapter already dropped it at runtime, since it is not in the catalog, so SIMPLE had silently been opening on lightning. Every slot now goes to nemotron-3-nano-omni-30b-a3b-reasoning: the gateway's own redirect target, the same family and size, and the model those slots were actually served by. The 2026-08-31 reason for excluding it (it answered as nano-30b) is gone with nano-30b. - router_core: re-synced to router-core e38958b (BlockRunAI/router-core#3). eco SIMPLE fallback[0] goes nano-30b -> nano-omni; nano-30b is dropped from model_capabilities; nano-omni gets the supportsVision:false override. The decision fixture is a verbatim copy of upstream's regenerated snapshot, and the parity test passes on it. - New test: FREE_TIERS may only name ids from the $0 chat set /v1/models published on 2026-09-29, and never nano-30b. The existing membership check compared the table against a hand-kept list and could not catch this. Docs - README "Available free models", the `free` routing-profile row, the hero line, both code examples, the NVIDIA section and the FAQ count now name the six live free models instead of the 08-12 lineup (step-3.7-flash, mistral-nemotron, nemotron-nano v2, gpt-oss). gpt-oss-120b/20b were retired 2026-09-03, so their "direct calls still work" rows and privacy note are gone. Counts use the models.free marker. - client.py streaming docstring: example ids moved off the retired deepseek-v4-flash / llama-4-maverick. Co-authored-by: 1bcMax <viewitter@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This repo allows squash merges only. Squash-merge this PR.
Depends on BlockRunAI/router-core#3. The
router_coreport is re-synced to that PR's heade38958b, and the decision fixture is its regenerated snapshot. Land router-core#3 first. If it lands under a different SHA, only the pin references in three docstrings change; the code and fixture do not.Sources, fetched 2026-09-29: the $0 chat set in
blockrun.ai/api/v1/modelsisnvidia/nemotron-3-nano-omni-30b-a3b-reasoning,nvidia/nemotron-3.5-lightning,nvidia/llama-3.2-11b-vision,nvidia/nemotron-3-ultra-550b,cohere/north-mini-codeandpoolside/laguna-xs-2.1.brand/numbers.jsonhasmodels.free = 6.nvidia/nemotron-3-nano-30bwas delisted on 2026-09-08, and the gateway redirects it to nano-omni.gpt-oss-120b/20bwere retired on 2026-09-03 (available:false, redirected).Code
router_adapter.FREE_TIERSSIMPLE primarynvidia/nemotron-3-nano-30bnvidia/nemotron-3-nano-omni-30b-a3b-reasoningFREE_TIERSMEDIUM / COMPLEX / REASONING fallbacksrouter_core/config.pyeco SIMPLEfallback[0]e38958brouter_core/model_capabilities.pysupports_vision: Truesupports_vision: Falsewith the upstream override commenttests/unit/router_core_decisions.snapshot.json5ee7c23e38958b, verbatimtest_router_core_snapshot.pypasses).router_core/__init__.py,config.py,test_router_core_snapshot.py5ee7c23e38958btests/unit/test_router_adapter.pyFREE_MODELSlists nano-30btests/unit/test_router_adapter.pytest_names_only_models_in_the_live_free_catalogFREE_TIERSwith the hand-keptFREE_MODELS, which can't catch a delisting. The new test pins the table to the published $0 set and names nano-30b. I checked that it fails on the old table.tests/unit/test_routing_parity.pycatalog fixtureclient.pychat_completion_streamdocstringnvidia/deepseek-v4-flash, fallbacknvidia/llama-4-mavericknvidia/nemotron-3.5-lightning, fallback nano-omniREADME
/v1/modelslists them, a note that neither vision-catalogued free model handles images reliably, and the retired list extended with the 08-30, 09-03 and 09-08 retirements. The gpt-oss privacy note is removed (those models are retired).freerouting-profile rowFREE_TIERSroutes across, plus Ultra 550B as direct-call only (it is deliberately not in the table)br:models.freemarker plus the six namesnvidia/step-3.7-flashnvidia/nemotron-3.5-lightning;result.modelexample is nano-omninvidia/*free models, with the non-NVIDIA pair cross-referencedbr:models.freemarkerLeft alone on purpose
nemotron-3-ultra-550bstays out ofFREE_TIERS. It was excluded on 2026-08-31 because it answered as another model, and today's probe got a 429, so there is no fresh evidence either way.examples/sweep_all_chat_models.pyhard-codes retired ids. It is a historical sweep script, not a list the SDK uses.openai/gpt-oss-*rows are paid testnet models, not the free tier.Checks (throwaway venv, Python 3.13,
.[dev,solana,anthropic]as in CI)pytest tests/unit: 984 passedblack --check .: passruff check .: passnode scripts/sync-brand-numbers.mjs --check: up to datemypy blockrun_llm: 256 errors, identical to main (mypy is not a CI step)🤖 Generated with Claude Code