Skip to content

fix(free): replace the delisted nemotron-3-nano-30b, and document the live free lineup - #73

Merged
VickyXAI merged 1 commit into
mainfrom
chore/savings-and-free-lineup
Sep 29, 2026
Merged

VickyXAI merged 1 commit into
mainfrom
chore/savings-and-free-lineup

Conversation

@VickyXAI

Copy link
Copy Markdown
Contributor

This repo allows squash merges only. Squash-merge this PR.

Depends on BlockRunAI/router-core#3. The router_core port is re-synced to that PR's head e38958b, and the decision fixture is its regenerated snapshot. Land router-core#3 first. If it lands under a different SHA, only the pin references in three docstrings change; the code and fixture do not.

Sources, fetched 2026-09-29: the $0 chat set in blockrun.ai/api/v1/models is nvidia/nemotron-3-nano-omni-30b-a3b-reasoning, nvidia/nemotron-3.5-lightning, nvidia/llama-3.2-11b-vision, nvidia/nemotron-3-ultra-550b, cohere/north-mini-code and poolside/laguna-xs-2.1. brand/numbers.json has models.free = 6. nvidia/nemotron-3-nano-30b was delisted on 2026-09-08, and the gateway redirects it to nano-omni. gpt-oss-120b/20b were retired on 2026-09-03 (available:false, redirected).

Code

Where Old New Why
router_adapter.FREE_TIERS SIMPLE primary nvidia/nemotron-3-nano-30b nvidia/nemotron-3-nano-omni-30b-a3b-reasoning nano-30b is no longer in the catalog. The adapter drops uncatalogued candidates, so SIMPLE was silently opening on its first fallback. nano-omni is the gateway's own redirect target for nano-30b and the same family and size (30B-A3B), so it is the model these slots were actually served by. The 2026-08-31 reason for excluding it (it answered as nano-30b) went away with nano-30b. On 2026-09-29 a model-echo probe got nano-omni's own NIM deployment back.
FREE_TIERS MEDIUM / COMPLEX / REASONING fallbacks nano-30b nano-omni Same. Each tier keeps its depth of ≥3 distinct models.
router_core/config.py eco SIMPLE fallback[0] nano-30b nano-omni Port of router-core e38958b
router_core/model_capabilities.py nano-30b entry; nano-omni supports_vision: True nano-30b removed; nano-omni supports_vision: False with the upstream override comment Port. Without it, eco SIMPLE image turns would start landing on a model whose image path fails real probes.
tests/unit/router_core_decisions.snapshot.json upstream 5ee7c23 upstream e38958b, verbatim The only decision change is one candidate id. The Python port reproduces it exactly (test_router_core_snapshot.py passes).
router_core/__init__.py, config.py, test_router_core_snapshot.py pin 5ee7c23 e38958b
tests/unit/test_router_adapter.py FREE_MODELS lists nano-30b nano-omni
tests/unit/test_router_adapter.py — test_names_only_models_in_the_live_free_catalog The existing check compared FREE_TIERS with the hand-kept FREE_MODELS, which can't catch a delisting. The new test pins the table to the published $0 set and names nano-30b. I checked that it fails on the old table.
tests/unit/test_routing_parity.py catalog fixture nano-30b row nano-omni row
client.py chat_completion_stream docstring nvidia/deepseek-v4-flash, fallback nvidia/llama-4-maverick nvidia/nemotron-3.5-lightning, fallback nano-omni Both old ids are retired

README

Section Old New
Available free models step-3.7-flash, mistral-nemotron, nano-omni, nemotron-nano-9b/12b, gpt-oss-120b/20b ("all NVIDIA-hosted") The six live ids exactly as /v1/models lists them, a note that neither vision-catalogued free model handles images reliably, and the retired list extended with the 08-30, 09-03 and 09-08 retirements. The gpt-oss privacy note is removed (those models are retired).
free routing-profile row "Step 3.7 Flash, Mistral Nemotron, Nemotron Nano Omni / 9B / 12B VL" The five models FREE_TIERS routes across, plus Ultra 550B as direct-call only (it is deliberately not in the table)
Hero line "8 fully-free NVIDIA-hosted models — DeepSeek V4 Flash … Llama 4 Maverick …" br:models.free marker plus the six names
Try-it-free and funding examples nvidia/step-3.7-flash nvidia/nemotron-3.5-lightning; result.model example is nano-omni
NVIDIA section 08-12 free rows The four live nvidia/* free models, with the non-NVIDIA pair cross-referenced
FAQ "11 NVIDIA-hosted models" br:models.free marker

Left alone on purpose

  • nemotron-3-ultra-550b stays out of FREE_TIERS. It was excluded on 2026-08-31 because it answered as another model, and today's probe got a 429, so there is no fresh evidence either way.
  • examples/sweep_all_chat_models.py hard-codes retired ids. It is a historical sweep script, not a list the SDK uses.
  • Testnet openai/gpt-oss-* rows are paid testnet models, not the free tier.
  • CHANGELOG is written at release time.

Checks (throwaway venv, Python 3.13, .[dev,solana,anthropic] as in CI)

  • pytest tests/unit: 984 passed
  • black --check .: pass
  • ruff check .: pass
  • node scripts/sync-brand-numbers.mjs --check: up to date
  • mypy blockrun_llm: 256 errors, identical to main (mypy is not a CI step)

🤖 Generated with Claude Code

… live free lineup

blockrun delisted nvidia/nemotron-3-nano-30b on 2026-09-08 (NVIDIA
deprovisioned it for the account). The gateway redirects the id to nano-omni.

Code
- router_adapter.FREE_TIERS: nano-30b was SIMPLE's primary and a fallback in
  MEDIUM, COMPLEX and REASONING. The adapter already dropped it at runtime,
  since it is not in the catalog, so SIMPLE had silently been opening on
  lightning. Every slot now goes to nemotron-3-nano-omni-30b-a3b-reasoning:
  the gateway's own redirect target, the same family and size, and the model
  those slots were actually served by. The 2026-08-31 reason for excluding it
  (it answered as nano-30b) is gone with nano-30b.
- router_core: re-synced to router-core e38958b (BlockRunAI/router-core#3).
  eco SIMPLE fallback[0] goes nano-30b -> nano-omni; nano-30b is dropped from
  model_capabilities; nano-omni gets the supportsVision:false override. The
  decision fixture is a verbatim copy of upstream's regenerated snapshot, and
  the parity test passes on it.
- New test: FREE_TIERS may only name ids from the $0 chat set /v1/models
  published on 2026-09-29, and never nano-30b. The existing membership check
  compared the table against a hand-kept list and could not catch this.

Docs
- README "Available free models", the `free` routing-profile row, the hero
  line, both code examples, the NVIDIA section and the FAQ count now name the
  six live free models instead of the 08-12 lineup (step-3.7-flash,
  mistral-nemotron, nemotron-nano v2, gpt-oss). gpt-oss-120b/20b were retired
  2026-09-03, so their "direct calls still work" rows and privacy note are
  gone. Counts use the models.free marker.
- client.py streaming docstring: example ids moved off the retired
  deepseek-v4-flash / llama-4-maverick.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@VickyXAI
VickyXAI merged commit 5915342 into main Sep 29, 2026
4 checks passed
VickyXAI added a commit that referenced this pull request Sep 30, 2026
… live free lineup (#73)

blockrun delisted nvidia/nemotron-3-nano-30b on 2026-09-08 (NVIDIA
deprovisioned it for the account). The gateway redirects the id to nano-omni.

Code
- router_adapter.FREE_TIERS: nano-30b was SIMPLE's primary and a fallback in
  MEDIUM, COMPLEX and REASONING. The adapter already dropped it at runtime,
  since it is not in the catalog, so SIMPLE had silently been opening on
  lightning. Every slot now goes to nemotron-3-nano-omni-30b-a3b-reasoning:
  the gateway's own redirect target, the same family and size, and the model
  those slots were actually served by. The 2026-08-31 reason for excluding it
  (it answered as nano-30b) is gone with nano-30b.
- router_core: re-synced to router-core e38958b (BlockRunAI/router-core#3).
  eco SIMPLE fallback[0] goes nano-30b -> nano-omni; nano-30b is dropped from
  model_capabilities; nano-omni gets the supportsVision:false override. The
  decision fixture is a verbatim copy of upstream's regenerated snapshot, and
  the parity test passes on it.
- New test: FREE_TIERS may only name ids from the $0 chat set /v1/models
  published on 2026-09-29, and never nano-30b. The existing membership check
  compared the table against a hand-kept list and could not catch this.

Docs
- README "Available free models", the `free` routing-profile row, the hero
  line, both code examples, the NVIDIA section and the FAQ count now name the
  six live free models instead of the 08-12 lineup (step-3.7-flash,
  mistral-nemotron, nemotron-nano v2, gpt-oss). gpt-oss-120b/20b were retired
  2026-09-03, so their "direct calls still work" rows and privacy note are
  gone. Counts use the models.free marker.
- client.py streaming docstring: example ids moved off the retired
  deepseek-v4-flash / llama-4-maverick.

Co-authored-by: 1bcMax <viewitter@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant