diff --git a/README.md b/README.md index c55f5fe..c43969a 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # BlockRun LLM SDK (Python) -> **blockrun-llm** is a Python SDK for accessing 82 large language models (GPT-5.x, Claude 4.x, Gemini 3.x, DeepSeek, Grok 4.x, GLM, MiniMax, Moonshot and more) plus image / video / music generation, Grok Live Search, prediction-market data (Predexon), Exa neural web search, and Pyth-backed market data. Every call is paid per request — no subscription, no seats, no minimum. Built for AI agents that need to operate autonomously. +> **blockrun-llm** is a Python SDK for accessing 82 large language models (GPT-6, GPT-5.x, Claude 5.x, Gemini 3.x, DeepSeek, Grok 4.x, GLM, MiniMax, Moonshot and more) plus image / video / music generation, Grok Live Search, prediction-market data (Predexon), Exa neural web search, and Pyth-backed market data. Every call is paid per request — no subscription, no seats, no minimum. Built for AI agents that need to operate autonomously. > > **Two ways to pay, same SDK, same catalogue.** Sign up at > **[user.blockrun.ai](https://user.blockrun.ai)** for an API key and prepaid @@ -412,127 +412,164 @@ One call, no separate pay step. ## Available Models -### OpenAI GPT-5.5 Family -Released 2026-04-23 — first fully retrained base since GPT-4.5. 1M context, 128K output, native agent + computer use. - -| Model | Input Price | Output Price | Context | -|-------|-------------|--------------|---------| -| `openai/gpt-5.5` | $5.00/M | $30.00/M | 1M | - -### OpenAI GPT-5.4 Family -| Model | Input Price | Output Price | Context | -|-------|-------------|--------------|---------| -| `openai/gpt-5.4` | $2.50/M | $15.00/M | 1M | -| `openai/gpt-5.4-pro` | $30.00/M | $180.00/M | 1M | -| `openai/gpt-5.4-mini` | $0.75/M | $4.50/M | 400K | -| `openai/gpt-5.4-nano` | $0.20/M | $1.25/M | 1M | - -### OpenAI GPT-5 Family -| Model | Input Price | Output Price | Context | -|-------|-------------|--------------|---------| -| `openai/gpt-5.3` | $1.75/M | $14.00/M | 128K | -| `openai/gpt-5.2` | $1.75/M | $14.00/M | 400K | -| `openai/gpt-5-mini` | $0.25/M | $2.00/M | 200K | -| `openai/gpt-5.2-pro` | $21.00/M | $168.00/M | 400K | -| `openai/gpt-5.3-codex` | $1.75/M | $14.00/M | 400K | - -### OpenAI GPT-4o Family -| Model | Input Price | Output Price | Context | -|-------|-------------|--------------|---------| -| `openai/gpt-4o` | $2.50/M | $10.00/M | 128K | -| `openai/gpt-4o-mini` | $0.15/M | $0.60/M | 128K | +**Prices are not listed here.** They change often, and a number copied into a +README is wrong the day after it lands. See **[blockrun.ai/models](https://blockrun.ai/models)** +for live rates, or read them from the catalog at runtime — `client.list_models()` +and `client.list_image_models()` return exactly what the gateway is charging. + +### OpenAI GPT-6 Family + +The GPT-6 generation: Astra is the flagship for long-horizon agentic work and +computer use, Sol the cost-efficient tier for complex coding, Luna the fast +low-cost tier. + +| Model | Context | +|---|---| +| `openai/gpt-6-astra` | 1.05M | +| `openai/gpt-6-sol` | 1.05M | +| `openai/gpt-6-luna` | 1.05M | + +### OpenAI GPT-5.6 Family + +Three tiers on one 1.05M-context base — Sol (deepest reasoning), Terra +(balanced), Luna (cheap and fast). Each has a `-pro` sibling that thinks +longer at the same token price. + +| Model | Context | +|---|---| +| `openai/gpt-5.6-sol` | 1.05M | +| `openai/gpt-5.6-sol-pro` | 1.05M | +| `openai/gpt-5.6-terra` | 1.05M | +| `openai/gpt-5.6-terra-pro` | 1.05M | +| `openai/gpt-5.6-luna` | 1.05M | +| `openai/gpt-5.6-luna-pro` | 1.05M | + +### OpenAI GPT-5.5 / 5.4 / 5.2 / 5.1 Families + +| Model | Context | Notes | +|---|---|---| +| `openai/gpt-5.5` | 1.05M | | +| `openai/gpt-5.5-pro` | 1.05M | | +| `openai/chat-latest` | 128K | ChatGPT Instant — the model behind chatgpt.com | +| `openai/gpt-5.4` | 1.05M | | +| `openai/gpt-5.4-pro` | 1.05M | | +| `openai/gpt-5.4-mini` | 400K | | +| `openai/gpt-5.4-nano` | 1.05M | | +| `openai/gpt-5.2` | 400K | | +| `openai/gpt-5.2-pro` | 400K | | +| `openai/gpt-5.1` | 400K | Configurable reasoning effort | +| `openai/gpt-5.3-codex` | 400K | Coding/agentic SKU | +| `openai/gpt-5-mini` | 200K | | + +### OpenAI GPT-4 Family + +| Model | Context | +|---|---| +| `openai/gpt-4.1` | 128K | +| `openai/gpt-4.1-mini` | 128K | +| `openai/gpt-4.1-nano` | 128K | +| `openai/gpt-4o` | 128K | +| `openai/gpt-4o-mini` | 128K | ### OpenAI O-Series (Reasoning) -| Model | Input Price | Output Price | Context | -|-------|-------------|--------------|---------| -| `openai/o1` | $15.00/M | $60.00/M | 200K | -| `openai/o3` | $2.00/M | $8.00/M | 200K | -| `openai/o3-mini` | $1.10/M | $4.40/M | 128K | + +| Model | Context | +|---|---| +| `openai/o1` | 200K | +| `openai/o3` | 200K | +| `openai/o3-mini` | 128K | +| `openai/o4-mini` | 128K | ### Anthropic Claude -| Model | Input Price | Output Price | Context | Notes | -|-------|-------------|--------------|---------|-------| -| `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | 1M | Most capable Claude — agentic coding + adaptive thinking, 128K output | -| `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | 1M | Agentic coding + adaptive thinking, 128K output | -| `anthropic/claude-opus-4.6` | $5.00/M | $25.00/M | 200K | Hidden from `/v1/models` (superseded by 4.7); direct calls still work | -| `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | | -| `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 200K | | -| `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | | + +| Model | Context | Notes | +|---|---|---| +| `anthropic/claude-fable-5.1` | 1M | Most capable — successor to Fable 5 at the same tier and price | +| `anthropic/claude-fable-5` | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output | +| `anthropic/claude-opus-5.5` | 1M | Newest Opus — Opus-class reasoning at a lower price than Opus 5 | +| `anthropic/claude-opus-5` | 1M | Flagship — the baseline the routing savings claim is measured against | +| `anthropic/claude-opus-4.8` | 1M | Agentic coding + adaptive thinking, 128K output | +| `anthropic/claude-opus-4.7` | 1M | | +| `anthropic/claude-opus-4.5` | 200K | | +| `anthropic/claude-sonnet-5.5` | 1M | Newest Sonnet — everyday coding and agent work, 128K output | +| `anthropic/claude-sonnet-5` | 1M | Best cost/quality balance for long-context agent turns | +| `anthropic/claude-sonnet-4.6` | 1M | | +| `anthropic/claude-sonnet-4.5` | 200K | | +| `anthropic/claude-haiku-4.5` | 200K | | ### Google Gemini -| Model | Input Price | Output Price | Context | -|-------|-------------|--------------|---------| -| `google/gemini-3.1-pro` | $2.00/M | $12.00/M | 1M | -| `google/gemini-3.5-flash` | $0.50/M | $3.00/M | 1M | -| `google/gemini-3-flash-preview` | $0.50/M | $3.00/M | 1M | -| `google/gemini-2.5-pro` | $1.25/M | $10.00/M | 1M | -| `google/gemini-2.5-flash` | $0.30/M | $2.50/M | 1M | -| `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M | 1M | -| `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M | 1M | -### DeepSeek +| Model | Context | +|---|---| +| `google/gemini-3.1-pro` | 1M | +| `google/gemini-3.8-flash` | 1M | +| `google/gemini-3.6-flash` | 1M | +| `google/gemini-3.5-flash` | 1M | +| `google/gemini-3-flash-preview` | 1M | +| `google/gemini-3.5-flash-lite` | 1M | +| `google/gemini-3.1-flash-lite` | 1M | +| `google/gemini-2.5-pro` | 1M | +| `google/gemini-2.5-flash` | 1M | +| `google/gemini-2.5-flash-lite` | 1M | -V4 family launched 2026-04-24. DeepSeek upstream now serves the legacy -`deepseek-chat` / `deepseek-reasoner` aliases as V4 Flash non-thinking / -thinking modes. V4 Pro is the new flagship paid SKU — 1.6T MoE / 49B active, -1M context, MMLU-Pro 87.5, GPQA 90.1, SWE-bench 80.6, LiveCodeBench 93.5. +### DeepSeek -| Model | Input Price | Output Price | Context | Notes | -|-------|-------------|--------------|---------|-------| -| `deepseek/deepseek-v4-pro` | $0.435/M | $0.87/M | 1M | V4 flagship — strongest open-weight reasoner. The 75% launch promo became the permanent list price after 2026-05-31 | -| `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking (paid endpoint with 5MB request bodies) | -| `deepseek/deepseek-reasoner` | $0.20/M | $0.40/M | 1M | V4 Flash thinking (same upstream as `deepseek-chat`, thinking enabled by default) | +DeepSeek upstream serves the legacy `deepseek-chat` / `deepseek-reasoner` +aliases as V4 Flash non-thinking / thinking modes. V4 Pro is the flagship +paid SKU; the vision SKU is an experimental preview. -### MiniMax -| Model | Input Price | Output Price | Context | Notes | -|-------|-------------|--------------|---------|-------| -| `minimax/minimax-m3` | $0.30/M | $1.20/M | 1M | M3 flagship — strong reasoning + coding, 1M context | -| `minimax/minimax-m2.7` | $0.30/M | $1.20/M | 200K | | +| Model | Context | Notes | +|---|---|---| +| `deepseek/deepseek-v4-pro` | 1M | V4 flagship — strongest open-weight reasoner | +| `deepseek/deepseek-v4-flash-vision-exp` | 1M | Experimental vision preview | +| `deepseek/deepseek-chat` | 1M | V4 Flash non-thinking | +| `deepseek/deepseek-reasoner` | 1M | V4 Flash thinking (same upstream, thinking on by default) | ### xAI Grok -Grok 4.3 and Grok Build are resold through BlockRun's OpenRouter credit pool -(same pattern as `deepseek/deepseek-v4-pro` and `minimax/minimax-m3`). Older -Grok chat SKUs (grok-3/4/4.1-fast families) are hidden from `/v1/models` but -direct calls by full ID still work. - -| Model | Input Price | Output Price | Context | Notes | -|-------|-------------|--------------|---------|-------| -| `xai/grok-4.3` | $1.25/M | $2.50/M | 1M | Reasoning model, vision-capable, tuned for agentic workflows | -| `xai/grok-build-0.1` | $1.00/M | $2.00/M | 256K | Fast agentic coding model — interactive software-engineering workflows | - -### ZAI - -The GLM flat-rate launch promos have fully ended (glm-5.1 on 2026-06-05; -glm-5 and glm-5-turbo on 2026-06-06) — the whole family now bills per-token. - -| Model | Input Price | Output Price | Context | Notes | -|-------|-------------|--------------|---------|-------| -| `zai/glm-5.2` | $1.40/M | $4.40/M | 1M | Z.AI's newest flagship — 1M-token context, top open-source on long-horizon coding | -| `zai/glm-5.1` | $1.40/M | $4.40/M | 200K | #1 open-source on SWE-Bench Pro, 8-hour autonomous execution | -| `zai/glm-5` | $0.60/M | $1.92/M | 200K | | -| `zai/glm-5-turbo` | $1.20/M | $4.00/M | 200K | | - -### NVIDIA (Free & Hosted) - -Free tier checked against `/v1/models` on 2026-09-29. NVIDIA has retired -(HTTP 410 end-of-life, or deprovisioned) most of its earlier free lineup — the -free DeepSeek family, `step-3.7-flash`, the `nemotron-nano` v2 SKUs, -`gpt-oss-120b/20b` (2026-09-03) and `nemotron-3-nano-30b` (2026-09-08) among -them. Retired models stay callable by ID: the gateway auto-redirects them to a -live model, so pinned callers still get a 200. The free tier also includes two -non-NVIDIA models, `cohere/north-mini-code` and `poolside/laguna-xs-2.1` (see -[Try It Free](#try-it-free-no-balance-required)). The live list is -`GET /v1/models` filtered on $0 pricing. - -| Model | Input Price | Output Price | Context | Notes | -|-------|-------------|--------------|---------|-------| -| `nvidia/nemotron-3.5-lightning` | **FREE** | **FREE** | 1M | Free default — thinking-mode reasoning | -| `nvidia/nemotron-3-ultra-550b` | **FREE** | **FREE** | 1M | Largest free model (550B / 55B active MoE) | -| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 256K | Fast general chat + reasoning | -| `nvidia/llama-3.2-11b-vision` | **FREE** | **FREE** | 128K | Meta Llama 3.2 11B | -| `moonshot/kimi-k2.5` | $0.60/M | $3.00/M | 262K | Kimi K2.5 direct from Moonshot (replaces `nvidia/kimi-k2.5`) | -| `moonshot/kimi-k2.6` | $0.95/M | $4.00/M | 256K | Moonshot flagship (vision + reasoning_content) | +The older Grok chat SKUs (grok-3/3-mini, the grok-4 fast families, +grok-code-fast-1, grok-2-vision) have left the catalog. Retired ids stay +callable — the gateway redirects them to a healthy model — but `smart_chat` +only ranks what `/v1/models` lists. + +| Model | Context | Notes | +|---|---|---| +| `xai/grok-4.7` | 500K | Flagship — reasoning + vision, selectable effort (low → xhigh), native Live Search (`search: true`) | +| `xai/grok-4.6` | 500K | Reasoning with selectable effort, native Live Search | +| `xai/grok-4.5` | 500K | Reasoning + vision, native Live Search | +| `xai/grok-4.3` | 1M | Reasoning + vision, tuned for agentic workflows | +| `xai/grok-build-0.1` | 256K | Fast agentic coding model | + +### Moonshot, MiniMax, Z.ai, Qwen + +| Model | Context | Notes | +|---|---|---| +| `moonshot/kimi-k3` | 1M | Replaces the retired `kimi-k2.5` / `k2.6` SKUs | +| `minimax/minimax-m3` | 1M | | +| `minimax/minimax-m2.7` | 200K | | +| `zai/glm-5.3` | 1M | | +| `zai/glm-5.3-flash` | 1M | Cheapest vision-capable paid SKU | +| `zai/glm-5.2` | 1M | | +| `zai/glm-5.1` | 200K | | +| `zai/glm-5` | 200K | | +| `zai/glm-5-turbo` | 200K | | +| `qwen/qwen3.7-max` | 1M | | +| `qwen/qwen3.7-plus` | 1M | | +| `qwen/qwen3.8-flash` | 1M | | +| `qwen/qwen3.7-flash` | 1M | Cheapest paid chat model in the catalog | + +### Xiaomi + +| Model | Context | +|---|---| +| `xiaomi/mimo-v2.5` | 1M | +| `xiaomi/mimo-v2.5-pro` | 1M | + +### Free Tier + +The 6 free models are listed under +[Try It Free](#try-it-free-no-balance-required). Pin them by full model id, or +use `routing_profile="free"`. ### Testnet Models (Base Sepolia) | Model | Price | diff --git a/blockrun_llm/router_core/model_capabilities.py b/blockrun_llm/router_core/model_capabilities.py index 6dcc6d2..dd8d97d 100644 --- a/blockrun_llm/router_core/model_capabilities.py +++ b/blockrun_llm/router_core/model_capabilities.py @@ -8,7 +8,7 @@ product catalog is temporarily unavailable, without importing product code. GENERATED upstream by ``scripts/sync-model-capabilities.mjs`` from the public -catalog (GET https://blockrun.ai/api/v1/models) on 2026-08-31; ``supports_tools`` +catalog (GET https://blockrun.ai/api/v1/models) on 2026-10-02; ``supports_tools`` comes from a live function-calling probe. Re-sync from ``model-capabilities.ts`` rather than editing by hand — a hand edit is lost on the next sync. """ @@ -28,6 +28,13 @@ "supports_tools": True, "supports_vision": True, }, + # supportsTools: not probed — fails closed + "anthropic/claude-fable-5.1": { + "context_window": 1_000_000, + "max_output_tokens": 128_000, + "supports_tools": False, + "supports_vision": True, + }, # override: The public catalog's `categories` omit "vision" for this Anthropic model even # though the gateway accepts image input for it (the prior hand-maintained snapshot had # it, and Anthropic's model card lists it). Without this the vision filter would silently @@ -62,6 +69,13 @@ "supports_tools": True, "supports_vision": True, }, + # supportsTools: not probed — fails closed + "anthropic/claude-opus-5.5": { + "context_window": 1_000_000, + "max_output_tokens": 128_000, + "supports_tools": False, + "supports_vision": True, + }, "anthropic/claude-sonnet-4.5": { "context_window": 200_000, "max_output_tokens": 64_000, @@ -85,6 +99,13 @@ "supports_vision": True, }, # supportsTools: not probed — fails closed + "anthropic/claude-sonnet-5.5": { + "context_window": 1_000_000, + "max_output_tokens": 128_000, + "supports_tools": False, + "supports_vision": True, + }, + # supportsTools: not probed — fails closed "cohere/north-mini-code": { "context_window": 256_000, "max_output_tokens": 16_384, @@ -103,6 +124,13 @@ "supports_tools": True, "supports_vision": False, }, + # supportsTools: not probed — fails closed + "deepseek/deepseek-v4-flash-vision-exp": { + "context_window": 1_048_576, + "max_output_tokens": 65_536, + "supports_tools": False, + "supports_vision": True, + }, "deepseek/deepseek-v4-pro": { "context_window": 1_048_576, "max_output_tokens": 65_536, @@ -163,6 +191,13 @@ "supports_tools": True, "supports_vision": True, }, + # supportsTools: not probed — fails closed + "google/gemini-3.8-flash": { + "context_window": 1_048_576, + "max_output_tokens": 65_536, + "supports_tools": False, + "supports_vision": True, + }, "minimax/minimax-m2.7": { "context_window": 204_800, "max_output_tokens": 16_384, @@ -257,6 +292,13 @@ "supports_tools": True, "supports_vision": False, }, + # supportsTools: not probed — fails closed + "openai/gpt-5.1": { + "context_window": 400_000, + "max_output_tokens": 128_000, + "supports_tools": False, + "supports_vision": True, + }, "openai/gpt-5.2": { "context_window": 400_000, "max_output_tokens": 128_000, @@ -356,6 +398,27 @@ "supports_tools": True, "supports_vision": True, }, + # supportsTools: not probed — fails closed + "openai/gpt-6-astra": { + "context_window": 1_050_000, + "max_output_tokens": 128_000, + "supports_tools": False, + "supports_vision": True, + }, + # supportsTools: not probed — fails closed + "openai/gpt-6-luna": { + "context_window": 1_050_000, + "max_output_tokens": 128_000, + "supports_tools": False, + "supports_vision": True, + }, + # supportsTools: not probed — fails closed + "openai/gpt-6-sol": { + "context_window": 1_050_000, + "max_output_tokens": 128_000, + "supports_tools": False, + "supports_vision": True, + }, "openai/o1": { "context_window": 200_000, "max_output_tokens": 100_000, @@ -405,11 +468,12 @@ "supports_tools": True, "supports_vision": False, }, - "tencent/hy3": { - "context_window": 262_144, - "max_output_tokens": 128_000, - "supports_tools": True, - "supports_vision": False, + # supportsTools: not probed — fails closed + "qwen/qwen3.8-flash": { + "context_window": 1_000_000, + "max_output_tokens": 131_072, + "supports_tools": False, + "supports_vision": True, }, "xai/grok-4.3": { "context_window": 1_000_000, @@ -423,12 +487,33 @@ "supports_tools": True, "supports_vision": True, }, + # supportsTools: not probed — fails closed + "xai/grok-4.6": { + "context_window": 500_000, + "max_output_tokens": 16_384, + "supports_tools": False, + "supports_vision": True, + }, + # supportsTools: not probed — fails closed + "xai/grok-4.7": { + "context_window": 500_000, + "max_output_tokens": 16_384, + "supports_tools": False, + "supports_vision": True, + }, "xai/grok-build-0.1": { "context_window": 256_000, "max_output_tokens": 16_384, "supports_tools": True, "supports_vision": False, }, + # supportsTools: not probed — fails closed + "xiaomi/mimo-v2.5": { + "context_window": 1_048_576, + "max_output_tokens": 131_072, + "supports_tools": False, + "supports_vision": True, + }, "xiaomi/mimo-v2.5-pro": { "context_window": 1_048_576, "max_output_tokens": 131_072,