Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 5 additions & 4 deletions docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@
"get-started",
"get-started/connect-to-runpod",
"get-started/concepts",
"get-started/credentials",
"get-started/credentials",
"get-started/agent-skills",
"get-started/mcp-servers",
"get-started/early-access"
Expand Down Expand Up @@ -238,7 +238,7 @@
}
]
},
{
{
"group": "Storage",
"pages": [
"storage/network-volumes",
Expand All @@ -248,7 +248,7 @@
"group": "Global volumes (Beta)",
"pages": [
"storage/globalvolume/overview",
"storage/globalvolume/globalvolume-pods"
"storage/globalvolume/globalvolume-pods"
]
}
]
Expand Down Expand Up @@ -724,7 +724,8 @@
"pages": [
"public-endpoints/models/granite-4",
"public-endpoints/models/moonshot-kimi",
"public-endpoints/models/qwen3-32b"
"public-endpoints/models/qwen3-32b",
"public-endpoints/models/microsoft-frognano-4b-2609"
]
},
{
Expand Down
43 changes: 43 additions & 0 deletions public-endpoints/models/microsoft-frognano-4b-2609.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
---
title: "FrogNano-4B-2609"
sidebarTitle: "FrogNano-4B-2609"
description: "Deploy Microsoft FrogNano-4B-2609 on Runpod Serverless with vLLM, fp8 precision, and an 8,192 token max model length."
---

## Overview

FrogNano-4B-2609 is a 4B parameter text model from Microsoft, published as `microsoft/frognano-4b-2609`. This recipe serves it on Runpod Serverless with vLLM on a single GPU, so you can run long multi-turn, tool-calling agent workloads yourself instead of paying per token.

## Recipe

| Setting | Value |
| --- | --- |
| Engine | vllm |
| Image | `runpod/worker-v1-vllm:v2.27.0` |
| GPU | 1x RTX 4090 |
| Precision | fp8 |
| Max model length | 8,192 |

## Deploy

Deploy the template from the console: [Deploy FrogNano-4B-2609](https://console.runpod.io/deploy?template=902q8tlzie&utm_source=docs&utm_medium=content&utm_campaign=202610_activation_indie-ml-dev_microsoft-frognano-4b-2609&utm_content=model-page)

```bash
runpodctl serverless create --hub-id <vllm listing> --model-reference hf://microsoft/FrogNano-4B-2609
```

Browse related listings in the [Hub](https://console.runpod.io/hub?utm_source=docs&utm_medium=content&utm_campaign=202610_activation_indie-ml-dev_microsoft-frognano-4b-2609&utm_content=model-page), or [create an account](https://console.runpod.io/signup?utm_source=docs&utm_medium=content&utm_campaign=202610_activation_indie-ml-dev_microsoft-frognano-4b-2609&utm_content=model-page) first.

## Measured performance

| GPU | Price | Throughput | Time to first token | Cold start | Cost per 1M output tokens |
| --- | --- | --- | --- | --- | --- |
| 1x RTX 4090, fp8 | $0.69/hr | 104 tokens per second | 582 ms | 870 s | $1.84 |

## Limits

- License: mit. Review the license terms before you deploy.
- Context length for the model is not stated. This recipe sets a max model length of 8,192, and requests above that limit are rejected.
- The recipe covers text only. It does not configure a network volume, quantization other than fp8, a custom handler, or multi-GPU serving.

See the [model page](https://www.runpod.io/models/microsoft-frognano-4b-2609?utm_source=docs&utm_medium=content&utm_campaign=202610_activation_indie-ml-dev_microsoft-frognano-4b-2609&utm_content=model-page) and the [deployment walkthrough](https://www.runpod.io/blog/microsoft-frognano-4b-2609-on-runpod-serverless?utm_source=docs&utm_medium=content&utm_campaign=202610_activation_indie-ml-dev_microsoft-frognano-4b-2609&utm_content=model-page).
Loading