A self-hosted AI API and MCP gateway for organizations. Every model request and every MCP tool call passes through one gateway, where it is authenticated against the organization's identity provider, checked against limits and budgets, inspected by security guards, priced, and written to the audit log. It plays the role for AI access that a bastion host plays for server access.
Sponsors: Want to appear here?
┌──────────────────────────────────────┐
Claude Code ──────>│ │──> OpenAI
Cursor ───────────>│ Gateway :3000 │──> Anthropic
Custom Agent ─────>│ AI API + MCP │──> Google Gemini
CI/CD Pipeline ───>│ │──> Azure OpenAI / AWS Bedrock
└──────────────────────────────────────┘
┌──────────────────────────────────────┐
Admin Browser ────>│ Console :3001 │
│ Management UI + Admin API │
└──────────────────────────────────────┘
- MCP tool calls run as the real user. Each user connects their own GitHub, Notion, Linear, Slack or Atlassian account through OAuth or a personal token, so the upstream's own audit log shows who acted. Tokens are encrypted at rest, tool lists are cached per user, and each tool can be granted per role and per API key.
- Security guards on every request. Outbound redaction replaces credentials and personal data anywhere in a request with placeholders such as
<<TW_EMAIL_1>>before it goes upstream, and restores them in the answer, streamed ones included. Tool-call inspection checks the tool calls in each response for dangerous commands, and the content filter looks for prompt-injection phrases and hidden characters in what the caller sent, then refuses the request, deletes them or records them. - Identity from the organization's directory. Sign-in works through any OIDC provider (Zitadel, Okta, Azure AD and others), with optional TOTP. Five built-in roles, from Super Admin to Viewer, and custom roles decide who may use which models, tools and admin pages.
- One key for AI and MCP. Users receive
tw-virtual keys that can be scoped to the AI gateway, the MCP gateway or both. Keys are stored only as hashes and rotate with a grace period. - Rate limits and budgets. Sliding windows from one minute to one week limit requests or tokens, and daily, weekly or monthly budgets cap spending. Both attach to users, API keys or roles, and rate limits apply to MCP tool calls as well as model requests.
- Cost accounting that finance can use. Spend is reported by model, user, provider and cost center, with CSV chargeback reports and a month-end forecast. Per-model weights make expensive models count for more against the same quota.
- Audit trail in ClickHouse. Every model request and tool call is recorded with user, parameters, response, latency and errors, and captured bodies can be redacted with the outbound redaction rules before storage (off by default). Events can be forwarded to a SIEM over Syslog, Kafka (through a REST proxy) or signed webhooks.
- One endpoint for every client. OpenAI Chat Completions, OpenAI Responses, Anthropic Messages and Gemini requests are served on one port and converted to whatever the upstream speaks. Routing spreads traffic by weight, latency or health, and a circuit breaker takes failing upstreams out of rotation.
# 1. Start infrastructure
make infra
# 2. Generate dev secrets + start backend (gateway :3000 + console :3001)
make dev-secrets # writes .env from .env.example with random secrets
make dev-backend
# 3. Start frontend dev server
cd web && pnpm install && pnpm dev
# 4. Complete the setup wizard at http://localhost:5173/setupThe setup wizard creates the first Super Admin account, sets the site name and issues a first API key; providers are added in the console afterwards. The console then has copy-paste setup instructions for Claude Code, Cursor, Continue, Cline, the OpenAI and Anthropic SDKs, and cURL.
| Option | Command | Notes |
|---|---|---|
| Docker Compose | make deploy |
Generates .env.production with random secrets on first run |
| Kubernetes | make helm-deploy |
Helm chart in deploy/helm/think-watch; secrets are generated on install and kept on upgrade |
The gateway (port 3000) is the only part that clients need to reach. The console (port 3001) serves the management UI and admin API and belongs behind a VPN or firewall. See the Deployment Guide for TLS, hardening and production settings.
MCP identity
- A user may connect several accounts to one server (work and personal, for example) and pin each
tw-key to one of them. - Adding a server takes its URL: the gateway discovers the OAuth endpoints and registers itself when the upstream supports dynamic client registration. The MCP Store ships 37 ready-made templates.
- A user who has not yet connected an account still sees the tool list; calling a tool returns JSON-RPC error
-32050with the authorization URL, which compliant MCP clients can show. - Responses from servers that use per-user credentials are cached per user and account, never shared.
Security guards
- There are three guards, each with three modes: off, observe, and one named for what it does — replace (outbound redaction), cut off (tool-call inspection) and enforce (content filter). A new installation starts all three in observe mode: hits go to the audit log and nothing is changed until a guard is switched to its third mode.
- Every rule is listed on the console's security page, built-in and custom. Built-in rules can be switched on or off, tool-call and content rules can take another action, custom rules can be added, and a sample can be tried against one rule or a whole guard first.
- Outbound redaction searches the whole request, system prompt and earlier answers included, but not base64 payloads. A match becomes
<<TW_LABEL_n>>—SECRETfor credentials,ID_NUMBER,CARD_NUMBER,EMAILandPHONEfor personal data, a label of its own for a custom rule — and is restored in the answer. E-mail addresses and phone numbers ship switched off. - A content rule matches a phrase, a regular expression or code points (
U+200B,U+E0000–U+E007F), and either refuses the request, deletes what it matched from the caller's messages and tool results, or records only. Hidden characters are content rules: Unicode tag characters and bidirectional controls ship on, zero-width and private-use characters off. - A tool-call rule cuts the response at the call or records it. Besides the dangerous-command rules, two built-in rules catch a credential sent to an unknown host and a local file uploaded to an external host.
- A model's maximum output tokens, set on the Models page, caps
max_tokenson every request to that model; it replaces the old output length guardrail.
Limits and budgets
- Request-count limits are checked before the request; token limits and budgets are counted after the response, so one request can cross a budget before the next is refused.
- If Redis is unavailable, limits fail open by default. Setting
security.rate_limit_fail_closedrefuses requests instead. - Budget alerts fire once per period at 50%, 80%, 95% and 100%.
Product page: thinkwat.ch/thinkwatch · Full documentation: thinkwat.ch/docs
| Document | Description |
|---|---|
| Architecture | System design, dual-port model, data flow |
| Deployment Guide | Docker Compose, Kubernetes, TLS, production hardening |
| Configuration | Environment variables and settings |
| API Reference | Gateway and console endpoints |
| Security | Auth model, encryption, RBAC, threat model |
| Secret Rotation | Rotating provider keys, JWT secrets and admin credentials |
ThinkWatch uses four crates from ThinkWatch Core (MIT): tw-dialect for conversion between API formats and usage parsing, tw-guard for redaction, tool-call inspection and the other guards, tw-breaker for the circuit-breaker state machine, and tw-bedrock for Amazon Bedrock signing, event streams and the model catalog.
ThinkWatch Lite is the desktop edition for individual developers, a local gateway for Claude Code, Codex and other clients on macOS, Windows and Linux (MIT).
Contributions are welcome. Please open an issue to discuss before submitting a PR for major changes.
ThinkWatch is source-available under the Business Source License 1.1.
Non-production use is free. Production use is free up to both 10,000,000
Billable Tokens and 10,000 MCP Tool Calls per UTC calendar month; above
either threshold, a commercial license is required and priced by usage tiers.
See LICENSING.md for the production-use thresholds, the
Billable Token and MCP Tool Call definitions, the tiering model, and the
changeover to GPL-2.0-or-later.