Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 8 additions & 4 deletions src/content/docs-lite/en/features.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ A request opens into its timeline, its routing (the rule it matched, the group i

## Client setup

The Clients page points Claude Code, Claude Desktop, Codex, opencode, Pi, oh-my-pi, Grok Build, Qwen Code, Hermes Agent, Zed, Aider and DeepSeek Harness at the gateway. Before anything is written, it lists the fields that change and what else the change affects (the ChatGPT desktop app, for instance, reads the same configuration file as Codex), shows the full diff and backs up the original file. Only the settings that point the client at the gateway change, and each client receives a key of its own. Claude Desktop is connected through its official third-party inference mode, and the page lists each of the files that change for it; a Claude Desktop managed by an organization is left as it is. A connected client can be restored at any time, on its own or together with all the others; a restored Codex keeps a plain OpenAI entry in place of the gateway's, so sessions started while it was connected can still be opened. opencode (v1 and v2), Pi, oh-my-pi, Grok Build and Qwen Code also get the list of models their key can use on the gateway; when that list changes, the page offers to update it, through the same diff. DeepSeek Harness is covered in its web app, its desktop app and headless runs, which all read the same configuration; models used through a DeepSeek account signed in to the desktop app still go to DeepSeek directly. Cursor, Continue and Antigravity CLI come with step-by-step instructions and a key created for them. For every client the page shows whether it is in use, waiting for its first request or not in effect, and its requests over the last 24 hours.
The Clients page points Claude Code, Claude Desktop, Codex, opencode, Pi, oh-my-pi, Grok Build, Qwen Code, Hermes Agent, Zed, Aider and DeepSeek Harness at the gateway. Before anything is written, it lists the fields that change and what else the change affects (the ChatGPT desktop app, for instance, reads the same configuration file as Codex), shows the full diff and backs up the original file. Only the settings that point the client at the gateway change, and each client receives a key of its own. Claude Desktop is connected through its official third-party inference mode, and the page lists each of the files that change for it; a Claude Desktop managed by an organization is left as it is. Claude Desktop accepts only Claude model names; when no upstream offers a Claude model, the takeover asks which model it should use and adds a routing rule for its key, which cancelling the takeover removes. A connected client can be restored at any time, on its own or together with all the others; a restored Codex keeps a plain OpenAI entry in place of the gateway's, so sessions started while it was connected can still be opened. opencode (v1 and v2), Pi, oh-my-pi, Grok Build and Qwen Code also get the list of models their key can use on the gateway; when that list changes, the page offers to update it, through the same diff. DeepSeek Harness is covered in its web app, its desktop app and headless runs, which all read the same configuration; models used through a DeepSeek account signed in to the desktop app still go to DeepSeek directly. Cursor, Continue and Antigravity CLI come with step-by-step instructions and a key created for them. For every client the page shows whether it is in use, waiting for its first request or not in effect, and its requests over the last 24 hours.

On Windows, Claude Code and Codex installed inside WSL appear in a group of their own for each distribution, next to the clients on the computer itself. They are pointed at the gateway on Windows, restored and diagnosed the same way, each with a key separate from the Windows copy, and their files are edited through `\\wsl.localhost`. They are given `127.0.0.1`, the same address as the clients on Windows, which WSL reaches in two setups:

Expand All @@ -35,7 +35,9 @@ Clients reach the gateway with a key, on the local machine as well. The Keys pag

## Upstreams

Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; proxies and price sheets have tabs on the same page. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs.
Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; aliases, proxies and price sheets have tabs on the same page. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs.

The Aliases tab gives a model the name clients use for it. An alias lists the names the same model has on different upstreams, such as `claude-sonnet-5` on Anthropic and `us.anthropic.claude-sonnet-5-v1:0` on Bedrock; every upstream that offers one of them serves the alias under its own name, and they back each other up. Clients see aliases in their model lists, and answers carry the name the client asked for, while the request log shows the model each upstream was sent. When the same Claude model has different names on the official API, Bedrock, Vertex or OpenRouter, the tab and the new-alias dialog suggest grouping them. An alias can also be created from an upstream's model list. The dialog shows which upstream receives which name and warns when the name would take over another upstream's model of the same name. A key whose model scope allows an upstream model can also use the aliases that list it.

The Check-up tab compares the upstreams over the last 24 hours, 7 days, 30 days or a custom range: requests and failure rate; whether the model named in each answer matches the one sent; the input tokens each upstream reports, as a multiple of the gateway's own estimate, against other upstreams serving the same model; the share of input read from the prompt cache in follow-up turns, also against other upstreams; and the median time to first token and generation speed. Each figure comes with its sample size, and a deviation is marked only when both sides have enough samples.

Expand All @@ -45,7 +47,7 @@ A relay or vendor can hand out an import link, `thinkwatch://import?…` or its

## Routing and failover

Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream or a group, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. A map at the top of the page traces every key through its route and groups to the upstreams.
Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream, a group or specified models, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Specified models name an upstream and one of its models, with backups tried in order, and send the model as written; to have one client use one model under another model's name, a rule on that client's key does it without changing what the name means for other clients. A model condition written for an upstream model also matches the aliases that list it. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. A map at the top of the page traces every key through its route and groups to the upstreams.

Auxiliary requests that clients send on their own (health checks, warm-ups, titles, topic detection and input suggestions) can be answered locally at no cost, or forwarded. Forwarded ones go through the routing rules like any other request, and a rule can send a given kind to a lower-cost upstream.

Expand All @@ -56,11 +58,13 @@ Every request records the rule it matched, the group it went through and each at
The Security page holds three protections: outbound redaction, tool-call inspection and the content filter. They apply to every upstream and every key alike, and each runs in one of three modes: Off, Observe (detect and record, change nothing) and a third mode named for what it does: Replace for outbound redaction, Cut off for tool-call inspection and Enforce for the content filter. All three start in Observe, so out of the box no request is changed or refused.

- **Outbound redaction** searches the whole request before it leaves, including the system prompt, earlier turns and tool calls, for credentials and personal information: API keys and tokens for Anthropic, OpenAI, GitHub, Slack, AWS, Google, GitLab, Stripe, npm, DigitalOcean and SendGrid, private keys, JWTs and passwords in connection strings, as well as Chinese resident ID numbers and bank card numbers, which count only when their structure and check digit are valid. In Replace mode they are replaced with placeholders such as `<<TW_SECRET_1>>` and restored where the response repeats them. Rules for email addresses, Chinese mainland mobile numbers, internal IP addresses and internal domains are included and start off, and a custom rule can name its own placeholder: `PROJECT` gives `<<TW_PROJECT_1>>`.
- **Tool-call inspection** checks the tool calls a model returns for commands that download or decode code and run it, send out environment variables or credential files, send a credential to a host that is neither local nor the credential's own provider, read private keys or cloud credentials, or install startup items and scheduled jobs. In Cut off mode such a call cuts the response off, so the client never receives a complete call to run. Deleting the home or root directory, making files world-writable and uploading a local file to an outside host are only recorded by default.
- **Tool-call inspection** checks the tool calls a model returns for commands that download or decode code and run it, send out environment variables or credential files, send a credential to a host that is neither local nor the credential's own provider, read private keys or cloud credentials, read or change ThinkWatch's own data directory, or install startup items and scheduled jobs. A path that only appears in text being written, such as a document that mentions the directory, does not count. In Cut off mode such a call cuts the response off, so the client never receives a complete call to run. Deleting the home or root directory, making files world-writable and uploading a local file to an outside host are only recorded by default.
- **Content filter** checks the user messages and tool results in each request, context compaction included. A rule matches a keyword, a regular expression or code points such as `U+E0000–U+E007F`, and in Enforce mode it refuses the request, deletes the matched text and sends the rest, or only records. Built-in rules delete Unicode tag characters and bidirectional controls, which can hide instructions from people but not from a model, and refuse explicit "ignore previous instructions" phrasing. Rules for zero-width and private-use characters, which emoji, Persian text and icon fonts also use, and for jailbreaks, persona manipulation, prompt extraction and their Chinese counterparts start off and can be switched on.

The page lists every rule, the built-in ones grouped by kind. Built-in rules can be switched on or off one at a time; a built-in tool-call rule can be set to cut off or only record, and a built-in content rule to refuse, delete or only record. Custom rules are regular expressions, or for the content filter also keywords or code points. Any rule can be tried on a sample text first; for redaction and the content filter, the test also shows the text as it would be sent. Everything the protections find is kept in the log on the first tab, together with the request it came from; when hidden characters spell out text, the log shows that text.

The protections act on requests and answers passing through the gateway, not on files on disk. The configuration file holds keys in plain text and only its owner can read it, which keeps out other users but not programs running as the same user; outbound redaction protects what leaves the machine, not the file.

## MCP

The MCP page covers what clients load from their own configuration files, which does not pass through the gateway.
Expand Down
Loading
Loading