diff --git a/src/content/docs-lite/en/features.md b/src/content/docs-lite/en/features.md index 0d17000..883bdb9 100644 --- a/src/content/docs-lite/en/features.md +++ b/src/content/docs-lite/en/features.md @@ -20,7 +20,7 @@ A request opens into its timeline, its routing (the rule it matched, the group i ## Client setup -The Clients page points Claude Code, Claude Desktop, Codex, opencode, Pi, oh-my-pi, Grok Build, Qwen Code, Hermes Agent, Zed, Aider and DeepSeek Harness at the gateway. Before anything is written, it lists the fields that change and what else the change affects (the ChatGPT desktop app, for instance, reads the same configuration file as Codex), shows the full diff and backs up the original file. Only the settings that point the client at the gateway change, and each client receives a key of its own. Claude Desktop is connected through its official third-party inference mode, and the page lists each of the files that change for it; a Claude Desktop managed by an organization is left as it is. A connected client can be restored at any time, on its own or together with all the others; a restored Codex keeps a plain OpenAI entry in place of the gateway's, so sessions started while it was connected can still be opened. opencode (v1 and v2), Pi, oh-my-pi, Grok Build and Qwen Code also get the list of models their key can use on the gateway; when that list changes, the page offers to update it, through the same diff. DeepSeek Harness is covered in its web app, its desktop app and headless runs, which all read the same configuration; models used through a DeepSeek account signed in to the desktop app still go to DeepSeek directly. Cursor, Continue and Antigravity CLI come with step-by-step instructions and a key created for them. For every client the page shows whether it is in use, waiting for its first request or not in effect, and its requests over the last 24 hours. +The Clients page points Claude Code, Claude Desktop, Codex, opencode, Pi, oh-my-pi, Grok Build, Qwen Code, Hermes Agent, Zed, Aider and DeepSeek Harness at the gateway. Before anything is written, it lists the fields that change and what else the change affects (the ChatGPT desktop app, for instance, reads the same configuration file as Codex), shows the full diff and backs up the original file. Only the settings that point the client at the gateway change, and each client receives a key of its own. Claude Desktop is connected through its official third-party inference mode, and the page lists each of the files that change for it; a Claude Desktop managed by an organization is left as it is. Claude Desktop accepts only Claude model names; when no upstream offers a Claude model, the takeover asks which model it should use and adds a routing rule for its key, which cancelling the takeover removes. A connected client can be restored at any time, on its own or together with all the others; a restored Codex keeps a plain OpenAI entry in place of the gateway's, so sessions started while it was connected can still be opened. opencode (v1 and v2), Pi, oh-my-pi, Grok Build and Qwen Code also get the list of models their key can use on the gateway; when that list changes, the page offers to update it, through the same diff. DeepSeek Harness is covered in its web app, its desktop app and headless runs, which all read the same configuration; models used through a DeepSeek account signed in to the desktop app still go to DeepSeek directly. Cursor, Continue and Antigravity CLI come with step-by-step instructions and a key created for them. For every client the page shows whether it is in use, waiting for its first request or not in effect, and its requests over the last 24 hours. On Windows, Claude Code and Codex installed inside WSL appear in a group of their own for each distribution, next to the clients on the computer itself. They are pointed at the gateway on Windows, restored and diagnosed the same way, each with a key separate from the Windows copy, and their files are edited through `\\wsl.localhost`. They are given `127.0.0.1`, the same address as the clients on Windows, which WSL reaches in two setups: @@ -35,7 +35,9 @@ Clients reach the gateway with a key, on the local machine as well. The Keys pag ## Upstreams -Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; proxies and price sheets have tabs on the same page. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs. +Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; aliases, proxies and price sheets have tabs on the same page. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs. + +The Aliases tab gives a model the name clients use for it. An alias lists the names the same model has on different upstreams, such as `claude-sonnet-5` on Anthropic and `us.anthropic.claude-sonnet-5-v1:0` on Bedrock; every upstream that offers one of them serves the alias under its own name, and they back each other up. Clients see aliases in their model lists, and answers carry the name the client asked for, while the request log shows the model each upstream was sent. When the same Claude model has different names on the official API, Bedrock, Vertex or OpenRouter, the tab and the new-alias dialog suggest grouping them. An alias can also be created from an upstream's model list. The dialog shows which upstream receives which name and warns when the name would take over another upstream's model of the same name. A key whose model scope allows an upstream model can also use the aliases that list it. The Check-up tab compares the upstreams over the last 24 hours, 7 days, 30 days or a custom range: requests and failure rate; whether the model named in each answer matches the one sent; the input tokens each upstream reports, as a multiple of the gateway's own estimate, against other upstreams serving the same model; the share of input read from the prompt cache in follow-up turns, also against other upstreams; and the median time to first token and generation speed. Each figure comes with its sample size, and a deviation is marked only when both sides have enough samples. @@ -45,7 +47,7 @@ A relay or vendor can hand out an import link, `thinkwatch://import?…` or its ## Routing and failover -Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream or a group, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. A map at the top of the page traces every key through its route and groups to the upstreams. +Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream, a group or specified models, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Specified models name an upstream and one of its models, with backups tried in order, and send the model as written; to have one client use one model under another model's name, a rule on that client's key does it without changing what the name means for other clients. A model condition written for an upstream model also matches the aliases that list it. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. A map at the top of the page traces every key through its route and groups to the upstreams. Auxiliary requests that clients send on their own (health checks, warm-ups, titles, topic detection and input suggestions) can be answered locally at no cost, or forwarded. Forwarded ones go through the routing rules like any other request, and a rule can send a given kind to a lower-cost upstream. @@ -56,11 +58,13 @@ Every request records the rule it matched, the group it went through and each at The Security page holds three protections: outbound redaction, tool-call inspection and the content filter. They apply to every upstream and every key alike, and each runs in one of three modes: Off, Observe (detect and record, change nothing) and a third mode named for what it does: Replace for outbound redaction, Cut off for tool-call inspection and Enforce for the content filter. All three start in Observe, so out of the box no request is changed or refused. - **Outbound redaction** searches the whole request before it leaves, including the system prompt, earlier turns and tool calls, for credentials and personal information: API keys and tokens for Anthropic, OpenAI, GitHub, Slack, AWS, Google, GitLab, Stripe, npm, DigitalOcean and SendGrid, private keys, JWTs and passwords in connection strings, as well as Chinese resident ID numbers and bank card numbers, which count only when their structure and check digit are valid. In Replace mode they are replaced with placeholders such as `<>` and restored where the response repeats them. Rules for email addresses, Chinese mainland mobile numbers, internal IP addresses and internal domains are included and start off, and a custom rule can name its own placeholder: `PROJECT` gives `<>`. -- **Tool-call inspection** checks the tool calls a model returns for commands that download or decode code and run it, send out environment variables or credential files, send a credential to a host that is neither local nor the credential's own provider, read private keys or cloud credentials, or install startup items and scheduled jobs. In Cut off mode such a call cuts the response off, so the client never receives a complete call to run. Deleting the home or root directory, making files world-writable and uploading a local file to an outside host are only recorded by default. +- **Tool-call inspection** checks the tool calls a model returns for commands that download or decode code and run it, send out environment variables or credential files, send a credential to a host that is neither local nor the credential's own provider, read private keys or cloud credentials, read or change ThinkWatch's own data directory, or install startup items and scheduled jobs. A path that only appears in text being written, such as a document that mentions the directory, does not count. In Cut off mode such a call cuts the response off, so the client never receives a complete call to run. Deleting the home or root directory, making files world-writable and uploading a local file to an outside host are only recorded by default. - **Content filter** checks the user messages and tool results in each request, context compaction included. A rule matches a keyword, a regular expression or code points such as `U+E0000–U+E007F`, and in Enforce mode it refuses the request, deletes the matched text and sends the rest, or only records. Built-in rules delete Unicode tag characters and bidirectional controls, which can hide instructions from people but not from a model, and refuse explicit "ignore previous instructions" phrasing. Rules for zero-width and private-use characters, which emoji, Persian text and icon fonts also use, and for jailbreaks, persona manipulation, prompt extraction and their Chinese counterparts start off and can be switched on. The page lists every rule, the built-in ones grouped by kind. Built-in rules can be switched on or off one at a time; a built-in tool-call rule can be set to cut off or only record, and a built-in content rule to refuse, delete or only record. Custom rules are regular expressions, or for the content filter also keywords or code points. Any rule can be tried on a sample text first; for redaction and the content filter, the test also shows the text as it would be sent. Everything the protections find is kept in the log on the first tab, together with the request it came from; when hidden characters spell out text, the log shows that text. +The protections act on requests and answers passing through the gateway, not on files on disk. The configuration file holds keys in plain text and only its owner can read it, which keeps out other users but not programs running as the same user; outbound redaction protects what leaves the machine, not the file. + ## MCP The MCP page covers what clients load from their own configuration files, which does not pass through the gateway. diff --git a/src/content/docs-lite/zh-CN/features.md b/src/content/docs-lite/zh-CN/features.md index f0a4919..5b38f1a 100644 --- a/src/content/docs-lite/zh-CN/features.md +++ b/src/content/docs-lite/zh-CN/features.md @@ -20,7 +20,7 @@ ## 客户端接管 -客户端页可以把 Claude Code、Claude Desktop、Codex、opencode、Pi、oh-my-pi、Grok Build、Qwen Code、Hermes Agent、Zed、Aider 与 DeepSeek Harness 指向网关。写入之前,页面列出将要修改的字段和这次接管的其他影响(例如 ChatGPT 桌面版与 Codex 读取同一份配置文件),给出完整的改动差异,并完整备份原文件。只修改指向网关所需的配置,每个客户端使用各自的密钥。Claude Desktop 通过官方的第三方推理模式接入,页面逐一列出要修改的各个文件;由组织统一管理的 Claude Desktop 不做修改。已接管的客户端可以随时单独还原或全部还原;Codex 还原后保留一项直连 OpenAI 的配置,接管期间的会话仍可打开。opencode(v1 与 v2)、Pi、oh-my-pi、Grok Build 与 Qwen Code 的配置中同时写入其密钥在网关上可用的模型列表;网关上可用的模型变化后,页面提示更新,更新同样先给出改动差异。DeepSeek Harness 的网页版、桌面版与 headless 模式读取同一份配置,都会经过网关;在桌面版中通过 DeepSeek 账号登录使用的模型仍直接连接 DeepSeek。Cursor、Continue 与 Antigravity CLI 提供逐步的配置方法,并为其创建密钥。页面列出每个客户端处于使用中、等待首个请求还是未生效,以及最近 24 小时的请求。 +客户端页可以把 Claude Code、Claude Desktop、Codex、opencode、Pi、oh-my-pi、Grok Build、Qwen Code、Hermes Agent、Zed、Aider 与 DeepSeek Harness 指向网关。写入之前,页面列出将要修改的字段和这次接管的其他影响(例如 ChatGPT 桌面版与 Codex 读取同一份配置文件),给出完整的改动差异,并完整备份原文件。只修改指向网关所需的配置,每个客户端使用各自的密钥。Claude Desktop 通过官方的第三方推理模式接入,页面逐一列出要修改的各个文件;由组织统一管理的 Claude Desktop 不做修改。Claude Desktop 只接受 Claude 的模型名;没有上游提供 Claude 模型时,接管会询问它使用哪个模型,并在它的密钥上添加一条路由规则,取消接管时一并删除。已接管的客户端可以随时单独还原或全部还原;Codex 还原后保留一项直连 OpenAI 的配置,接管期间的会话仍可打开。opencode(v1 与 v2)、Pi、oh-my-pi、Grok Build 与 Qwen Code 的配置中同时写入其密钥在网关上可用的模型列表;网关上可用的模型变化后,页面提示更新,更新同样先给出改动差异。DeepSeek Harness 的网页版、桌面版与 headless 模式读取同一份配置,都会经过网关;在桌面版中通过 DeepSeek 账号登录使用的模型仍直接连接 DeepSeek。Cursor、Continue 与 Antigravity CLI 提供逐步的配置方法,并为其创建密钥。页面列出每个客户端处于使用中、等待首个请求还是未生效,以及最近 24 小时的请求。 在 Windows 上,安装在 WSL 中的 Claude Code 与 Codex 按发行版单独成组,列在这台电脑的客户端之后。它们同样可以指向 Windows 上的网关、还原和检查,使用与 Windows 上那一份分开的密钥,配置文件经由 `\\wsl.localhost` 修改。写入的地址与 Windows 上的客户端相同,是 `127.0.0.1`,WSL 在以下两种情况下可以访问: @@ -35,7 +35,9 @@ WSL 2 默认使用 NAT 网络,此时 Windows 上的网关无法从 WSL 内访 ## 上游 -上游是网关转发请求的目标:Anthropic、OpenAI、Google Gemini、DeepSeek 或任何兼容接口的 API 密钥,Amazon Bedrock(API 密钥、访问密钥或 AWS 配置文件),在应用内登录的 ChatGPT 账号或 Z.ai / BigModel 账号,OpenRouter 等中转服务,以及 Ollama 等本机模型。ChatGPT 账号显示订阅额度与重置时间;GLM Coding Plan 的上游(地址在 `api.z.ai` 或 `open.bigmodel.cn` 上,在应用内登录或手动填写密钥均可)同样显示:5 小时与每周额度,积分制套餐另外显示剩余积分(「剩余 1,976 / 2,000 积分」)。客户端与上游的 API 格式不同时,请求在 Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 与 Gemini 之间自动转换,无法转换的字段会在请求上逐一列出。上游可以经出站代理访问,也可以使用单独的价目表计价,代理与价目表在同一页的标签中管理。链路测速测量 DNS 解析以及 TCP、TLS、代理握手的耗时,不产生费用;推理测速测量首个 token 的时间,运行前先给出费用预估。 +上游是网关转发请求的目标:Anthropic、OpenAI、Google Gemini、DeepSeek 或任何兼容接口的 API 密钥,Amazon Bedrock(API 密钥、访问密钥或 AWS 配置文件),在应用内登录的 ChatGPT 账号或 Z.ai / BigModel 账号,OpenRouter 等中转服务,以及 Ollama 等本机模型。ChatGPT 账号显示订阅额度与重置时间;GLM Coding Plan 的上游(地址在 `api.z.ai` 或 `open.bigmodel.cn` 上,在应用内登录或手动填写密钥均可)同样显示:5 小时与每周额度,积分制套餐另外显示剩余积分(「剩余 1,976 / 2,000 积分」)。客户端与上游的 API 格式不同时,请求在 Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 与 Gemini 之间自动转换,无法转换的字段会在请求上逐一列出。上游可以经出站代理访问,也可以使用单独的价目表计价,别名、代理与价目表在同一页的标签中管理。链路测速测量 DNS 解析以及 TCP、TLS、代理握手的耗时,不产生费用;推理测速测量首个 token 的时间,运行前先给出费用预估。 + +别名标签为模型设定客户端使用的名称。一个别名列出同一个模型在各家上游的名称,例如 Anthropic 上的 `claude-sonnet-5` 和 Bedrock 上的 `us.anthropic.claude-sonnet-5-v1:0`;提供其中任一名称的上游都能以自己的名称服务这个别名,并互为备用。客户端的模型列表里能看到别名,回答里的模型名写成客户端请求的名称,请求记录则保留每家上游实际收到的模型。同一个 Claude 模型在官方 API、Bedrock、Vertex 或 OpenRouter 上名称不同时,别名标签和新建别名对话框会建议合并。也可以在上游的模型列表里直接起别名。对话框列出每家上游将收到的名称,名称会接管另一家上游的同名模型时给出提示。密钥的可见模型允许某个上游模型时,列有它的别名也可以使用。 「体检」标签按最近 24 小时、7 天、30 天或自定义区间对照各个上游:请求数与失败率;回答中的模型名与发出的是否相同;各上游报告的输入 token 相对网关本地估算的倍数,与服务同一模型的其他上游对照;后续轮次中输入从提示缓存读取的比例,同样与其他上游对照;以及首 token 时间与生成速度的中位数。每项数字都附样本数,两边样本都足够时才标出偏差。 @@ -45,7 +47,7 @@ API 密钥和请求头的值可以写成 `${变量名}`,读取系统环境变 ## 路由与故障转移 -每把密钥使用一条路由,未指定的使用默认路由。路由由按顺序匹配的规则组成。规则的条件包括模型、密钥、客户端的 API 格式、输入 token 数、`max_tokens`、工具数量、图片、扩展思考、流式、提示缓存以及辅助请求的类型;命中后把请求交给某个上游或策略组,或拒绝请求,也可以改写模型、`max_tokens` 或扩展思考。设置 `max_tokens` 可以限制回答的长度:上游到达上限时自行停止,不会报错。策略组把多个上游放在同一个名字下,并决定尝试的先后:按顺序、手动选择、轮询、延迟最低优先或费用最低优先。一次尝试失败时,请求转到下一个上游;同一会话默认保持在同一个上游上,以便提示缓存持续命中。页面顶部的路由图显示每把密钥经过的路由、策略组和上游。 +每把密钥使用一条路由,未指定的使用默认路由。路由由按顺序匹配的规则组成。规则的条件包括模型、密钥、客户端的 API 格式、输入 token 数、`max_tokens`、工具数量、图片、扩展思考、流式、提示缓存以及辅助请求的类型;命中后把请求交给某个上游、策略组或指定模型,或拒绝请求,也可以改写模型、`max_tokens` 或扩展思考。指定模型写明上游和它的某个模型,可按顺序列出备用,模型名原样发出;要让某个客户端把一个模型当另一个模型的名称使用,在这个客户端的密钥上写规则即可,不会改变这个名称对其他客户端的含义。条件写上游模型名时,也匹配列有它的别名。设置 `max_tokens` 可以限制回答的长度:上游到达上限时自行停止,不会报错。策略组把多个上游放在同一个名字下,并决定尝试的先后:按顺序、手动选择、轮询、延迟最低优先或费用最低优先。一次尝试失败时,请求转到下一个上游;同一会话默认保持在同一个上游上,以便提示缓存持续命中。页面顶部的路由图显示每把密钥经过的路由、策略组和上游。 客户端自行发出的辅助请求(连通性检查、预热、生成标题、话题识别、输入建议)可以由网关在本地应答而不产生费用,也可以转发。转发的辅助请求和普通请求一样经过路由规则,规则可以按辅助请求的类型把它们分流到费用更低的上游。 @@ -56,11 +58,13 @@ API 密钥和请求头的值可以写成 `${变量名}`,读取系统环境变 安全页有三项防护:出站脱敏、工具调用审查和内容过滤,对所有上游和所有密钥统一生效。每项都有三档:「关闭」「观察」,以及按作用命名的第三档,出站脱敏为「替换」,工具调用审查为「切断」,内容过滤为「处置」;其中「观察」只检测和记录,不做任何改动。三项出厂均为「观察」,因此默认不会改动或拒绝任何请求。 - **出站脱敏**:请求发出之前查找整个请求(含系统提示、之前的对话和工具调用)中的凭据和个人信息,包括 Anthropic、OpenAI、GitHub、Slack、AWS、Google、GitLab、Stripe、npm、DigitalOcean、SendGrid 的 API 密钥与令牌,以及私钥、JWT 和连接串中的口令,还有身份证号与银行卡号(号码结构与校验位都正确才算)。「替换」档下把它们替换为 `<>` 这样的占位符,响应中回显时再还原。邮箱地址、中国大陆手机号、内网 IP 地址和内网域名四条规则出厂为停用,可以按需启用;自定义规则可以指定占位符名称,如 `PROJECT` 替换为 `<>`。 -- **工具调用审查**:检查模型返回的工具调用中是否含有下载或解码后执行代码、外发环境变量或凭据文件、把凭据发往本机和其服务商以外的主机、读取私钥或云服务凭据、写入启动项或定时任务等命令。「切断」档下命中即切断响应,客户端收不到一个完整、可执行的调用。删除主目录或根目录、设置全员可写权限、把本地文件上传到外部主机三条规则出厂只记录。 +- **工具调用审查**:检查模型返回的工具调用中是否含有下载或解码后执行代码、外发环境变量或凭据文件、把凭据发往本机和其服务商以外的主机、读取私钥或云服务凭据、读写 ThinkWatch 自己的数据目录、写入启动项或定时任务等命令。只出现在要写入的文字里的路径(例如提到这个目录的文档)不算。「切断」档下命中即切断响应,客户端收不到一个完整、可执行的调用。删除主目录或根目录、设置全员可写权限、把本地文件上传到外部主机三条规则出厂只记录。 - **内容过滤**:检查每个请求中的用户消息和工具结果,压缩上下文的请求也在其内。规则按关键词、正则表达式或码位(如 `U+E0000–U+E007F`)匹配,「处置」档下按规则拒绝请求、删除命中的文字后发出,或仅记录。内置规则删除 Unicode 标签字符和双向控制符(人看不见、模型读得到,可用于隐藏指令),拒绝明确要求「忽略先前指令」的请求。零宽字符、私用区字符两条规则(表情符号、波斯文和图标字体也会用到这些字符),以及越狱、身份操纵、套取提示词等规则及其中文版本出厂为停用,可以按需启用。 安全页列出全部规则,内置规则按类别分组。内置规则可以逐条启用或停用;工具调用审查的内置规则可以设为切断或仅记录,内容过滤的内置规则可以设为拒绝、删除或仅记录。自定义规则为正则表达式,内容过滤也可以使用关键词或码位。任何规则都可以先用一段文本测试,出站脱敏和内容过滤的测试还会显示发出时的文本。各项防护检出的内容都记入第一个标签页的日志,并注明所属的请求;隐藏字符拼出文字时,日志中一并显示这段文字。 +各项防护作用于经过网关的请求和回答,不作用于磁盘上的文件。配置文件以明文保存密钥,只有所有者可读:挡得住其他用户,挡不住以同一用户身份运行的程序;出站脱敏保护的是带出本机的内容,不是这份文件。 + ## MCP MCP 页管理客户端从自己的配置文件中加载的内容,这些内容不经过网关。 diff --git a/src/data/core-docs/config.md b/src/data/core-docs/config.md index f542b4d..9066764 100644 --- a/src/data/core-docs/config.md +++ b/src/data/core-docs/config.md @@ -25,6 +25,25 @@ control socket (`twcore.sock`; on Windows a loopback port recorded in `control.port`). The directory is private to its owner (`0700`), the file is `0600`: it holds keys in plain text. +These permissions are the file's only protection on this machine. They keep +out other users, not programs running as the same user: such a program can +read every key in the file, and with the control key it holds, change the +configuration through the control plane. Outbound redaction does not change +that; it protects what a request carries off the machine, not what is on disk. +The control key is masked where the configuration is shown so that a write +through the control plane cannot change it +([`listen.control`](#cfg-listen-control)), not to keep it from local +programs; upstream and gateway keys are shown as written. + +Tool-call inspection has a built-in rule for this directory, +`thinkwatch-data`. A tool call whose path or command points into one of the +default locations above, or into `/var/lib/thinkwatch` or `/etc/thinkwatch` +on a server, is recorded, and under `enforce` the response is cut off, so a +model cannot be steered into reading these keys or rewriting its own +protections. Mentioning the path, as in a document being edited, does not +count. With `THINKWATCH_HOME` elsewhere, a custom rule in +[`security.inspect_tools`](#cfg-security-inspect_tools) can cover that path. + `twcore serve` writes a starting configuration when there is none, and `twcore init` writes one on request. Both produce this: @@ -153,6 +172,7 @@ means. | `security` | object, [`security`](#cfg-security) | — | The three guards. All of them start in `observe`, so out of the box nothing is changed or refused. | | `retention` | object, [`retention`](#cfg-retention) | — | How long request logs are kept. | | `failover` | object, [`failover`](#cfg-failover) | — | How long an upstream is set aside after it fails, and how long the start of a stream is awaited. | +| `aliases` | map of alias → string or list of strings | `{}` | Model aliases: one name for the same model across upstreams, mapped to the name each upstream uses, in order. A request for an alias goes to any upstream offering one of the listed names, under the first of them it offers. An alias cannot list another alias. | | `groups` | list of [`groups[]`](#cfg-groups) | `[]` | Strategy groups: several upstreams behind one name, with a way to pick among them. | | `routes` | list of [`routes[]`](#cfg-routes) | `[]` | Routes. Without any, requests fail over across all upstreams in the order they are declared. | | `default_route` | string | — | The route for keys that do not name one. Unset: the route named `default`, or the built-in failover when there is none. | @@ -284,7 +304,7 @@ gateway. A key is an identity. Limits, model scope and route are per key. | `name` | string | **required** | Name of the key; unique. Routing rules match it with `when.client`. | | `key` | string | **required** | The key clients send (as `x-api-key` or `Authorization: Bearer`). Generated keys start with `tw-` so they are not mistaken for an upstream's key. Unique. | | `max_concurrent` | integer | — | Requests with this key that may run at once; the rest wait. Unset: no limit. `0` is refused. | -| `allow` | list of strings | — | Models this key may use, as model ids or globs (`claude-*`). Unset: every model. `[]`: none at all. | +| `allow` | list of strings | — | Models this key may use, as model ids or globs (`claude-*`). An upstream model name also allows the aliases that list it; an alias allows only the alias. Unset: every model. `[]`: none at all. | | `route` | string | — | Name of the route requests with this key take. Unset: `default_route`. | | `client` | string | — | The client this key was made for (`claude-code`, `codex`, …), recorded when the desktop app points a client at the gateway. A client has at most one. | | `disabled` | bool | `false` | Refuse every request made with this key, and keep the key. | @@ -736,6 +756,7 @@ Built-in rules: | `exfil-credentials-reversed` | Send out a credential file (verb first) | `cut` | | `ssh-key-read` | Read a private key or cloud credential | `cut` | | `secret-to-unknown-host` | Send a credential to an unknown host | `cut` | +| `thinkwatch-data` | Read or change ThinkWatch's own data | `cut` | | `write-startup-item` | Write a startup item | `cut` | | `crontab-install` | Install a scheduled job | `cut` | | `rm-rf-root` | Delete home or root | `record` | @@ -886,6 +907,40 @@ the same as an error status would. | `stream_start_wait_secs` | integer | `15` | Seconds to hold a streamed answer until its first content arrives. An error before then moves the request to the next upstream; after this long, what has arrived is passed on. From 1 to 120. | +### `aliases` + +A model alias is another name for the same model. Clients request it like any +model; each upstream receives the name it uses for that model. + +```yaml +aliases: + deepseek-v4.1: DeepSeek-v4.1-flash + claude-sonnet-5: + - claude-sonnet-5 + - us.anthropic.claude-sonnet-5-v1:0 + - anthropic/claude-sonnet-5 +``` + +- An upstream serves an alias when it offers one of the listed names within + its `models_only`, and receives the first such name in the list. Every + upstream that serves an alias takes part in failover for it. +- `GET /v1/models` lists aliases next to the upstream models, and the + original names stay available. +- An alias takes precedence over an upstream model of the same name: a + request for `claude-sonnet-5` above goes only to upstreams offering one of + the three names. List an upstream's own name when it should keep serving it. +- A key's `allow` and a rule's `when.model` written for an upstream model name + also cover the aliases that list it; written for an alias, they cover only + the alias. +- When the upstream answers with the model it was sent, the answer carries the + name the client asked for. An answer naming a different model is passed on + as it is. +- Prices and the upstream check-up use the name sent upstream. The request log + keeps the name the client asked for, and each attempt the name it sent. +- To send one key's requests to a particular upstream model, use a rule with + pinned models ([`routes[].rules[].to`](#cfg-routes-rules-to)) rather than an + alias: an alias changes what the name means for every client. + ### `groups` A group puts several upstreams behind one name. Rules send requests to a @@ -942,11 +997,38 @@ order they are declared. |---|---|---|---| | `name` | string | **required** | Name shown in logs and in the traffic view. | | `when` | object, [`routes[].rules[].when`](#cfg-routes-rules-when) | — | Conditions, all of which have to hold. Unset: matches every request. | -| `to` | string | — | An upstream or a group, by name; `__all__` is every upstream in declared order. Not allowed together with `when.provider_would_be`. | +| `to` | string, or list of [`routes[].rules[].to[]`](#cfg-routes-rules-to) | — | An upstream or a group, by name; `__all__` is every upstream in declared order. Or pinned models: a list of upstreams with the model to send to each, tried in order. Not allowed together with `when.provider_would_be`. | | `set` | object, [`routes[].rules[].set`](#cfg-routes-rules-set) | — | Parameters to rewrite. Collected from every matching rule, not only the first. | | `deny` | string | — | Refuse the request with this reason. | +#### `routes[].rules[].to` + +`to` names an upstream or a group, or pins models: a list of upstreams, each +with the model sent to it, tried in order. A pinned model is sent as written, +without aliases or `set.model`, so a rule can send a key's requests to one +upstream's model even when an alias of the same name points elsewhere. + +```yaml +routes: + - name: default + rules: + - name: Opus on Bedrock + when: { model: claude-opus-5 } + to: + - { provider: bedrock, model: us.anthropic.claude-opus-5-v1:0 } + - { provider: anthropic, model: claude-opus-5 } +``` + + + + +| Field | Type | Default | Description | +|---|---|---|---| +| `provider` | string | **required** | An upstream, by name; not a group. Each upstream appears once in the list. | +| `model` | string | **required** | The model name sent to that upstream, as written: aliases do not apply, and no `set.model` changes it. Not allowed together with `set.model` in the same rule. | + + #### `routes[].rules[].when` @@ -954,7 +1036,7 @@ order they are declared. | Field | Type | Default | Description | |---|---|---|---| -| `model` | string | — | Requested model, glob (`claude-opus-*`). | +| `model` | string | — | Requested model, glob (`claude-opus-*`). An upstream model name also matches requests for the aliases that list it; an alias matches only requests for the alias. | | `client` | string | — | Name of the gateway key the request used, exactly. | | `dialect` | string | — | API format the client spoke: `anthropic`, `openai-chat`, `openai-responses`, `gemini`. | | `input_tokens` | comparison (`>200k`, `<=4k`, `==3`) | — | Estimated input tokens. | @@ -975,6 +1057,9 @@ not an equality: `"200k"` alone is refused. #### `routes[].rules[].set` +`set.model` may name an alias; each upstream then receives its own name for +it. Answers carry the name the client asked for, as with aliases. + @@ -1064,7 +1149,7 @@ Requests of a kind a plugin does not handle pass without it, whatever its | Field | Type | Default | Description | |---|---|---|---| -| `id` | string | **required** | Lowercase letters, digits and hyphens, 1 to 40 characters; unique. `order`, `inspect`, `rewrite` and `confirmed` are taken by the control plane. | +| `id` | string | **required** | Lowercase letters, digits and hyphens, 1 to 40 characters; unique. | | `file` | string | **required** | The plugin's code, relative to this file's directory. It is always `plugins/.js`; the app writes it. | | `sha256` | string | **required** | SHA-256 of the approved code, 64 lowercase hexadecimal characters. When the file no longer has this hash, the plugin stops running until the change is approved in the app. The approved code is kept in `plugins/.approved/.js`. | | `enabled` | bool | `true` | Run the plugin. `false` keeps it installed and out of every request. | diff --git a/src/data/core-docs/config.zh-CN.md b/src/data/core-docs/config.zh-CN.md index 3e0307d..a0b9ad7 100644 --- a/src/data/core-docs/config.zh-CN.md +++ b/src/data/core-docs/config.zh-CN.md @@ -15,6 +15,10 @@ ThinkWatch Core 只读一个文件:`config.yaml`。本文逐项说明其中每 `THINKWATCH_HOME` 替换整个目录;`--config <路径>` 为单条命令指定文件。core 的其余数据也在这个目录里:请求数据库(`data.db`)、配置历史(`history/`)、下载的价目表(`model_prices.json`),以及本地控制通道的 socket 文件(`twcore.sock`;Windows 上是回环端口,记录在 `control.port` 中)。目录只有所有者可访问(`0700`),配置文件权限为 `0600`:其中以明文保存密钥。 +这两项权限就是这份文件在本机上的全部保护:挡得住其他用户,挡不住以同一用户身份运行的程序。这样的程序能读到文件里的每一把密钥,也能用其中的控制通道密钥经控制通道修改配置。出站脱敏不改变这一点:它保护的是请求带出本机的内容,不是磁盘上的文件。显示配置时给控制通道密钥打码,是为了让经控制通道的写入改不了它(见 [`listen.control`](#cfg-listen-control)),而不是对本机程序隐藏它;上游密钥和网关密钥按原样显示。 + +工具调用审查为这个目录内置了一条规则 `thinkwatch-data`:工具调用的路径或命令指向上述默认位置之一,或服务器上的 `/var/lib/thinkwatch`、`/etc/thinkwatch` 时记录下来,`enforce` 下切断响应。这样模型无法被诱导去读取这些密钥、改写约束它自己的防护。只是提到这个路径(例如在正在编辑的文档里)不算。`THINKWATCH_HOME` 指向别处时,可以在 [`security.inspect_tools`](#cfg-security-inspect_tools) 里为那个路径加一条自定义规则。 + 没有配置文件时,`twcore serve` 会写入一份初始配置;`twcore init` 也可以按需生成。两者生成的内容如下: ```yaml @@ -102,6 +106,7 @@ twcore config set /listen/gateway/port 8790 --int | `security` | 对象,见 [`security`](#cfg-security) | — | 三项防护。出厂时都处在 `observe`,不改变、不拒绝任何请求。 | | `retention` | 对象,见 [`retention`](#cfg-retention) | — | 请求日志保留多久。 | | `failover` | 对象,见 [`failover`](#cfg-failover) | — | 上游失败后停用多久,以及流式回答的开头最多等多久。 | +| `aliases` | 映射: 别名 → 字符串或字符串列表 | `{}` | 模型别名:同一个模型在各上游的名称合用一个名字,按顺序列出各上游的名称。请求别名时,提供其中任一名称的上游都能服务,发给它的是列表中它提供的第一个名称。别名不能列出别的别名。 | | `groups` | 对象列表,见 [`groups[]`](#cfg-groups) | `[]` | 策略组:多个上游合用一个名字,并规定如何在其中选择。 | | `routes` | 对象列表,见 [`routes[]`](#cfg-routes) | `[]` | 路由。一条都不写时,请求按上游的声明顺序故障转移。 | | `default_route` | 字符串 | — | 未指定路由的密钥走哪条路由。不写:名为 `default` 的路由;没有这条路由时走内置的故障转移。 | @@ -208,7 +213,7 @@ listen: | `name` | 字符串 | **必填** | 密钥的名字,不能重复。路由规则用 `when.client` 匹配它。 | | `key` | 字符串 | **必填** | 客户端发送的密钥(放在 `x-api-key` 或 `Authorization: Bearer` 中)。生成的密钥以 `tw-` 开头,以免被误认作上游的密钥。不能重复。 | | `max_concurrent` | 整数 | — | 用这把密钥同时进行的请求数上限,超出的排队等待。不写:不限。`0` 会被拒绝。 | -| `allow` | 字符串列表 | — | 这把密钥可用的模型,写模型 ID 或通配(`claude-*`)。不写:全部模型。`[]`:一个都不给。 | +| `allow` | 字符串列表 | — | 这把密钥可用的模型,写模型 ID 或通配(`claude-*`)。写上游模型名,列有它的别名一并可用;写别名只放行别名。不写:全部模型。`[]`:一个都不给。 | | `route` | 字符串 | — | 这把密钥的请求走哪条路由。不写:`default_route`。 | | `client` | 字符串 | — | 这把密钥是为哪个客户端生成的(`claude-code`、`codex` 等),由桌面应用接管客户端时写入。一个客户端最多一把。 | | `disabled` | 布尔 | `false` | 拒绝使用这把密钥的所有请求,密钥本身保留。 | @@ -581,6 +586,7 @@ pricing: | `exfil-credentials-reversed` | Send out a credential file (verb first) | `cut` | | `ssh-key-read` | Read a private key or cloud credential | `cut` | | `secret-to-unknown-host` | Send a credential to an unknown host | `cut` | +| `thinkwatch-data` | Read or change ThinkWatch's own data | `cut` | | `write-startup-item` | Write a startup item | `cut` | | `crontab-install` | Install a scheduled job | `cut` | | `rm-rf-root` | Delete home or root | `record` | @@ -710,6 +716,27 @@ security: | `stream_start_wait_secs` | 整数 | `15` | 流式回答在第一段内容到达前最多暂存的秒数。在此之前上游报错,请求换到下一家;超过这个时间,已收到的部分照常交给客户端。取值 1 到 120。 | +### `aliases` + +模型别名是同一个模型的另一个名称。客户端像请求其他模型一样请求它,每家上游收到的是自己对这个模型的叫法。 + +```yaml +aliases: + deepseek-v4.1: DeepSeek-v4.1-flash + claude-sonnet-5: + - claude-sonnet-5 + - us.anthropic.claude-sonnet-5-v1:0 + - anthropic/claude-sonnet-5 +``` + +- 上游在 `models_only` 范围内提供列出的任一名称,就能服务这个别名,收到的是列表里它提供的第一个名称。能服务同一别名的上游互为备用。 +- `GET /v1/models` 在上游模型旁一并列出别名,原来的名称照常可用。 +- 别名优先于同名的上游模型:上例中请求 `claude-sonnet-5` 只会发给提供这三个名称之一的上游。要让某家上游继续用自己的同名模型提供服务,把它的名称列进去。 +- 密钥的 `allow` 和规则的 `when.model` 写上游模型名时,对列有这个名称的别名同样有效;写别名时只对别名有效。 +- 上游答的是发给它的那个模型时,回答里的模型名写成客户端请求的名称;答的是别的模型,原样转发。 +- 计价和上游体检按发给上游的名称;请求记录保留客户端请求的名称,每次尝试记录实际发出的名称。 +- 要把某把密钥的请求发到某家上游的某个模型,用带指定模型的规则([`routes[].rules[].to`](#cfg-routes-rules-to)),不要用别名:别名会改变这个名称对所有客户端的含义。 + ### `groups` 策略组让多个上游合用一个名字。规则用 `to` 把请求交给策略组。 @@ -751,11 +778,35 @@ security: |---|---|---|---| | `name` | 字符串 | **必填** | 日志和流量详情中显示的名字。 | | `when` | 对象,见 [`routes[].rules[].when`](#cfg-routes-rules-when) | — | 条件,须全部满足。不写:匹配所有请求。 | -| `to` | 字符串 | — | 上游或策略组的名字;`__all__` 表示按声明顺序的全部上游。不能与 `when.provider_would_be` 同时写。 | +| `to` | 字符串,或对象列表,见 [`routes[].rules[].to[]`](#cfg-routes-rules-to) | — | 上游或策略组的名字;`__all__` 表示按声明顺序的全部上游。也可以指定模型:列出上游及发给它的模型,按顺序备用。不能与 `when.provider_would_be` 同时写。 | | `set` | 对象,见 [`routes[].rules[].set`](#cfg-routes-rules-set) | — | 改写请求参数。从所有匹配的规则累积,不只第一条。 | | `deny` | 字符串 | — | 以这句原因拒绝请求。 | +#### `routes[].rules[].to` + +`to` 写上游或策略组的名称,或者指定模型:一组上游,每家写明发给它的模型,按顺序备用。指定模型原样发出,不经过别名,也不受 `set.model` 影响;所以即使同名别名指向别处,规则仍能把某把密钥的请求发到某家上游的这个模型。 + +```yaml +routes: + - name: default + rules: + - name: Opus 走 Bedrock + when: { model: claude-opus-5 } + to: + - { provider: bedrock, model: us.anthropic.claude-opus-5-v1:0 } + - { provider: anthropic, model: claude-opus-5 } +``` + + + + +| 字段 | 类型 | 默认值 | 说明 | +|---|---|---|---| +| `provider` | 字符串 | **必填** | 上游的名字,不能是策略组。同一个上游在列表中只出现一次。 | +| `model` | 字符串 | **必填** | 发给这个上游的模型名,原样发出:不经过别名,也不受 `set.model` 改写。同一条规则里不能再写 `set.model`。 | + + #### `routes[].rules[].when` @@ -763,7 +814,7 @@ security: | 字段 | 类型 | 默认值 | 说明 | |---|---|---|---| -| `model` | 字符串 | — | 请求的模型,可用通配(`claude-opus-*`)。 | +| `model` | 字符串 | — | 请求的模型,可用通配(`claude-opus-*`)。写上游模型名,也匹配请求列有它的别名的请求;写别名只匹配请求这个别名的。 | | `client` | 字符串 | — | 请求所用网关密钥的名字,精确匹配。 | | `dialect` | 字符串 | — | 客户端使用的接口格式:`anthropic`、`openai-chat`、`openai-responses`、`gemini`。 | | `input_tokens` | 比较式(`>200k`、`<=4k`、`==3`) | — | 估算的输入 token 数。 | @@ -782,6 +833,8 @@ security: #### `routes[].rules[].set` +`set.model` 可以写别名,每家上游收到各自对它的叫法。和别名一样,回答里的模型名写成客户端请求的名称。 + @@ -834,7 +887,7 @@ default_route: default | 字段 | 类型 | 默认值 | 说明 | |---|---|---|---| -| `id` | 字符串 | **必填** | 小写字母、数字和连字符,1 到 40 个字符,不能重复。`order`、`inspect`、`rewrite` 和 `confirmed` 被控制面占用。 | +| `id` | 字符串 | **必填** | 小写字母、数字和连字符,1 到 40 个字符,不能重复。 | | `file` | 字符串 | **必填** | 插件的代码,相对本文件所在的目录。只能是 `plugins/.js`,由应用写入。 | | `sha256` | 字符串 | **必填** | 批准过的代码的 SHA-256,64 个小写十六进制字符。文件的哈希与它不符时插件停止运行,直到在应用里批准这次改动。批准过的代码另存在 `plugins/.approved/.js`。 | | `enabled` | 布尔 | `true` | 是否运行这个插件。`false`:插件保留,不参与任何请求。 | diff --git a/src/data/core-docs/manifest.json b/src/data/core-docs/manifest.json index afe5417..f7459d4 100644 --- a/src/data/core-docs/manifest.json +++ b/src/data/core-docs/manifest.json @@ -1,9 +1,9 @@ { "repository": "ThinkWatchProject/ThinkWatch-Core", - "ref": "v0.60.0", + "ref": "v0.62.0", "files": { - "docs/config.md": "c5ccc802e4f502e28b468eab6b995ae737cc2a04a46bfe8012c89560bdf4ade1", - "docs/config.zh-CN.md": "feba4c83df62f5fcf8c8dfad6b3c70d6f6ebc8e3e6c05f6b428d6ec80351da9a", + "docs/config.md": "e916418233efc68eb2a7883c22463b3cdeee76b0ac034cb7bdc1516d4223ae07", + "docs/config.zh-CN.md": "46c3f4bcddffd7f5d74d4f066688ffba9d9d0d4d8c969b5548356c14b918fd57", "docs/server.md": "5e1e9b901bef2b46d417aea057a1db24c78457d3c1938b8540ecb4766515b68f", "docs/server.zh-CN.md": "1f508e7c39b8ded28653773ca4d8701e6bdc2247bcd359c3a8fd00fd1401a6ff" }