diff --git a/src/content/docs-lite/en/failover-and-load-balancing.md b/src/content/docs-lite/en/failover-and-load-balancing.md index 823c26d..47d93aa 100644 --- a/src/content/docs-lite/en/failover-and-load-balancing.md +++ b/src/content/docs-lite/en/failover-and-load-balancing.md @@ -4,13 +4,13 @@ ThinkWatch Lite turns each relay key into an upstream and puts several upstreams ## Before you start -- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.4. +- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.6. - The base URL and API keys of each relay. ## Steps 1. On the Upstreams page, choose **New upstream**. Under **Connection**, enter a **Name**, the **Base URL** and one **API key**, choose **Check connection**, then **Next** until **Create**. Repeat for each key: an upstream holds one key, and failover and pauses work per upstream. -2. On the Routing page, choose **New group** under **Groups**. Enter a **Name**, choose a **Strategy**, tick the upstreams under **Members**, drag them into order and choose **Create**. +2. On the Routing page, choose **New group** under **Groups**. Enter a **Name**, choose a **Strategy**, tick the upstreams under **Members**, drag them into order and choose **Create**. For **Round robin**, also enter each member's **Weight** and choose **Distribute by**. 3. Under **Routes**, open the route, choose **Edit…** on the rule that should use the group (usually **All requests (catch-all)**), select the group in **Forward to** and choose **Save** in both dialogs. 4. **Dry run** checks the result without sending anything. On the Traffic page, a request's **Routing** tab lists the upstreams it tried under **Attempts**. @@ -19,14 +19,17 @@ ThinkWatch Lite turns each relay key into an upstream and puts several upstreams | Strategy | Order of the members | |---|---| | **In order** | As listed. | -| **Round robin** | Requests are distributed across the members in turn. | -| **Lowest latency** | Median time to first byte over each upstream's last 32 requests; an upstream with fewer than 3 samples ranks after. | +| **Round robin** | Requests are distributed across the members in turn, in proportion to their weights. | +| **Lowest latency** | Median time to first token over each upstream's last 32 streamed answers; until an upstream has 3, a connection timing taken at startup stands in, and an upstream with neither ranks after. | | **Lowest cost** | Input price of the requested model in each upstream's price sheet; free upstreams first, unpriced ones last. | | **Manual** | The preferred upstream first, then the rest as listed. | A strategy only sets the order; every member remains available for failover. - **When the next upstream is tried.** Before anything has reached the client, the gateway moves on when the upstream cannot be reached or its credential cannot be read; answers 5xx, 429, 401, 403, 402 or 404; answers 400 or 422 with an error about an insufficient balance, a used-up quota or an unavailable model; or, in a streamed answer, reports an error before the first content, such as an overload. The gateway waits for that first content for up to **Wait for the answer to start** in Settings › Failover, 15 seconds by default. Other 4xx responses, and the last member's 4xx other than 429, go back to the client unchanged. +- **Weights and distribution.** In a **Round robin** group, a member's weight, from 1 to 100, is its long-run share of requests: weights 7 and 3 send seven requests in ten to the first, interleaved with the other three. **By ratio** uses the weights alone; **By speed** multiplies each by how much faster than the members' median the upstream starts answering (time to first token, squared, kept between 0.1 and 10), **By reliability** by its success rate over its last 50 outcomes within 30 minutes (squared, at least 0.05), and **By speed and reliability** by both. An upstream with too few samples counts as average, requests that stay with an upstream for their conversation count toward its share, and paused or full members sit out. +- **Slow starts.** With **Move to the next upstream when the start times out** on in Settings › Failover, a streamed answer with no content after **Wait for the answer to start** is cancelled and the request goes to the next upstream, if one can take it at that moment; the last upstream always waits, and the slow one is not paused. The upstream may have billed the input of the abandoned attempt, which the request's **Routing** tab marks; with the switch on, the wait must be at least 5 seconds, and 30 or more suits models that think before they answer. +- **Concurrency limits.** An upstream with a **Concurrency limit** takes at most that many requests at once. When it is full, a conversation that stays on it waits for a free slot and then moves on, and other requests go straight to the next member; when every member is full, a request waits for the first free slot for up to **Wait for a free slot at most** in Settings › Failover, 30 seconds by default, and then receives a 429 with `Retry-After`, or the error of an earlier attempt when one was sent. The same wait covers a key's per-minute and per-hour usage limits, and the Traffic page shows each skip and the time queued. - **Pauses.** A failing upstream is paused for as long as Settings › Failover sets: by default 60 seconds after 3 **Consecutive failures**, doubling up to 600; 30 minutes for **Insufficient balance**; until the reset time, or 60 minutes, for **Quota used up**; the wait a rate-limited upstream asks for, up to 60 minutes. A rule with a single upstream is never held back, and when every member is paused they are tried anyway. - **Sessions and prompt cache.** Within a turn, while the client sends tool results back, requests keep the rule chosen at the start of the turn and the upstream that answered. Across turns, a conversation stays with the upstream that answered last if that answer read or wrote at least 1,024 cached tokens within the last five minutes; moving would rebuild the cache at full price. Otherwise, or when that upstream is paused, the strategy orders the members again, which is when **Round robin** moves on. The **Conversation** line on a request's **Routing** tab shows when a request stayed. - **Same model name.** Every member is asked for the model the client sent, or the name a rule rewrote it to. A member whose model list lacks it is skipped; one without a list is tried, and its 404 moves the request on. diff --git a/src/content/docs-lite/en/features.md b/src/content/docs-lite/en/features.md index 883bdb9..fc774e3 100644 --- a/src/content/docs-lite/en/features.md +++ b/src/content/docs-lite/en/features.md @@ -18,6 +18,8 @@ The Traffic page lists requests as they arrive: status, key, model, upstream, ti A request opens into its timeline, its routing (the rule it matched, the group it went through and each attempt with its status and duration), the request and response bodies, and its usage and cost. A request from DeepSeek Harness also shows the size of the session log it carried, the whole conversation the client attaches to every request; the gateway removes it before a request goes to an upstream other than DeepSeek. A finished request can be sent again, unchanged, to another upstream after an estimate of its cost, and the two responses are shown side by side. +The attempts also show an upstream passed over at its concurrency limit (**At its concurrency limit**), the time spent waiting for a free slot, and a streamed answer given up because its start timed out (**Start timed out**), with the input tokens that upstream may have billed for it. A request refused by its key's usage limit, or because every upstream stayed at its concurrency limit, shows **Usage limit** or **At capacity** in place of an upstream, and a connection in the Responses API's WebSocket mode is listed one request per turn, each with its own usage and cost. + ## Client setup The Clients page points Claude Code, Claude Desktop, Codex, opencode, Pi, oh-my-pi, Grok Build, Qwen Code, Hermes Agent, Zed, Aider and DeepSeek Harness at the gateway. Before anything is written, it lists the fields that change and what else the change affects (the ChatGPT desktop app, for instance, reads the same configuration file as Codex), shows the full diff and backs up the original file. Only the settings that point the client at the gateway change, and each client receives a key of its own. Claude Desktop is connected through its official third-party inference mode, and the page lists each of the files that change for it; a Claude Desktop managed by an organization is left as it is. Claude Desktop accepts only Claude model names; when no upstream offers a Claude model, the takeover asks which model it should use and adds a routing rule for its key, which cancelling the takeover removes. A connected client can be restored at any time, on its own or together with all the others; a restored Codex keeps a plain OpenAI entry in place of the gateway's, so sessions started while it was connected can still be opened. opencode (v1 and v2), Pi, oh-my-pi, Grok Build and Qwen Code also get the list of models their key can use on the gateway; when that list changes, the page offers to update it, through the same diff. DeepSeek Harness is covered in its web app, its desktop app and headless runs, which all read the same configuration; models used through a DeepSeek account signed in to the desktop app still go to DeepSeek directly. Cursor, Continue and Antigravity CLI come with step-by-step instructions and a key created for them. For every client the page shows whether it is in use, waiting for its first request or not in effect, and its requests over the last 24 hours. @@ -31,15 +33,17 @@ WSL 2 uses NAT networking by default, and the gateway on Windows cannot be reach ## Keys -Clients reach the gateway with a key, on the local machine as well. The Keys page lists the default key, used by clients that were not given one of their own, and a key for each connected client, labelled with the client it belongs to so that its requests can be told apart in Traffic. Each key has a route, the models it may use (all, none, or chosen models and patterns such as `gpt-5*`), an optional limit on concurrent requests, and its requests and cost over the last 24 hours. A key can be disabled, which rejects every request made with it, or rotated; rotating writes the new key into the configuration of the client that uses it. +Clients reach the gateway with a key, on the local machine as well. The Keys page lists the default key, used by clients that were not given one of their own, and a key for each connected client, labelled with the client it belongs to so that its requests can be told apart in Traffic. Each key has a route, the models it may use (all, none, or chosen models and patterns such as `gpt-5*`), an optional limit on concurrent requests, optional usage limits, and its requests and cost over the last 24 hours. A key can be disabled, which rejects every request made with it, or rotated; rotating writes the new key into the configuration of the client that uses it. + +Usage limits keep one client from using up a budget: each caps the key at a number of requests, tokens or a cost in USD per minute, hour, day, week or month, and once one is used up the key's requests are refused until it resets, with the key marked **Limit reached**. Token limits can count cache reads, models without a price count as zero toward a cost limit and are listed in the key dialog, and a daily, weekly or monthly limit raises a notification at 80% and when it is reached. ## Upstreams -Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; aliases, proxies and price sheets have tabs on the same page. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs. +Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; aliases, proxies and price sheets have tabs on the same page. An optional **Concurrency limit** suits relays and accounts that allow only so many requests at once: when the upstream is full, a conversation that stays on it waits for a free slot, other requests go to the next upstream, and when every upstream is full a request waits and then receives a busy error. **Specs…** in an upstream's model list sets a model's context window and maximum output by hand, for a model the price table lacks or gets wrong, and the values set there are used in place of the price table's. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs. The Aliases tab gives a model the name clients use for it. An alias lists the names the same model has on different upstreams, such as `claude-sonnet-5` on Anthropic and `us.anthropic.claude-sonnet-5-v1:0` on Bedrock; every upstream that offers one of them serves the alias under its own name, and they back each other up. Clients see aliases in their model lists, and answers carry the name the client asked for, while the request log shows the model each upstream was sent. When the same Claude model has different names on the official API, Bedrock, Vertex or OpenRouter, the tab and the new-alias dialog suggest grouping them. An alias can also be created from an upstream's model list. The dialog shows which upstream receives which name and warns when the name would take over another upstream's model of the same name. A key whose model scope allows an upstream model can also use the aliases that list it. -The Check-up tab compares the upstreams over the last 24 hours, 7 days, 30 days or a custom range: requests and failure rate; whether the model named in each answer matches the one sent; the input tokens each upstream reports, as a multiple of the gateway's own estimate, against other upstreams serving the same model; the share of input read from the prompt cache in follow-up turns, also against other upstreams; and the median time to first token and generation speed. Each figure comes with its sample size, and a deviation is marked only when both sides have enough samples. +Each upstream is also compared over the last 7 days with the others serving the same model, and a deviation is marked next to its name in the list: **Model name differs** when answers name a model other than the one sent, **Input reported high** or **Input reported low** when the input tokens it reports, as a multiple of the gateway's own estimate, are well off the other upstreams', and **Low cache reads** when follow-up turns read a smaller share of their input from the prompt cache. Hovering over a mark shows the evidence and the sample sizes, and a deviation is marked only when both sides have enough samples. API keys and header values can be written as `${NAME}` to read a system environment variable. On macOS these come from the login shell, so variables exported in `~/.zshrc` and similar files apply, and the same holds on Linux (`~/.bashrc`, `~/.profile` and so on); on Windows they are the environment variables configured in system settings. After a variable changes, reopening the app picks it up. Proxy variables such as `HTTPS_PROXY`, and `PATH`, are not read. @@ -47,7 +51,7 @@ A relay or vendor can hand out an import link, `thinkwatch://import?…` or its ## Routing and failover -Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream, a group or specified models, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Specified models name an upstream and one of its models, with backups tried in order, and send the model as written; to have one client use one model under another model's name, a rule on that client's key does it without changing what the name means for other clients. A model condition written for an upstream model also matches the aliases that list it. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. A map at the top of the page traces every key through its route and groups to the upstreams. +Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream, a group or specified models, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Specified models name an upstream and one of its models, with backups tried in order, and send the model as written; to have one client use one model under another model's name, a rule on that client's key does it without changing what the name means for other clients. A model condition written for an upstream model also matches the aliases that list it. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. In a round-robin group each member has a weight from 1 to 100, its long-run share of requests, and **Distribute by** can multiply the weights by each upstream's speed (time to first token), its reliability (recent success rate) or both; conversations in progress stay on their upstream. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. With **Move to the next upstream when the start times out** on in Settings › Failover, a streamed answer that has sent nothing within the wait also moves on, when another upstream can take it; the last upstream always waits. A map at the top of the page traces every key through its route and groups to the upstreams. Auxiliary requests that clients send on their own (health checks, warm-ups, titles, topic detection and input suggestions) can be answered locally at no cost, or forwarded. Forwarded ones go through the routing rules like any other request, and a rule can send a given kind to a lower-cost upstream. @@ -81,7 +85,7 @@ Plugins are short JavaScript files that change requests before they go to an ups ## Settings -Settings has seven sections. Connection lists the local core and the saved remote cores, described in [Connecting to a remote core](/docs/lite/remote-core). General sets the language, the appearance, what the menu bar item shows on macOS, launch at login, whether notices arrive as system notifications, in the app only or not at all, and shows hidden guidance hints again. Listening sets who can reach the gateway (this machine only, the local network of a chosen interface, or every interface), its port and the allowed address ranges. Failover sets how long a failing upstream is paused before requests go to the next one: after how many consecutive failures, for how long, and separate pauses for an insufficient balance, a used-up quota and rate limits, and how long to wait for a streamed answer to start before trying the next upstream. Log retention sets how long request payloads and request records are kept, and a size cap for payloads. About shows the version, checks for updates and produces a diagnostics bundle with keys and addresses masked. Uninstall restores every connected client and removes the autostart entry, and is meant to be run before the app is deleted. +Settings has seven sections. Connection lists the local core and the saved remote cores, described in [Connecting to a remote core](/docs/lite/remote-core). General sets the language, the appearance, what the menu bar item shows on macOS, launch at login, whether notices arrive as system notifications, in the app only or not at all, and shows hidden guidance hints again. Listening sets who can reach the gateway (this machine only, the local network of a chosen interface, or every interface), its port and the allowed address ranges. Failover sets how long a failing upstream is paused before requests go to the next one: after how many consecutive failures, for how long, and separate pauses for an insufficient balance, a used-up quota and rate limits. It also sets how long to wait for a streamed answer to start, whether a stream that has not started by then moves to the next upstream (off by default), and how long in total a request may wait for a free slot on a full upstream or for a key's per-minute or per-hour limit, 30 seconds by default. Log retention sets how long request payloads and request records are kept, and a size cap for payloads. About shows the version, checks for updates and produces a diagnostics bundle with keys and addresses masked. Uninstall restores every connected client and removes the autostart entry, and is meant to be run before the app is deleted. ## Menu bar, system tray and notifications @@ -93,4 +97,4 @@ On Windows the icon sits in the notification area. Hovering over it shows the ga On Linux the icon sits in the system tray. Clicking it opens the same menu, with Open ThinkWatch Lite as its first item and quota bars written out as text. -System notifications, native on macOS and Windows and sent through the desktop's notification service on Linux, report when the gateway stops forwarding or keeps restarting, the connection to a remote core drops, a subscription quota runs out, a sign-in expires or an upstream rejects its credential, a proxy cannot be reached, the configuration file fails validation, a tool call matches a rule that cuts the response off, or suspicious content appears in a client's configuration. A new version found by the automatic check is announced the same way (see [Updates](/docs/lite/install#updates)). An unreachable upstream, which a fallback usually covers, is only listed in the app. Notices as a whole can be set to system notifications, in-app only, or off. Marking a notice as read stops the bell from counting it; the notice stays in the list until the problem behind it clears or the list is cleared. +System notifications, native on macOS and Windows and sent through the desktop's notification service on Linux, report when the gateway stops forwarding or keeps restarting, the connection to a remote core drops, a subscription quota runs out, a sign-in expires or an upstream rejects its credential, a proxy cannot be reached, the configuration file fails validation, a tool call matches a rule that cuts the response off, a key nears or reaches a daily, weekly or monthly usage limit, or suspicious content appears in a client's configuration. A new version found by the automatic check is announced the same way (see [Updates](/docs/lite/install#updates)). An unreachable upstream, which a fallback usually covers, is only listed in the app. Notices as a whole can be set to system notifications, in-app only, or off. Marking a notice as read stops the bell from counting it; the notice stays in the list until the problem behind it clears or the list is cleared. diff --git a/src/content/docs-lite/en/keep-api-keys-from-relays.md b/src/content/docs-lite/en/keep-api-keys-from-relays.md index dd5f0e5..0ddc2e3 100644 --- a/src/content/docs-lite/en/keep-api-keys-from-relays.md +++ b/src/content/docs-lite/en/keep-api-keys-from-relays.md @@ -4,7 +4,7 @@ A relay receives every request in full, including keys that end up in the conver ## Before you start -- ThinkWatch Lite, [installed](/lite/#install), with clients connected to the gateway. This guide follows version 2026.10.4. +- ThinkWatch Lite, [installed](/lite/#install), with clients connected to the gateway. This guide follows version 2026.10.6. - Both protections start in **Observe**: they record what they find and change nothing. **Off** checks nothing. ## Steps @@ -28,7 +28,7 @@ A relay receives every request in full, including keys that end up in the conver - **What is cut off.** Calls that download or decode code and run it, send out environment variables or credential files, read private keys or cloud credentials, send a credential to an unknown host, or install startup items or scheduled jobs. Built-in rules for deleting home or root, world-writable permissions and uploading a local file only record. - **What the relay still sees.** The key configured for its upstream in the app, and the rest of the request as written: code, file contents, the conversation. The key a client uses for the gateway is not forwarded, and neither is the client's identity unless **Forward client identity** is on for that upstream. - **Rules, not judgement.** Only values that match a rule are replaced; a credential without a recognizable prefix, such as an AWS secret access key, needs a custom rule. A dangerous command written in a form no rule matches passes tool-call inspection. -- **Nothing beyond the gateway.** What a relay does on its own servers, such as keeping requests or answering with another model, is out of sight; the **Check-up** tab on the Upstreams page can show signs of the latter, not prove it. +- **Nothing beyond the gateway.** What a relay does on its own servers, such as keeping requests or answering with another model, is out of sight; a mark such as **Model name differs** next to an upstream on the Upstreams page can show signs of the latter, not prove it. - **The cost of acting.** Changed request content may miss the upstream's prompt cache, and a false match in **Cut off** stops the answer at that call. Related: [Features](/docs/lite/features/#security), including the third protection, the content filter; [Install and update](/docs/lite/install/). diff --git a/src/content/docs-lite/zh-CN/failover-and-load-balancing.md b/src/content/docs-lite/zh-CN/failover-and-load-balancing.md index 28418e9..431d7a5 100644 --- a/src/content/docs-lite/zh-CN/failover-and-load-balancing.md +++ b/src/content/docs-lite/zh-CN/failover-and-load-balancing.md @@ -4,13 +4,13 @@ ThinkWatch Lite 把每个中转站的每把密钥建成一个上游,再把多 ## 准备 -- 已[安装](/zh-CN/lite/#install) ThinkWatch Lite,并已在客户端页接管客户端。本文按 2026.10.4 版编写。 +- 已[安装](/zh-CN/lite/#install) ThinkWatch Lite,并已在客户端页接管客户端。本文按 2026.10.6 版编写。 - 各中转站的接口地址和 API 密钥。 ## 步骤 1. 在上游页选择「新建上游」。在「连接」中填写「名称」「接口地址」和一把「API 密钥」,选择「检测连接」,再依次选择「下一步」直到「创建」。每把密钥重复一次:一个上游只有一把密钥,故障转移和暂停都以上游为单位。 -2. 在路由页的「策略组」标签中选择「新建策略组」。填写「名称」,选择「策略」,在「成员」中勾选上游并拖动排序,然后选择「创建」。 +2. 在路由页的「策略组」标签中选择「新建策略组」。填写「名称」,选择「策略」,在「成员」中勾选上游并拖动排序,然后选择「创建」。策略为「轮询」时,还要填写每个成员的「权重」并选择「分配依据」。 3. 在「路由」标签中打开路由,对要使用该策略组的规则选择「编辑…」(通常是「全部请求(兜底)」那一条),在「转发至」中选择该策略组,并在两个对话框中各选择一次「保存」。 4. 「试算」可以检查结果,不会发出请求。在流量页打开请求,「路由」标签的「尝试链」列出尝试过的每个上游。 @@ -19,14 +19,17 @@ ThinkWatch Lite 把每个中转站的每把密钥建成一个上游,再把多 | 策略 | 成员的顺序 | |---|---| | 「按顺序」 | 按列表顺序。 | -| 「轮询」 | 在成员之间轮流分配请求。 | -| 「延迟最低」 | 按各上游最近 32 个请求首字节时间的中位数排序;样本少于 3 个的排在后面。 | +| 「轮询」 | 在成员之间轮流分配请求,比例与权重相同。 | +| 「延迟最低」 | 按各上游最近 32 个流式回答首 token 时间的中位数排序;样本不足 3 个时以启动时的链路测速代替,两者都没有的排在后面。 | | 「费用最低」 | 按各上游价目表中所请求模型的输入单价排序;设为不计费的排在最前,无法计价的排在最后。 | | 「手动选择」 | 优先使用的上游排在最前,其余按列表顺序。 | 策略只决定顺序,所有成员都留作故障转移的后备。 - **何时换到下一个上游。**在尚未向客户端发出任何内容时,上游出现以下情况,网关即改用下一个成员:无法连接或读不到它的凭据;返回 5xx、429、401、403、402 或 404;返回 400 或 422,且错误信息说明余额不足、额度用完或模型不可用;流式回答在第一段内容之前报错,例如过载。等待第一段内容的时长为「设置 › 故障转移」中的「等待回答开头」,默认 15 秒。其他 4xx,以及最后一个成员返回的 4xx(429 除外),原样返回客户端。 +- **权重与分配依据。**「轮询」策略组中,成员的权重(1 到 100)就是它长期分到的请求比例:权重 7 与 3 时,每十个请求中七个交给前者,与另外三个交错分配。「按比例」只看权重;「按速度」再乘以该上游开始回答比成员中位数快多少(首 token 时间之比的平方,限定在 0.1 到 10 之间);「按稳定性」再乘以它 30 分钟内最近 50 次结果的成功率(取平方,最低 0.05);「按速度和稳定性」两者都乘。样本不足的上游按平均水平计算;因对话延续而留在某个上游的请求计入它的份额;暂停中或并发已满的成员本轮不参与。 +- **开头超时。**在「设置 › 故障转移」中打开「开头超时时转到下一个上游」后,流式回答等过「等待回答开头」仍没有内容时,只要此刻另有上游可以接下,网关就取消这次尝试,把请求交给下一个上游;最后一个上游照常等待,慢的上游也不会被暂停。被放弃的那次尝试的输入可能已被上游计费,请求的「路由」标签会标出。开启时等待时长至少为 5 秒;先思考再输出的模型宜设为 30 秒以上。 +- **并发上限。**设置了「并发上限」的上游同时最多接收这么多请求。上游已满时,留在它上面的对话等待空位,等不到再换下一个,其他请求直接交给下一个成员;所有成员都满时,请求等待最先空出的位置,最长为「设置 › 故障转移」中的「最多等待空位」(默认 30 秒),仍无空位则返回 429 并带 `Retry-After`;此前已有尝试发出时,返回那次尝试的错误。密钥的分钟、小时用量上限共用这段等待;流量页会标出每一次跳过和排队时间。 - **暂停。**失败的上游按「设置 › 故障转移」暂停使用,默认值为:「连续失败」3 次后暂停 60 秒,此后每次加倍,最长 600 秒;「余额不足」暂停 30 分钟;「额度用完」暂停到重置时刻,未给出时暂停 60 分钟;「限流」按上游要求的等待时间暂停,最长 60 分钟。只有一个上游的规则不受暂停影响;所有成员都在暂停时,网关仍会逐个尝试。 - **会话与提示缓存。**同一轮之内(客户端回传工具结果期间),请求沿用这一轮开头确定的规则和回答它的上游。跨轮时,如果上次回答在五分钟以内、且读写了至少 1,024 个缓存 token,对话继续使用该上游,因为换到别处要按全价重建缓存;否则,或者该上游正在暂停时,由策略重新排序,「轮询」正是在这时轮到下一个成员。请求「路由」标签中的「对话延续」一行说明请求是否因此留在原上游。 - **模型名相同。**每个成员收到的都是客户端请求的模型名,或者规则改写之后的模型名。模型列表中没有该模型的成员会被跳过;没有模型列表的成员照常尝试,它返回 404 时请求换到下一个。 diff --git a/src/content/docs-lite/zh-CN/features.md b/src/content/docs-lite/zh-CN/features.md index 5b38f1a..d93d3f3 100644 --- a/src/content/docs-lite/zh-CN/features.md +++ b/src/content/docs-lite/zh-CN/features.md @@ -18,6 +18,8 @@ 打开一个请求可以查看时间线、路由(命中的规则、经过的策略组,以及每一次尝试的状态与耗时)、请求与响应正文、用量与费用。DeepSeek Harness 发出的请求还会显示所带会话日志的大小:这是客户端随每个请求附带的整段对话记录,发往 DeepSeek 以外的上游之前由网关去除。已结束的请求可以在预估费用后原样发送到另一个上游,两次的响应并排对照。 +尝试链还会标出因并发已满而跳过的上游、等待空位的排队时间,以及因开头超时而放弃的流式回答(「开头超时」),并注明该上游可能已计费的输入 token。因密钥的用量上限或上游并发均已满而未被接下的请求,在上游的位置显示「用量上限」或「并发已满」;以 Responses API 的 WebSocket 模式建立的连接按轮记为请求,每一轮单独统计用量与费用。 + ## 客户端接管 客户端页可以把 Claude Code、Claude Desktop、Codex、opencode、Pi、oh-my-pi、Grok Build、Qwen Code、Hermes Agent、Zed、Aider 与 DeepSeek Harness 指向网关。写入之前,页面列出将要修改的字段和这次接管的其他影响(例如 ChatGPT 桌面版与 Codex 读取同一份配置文件),给出完整的改动差异,并完整备份原文件。只修改指向网关所需的配置,每个客户端使用各自的密钥。Claude Desktop 通过官方的第三方推理模式接入,页面逐一列出要修改的各个文件;由组织统一管理的 Claude Desktop 不做修改。Claude Desktop 只接受 Claude 的模型名;没有上游提供 Claude 模型时,接管会询问它使用哪个模型,并在它的密钥上添加一条路由规则,取消接管时一并删除。已接管的客户端可以随时单独还原或全部还原;Codex 还原后保留一项直连 OpenAI 的配置,接管期间的会话仍可打开。opencode(v1 与 v2)、Pi、oh-my-pi、Grok Build 与 Qwen Code 的配置中同时写入其密钥在网关上可用的模型列表;网关上可用的模型变化后,页面提示更新,更新同样先给出改动差异。DeepSeek Harness 的网页版、桌面版与 headless 模式读取同一份配置,都会经过网关;在桌面版中通过 DeepSeek 账号登录使用的模型仍直接连接 DeepSeek。Cursor、Continue 与 Antigravity CLI 提供逐步的配置方法,并为其创建密钥。页面列出每个客户端处于使用中、等待首个请求还是未生效,以及最近 24 小时的请求。 @@ -31,15 +33,17 @@ WSL 2 默认使用 NAT 网络,此时 Windows 上的网关无法从 WSL 内访 ## 密钥 -客户端连接网关必须携带密钥,本机也不例外。密钥页列出默认密钥(供未单独分配密钥的客户端使用)和每个已接管客户端的专用密钥,并注明所属客户端,便于在流量页中区分各客户端的请求。每把密钥有各自的路由、可用模型(全部、无,或指定的模型与 `gpt-5*` 这类通配模式)、可选的并发上限,以及最近 24 小时的请求数与费用。密钥可以停用,停用后使用它的请求一律被拒绝;也可以更换,新密钥会写入使用它的客户端的配置。 +客户端连接网关必须携带密钥,本机也不例外。密钥页列出默认密钥(供未单独分配密钥的客户端使用)和每个已接管客户端的专用密钥,并注明所属客户端,便于在流量页中区分各客户端的请求。每把密钥有各自的路由、可用模型(全部、无,或指定的模型与 `gpt-5*` 这类通配模式)、可选的并发上限与用量上限,以及最近 24 小时的请求数与费用。密钥可以停用,停用后使用它的请求一律被拒绝;也可以更换,新密钥会写入使用它的客户端的配置。 + +用量上限防止单个客户端用光预算:每条上限限定这把密钥每分钟、小时、天、周或月最多使用的请求次数、token 数或费用(美元),任一上限用满后,使用这把密钥的请求一律被拒绝,直到重置,密钥列表中标出「已达上限」。token 上限可以计入缓存读取;没有价格的模型按 0 计入费用上限,密钥对话框会列出这些模型;每天、每周或每月的上限用到八成和用满时各发送一次通知。 ## 上游 -上游是网关转发请求的目标:Anthropic、OpenAI、Google Gemini、DeepSeek 或任何兼容接口的 API 密钥,Amazon Bedrock(API 密钥、访问密钥或 AWS 配置文件),在应用内登录的 ChatGPT 账号或 Z.ai / BigModel 账号,OpenRouter 等中转服务,以及 Ollama 等本机模型。ChatGPT 账号显示订阅额度与重置时间;GLM Coding Plan 的上游(地址在 `api.z.ai` 或 `open.bigmodel.cn` 上,在应用内登录或手动填写密钥均可)同样显示:5 小时与每周额度,积分制套餐另外显示剩余积分(「剩余 1,976 / 2,000 积分」)。客户端与上游的 API 格式不同时,请求在 Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 与 Gemini 之间自动转换,无法转换的字段会在请求上逐一列出。上游可以经出站代理访问,也可以使用单独的价目表计价,别名、代理与价目表在同一页的标签中管理。链路测速测量 DNS 解析以及 TCP、TLS、代理握手的耗时,不产生费用;推理测速测量首个 token 的时间,运行前先给出费用预估。 +上游是网关转发请求的目标:Anthropic、OpenAI、Google Gemini、DeepSeek 或任何兼容接口的 API 密钥,Amazon Bedrock(API 密钥、访问密钥或 AWS 配置文件),在应用内登录的 ChatGPT 账号或 Z.ai / BigModel 账号,OpenRouter 等中转服务,以及 Ollama 等本机模型。ChatGPT 账号显示订阅额度与重置时间;GLM Coding Plan 的上游(地址在 `api.z.ai` 或 `open.bigmodel.cn` 上,在应用内登录或手动填写密钥均可)同样显示:5 小时与每周额度,积分制套餐另外显示剩余积分(「剩余 1,976 / 2,000 积分」)。客户端与上游的 API 格式不同时,请求在 Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 与 Gemini 之间自动转换,无法转换的字段会在请求上逐一列出。上游可以经出站代理访问,也可以使用单独的价目表计价,别名、代理与价目表在同一页的标签中管理。可选的「并发上限」适用于限制并发的中转站或账号:上游已满时,留在它上面的对话等待空位,其他请求交给下一个上游;所有上游都满时,请求先等待,仍无空位则返回繁忙错误。在上游的模型列表中选择「规格…」,可以手动设置模型的上下文窗口与输出上限,适用于价目表中没有或数值有误的模型,手动设置的值优先于价目表。链路测速测量 DNS 解析以及 TCP、TLS、代理握手的耗时,不产生费用;推理测速测量首个 token 的时间,运行前先给出费用预估。 别名标签为模型设定客户端使用的名称。一个别名列出同一个模型在各家上游的名称,例如 Anthropic 上的 `claude-sonnet-5` 和 Bedrock 上的 `us.anthropic.claude-sonnet-5-v1:0`;提供其中任一名称的上游都能以自己的名称服务这个别名,并互为备用。客户端的模型列表里能看到别名,回答里的模型名写成客户端请求的名称,请求记录则保留每家上游实际收到的模型。同一个 Claude 模型在官方 API、Bedrock、Vertex 或 OpenRouter 上名称不同时,别名标签和新建别名对话框会建议合并。也可以在上游的模型列表里直接起别名。对话框列出每家上游将收到的名称,名称会接管另一家上游的同名模型时给出提示。密钥的可见模型允许某个上游模型时,列有它的别名也可以使用。 -「体检」标签按最近 24 小时、7 天、30 天或自定义区间对照各个上游:请求数与失败率;回答中的模型名与发出的是否相同;各上游报告的输入 token 相对网关本地估算的倍数,与服务同一模型的其他上游对照;后续轮次中输入从提示缓存读取的比例,同样与其他上游对照;以及首 token 时间与生成速度的中位数。每项数字都附样本数,两边样本都足够时才标出偏差。 +每个上游还会与服务同一模型的其他上游对照最近 7 天的数据,偏差直接标在上游列表中它的名称旁:回答中的模型名与发出的不同,标为「模型名不符」;上游报告的输入 token 相对网关本地估算的倍数明显高于或低于其他上游,标为「输入 token 偏多」或「输入 token 偏少」;后续轮次中从提示缓存读取的输入比例偏低,标为「缓存读取偏低」。悬停可以查看依据与样本数,两边样本都足够时才标出偏差。 API 密钥和请求头的值可以写成 `${变量名}`,读取系统环境变量。macOS 上读的是登录 shell 里的环境变量,`~/.zshrc` 等文件中 `export` 的变量都会生效,Linux 同理(`~/.bashrc`、`~/.profile` 等);Windows 上读的是系统设置里配置的环境变量。修改变量后,重新打开应用即可生效。代理相关的变量(`HTTPS_PROXY` 等)和 `PATH` 不会被读取。 @@ -47,7 +51,7 @@ API 密钥和请求头的值可以写成 `${变量名}`,读取系统环境变 ## 路由与故障转移 -每把密钥使用一条路由,未指定的使用默认路由。路由由按顺序匹配的规则组成。规则的条件包括模型、密钥、客户端的 API 格式、输入 token 数、`max_tokens`、工具数量、图片、扩展思考、流式、提示缓存以及辅助请求的类型;命中后把请求交给某个上游、策略组或指定模型,或拒绝请求,也可以改写模型、`max_tokens` 或扩展思考。指定模型写明上游和它的某个模型,可按顺序列出备用,模型名原样发出;要让某个客户端把一个模型当另一个模型的名称使用,在这个客户端的密钥上写规则即可,不会改变这个名称对其他客户端的含义。条件写上游模型名时,也匹配列有它的别名。设置 `max_tokens` 可以限制回答的长度:上游到达上限时自行停止,不会报错。策略组把多个上游放在同一个名字下,并决定尝试的先后:按顺序、手动选择、轮询、延迟最低优先或费用最低优先。一次尝试失败时,请求转到下一个上游;同一会话默认保持在同一个上游上,以便提示缓存持续命中。页面顶部的路由图显示每把密钥经过的路由、策略组和上游。 +每把密钥使用一条路由,未指定的使用默认路由。路由由按顺序匹配的规则组成。规则的条件包括模型、密钥、客户端的 API 格式、输入 token 数、`max_tokens`、工具数量、图片、扩展思考、流式、提示缓存以及辅助请求的类型;命中后把请求交给某个上游、策略组或指定模型,或拒绝请求,也可以改写模型、`max_tokens` 或扩展思考。指定模型写明上游和它的某个模型,可按顺序列出备用,模型名原样发出;要让某个客户端把一个模型当另一个模型的名称使用,在这个客户端的密钥上写规则即可,不会改变这个名称对其他客户端的含义。条件写上游模型名时,也匹配列有它的别名。设置 `max_tokens` 可以限制回答的长度:上游到达上限时自行停止,不会报错。策略组把多个上游放在同一个名字下,并决定尝试的先后:按顺序、手动选择、轮询、延迟最低优先或费用最低优先。轮询策略组的每个成员有 1 到 100 的权重,即长期分到的请求比例;「分配依据」可以在权重之上再乘以各上游的速度(首 token 时间)、稳定性(最近的成功率)或两者,进行中的对话仍留在原上游。一次尝试失败时,请求转到下一个上游;同一会话默认保持在同一个上游上,以便提示缓存持续命中。在「设置 › 故障转移」中打开「开头超时时转到下一个上游」后,流式回答在等待时长内没有任何内容时,只要另有上游可以接下,请求也会转过去;最后一个上游照常等待。页面顶部的路由图显示每把密钥经过的路由、策略组和上游。 客户端自行发出的辅助请求(连通性检查、预热、生成标题、话题识别、输入建议)可以由网关在本地应答而不产生费用,也可以转发。转发的辅助请求和普通请求一样经过路由规则,规则可以按辅助请求的类型把它们分流到费用更低的上游。 @@ -81,7 +85,7 @@ MCP 页管理客户端从自己的配置文件中加载的内容,这些内容 ## 设置 -设置页分为七节。「连接」列出本机 core 和已保存的远程 core,详见[连接远程 core](/zh-CN/docs/lite/remote-core)。「通用」设置语言、外观、菜单栏显示的内容(仅 macOS)、开机启动、提醒以系统通知发送、仅在应用内显示还是关闭,并可让设为不再显示的引导提示重新显示。「网关监听」设置网关的访问范围(仅本机、所选网卡所在的局域网或所有网卡)、端口和放行网段。「故障转移」设置上游失败后暂停多久、请求交给下一个上游:连续失败几次后暂停、暂停多长,以及余额不足、额度用完和限流时各自的暂停时长,以及流式回答等待开头的时长。「日志保留」分别设置请求报文与请求记录的保留天数,以及报文的空间上限。「关于」显示版本、检查更新,并可生成诊断包,其中的密钥与地址均已脱敏。「卸载」还原所有已接管的客户端并取消开机启动,应在删除应用之前执行。 +设置页分为七节。「连接」列出本机 core 和已保存的远程 core,详见[连接远程 core](/zh-CN/docs/lite/remote-core)。「通用」设置语言、外观、菜单栏显示的内容(仅 macOS)、开机启动、提醒以系统通知发送、仅在应用内显示还是关闭,并可让设为不再显示的引导提示重新显示。「网关监听」设置网关的访问范围(仅本机、所选网卡所在的局域网或所有网卡)、端口和放行网段。「故障转移」设置上游失败后暂停多久、请求交给下一个上游:连续失败几次后暂停、暂停多长,以及余额不足、额度用完和限流时各自的暂停时长;还设置流式回答等待开头的时长、开头超时时是否转到下一个上游(默认关闭),以及上游并发已满或密钥的分钟、小时上限用满时请求合计最多等待的时长(默认 30 秒)。「日志保留」分别设置请求报文与请求记录的保留天数,以及报文的空间上限。「关于」显示版本、检查更新,并可生成诊断包,其中的密钥与地址均已脱敏。「卸载」还原所有已接管的客户端并取消开机启动,应在删除应用之前执行。 ## 菜单栏、系统托盘与通知 @@ -93,4 +97,4 @@ Windows 上图标位于通知区域:悬停显示网关状态与今日 token、 Linux 上图标位于系统托盘:点击打开同一份菜单,第一项为「打开主界面」,额度条同样改为文字。 -以下情况会发送系统通知(macOS 与 Windows 使用原生通知,Linux 通过桌面环境的通知服务):网关停止转发或反复重启、与远程 core 的连接断开、订阅额度用完、账号登录失效或上游拒绝当前凭据、代理不通、配置文件未通过校验、工具调用命中切断类规则、客户端配置中出现可疑内容。自动检查到新版本时也以系统通知告知(见[更新](/zh-CN/docs/lite/install#更新))。上游无法连接时通常由回退上游承接,因此只记录在应用内。提醒可以整体设为系统通知、仅在应用内显示或关闭。标为已读的提醒不再计入铃铛上的数字,但在问题解决或清空列表之前仍留在列表中。 +以下情况会发送系统通知(macOS 与 Windows 使用原生通知,Linux 通过桌面环境的通知服务):网关停止转发或反复重启、与远程 core 的连接断开、订阅额度用完、账号登录失效或上游拒绝当前凭据、代理不通、配置文件未通过校验、工具调用命中切断类规则、密钥的每日、每周或每月用量接近或达到上限、客户端配置中出现可疑内容。自动检查到新版本时也以系统通知告知(见[更新](/zh-CN/docs/lite/install#更新))。上游无法连接时通常由回退上游承接,因此只记录在应用内。提醒可以整体设为系统通知、仅在应用内显示或关闭。标为已读的提醒不再计入铃铛上的数字,但在问题解决或清空列表之前仍留在列表中。 diff --git a/src/content/docs-lite/zh-CN/keep-api-keys-from-relays.md b/src/content/docs-lite/zh-CN/keep-api-keys-from-relays.md index 755237b..51e0008 100644 --- a/src/content/docs-lite/zh-CN/keep-api-keys-from-relays.md +++ b/src/content/docs-lite/zh-CN/keep-api-keys-from-relays.md @@ -4,7 +4,7 @@ ## 准备 -- 已[安装](/zh-CN/lite/#install) ThinkWatch Lite,并已接管客户端,请求经过网关。本文按 2026.10.4 版编写。 +- 已[安装](/zh-CN/lite/#install) ThinkWatch Lite,并已接管客户端,请求经过网关。本文按 2026.10.6 版编写。 - 两项防护出厂均为「观察」:只记录检出的内容,不做任何改动。「关闭」档不检查也不记录。 ## 步骤 @@ -28,7 +28,7 @@ - **哪些调用会被切断。**下载或解码后执行代码、外发环境变量或凭据文件、读取私钥或云凭据、把凭据发往陌生主机、写入启动项或安装定时任务的调用。删除主目录或根目录、开放全部写权限、上传本地文件三类内置规则出厂只记录。 - **中转站仍能看到的内容。**应用中为该上游配置的密钥,以及请求的其余内容,例如代码、文件内容和对话,均按原样发给中转站。客户端连接网关所用的密钥不会转发;客户端的身份信息也不转发,除非在该上游上打开「转发客户端身份」。 - **只按规则匹配。**只有命中规则的值才会被替换;没有固定前缀的凭据(例如 AWS 私有访问密钥)需要自定义规则。写法不在任何规则之内的危险命令,工具调用审查拦不住。 -- **看不到网关之外。**中转站在自己的服务器上做了什么,例如留存请求、换用其他模型作答,各项防护无从得知。上游页的「体检」标签可以显示后者的迹象,但不能证明。 +- **看不到网关之外。**中转站在自己的服务器上做了什么,例如留存请求、换用其他模型作答,各项防护无从得知。上游页在上游名称旁标出的「模型名不符」等偏差可以显示后者的迹象,但不能证明。 - **动手的代价。**请求内容改变后,上游的提示缓存可能无法命中;「切断」档下误判时,回答会在该调用处中断。 相关文档:[功能详解](/zh-CN/docs/lite/features/#安全)(含第三项防护内容过滤)、[安装与更新](/zh-CN/docs/lite/install/)。 diff --git a/src/content/docs/en/configuration.md b/src/content/docs/en/configuration.md index d9be525..fa39414 100644 --- a/src/content/docs/en/configuration.md +++ b/src/content/docs/en/configuration.md @@ -22,6 +22,7 @@ PostgreSQL connection string. ThinkWatch requires PostgreSQL 15 or later. - In production, use `sslmode=require` to enforce encrypted connections: `postgres://user:pass@host:5432/db?sslmode=require` - Avoid embedding passwords in URLs checked into version control. Use a secrets manager. - The database user needs permissions to create tables (for migrations) or should have migrations applied separately. +- Connect directly or through a pooler in session mode, not transaction mode: schema setup at startup holds a session-level advisory lock. See [External PostgreSQL and Redis](/docs/deployment-guide#46-external-postgresql-and-redis). --- @@ -38,6 +39,7 @@ Redis connection string. Used for rate limiting, OIDC state/nonce storage, and s **Security notes:** - In production, enable Redis authentication: `redis://:yourpassword@host:6379` - For TLS-enabled Redis, use the `rediss://` scheme: `rediss://:password@host:6380` +- For a Redis Cluster, use `redis-cluster://` (`rediss-cluster://` over TLS). For a certificate signed by a private CA, set `REDIS_CA_CERT` to the CA's PEM file. See [External PostgreSQL and Redis](/docs/deployment-guide#46-external-postgresql-and-redis). --- diff --git a/src/content/docs/en/deployment-guide.md b/src/content/docs/en/deployment-guide.md index 8e2e99d..fefdd27 100644 --- a/src/content/docs/en/deployment-guide.md +++ b/src/content/docs/en/deployment-guide.md @@ -363,11 +363,11 @@ No authentication is needed — the packages are public. helm install think-watch deploy/helm/think-watch \ --set secrets.jwtSecret=$(openssl rand -hex 32) \ --set secrets.encryptionKey=$(openssl rand -hex 32) \ - --set secrets.databaseUrl="postgres://thinkwatch:password@postgres:5432/think_watch" \ - --set secrets.redisUrl="redis://:password@redis:6379" \ --set config.corsOrigins="https://console.internal.example.com" ``` +The chart runs PostgreSQL, Redis and ClickHouse alongside the server. Secrets left empty are generated on the first install and kept across upgrades. To use managed databases instead, see [4.6 External PostgreSQL and Redis](#46-external-postgresql-and-redis). + To deploy a specific image tag: ```bash @@ -410,10 +410,10 @@ ingress: ### 4.4 External Secrets -For production, use the External Secrets Operator instead of passing secrets via `--set`: +For production, use the External Secrets Operator instead of passing secrets via `--set`. The chart always creates its own Secret, `-secrets` (`think-watch-secrets` for the release above), and the server reads its variables from it, so the ExternalSecret merges values into that Secret instead of creating one: install the chart first, then apply: ```yaml -apiVersion: external-secrets.io/v1beta1 +apiVersion: external-secrets.io/v1 kind: ExternalSecret metadata: name: think-watch @@ -423,22 +423,19 @@ spec: name: vault-backend # or aws-secrets-manager, etc. kind: SecretStore target: - name: think-watch-secrets + name: think-watch-secrets # -secrets, created by the chart + creationPolicy: Merge data: - - secretKey: jwt-secret + - secretKey: JWT_SECRET remoteRef: key: think-watch/jwt-secret - - secretKey: encryption-key + - secretKey: ENCRYPTION_KEY remoteRef: key: think-watch/encryption-key - - secretKey: database-url - remoteRef: - key: think-watch/database-url - - secretKey: redis-url - remoteRef: - key: think-watch/redis-url ``` +On later upgrades the chart reads `JWT_SECRET` and `ENCRYPTION_KEY` back from the Secret, so the values from the store are kept; leave `secrets.jwtSecret` and `secrets.encryptionKey` unset, since values set there take precedence over the Secret. The server reads the Secret only at start, so restart it once after the first sync: `kubectl rollout restart deployment/think-watch-server`. `DATABASE_URL` and `REDIS_URL` cannot come from the store, because the chart writes them from `postgres.externalUrl` and `redis.externalUrl` on every upgrade (see [4.6](#46-external-postgresql-and-redis)). + ### 4.5 Horizontal Pod Autoscaling Enable HPA in `values.yaml`: @@ -453,6 +450,48 @@ autoscaling: The server is stateless (all state is in PostgreSQL and Redis), so it scales horizontally without issue. +### 4.6 External PostgreSQL and Redis + +A managed PostgreSQL or Redis can replace the bundled one. In the Helm chart, set `bundled: false` and `externalUrl` for that service; outside the chart, the URLs go in `DATABASE_URL` and `REDIS_URL`. + +```yaml +postgres: + bundled: false + externalUrl: postgres://user:pass@pg.example.com:5432/think_watch?sslmode=require +redis: + bundled: false + externalUrl: rediss://:pass@my-cache.example.com:6379 +``` + +**PostgreSQL behind a connection pooler.** Each server instance sets up the schema when it starts, holding a session-level advisory lock so that instances starting together take turns. The URL must therefore reach PostgreSQL directly or through a pooler in session mode, never one in transaction mode (PgBouncer `pool_mode = transaction`, or the transaction-mode port of a managed pooler): there the lock can stay held on a server connection the pooler keeps, and every later start waits for it indefinitely. Schema setup uses the connections of `DATABASE_URL`; there is no separate URL for it. + +**Redis Cluster.** Give the URL the `redis-cluster://` scheme and name one node or more; the server finds the rest of the cluster from them: + +```text +redis-cluster://:pass@redis-0.redis:6379?node=redis-1.redis:6379&node=redis-2.redis:6379 +``` + +Every node must be reachable from the server at the address it announces to the cluster (`cluster-announce-ip` / `-port`). A cluster has only database 0, so the URL names no `/`. + +**Redis over TLS.** Managed Redis services usually require TLS, for example ElastiCache with in-transit encryption, Upstash, Azure Cache for Redis and Redis Cloud. Use the `rediss://` scheme, or `rediss-cluster://` for a cluster, with the host name the service gives and the port it uses for TLS (a URL without a port means `6379`; Azure Cache for Redis uses `6380`). The certificate is checked against the public CAs and the host name in the URL; in a cluster, each node's certificate must name the address that node announces. + +A Redis whose certificate a private CA signed needs that CA: set `REDIS_CA_CERT` to the path of its PEM certificate, and the server then trusts only the certificates in that file for Redis. With the Helm chart, put the certificate in a Secret and name it under `redis.caSecret`; the chart mounts it and sets `REDIS_CA_CERT`: + +```bash +kubectl -n thinkwatch create secret generic redis-ca --from-file=ca.crt=./ca.crt +``` + +```yaml +redis: + bundled: false + externalUrl: rediss://:pass@redis.internal:6379 + caSecret: + name: redis-ca + key: ca.crt +``` + +The file is read at start, so restart the server after changing it. Client certificates (mutual TLS) are not supported: such a Redis needs `tls-auth-clients no` and password authentication. With `networkPolicy.enabled`, ports that no URL names, such as cluster nodes announcing other ports, go in `networkPolicy.extraEgress`. + --- ## 5. SSL/TLS diff --git a/src/content/docs/zh-CN/configuration.md b/src/content/docs/zh-CN/configuration.md index d195c02..f62d4d0 100644 --- a/src/content/docs/zh-CN/configuration.md +++ b/src/content/docs/zh-CN/configuration.md @@ -22,6 +22,7 @@ PostgreSQL 连接字符串。ThinkWatch 需要 PostgreSQL 15 或更高版本。 - 在生产环境中,使用 `sslmode=require` 强制加密连接:`postgres://user:pass@host:5432/db?sslmode=require` - 避免在提交到版本控制的 URL 中嵌入密码。请使用密钥管理器。 - 数据库用户需要创建表的权限(用于迁移),或者应单独应用迁移。 +- 直连数据库,或经由会话模式的连接池,不能使用事务模式:启动时建立数据库结构需要持有会话级 advisory lock。详见[外部 PostgreSQL 与 Redis](/zh-CN/docs/deployment-guide#46-外部-postgresql-与-redis)。 --- @@ -38,6 +39,7 @@ Redis 连接字符串。用于速率限制、OIDC 状态/nonce 存储和会话 **安全说明:** - 在生产环境中,启用 Redis 认证:`redis://:yourpassword@host:6379` - 对于启用了 TLS 的 Redis,使用 `rediss://` 协议:`rediss://:password@host:6380` +- Redis Cluster 使用 `redis-cluster://` 协议(TLS 为 `rediss-cluster://`)。证书由私有 CA 签发时,把 `REDIS_CA_CERT` 设为该 CA 的 PEM 文件路径。详见[外部 PostgreSQL 与 Redis](/zh-CN/docs/deployment-guide#46-外部-postgresql-与-redis)。 --- diff --git a/src/content/docs/zh-CN/deployment-guide.md b/src/content/docs/zh-CN/deployment-guide.md index 9b16f4e..f734d19 100644 --- a/src/content/docs/zh-CN/deployment-guide.md +++ b/src/content/docs/zh-CN/deployment-guide.md @@ -363,11 +363,11 @@ ghcr.io/thinkwatch/think-watch-web: helm install think-watch deploy/helm/think-watch \ --set secrets.jwtSecret=$(openssl rand -hex 32) \ --set secrets.encryptionKey=$(openssl rand -hex 32) \ - --set secrets.databaseUrl="postgres://thinkwatch:password@postgres:5432/think_watch" \ - --set secrets.redisUrl="redis://:password@redis:6379" \ --set config.corsOrigins="https://console.internal.example.com" ``` +Chart 会在服务器旁一并运行 PostgreSQL、Redis 和 ClickHouse。留空的密钥在首次安装时自动生成,升级时保留。改用托管数据库见 [4.6 外部 PostgreSQL 与 Redis](#46-外部-postgresql-与-redis)。 + 如需部署特定版本: ```bash @@ -410,10 +410,10 @@ ingress: ### 4.4 外部密钥管理 -在生产环境中,使用 External Secrets Operator 而非通过 `--set` 传递密钥: +在生产环境中,使用 External Secrets Operator 而非通过 `--set` 传递密钥。Chart 总会创建自己的 Secret,即 `-secrets`(上文的 release 对应 `think-watch-secrets`),服务器从中读取环境变量,因此 ExternalSecret 应把值合并进这个 Secret,而不是另建一个:先安装 Chart,再应用: ```yaml -apiVersion: external-secrets.io/v1beta1 +apiVersion: external-secrets.io/v1 kind: ExternalSecret metadata: name: think-watch @@ -423,22 +423,19 @@ spec: name: vault-backend # or aws-secrets-manager, etc. kind: SecretStore target: - name: think-watch-secrets + name: think-watch-secrets # -secrets, created by the chart + creationPolicy: Merge data: - - secretKey: jwt-secret + - secretKey: JWT_SECRET remoteRef: key: think-watch/jwt-secret - - secretKey: encryption-key + - secretKey: ENCRYPTION_KEY remoteRef: key: think-watch/encryption-key - - secretKey: database-url - remoteRef: - key: think-watch/database-url - - secretKey: redis-url - remoteRef: - key: think-watch/redis-url ``` +此后升级时,Chart 会从该 Secret 读回 `JWT_SECRET` 和 `ENCRYPTION_KEY`,密钥库中的值得以保留;`secrets.jwtSecret` 与 `secrets.encryptionKey` 须留空,因为在其中设置的值优先于 Secret。服务器只在启动时读取 Secret,因此首次同步后需要重启一次:`kubectl rollout restart deployment/think-watch-server`。`DATABASE_URL` 和 `REDIS_URL` 不能来自密钥库:每次升级时 Chart 都按 `postgres.externalUrl` 和 `redis.externalUrl` 重新写入(见 [4.6](#46-外部-postgresql-与-redis))。 + ### 4.5 水平 Pod 自动伸缩 在 `values.yaml` 中启用 HPA: @@ -453,6 +450,48 @@ autoscaling: 服务器是无状态的(所有状态存储在 PostgreSQL 和 Redis 中),因此可以无障碍地水平扩展。 +### 4.6 外部 PostgreSQL 与 Redis + +托管的 PostgreSQL 或 Redis 可以代替 Chart 自带的实例。使用 Helm Chart 时,为该服务设置 `bundled: false` 和 `externalUrl`;不使用 Chart 时,地址写在 `DATABASE_URL` 和 `REDIS_URL` 中。 + +```yaml +postgres: + bundled: false + externalUrl: postgres://user:pass@pg.example.com:5432/think_watch?sslmode=require +redis: + bundled: false + externalUrl: rediss://:pass@my-cache.example.com:6379 +``` + +**经连接池访问 PostgreSQL。**每个服务器实例启动时都会建立数据库结构,期间持有一把会话级 advisory lock,使同时启动的实例依次进行。因此地址必须直连 PostgreSQL,或经由会话模式的连接池,不能使用事务模式的连接池(PgBouncer 的 `pool_mode = transaction`,或托管连接池的事务模式端口):在事务模式下,这把锁可能留在连接池保留的服务端连接上,此后每次启动都会一直等待。建立数据库结构使用的就是 `DATABASE_URL` 的连接,没有单独的地址。 + +**Redis Cluster。**地址使用 `redis-cluster://` 协议,列出一个或多个节点,服务器由此找到集群的其余节点: + +```text +redis-cluster://:pass@redis-0.redis:6379?node=redis-1.redis:6379&node=redis-2.redis:6379 +``` + +每个节点都必须能从服务器按它向集群公布的地址(`cluster-announce-ip` / `-port`)访问。集群只有 0 号数据库,因此地址中不写 `/`。 + +**经 TLS 访问 Redis。**托管 Redis 服务通常要求 TLS,例如开启传输加密的 ElastiCache、Upstash、Azure Cache for Redis 和 Redis Cloud。地址使用 `rediss://` 协议,集群使用 `rediss-cluster://`,主机名写服务给出的名称,端口写它的 TLS 端口(地址中不写端口即为 `6379`;Azure Cache for Redis 为 `6380`)。证书按公共 CA 和地址中的主机名校验;集群中每个节点的证书都必须包含该节点公布的地址。 + +证书由私有 CA 签发的 Redis 需要提供该 CA:把 `REDIS_CA_CERT` 设为它的 PEM 证书路径,服务器连接 Redis 时便只信任该文件中的证书。使用 Helm Chart 时,把证书放进 Secret,并在 `redis.caSecret` 中指定;Chart 会挂载它并设置 `REDIS_CA_CERT`: + +```bash +kubectl -n thinkwatch create secret generic redis-ca --from-file=ca.crt=./ca.crt +``` + +```yaml +redis: + bundled: false + externalUrl: rediss://:pass@redis.internal:6379 + caSecret: + name: redis-ca + key: ca.crt +``` + +该文件在启动时读取,修改后需要重启服务器。不支持客户端证书(双向 TLS):这类 Redis 需设置 `tls-auth-clients no`,并使用密码认证。启用 `networkPolicy.enabled` 时,地址中没有写到的端口(例如集群节点公布的其他端口)写在 `networkPolicy.extraEgress` 中。 + --- ## 5. SSL/TLS diff --git a/src/data/core-docs/config.md b/src/data/core-docs/config.md index 9066764..a84134c 100644 --- a/src/data/core-docs/config.md +++ b/src/data/core-docs/config.md @@ -308,6 +308,7 @@ gateway. A key is an identity. Limits, model scope and route are per key. | `route` | string | — | Name of the route requests with this key take. Unset: `default_route`. | | `client` | string | — | The client this key was made for (`claude-code`, `codex`, …), recorded when the desktop app points a client at the gateway. A client has at most one. | | `disabled` | bool | `false` | Refuse every request made with this key, and keep the key. | +| `limits` | list of [`clients[].limits[]`](#cfg-clients-limits) | — | Usage limits: requests, tokens or cost per minute, hour, day, week or month. A request has to pass every one. Unset: no limit. | ```yaml @@ -321,6 +322,63 @@ clients: route: cheap ``` +#### `clients[].limits` + +Usage limits for a key: at most so many requests, tokens or dollars per +minute, hour, day, week or month. Each entry counts one of the three; a key +can have several, and a request has to pass every one. + +`minute` and `hour` are rolling: the last 60 seconds, the last 60 minutes. +When one is used up, a request waits for the next free slot if it frees +within `failover.slot_wait_secs` (30 seconds by default), and is refused +otherwise. Any later wait for a busy upstream comes out of the same time. +`day`, `week` and `month` follow the calendar in the time zone of the machine +twcore runs on and start again at midnight, on Monday and on the 1st. When one is used up, requests are refused until it starts again. If the +machine's time zone changes, the current day, week and month are added up +again from the request records. + +A refused request gets HTTP 429 in the client's own error format, naming the +key, the limit, the amount used and when it resets, and it shows in the +traffic list. A request that never reaches an upstream counts toward no +limit: one refused by a rule, the content filter or a limit, or turned away +because every upstream was at its `max_concurrent`. Cost is what is recorded +for each request, so a model without a price and an upstream with +`billing: free` count as $0. A request still +running counts with an estimate of its input until it is recorded. After a +restart, the day, week and month are added up again from the request records, +so the records have to cover the period: `retention.row_days` of at least 1 +for a daily limit, 7 for a weekly one and 31 for a monthly one. Minute and +hour limits start empty. + +On a Responses WebSocket connection, each `response.create` is a request of +its own: it is recorded with its usage and cost and checked against these +limits, and a refused one is answered with `response.failed` while the +connection stays open. A Realtime connection (`/v1/realtime`) is one request: +it is checked against the limits when it opens, and the tokens and cost of +all its answers count when it closes. + + + + +| Field | Type | Default | Description | +|---|---|---|---| +| `per` | `minute` \| `hour` \| `day` \| `week` \| `month` | **required** | The period. `minute` and `hour` are rolling (the last 60 seconds, the last 60 minutes); `day`, `week` and `month` start again at local midnight, on Monday and on the 1st. | +| `requests` | integer | — | At most this many requests. Token counts, answers the gateway gives itself and requests that never reach an upstream do not count. | +| `tokens` | integer | — | At most this many tokens: uncached input, cache writes and output. | +| `cost` | number | — | At most this much, in US dollars, as recorded for each request; at least 0.01. Models without a price and upstreams with `billing: free` count as 0. | +| `cache_reads` | bool | `false` | Count cache reads too. Only for a `tokens` limit. | + + +```yaml +clients: + - name: build-server + key: tw-q8r2s4t6u8v2w4x6y8z2a4b6 + limits: + - { per: minute, requests: 30 } + - { per: day, cost: 5 } + - { per: month, tokens: 20000000, cache_reads: true } +``` + ### `providers` Upstreams: the APIs requests are forwarded to. @@ -344,6 +402,8 @@ Upstreams: the APIs requests are forwarded to. | `models_only` | list of strings | — | Use only these of the upstream's models, as ids or globs. Others are not listed and are not routed here. Unset: all of them. Empty is refused; use `disabled`. | | `billing` | `per-token` \| `free` | `per-token` | `per-token`: cost is usage times the price in the upstream's price sheet, subscription accounts included. `free`: cost is recorded as 0. | | `pricing` | string | — | Name of a price sheet under `pricing.sheets`. Unset: the default price table. | +| `model_specs` | map of model id → [`providers[].model_specs.*`](#cfg-providers-model_specs) | `{}` | Context window and output limit of single models of this upstream, written by hand, by exact model id. They take precedence over the price table: for models it does not know, or gets wrong. | +| `max_concurrent` | integer | — | Most requests sent to this upstream at the same time, from 1 to 1000. When it is full, a conversation that stays on it waits for a free slot and other requests go to the next upstream; see `failover.slot_wait_secs`. Unset: no limit. | | `disabled` | bool | `false` | Take the upstream out of routing and out of the model list, and keep its configuration. | @@ -367,6 +427,7 @@ providers: proxy: office models_only: [gpt-4.1*, o3] pricing: relay-discount + max_concurrent: 4 - name: local base_url: http://127.0.0.1:11434/v1 @@ -387,6 +448,18 @@ A ChatGPT account upstream (`protocol: chatgpt`) takes only the credential the desktop app obtains by signing in; it cannot be written by hand. Claude and Google subscription sign-ins are not supported; use an API key. +Some relays and accounts accept only a few requests at a time and refuse the +rest. `max_concurrent` keeps the gateway within that number: a request takes a +slot on the upstream when it is sent, and gives it back when the answer has +been passed on in full or the client has gone. When the upstream is full, a +conversation that stays on it to reuse its prompt cache waits for a slot; +any other request goes straight to the next upstream. How long a request +waits is `failover.slot_wait_secs`. Waiting is not a failure: the upstream is +not set aside. Requests that only count tokens do not take a slot. On a +Responses WebSocket connection, each `response.create` takes a slot from when +it is sent until its answer ends, and so does a key's `max_concurrent`; an +idle connection takes none. + #### `providers[].oauth` @@ -470,6 +543,40 @@ without them requests are still forwarded, and `models` can list the models by hand. For a VPC endpoint or a proxy, write its address in `base_url` and the region in `aws.region`; the model list is asked of that address too. +#### `providers[].model_specs` + +A model's context window and output limit come from the price table. A relay's +own models are often missing from it, and now and then it is wrong. Write the +numbers here, for this upstream and by the exact id in its model list. A value +written here takes precedence over the price table; one left out still comes +from it. At least one of the two is written, and neither can be 0. + +The same numbers are used everywhere: in `/v1/models` for every client format, +for an alias this upstream serves, when the gateway judges whether a +conversation still fits the model it is on, and as the output limit of a +request converted to Anthropic that does not set one. When several upstreams +offer the same model, `/v1/models` describes it by the first of them in +`providers`. + + + + +| Field | Type | Default | Description | +|---|---|---|---| +| `context_window` | integer | — | Context window: the most tokens a request can take in. Unset: the price table's. | +| `max_output_tokens` | integer | — | The most tokens an answer can have. Unset: the price table's. | + + +```yaml +providers: + - name: relay + base_url: https://relay.example.com/v1 + protocol: openai-chat + model_specs: + glm-5-air: { context_window: 128000, max_output_tokens: 16384 } + claude-sonnet-4-5: { context_window: 1000000 } +``` + ### `proxies` Outbound proxies. Different upstreams often need different ones, so there @@ -888,11 +995,41 @@ go straight to the next candidate. How long depends on the reason the upstream gives: an insufficient balance waits for a top-up, a used-up quota waits until the moment the upstream says it resets, and a rate limit usually passes within seconds. A request with a single candidate is never affected. +On a Responses WebSocket connection, each `response.create` counts here like +a request: one that fails because of the upstream before any content arrives +counts as a failure, and so does a connection the upstream refuses or that +cannot be made. Before the first content of a streamed answer reaches the client, an error the upstream sends in the stream moves the request to the next candidate, the same as an error status would. +An upstream can also be slow to start: it accepts the request and then sends +nothing for a long time. With `next_on_slow_start`, the request moves on to the +next candidate when no content has arrived `stream_start_wait_secs` after it +was sent. It is off by default, because models that think before they write +can take long to start; with it on, wait 30 seconds or more. The last +candidate always waits, and the upstream given up on is not set aside. A +candidate that is at its `max_concurrent` at that moment does not count as a +next one: the slow upstream keeps the request. + +```yaml +failover: + stream_start_wait_secs: 30 + next_on_slow_start: true +``` + +When upstreams are at their `max_concurrent`, a request waits for a free slot +for at most `slot_wait_secs` in all. The same time also covers waiting for a +key's `minute` or `hour` limit, so a request never waits longer than +`slot_wait_secs` for the two together. An ongoing conversation waits for the +upstream it stays on and, if no slot frees in time, moves on to the next one, +where its cache starts over. A new conversation skips a full upstream at once. +When every candidate is full, the request waits for whichever frees first; if +none does, the client gets a 429 with `Retry-After` saying the upstreams are +busy. If an upstream did receive the request and failed, the client gets that +failure instead. + @@ -905,6 +1042,8 @@ the same as an error status would. | `quota_pause_secs` | integer | `3600` | Seconds to set aside an upstream whose quota is used up when it does not say when the quota resets. When it does, the upstream is set aside until then. | | `rate_limit_max_pause_secs` | integer | `3600` | A rate-limited upstream is set aside for the time its `Retry-After` gives, at most this many seconds. Without `Retry-After` it counts as a failure without a stated reason. | | `stream_start_wait_secs` | integer | `15` | Seconds to hold a streamed answer until its first content arrives. An error before then moves the request to the next upstream; after this long, what has arrived is passed on. From 1 to 120. | +| `next_on_slow_start` | bool | `false` | When a streamed answer still has no content `stream_start_wait_secs` after the request was sent, give up on that upstream and send the request to the next one. The last upstream always waits. The upstream given up on is not set aside. Needs `stream_start_wait_secs` of at least 5. | +| `slot_wait_secs` | integer | `30` | Seconds a request waits in all, counted once the key's own `max_concurrent` lets it in: for a key's `minute` or `hour` limit to free up, and for a free slot on upstreams at their `max_concurrent`. A key limit that does not free up in time refuses the request; without an upstream slot in time it goes to the next upstream, or, when every candidate is full, is answered with 429. `0`: never wait. From 0 to 300. | ### `aliases` @@ -952,14 +1091,45 @@ group with `to`. | Field | Type | Default | Description | |---|---|---|---| | `name` | string | **required** | Name of the group; unique, and not the name of an upstream. | -| `type` | `fallback` \| `select` \| `load-balance` \| `url-test` \| `cheapest` | `fallback` | `fallback`: the first healthy member, in order. `select`: the member named in `selected`. `load-balance`: take turns between new conversations. `url-test`: the fastest by measured time to first byte. `cheapest`: the lowest input price. | -| `providers` | list of strings | **required** | Member upstreams, by name. | +| `type` | `fallback` \| `select` \| `load-balance` \| `url-test` \| `cheapest` | `fallback` | `fallback`: the first healthy member, in order. `select`: the member named in `selected`. `load-balance`: requests are shared out in proportion to the members' weights; a conversation in progress stays where it is. `url-test`: the fastest by measured time from sending a request to the first content of the answer. `cheapest`: the lowest input price. | +| `providers` | list of strings or [`groups[].providers[]`](#cfg-groups-providers) | **required** | Member upstreams, by name; not groups. Each upstream appears once in a group. In a `load-balance` group, a member can be written as `{name, weight}`. | | `selected` | string | — | For `select`: the chosen member. | +| `balance_by` | `weights` \| `latency` \| `health` \| `latency-health` | `weights` | For `load-balance`: what the members' weights are multiplied by. `weights`: nothing; the weights alone. `latency`: faster upstreams get more. `health`: upstreams that fail less get more. `latency-health`: both. Other group types take only `weights`. | `fallback` is the default because a single user's machine has no load to spread. +In a `load-balance` group, a member can carry a weight, from 1 to 100; a +member written as just its name has weight 1. Weights set how the group's +requests are shared out: with `{ name: anthropic, weight: 7 }` and `relay`, +the official API serves seven requests in ten. Conversations in progress +stay on the upstream that answers them (see below) and count toward its +share, so the balance is kept by where new conversations start. An upstream +that is cooling down after failures, is at its `max_concurrent`, or cannot +serve a request, sits that request out, and the others share it by their +weights. A new WebSocket connection is placed the same way and counts as one +request; everything sent on it then goes to the upstream it connected to. +Other group types take no weights. + + + + +| Field | Type | Default | Description | +|---|---|---|---| +| `name` | string | **required** | The upstream, by name. A member written as just its name has weight 1. | +| `weight` | integer | `1` | The member's share of a `load-balance` group's requests, in proportion to the other members' weights. From 1 to 100. Other group types take no weight other than 1. | + + +```yaml +groups: + - name: pool + type: load-balance + providers: + - { name: anthropic, weight: 7 } + - relay +``` + Whatever the type, a conversation stays on the upstream that last answered it, so that what the upstream holds of it in its prompt cache is read again rather than paid for in full elsewhere. Within a turn (while the client sends @@ -970,7 +1140,53 @@ the conversation, and whichever upstream answered after a failover is the one it stays on. The rule a turn matched at its start also holds for the rest of that turn: rules keyed on input size or images do not move a turn halfway, unless its input no longer fits the context window of a model the rule sends -it to. `load-balance` therefore takes turns between new conversations. +it to. A `load-balance` weight is therefore the long-run share of requests: +conversations in progress stay where they are and count toward that +upstream's share. + +`balance_by` lets a `load-balance` group also look at how each upstream has +been doing lately. Each member's weight is multiplied by a factor, and the +group shares out requests by the result in the same way as above. + +- `weights` (the default): the weights alone. +- `latency`: faster upstreams get a larger share. Speed is the typical time + from sending a request to the first content of the answer, the same + measurement `url-test` uses. An upstream twice as fast as the middle of the + group has its weight multiplied by four, by at most ten and at least a + tenth. +- `health`: upstreams that fail less get a larger share. It looks at the + last 50 requests within the past 30 minutes. Server errors, rate limits, + used-up quota or balance, rejected credentials, timeouts and connection + errors count as failures; errors caused by the request itself do not, and + neither does a client that cancels, a switch away from a stream that is + slow to start, or an upstream skipped because it is at its + `max_concurrent`. An upstream that keeps failing keeps a twentieth of its + weight, so it still gets the occasional request and its recovery + is noticed; one that fails outright is set aside by + [`failover`](#cfg-failover) as before. +- `latency-health`: both factors, multiplied. + +Speed is measured on streamed answers only, from the moment the request is +sent to that upstream, so waiting and upstreams that failed before it do not +count. An upstream given up on because its stream was slow to start +(`failover.next_on_slow_start`) counts as taking the whole wait. On a +Responses WebSocket connection, each `response.create` counts as one request +for both speed and failures, its speed measured from the moment the upstream +starts answering it. + +An upstream without enough measurements yet counts as average. As with +weights alone, conversations in progress stay where they are, and new +conversations make up the difference. + +```yaml +groups: + - name: pool + type: load-balance + balance_by: latency-health + providers: + - { name: official, weight: 3 } + - relay +``` ### `routes` diff --git a/src/data/core-docs/config.zh-CN.md b/src/data/core-docs/config.zh-CN.md index a0b9ad7..49507c7 100644 --- a/src/data/core-docs/config.zh-CN.md +++ b/src/data/core-docs/config.zh-CN.md @@ -217,6 +217,7 @@ listen: | `route` | 字符串 | — | 这把密钥的请求走哪条路由。不写:`default_route`。 | | `client` | 字符串 | — | 这把密钥是为哪个客户端生成的(`claude-code`、`codex` 等),由桌面应用接管客户端时写入。一个客户端最多一把。 | | `disabled` | 布尔 | `false` | 拒绝使用这把密钥的所有请求,密钥本身保留。 | +| `limits` | 对象列表,见 [`clients[].limits[]`](#cfg-clients-limits) | — | 用量上限:每分钟、每小时、每天、每周或每月的请求数、token 数或费用。请求要通过每一条。不写:不限。 | ```yaml @@ -230,6 +231,52 @@ clients: route: cheap ``` +#### `clients[].limits` + +一把密钥的用量上限:每分钟、每小时、每天、每周或每月最多多少个请求、多少 token、多少美元。 +每一条只数其中一种;一把密钥可以有好几条,请求要通过每一条。 + +`minute`、`hour` 是滚动的:最近 60 秒、最近 60 分钟。用满之后,下一个空位在 +`failover.slot_wait_secs`(默认 30 秒)之内空出来,请求就等它;等不到就拒绝。之后若还要等 +上游空出位置,用的是同一段时间里剩下的部分。`day`、`week`、`month` 按 twcore 所在机器的 +本地时区算,在零点、周一零点、每月一号零点重新开始;用满之后,到重新开始之前的请求一律拒绝。 +机器换了时区时,当前这一天、这一周、这个月的用量按新的时区从请求记录里重新加起来。 + +被拒的请求收到 HTTP 429,错误格式和客户端自己的一致,写明是哪把密钥、哪一条上限、用了多少、 +什么时候重置;流量列表里也有这一条。没有发到上游的请求不计入任何上限:被规则、内容过滤或 +上限拒绝的,以及因为上游都已达到 `max_concurrent` 而被退回的。费用按每个请求记下的费用算,所以没有价格的模型、 +`billing: free` 的上游算 0。还在进行的请求先按输入的估算计入,记下之后换成实际用量。 +重启之后,这一天、这一周、这个月的用量从请求记录里重新加起来,所以请求记录要留够那一期: +设了按天的上限时 `retention.row_days` 至少 1,按周至少 7,按月至少 31;按分钟、按小时的 +上限从零开始。 + +Responses 的 WebSocket 连接上,每个 `response.create` 各算一个请求:带着各自的用量和费用记下, +按这些上限检查;被拒的那一个收到 `response.failed`,连接保持不断。Realtime 的连接 +(`/v1/realtime`)整条算一个请求:连上时按这些上限检查,这条连接上所有回答的 token 和费用在 +断开时计入。 + + + + +| 字段 | 类型 | 默认值 | 说明 | +|---|---|---|---| +| `per` | `minute` \| `hour` \| `day` \| `week` \| `month` | **必填** | 按多长一段时间算。`minute`、`hour` 是滚动的(最近 60 秒、最近 60 分钟);`day`、`week`、`month` 在本地时间的零点、周一零点、每月一号零点重新算。 | +| `requests` | 整数 | — | 最多这么多个请求。数 token 的请求、网关自己答的、没有发到上游的不算。 | +| `tokens` | 整数 | — | 最多这么多 token:未命中缓存的输入、写入缓存的和输出。 | +| `cost` | 数字 | — | 最多花这么多美元,按每个请求记下的费用算;至少 0.01。没有价格的模型、`billing: free` 的上游算 0。 | +| `cache_reads` | 布尔 | `false` | 把读取缓存的 token 也算进去。只有 `tokens` 上限能写。 | + + +```yaml +clients: + - name: build-server + key: tw-q8r2s4t6u8v2w4x6y8z2a4b6 + limits: + - { per: minute, requests: 30 } + - { per: day, cost: 5 } + - { per: month, tokens: 20000000, cache_reads: true } +``` + ### `providers` 上游,即请求被转发到的接口。 @@ -253,6 +300,8 @@ clients: | `models_only` | 字符串列表 | — | 只使用这家的这些模型,写 ID 或通配。范围外的模型不出现在模型列表里,也不会路由到这家。不写:全部。写空列表会被拒绝,暂停使用请用 `disabled`。 | | `billing` | `per-token` \| `free` | `per-token` | `per-token`:费用为用量乘以所选价目表中的单价,订阅账号同样如此。`free`:费用记为 0。 | | `pricing` | 字符串 | — | `pricing.sheets` 中某张价目表的名字。不写:默认价目表。 | +| `model_specs` | 映射: 模型 ID → [`providers[].model_specs.*`](#cfg-providers-model_specs) | `{}` | 手写这家上游某些模型的上下文窗口和输出上限,按模型 ID 完全匹配。写了就优先于价目表,用于价目表里没有或写错的模型。 | +| `max_concurrent` | 整数 | — | 同时发给这家的请求最多几个,取值 1 到 1000。满了的时候,留在这家的对话等空位,别的请求换下一家;等多久见 `failover.slot_wait_secs`。不写:不限。 | | `disabled` | 布尔 | `false` | 不参与路由,模型也不出现在模型列表里;配置原样保留。 | @@ -272,6 +321,7 @@ providers: proxy: office models_only: [gpt-4.1*, o3] pricing: relay-discount + max_concurrent: 4 - name: local base_url: http://127.0.0.1:11434/v1 @@ -283,6 +333,8 @@ providers: ChatGPT 账号上游(`protocol: chatgpt`)只接受桌面应用登录得到的凭据,不能手写。不支持 Claude 和 Google 的订阅登录,请使用 API 密钥。 +有的中转站和账号同时只接受几个请求,多出来的直接拒绝。`max_concurrent` 让网关守住这个数:请求发出时占用这家的一个位置,回答完整交给客户端、或者客户端断开时归还。这家满了的时候,为复用提示缓存而留在这家的对话等空位,别的请求直接换下一家。最多等多久由 `failover.slot_wait_secs` 决定。等待不算失败,这家不会因此停用。只计算 token 数的请求不占位置。Responses 的 WebSocket 连接上,每个 `response.create` 从发出起占一个位置,直到它的回答结束,密钥的 `max_concurrent` 也一样;空闲的连接不占位置。 + #### `providers[].oauth` @@ -347,6 +399,31 @@ providers: 请求转换为 Converse 格式。模型清单取自所在区域的控制面:可按需调用的基础模型、AWS 预设的推理配置(`us.anthropic.claude-…`),以及账号自己创建的应用推理配置(按调用时使用的 ARN 列出)。列出清单需要 `bedrock:ListFoundationModels` 和 `bedrock:ListInferenceProfiles` 权限;没有这两项权限时请求照常转发,可以在 `models` 中手动列出模型。使用 VPC 端点或代理时,在 `base_url` 中写它的地址,在 `aws.region` 中写区域;模型清单也向该地址获取。 +#### `providers[].model_specs` + +模型的上下文窗口和输出上限取自价目表。中转站自有的模型常常不在价目表里,价目表偶尔也会写错。这时在这里按这家上游、按它模型清单里的 ID(完全匹配)手写。写了的一项优先于价目表,没写的一项仍取价目表。两项至少写一项,都不能是 0。 + +各处用的是同一个数:各种客户端格式的 `/v1/models`、由这家上游服务的别名、网关判断一段对话是否还装得下当前模型,以及请求转换为 Anthropic 格式且没有写输出上限时补上的值。同一个模型由几家上游提供时,`/v1/models` 按 `providers` 中排在最前的那一家给出。 + + + + +| 字段 | 类型 | 默认值 | 说明 | +|---|---|---|---| +| `context_window` | 整数 | — | 上下文窗口,即一次请求最多输入多少 token。不写:取价目表的。 | +| `max_output_tokens` | 整数 | — | 一次回答最多输出多少 token。不写:取价目表的。 | + + +```yaml +providers: + - name: relay + base_url: https://relay.example.com/v1 + protocol: openai-chat + model_specs: + glm-5-air: { context_window: 128000, max_output_tokens: 16384 } + claude-sonnet-4-5: { context_window: 1000000 } +``` + ### `proxies` 出站代理。不同上游需要的代理往往不同,因此没有全局开关:由每个上游用 `proxy` 选择。 @@ -697,11 +774,32 @@ security: 上游失败后会停用一段时间,接下来的请求直接交给下一个候选。停用多久取决于上游 给出的原因:余额不足要等充值,额度用完要等到上游说的重置时刻,限流通常几秒钟就 -过去。只有一个候选的请求不受影响。 +过去。只有一个候选的请求不受影响。Responses 的 WebSocket 连接上,每个 `response.create` +在这里和一个请求一样算:第一段内容到达之前因为上游出错而失败的,算作一次失败;上游 +拒绝或连不上的连接也一样。 流式回答在第一段内容交给客户端之前,上游在流里报的错误和错误状态码一样,会把 请求换到下一个候选。 +上游也可能开头很慢:收下请求之后很久都不发内容。开启 `next_on_slow_start` 后,请求 +发出 `stream_start_wait_secs` 秒仍没有内容,就换到下一个候选。默认关闭,因为先思考 +再输出的模型本来就可能很久才开始;开启时建议等 30 秒以上。最后一个候选总是等下去, +被放弃的上游不会停用。到点那一刻并发数已满(`max_concurrent`)的候选不算下一个:请求 +留在慢的那一家。 + +```yaml +failover: + stream_start_wait_secs: 30 + next_on_slow_start: true +``` + +上游的并发数满了(`max_concurrent`)时,一个请求等空位合计最多 `slot_wait_secs` +秒。等密钥的 `minute`、`hour` 上限也算在这段时间里,两样加起来不超过 `slot_wait_secs`。 +进行中的对话等它留在的那一家,到时还没有空位就换下一家,缓存在那边从头建; +新的对话遇到满着的上游直接跳过。候选全满时,请求等先空出来的那一家;都没有空出来, +客户端收到 429 和 `Retry-After`,说明上游都忙。如果有上游收到过这个请求并且失败了, +客户端收到的是那次失败。 + @@ -714,6 +812,8 @@ security: | `quota_pause_secs` | 整数 | `3600` | 上游报告额度用完、但没有给出重置时间时停用的秒数。给出了重置时间的,停用到那一刻。 | | `rate_limit_max_pause_secs` | 整数 | `3600` | 被限流的上游按它给的 `Retry-After` 停用,最多这么多秒。没有 `Retry-After` 的按没有说明原因的失败计。 | | `stream_start_wait_secs` | 整数 | `15` | 流式回答在第一段内容到达前最多暂存的秒数。在此之前上游报错,请求换到下一家;超过这个时间,已收到的部分照常交给客户端。取值 1 到 120。 | +| `next_on_slow_start` | 布尔 | `false` | 流式回答在请求发出 `stream_start_wait_secs` 秒后仍没有内容时,放弃这家上游,把请求交给下一家。最后一家总是等下去。被放弃的上游不会停用。开启时 `stream_start_wait_secs` 至少为 5。 | +| `slot_wait_secs` | 整数 | `30` | 一个请求合计最多等的秒数,从过了密钥自己的 `max_concurrent` 时算起:等密钥的 `minute`、`hour` 上限空出名额,和等并发数满了(`max_concurrent`)的上游空出位置,都算在里面。密钥的上限到时空不出来就拒绝;等不到上游的空位就换下一家,候选全满时回 429。`0`:不等。取值 0 到 300。 | ### `aliases` @@ -747,14 +847,56 @@ aliases: | 字段 | 类型 | 默认值 | 说明 | |---|---|---|---| | `name` | 字符串 | **必填** | 策略组的名字,不能重复,也不能和上游同名。 | -| `type` | `fallback` \| `select` \| `load-balance` \| `url-test` \| `cheapest` | `fallback` | `fallback`:按顺序取第一个健康的。`select`:取 `selected` 指定的那个。`load-balance`:新对话轮流。`url-test`:按实测首字节时间取最快的。`cheapest`:取输入单价最低的。 | -| `providers` | 字符串列表 | **必填** | 成员上游的名字。 | +| `type` | `fallback` \| `select` \| `load-balance` \| `url-test` \| `cheapest` | `fallback` | `fallback`:按顺序取第一个健康的。`select`:取 `selected` 指定的那个。`load-balance`:请求按成员的权重分;进行中的对话留在原来那一家。`url-test`:按实测从发出请求到回答第一段内容的时间取最快的。`cheapest`:取输入单价最低的。 | +| `providers` | 列表,每项是字符串或对象,对象见 [`groups[].providers[]`](#cfg-groups-providers) | **必填** | 成员上游的名字,不能是策略组。同一个上游在一个策略组中只出现一次。`load-balance` 组的成员可以写成 `{name, weight}`。 | | `selected` | 字符串 | — | `select` 类型选中的成员。 | +| `balance_by` | `weights` \| `latency` \| `health` \| `latency-health` | `weights` | `load-balance` 类型用:成员的权重再乘上什么。`weights`:不乘,只按权重。`latency`:越快的上游分得越多。`health`:越少失败的上游分得越多。`latency-health`:两者都看。其他类型只能是 `weights`。 | 默认类型为 `fallback`:单个使用者的机器上没有需要分散的负载。 -无论哪种类型,一段对话都留在上次回答它的那一家上游,让上游缓存着的那部分被再次读取,而不是换一家全价重算。同一轮之内(客户端正在回传工具结果)一律不换;跨轮时,上一次回答读或写了至少 1024 个 token 的 prompt cache、且距今不到五分钟,才继续留下。上游因失败进入冷却时,对话随之放开;故障转移之后接下回答的那一家,就是之后留下的那一家。一轮开始时命中的规则也沿用到这一轮结束:按输入大小或图片分流的规则不会让一轮半路换家,除非输入已经超出规则所指模型的上下文窗口。因此 `load-balance` 轮流的是新对话。 +`load-balance` 组的成员可以带权重,取值 1 到 100;只写名字的成员权重为 1。权重决定组内请求怎么分:写成 `{ name: anthropic, weight: 7 }` 和 `relay` 时,每十个请求有七个由官方 API 服务。进行中的对话留在回答它的那一家(见下文),也算进那一家的份额,因此份额靠新对话从哪一家开始来补齐。因失败处于冷却、并发数已满(`max_concurrent`)、或服务不了某个请求的上游不参与这一次分配,其余成员按各自的权重分。新的 WebSocket 连接也这样分配,算作一个请求;之后在这条连接上发的都交给它连上的那一家。其他类型的策略组不用权重。 + + + + +| 字段 | 类型 | 默认值 | 说明 | +|---|---|---|---| +| `name` | 字符串 | **必填** | 上游的名字。只写名字的成员权重为 1。 | +| `weight` | 整数 | `1` | 成员在 `load-balance` 组中分到的请求份额,与其他成员的权重成比例。取值 1 到 100。其他类型的策略组只能写 1。 | + + +```yaml +groups: + - name: pool + type: load-balance + providers: + - { name: anthropic, weight: 7 } + - relay +``` + +无论哪种类型,一段对话都留在上次回答它的那一家上游,让上游缓存着的那部分被再次读取,而不是换一家全价重算。同一轮之内(客户端正在回传工具结果)一律不换;跨轮时,上一次回答读或写了至少 1024 个 token 的 prompt cache、且距今不到五分钟,才继续留下。上游因失败进入冷却时,对话随之放开;故障转移之后接下回答的那一家,就是之后留下的那一家。一轮开始时命中的规则也沿用到这一轮结束:按输入大小或图片分流的规则不会让一轮半路换家,除非输入已经超出规则所指模型的上下文窗口。因此 `load-balance` 的权重是长期看各家分到的请求的比例:进行中的对话留在原来的上游,也算进那一家的份额。 + +`balance_by` 让 `load-balance` 组再看各上游最近的表现:每个成员的权重乘上一个系数,组内请求按乘出来的结果照上文的方式分。 + +- `weights`(默认):只按权重。 +- `latency`:越快的上游分得越多。快慢看典型的从发出请求到回答第一段内容的时间,与 `url-test` 使用同一份测量。比组内居中者快一倍的上游,权重乘以四;最多乘以十,最少乘以十分之一。 +- `health`:越少失败的上游分得越多。依据是最近 30 分钟内的最近 50 次请求:服务器错误、限流、额度或余额用尽、凭据被拒、超时和连接失败算作失败;请求本身导致的错误不算,客户端取消、因开头太慢而换走、因并发数满了(`max_concurrent`)而跳过也不算。经常失败的上游至少保留权重的二十分之一,仍会偶尔分到请求,以便发现它已经恢复;完全失败的上游照旧由 [`failover`](#cfg-failover) 暂停。 +- `latency-health`:两个系数相乘。 + +快慢只在流式回答上测,从请求发给这家上游的那一刻算起:之前的等待、之前失败的上游都不算在内。因开头太慢而被放弃的上游(`failover.next_on_slow_start`),按等满的那段时间计。Responses 的 WebSocket 连接上,每个 `response.create` 在快慢和成败上都算一个请求,快慢从上游开始回答它的那一刻算起。 + +测量还不够的上游按中等对待。与只按权重时一样,进行中的对话留在原来的上游,差额由新对话补齐。 + +```yaml +groups: + - name: 均摊 + type: load-balance + balance_by: latency-health + providers: + - { name: 官方, weight: 3 } + - 中转 +``` ### `routes` diff --git a/src/data/core-docs/manifest.json b/src/data/core-docs/manifest.json index f7459d4..82745bb 100644 --- a/src/data/core-docs/manifest.json +++ b/src/data/core-docs/manifest.json @@ -1,9 +1,9 @@ { "repository": "ThinkWatchProject/ThinkWatch-Core", - "ref": "v0.62.0", + "ref": "v0.63.0", "files": { - "docs/config.md": "e916418233efc68eb2a7883c22463b3cdeee76b0ac034cb7bdc1516d4223ae07", - "docs/config.zh-CN.md": "46c3f4bcddffd7f5d74d4f066688ffba9d9d0d4d8c969b5548356c14b918fd57", + "docs/config.md": "66ac86b77c3120da283280a61c2f4c88cc9775b72b02c24cf763643adc37568f", + "docs/config.zh-CN.md": "58e23593663425c45bd975ba978d2322a7837faea9ec36da2d0064da5003230e", "docs/server.md": "5e1e9b901bef2b46d417aea057a1db24c78457d3c1938b8540ecb4766515b68f", "docs/server.zh-CN.md": "1f508e7c39b8ded28653773ca4d8701e6bdc2247bcd359c3a8fd00fd1401a6ff" } diff --git a/src/i18n/pages/lite.ts b/src/i18n/pages/lite.ts index 66dd58a..ac38dc0 100644 --- a/src/i18n/pages/lite.ts +++ b/src/i18n/pages/lite.ts @@ -166,7 +166,7 @@ export const liteCopy = { { id: "keys", title: "A key for each client", - body: "Connecting a client gives it a key of its own, so traffic and cost are counted per client. Each key has its own route, visible models and concurrency limit, and a rotated key is written into its client's configuration.", + body: "Connecting a client gives it a key of its own, so traffic and cost are counted per client. Each key has its own route, visible models, concurrency limit and usage limits on requests, tokens or cost, and a rotated key is written into its client's configuration.", alt: "The Keys page: the default key and one key each for Claude Code, Codex and Cursor, with the route each key uses, the models it may use, and its requests and cost over the last 24 hours", }, // No screenshot: the bento shows a short plugin instead, so there is no alt. @@ -409,7 +409,7 @@ export const liteCopy = { { id: "keys", title: "每个客户端一把密钥", - body: "接入客户端时为它生成专用密钥,流量与费用按客户端分开统计。每把密钥可单独设置路由、可见模型与并发上限,更换后的新密钥自动写入客户端配置。", + body: "接入客户端时为它生成专用密钥,流量与费用按客户端分开统计。每把密钥可单独设置路由、可见模型、并发上限,以及按请求次数、token 或费用计的用量上限,更换后的新密钥自动写入客户端配置。", alt: "密钥页:默认密钥与 Claude Code、Codex、Cursor 各自的密钥,列出各自使用的路由、可用模型,以及最近 24 小时的请求数与费用", }, {