diff --git a/src/content/docs-lite/en/codex-other-models.md b/src/content/docs-lite/en/codex-other-models.md index 43e2a0a..01f31a7 100644 --- a/src/content/docs-lite/en/codex-other-models.md +++ b/src/content/docs-lite/en/codex-other-models.md @@ -10,7 +10,7 @@ ThinkWatch Lite connects Codex to a gateway on the same computer that accepts th ## Steps -1. On the Upstreams page, choose **New upstream** and pick a **Service**: **Anthropic** for Claude, or **Google Gemini**; each fills in the address and protocol. For a relay, choose **Custom**, enter its **Base URL** without an endpoint path such as `/chat/completions`, and set **Protocol** to **OpenAI Chat Completions**; for GLM, for example, `https://api.z.ai/api/paas/v4`, or `…/api/coding/paas/v4` on a GLM Coding Plan. A base URL that ends with its own version, such as `/v4` or Volcengine Ark's `/api/v3`, is used as written from ThinkWatch Lite 2026.10.5. Enter the **API key**, choose **Check connection**, then **Next**. If the relay does not list its models, enter them one per line under **Manual list**. Choose **Next**, then **Create**. +1. On the Upstreams page, choose **New upstream** and pick a **Service**: **Anthropic** for Claude, or **Google Gemini**; each fills in the address and protocol. For a relay, choose **Custom**, enter its **Base URL** without an endpoint path such as `/chat/completions`, and set **Protocol** to **OpenAI Chat Completions**; for GLM, for example, `https://api.z.ai/api/paas/v4`, or `…/api/coding/paas/v4` on a GLM Coding Plan. A base URL that ends with its own version, such as `/v4` or Volcengine Ark's `/api/v3`, is used as written from ThinkWatch Lite 2026.10.5. Enter the **API key**, choose **Check connection**, then **Next**. If the relay lists no models, or leaves some out, type each missing model ID in the box at the end of **Models** and press Enter. Choose **Next**, then **Create**. 2. On the Clients page, choose **Connect…** on the Codex row. The dialog shows the change to `~/.codex/config.toml`: | Field | Value | @@ -31,8 +31,9 @@ ThinkWatch Lite connects Codex to a gateway on the same computer that accepts th - **Codex's model table.** Codex carries metadata for its own models, such as the context window, inside the program. A model it does not know, such as a Claude or Gemini model, runs on fallback metadata with a 272,000-token context window, and Codex warns: "Model metadata for `` not found. Defaulting to fallback metadata; this can degrade performance and cause issues." For a model with a smaller window, `model_context_window` in `config.toml` sets the window Codex assumes. The gateway's model list does not appear in Codex's model picker, as Codex expects a catalog in its own format. - **Conversion.** Requests and streamed answers are converted in both directions. Traffic marks such requests **Converted**; fields the target format cannot carry are dropped and listed in the request details. Server-side tools such as web search run only at the provider they belong to and are dropped in conversion. Codex accepts only `responses` for `wire_api`, so this conversion is what makes a Chat Completions-only relay usable. +- **Tools, history and compaction.** From ThinkWatch Lite 2026.10.11, every tool Codex declares reaches the upstream, including those it lists in its input rather than in `tools`, and tool calls come back under the names Codex gave them. Local shell calls, tool search and reasoning-effort changes in the history are converted as well. Compacting a long session works: the upstream writes a summary, which Codex keeps as its compaction and sends back in later requests. A compaction encrypted by OpenAI cannot be read by another upstream, and a request that carries one is refused with an error saying so. Instructions Codex adds in the middle of a conversation stay where they are, so the system prompt stays the same from turn to turn and the prompt cache keeps hitting. - **Credentials.** With `requires_openai_auth = false`, Codex authenticates to the gateway with its own key and sends no OpenAI key or ChatGPT token. A ChatGPT account signed in from the app can still serve the OpenAI models next to Claude or Gemini: each request goes to an upstream that lists the model asked for. - **Sessions.** Codex lists sessions started before and after connecting separately. `codex resume -c model_provider=thinkwatch` continues an earlier session through the gateway. After **Restore…**, sessions started while connected can still be opened, and go straight to OpenAI. -- **Cost.** A rewritten request is priced as the model actually sent. +- **Cost.** A rewritten request is priced as the model actually sent. Codex marks no prompt-cache breakpoints, so on a route to an upstream in Anthropic Messages format, or to a Claude model on Bedrock that AWS lists for prompt caching (Claude 3.5 Sonnet v2, Claude 3.7 Sonnet, and Claude 4.5 and later), the gateway marks the tools, the system prompt and the last two user turns, and each turn reads what the previous one cached. Older Claude models on Bedrock, such as Sonnet 4, get no marks, and an upstream that refuses them receives the request again without them. Cache writes and reads are priced at the price sheet's cache rates; at Anthropic, writes cost 25% more than input and reads a tenth of it. Related: [Features](/docs/lite/features/), [Use Claude Code with GLM, DeepSeek or Kimi](/docs/lite/claude-code-other-models/), [Install and update](/docs/lite/install/). diff --git a/src/content/docs-lite/en/failover-and-load-balancing.md b/src/content/docs-lite/en/failover-and-load-balancing.md index 8d1c5c6..095719f 100644 --- a/src/content/docs-lite/en/failover-and-load-balancing.md +++ b/src/content/docs-lite/en/failover-and-load-balancing.md @@ -4,7 +4,7 @@ ThinkWatch Lite turns each relay key into an upstream and puts several upstreams ## Before you start -- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.10. +- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.11. - The base URL and API keys of each relay. ## Steps @@ -26,13 +26,13 @@ ThinkWatch Lite turns each relay key into an upstream and puts several upstreams A strategy only sets the order; every member remains available for failover. -- **When the next upstream is tried.** Before anything has reached the client, the gateway moves on when the upstream cannot be reached or its credential cannot be read; answers 5xx, 429, 401, 403, 402 or 404; answers 400 or 422 with an error about an insufficient balance, a used-up quota or an unavailable model; or, in a streamed answer, reports an error before the first content, such as an overload. The gateway waits for that first content for up to **Wait for the answer to start** in Settings › Failover, 15 seconds by default. Other 4xx responses, and the last member's 4xx other than 429, go back to the client unchanged. +- **When the next upstream is tried.** Before anything has reached the client, the gateway moves on when the upstream cannot be reached or its credential cannot be read; answers 5xx, 429, 401, 403, 402 or 404; answers 400 or 422 with an error about an insufficient balance, a used-up quota or an unavailable model; in a streamed answer, reports an error before the first content, such as an overload; or sends no content within the no-response timeout. Other 4xx responses, and the last member's 4xx other than 429, go back to the client unchanged. - **Weights and distribution.** In a **Round robin** group, a member's weight, from 1 to 100, is its long-run share of requests: weights 7 and 3 send seven requests in ten to the first, interleaved with the other three. **By ratio** uses the weights alone; **By speed** multiplies each by how much faster than the members' median the upstream starts answering (time to first token, squared, kept between 0.1 and 10), **By reliability** by its success rate over its last 50 outcomes within 30 minutes (squared, at least 0.05), and **By speed and reliability** by both. An upstream with too few samples counts as average, requests that stay with an upstream for their conversation count toward its share, and paused or full members sit out. -- **Slow starts.** With **Move to the next upstream when the start times out** on in Settings › Failover, a streamed answer with no content after **Wait for the answer to start** is cancelled and the request goes to the next upstream, if one can take it at that moment; the last upstream always waits, and the slow one is not paused. The upstream may have billed the input of the abandoned attempt, which the request's **Routing** tab marks; with the switch on, the wait must be at least 5 seconds, and 30 or more suits models that think before they answer. +- **No response.** An upstream that sends no content for the **No-response timeout** in Settings › Failover, 300 seconds by default (30 to 3,600), is given up on. Only content counts: text, reasoning or a tool call, not keep-alive pings. The clock starts when the request is sent to that upstream, so time spent waiting for a free slot does not count, and starts again with every piece of content. Before anything has reached the client, the request moves to the next upstream and the silent one counts as a failure toward its pause; when none is left, the client receives a timeout error in its own format. Once the answer has started, it ends with an error event the client understands. A whole, non-streamed answer has the same time to arrive. A streamed answer is held back for at most 15 seconds; after that the client receives the response headers and keep-alive comments (Gemini clients get the headers only), so it does not give up while the gateway can still move to the next upstream. The silent upstream's time to first token is recorded as the full wait, which ranks it lower for **Lowest latency** and **By speed**. The request's **Routing** tab marks the attempt **No response · 300 s** with the input the upstream may have billed for it. A model that thinks for a long time without sending anything may need a longer timeout. - **Concurrency limits.** An upstream with a **Concurrency limit** takes at most that many requests at once. When it is full, a conversation that stays on it waits for a free slot and then moves on, and other requests go straight to the next member; when every member is full, a request waits for the first free slot for up to **Wait for a free slot at most** in Settings › Failover, 30 seconds by default, and then receives a 429 with `Retry-After`, or the error of an earlier attempt when one was sent. The same wait covers a key's per-minute and per-hour usage limits, and the Traffic page shows each skip and the time queued. -- **Pauses.** A failing upstream is paused for as long as Settings › Failover sets: by default 60 seconds after 3 **Consecutive failures**, doubling up to 600; 30 minutes for **Insufficient balance**; until the reset time, or 60 minutes, for **Quota used up**; the wait a rate-limited upstream asks for, up to 60 minutes. A rule with a single upstream is never held back, and when every member is paused they are tried anyway. +- **Pauses.** A failing upstream is paused for as long as Settings › Failover sets: by default 60 seconds after 3 **Consecutive failures**, doubling up to 600; 30 minutes for **Insufficient balance**; until the reset time, or 60 minutes, for **Quota used up**; the wait a rate-limited upstream asks for, up to 60 minutes. A rule with a single upstream is never held back, and when every member is paused they are tried anyway. A request aborted on the Traffic page does not count against its upstream. - **Sessions and prompt cache.** Within a turn, while the client sends tool results back, requests keep the rule chosen at the start of the turn and the upstream that answered. Across turns, a conversation stays with the upstream that answered last if that answer read or wrote at least 1,024 cached tokens within the last five minutes; moving would rebuild the cache at full price. Otherwise, or when that upstream is paused, the strategy orders the members again, which is when **Round robin** moves on. The **Conversation** line on a request's **Routing** tab shows when a request stayed. -- **Same model name.** Every member is asked for the model the client sent, or the name a rule rewrote it to. A member whose model list lacks it is skipped; one without a list is tried, and its 404 moves the request on. A model that a member serves but leaves out of its list can be added with **Add models…** in the member's model list on the Upstreams page. +- **Same model name.** Every member is asked for the model the client sent, or the name a rule rewrote it to. A member whose model list lacks it is skipped; one without a list is tried, and its 404 moves the request on. A model that a member serves but leaves out of its list can be added with **Add models…** in the member's model list on the Upstreams page, or in the **Models** section of its edit dialog. - **Falling back to another model.** Add a rule with the condition **Selected upstream** set to the backup upstream, **On match** set to **Continue matching**, and its model in **Change model to** under **Parameter rewrites**. Such a rule is evaluated for each upstream as it is tried, failover included. The changed model no longer hits the cached prompt, and the request is priced by the name sent. Related: [Switch relays, upstreams or models without restarting Claude Code or Codex](/docs/lite/switch-upstreams-without-restart/), [Features](/docs/lite/features/#routing-and-failover), [Install and update](/docs/lite/install/). diff --git a/src/content/docs-lite/en/features.md b/src/content/docs-lite/en/features.md index 4166789..339b713 100644 --- a/src/content/docs-lite/en/features.md +++ b/src/content/docs-lite/en/features.md @@ -14,11 +14,13 @@ Until the first request has gone to an upstream, the Overview page shows a get-s ## Traffic and sessions -The Traffic page lists requests as they arrive: status, key, model, upstream, time to first token and total time (with the generation speed on hover), tokens and cost, with marks for a converted API format, redacted keys, text deleted by the content filter and a blocked or suspicious tool call. The list can be searched (path, key, upstream, model and error message), filtered by key, upstream and model, or narrowed to failed or unpriced requests. It holds the latest 2,000 requests, and a search or filter also runs over the stored history, further back on request. With content search on, it also covers what each request newly sent (the last user turn, tool results included) and its answer (tool calls included), for as long as payloads are kept; a match shows its excerpt under the request. The Sessions view groups the requests of one conversation into turns, with the input tokens and the cost of each turn. A session's Conversation tab replays it turn by turn: the messages each turn added and its answer, with text, folded thinking, tool calls beside their results, and images by type and size. Changes to the system prompt and restarts after compaction are marked, and a turn whose bodies are past retention, were too large to keep whole or cannot be read says so. Each turn opens its request. +The Traffic page lists requests as they arrive: status, key, model, upstream, time to first token and total time (with the generation speed on hover), tokens and cost, with marks for a converted API format, redacted keys, text deleted by the content filter and a blocked or suspicious tool call. Long upstream, model and key names are cut short, and hovering shows them in full. The list can be searched (path, key, upstream, model and error message), filtered by key, upstream and model, or narrowed to failed or unpriced requests. It holds the latest 2,000 requests, and a search or filter also runs over the stored history, further back on request. With content search on, it also covers what each request newly sent (the last user turn, tool results included) and its answer (tool calls included), for as long as payloads are kept; a match shows its excerpt under the request. The Sessions view groups the requests of one conversation into turns, with the input tokens and the cost of each turn. A session's Conversation tab replays it turn by turn: the messages each turn added and its answer, with text, folded thinking, tool calls beside their results, and images by type and size. Changes to the system prompt and restarts after compaction are marked, and a turn whose bodies are past retention, were too large to keep whole or cannot be read says so. Each turn opens its request. A request opens into its timeline, its routing (the rule it matched, the group it went through and each attempt with its status and duration), the request and response bodies, and its usage and cost. A request from DeepSeek Harness also shows the size of the session log it carried, the whole conversation the client attaches to every request; the gateway removes it before a request goes to an upstream other than DeepSeek. A finished request can be sent again, unchanged, to another upstream after an estimate of its cost, and the two responses are shown side by side. -The attempts also show an upstream passed over at its concurrency limit (**At its concurrency limit**), the time spent waiting for a free slot, and a streamed answer given up because its start timed out (**Start timed out**), with the input tokens that upstream may have billed for it. A request refused by its key's usage limit, or because every upstream stayed at its concurrency limit, shows **Usage limit** or **At capacity** in place of an upstream, and a connection in the Responses API's WebSocket mode is listed one request per turn, each with its own usage and cost. +The attempts also show an upstream passed over at its concurrency limit (**At its concurrency limit**), the time spent waiting for a free slot, an upstream given up because it sent no content within the no-response timeout (**No response · 300 s** at the default), with the input tokens that upstream may have billed for it, and the attempt under way when a request was aborted (**Aborted**). A request refused by its key's usage limit, or because every upstream stayed at its concurrency limit, shows **Usage limit** or **At capacity** in place of an upstream, and a connection in the Responses API's WebSocket mode is listed one request per turn, each with its own usage and cost. + +A request in progress can be stopped with **Abort request** in its menu. **Abort session**, in a session's menu or at the top of its panel, stops every request in progress in that session after a confirmation. The upstream call is dropped at once, the client receives an error, and the request is marked **Aborted**, in grey; its upstream is not paused for it. Turns on a Responses API WebSocket connection cannot be aborted. ## Client setup @@ -39,9 +41,9 @@ Usage limits keep one client from using up a budget: each caps the key at a numb ## Upstreams -Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; aliases, proxies and price sheets have tabs on the same page. An optional **Concurrency limit** suits relays and accounts that allow only so many requests at once: when the upstream is full, a conversation that stays on it waits for a free slot, other requests go to the next upstream, and when every upstream is full a request waits and then receives a busy error. **Specs…** in an upstream's model list sets by hand a model's context window, maximum output, and whether it reasons and takes images, for a model the price table lacks or gets wrong; the values set there are used in place of the price table's. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs. +Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Clients such as Codex mark no prompt-cache breakpoints; when their requests are converted for an upstream in Anthropic Messages format, or for a Claude model on Bedrock that AWS lists for prompt caching (Claude 3.5 Sonnet v2, Claude 3.7 Sonnet, and Claude 4.5 and later), the gateway marks them, so each turn reads what the previous one cached. An upstream that refuses the marks receives the request again without them. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; aliases, proxies and price sheets have tabs on the same page. An optional **Concurrency limit** suits relays and accounts that allow only so many requests at once: when the upstream is full, a conversation that stays on it waits for a free slot, other requests go to the next upstream, and when every upstream is full a request waits and then receives a busy error. **Specs…** in an upstream's model list sets by hand a model's context window, maximum output, and whether it reasons and takes images, for a model the price table lacks or gets wrong; the values set there are used in place of the price table's. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs. -An upstream's models are the ones it lists plus any added by hand. Upstreams often leave out models they serve: an account backend hides newer models, a relay lists only some. **Add models…** in an upstream's model list adds such models by their exact IDs, several at a time. An ID cannot contain `*` or `?` or have spaces around it, is at most 256 characters, and is added once; the dialog points out an ID that breaks these rules, or one the upstream already lists, before saving. A model added by hand is marked **Manual**, and routing and failover, aliases, routing rules that name a model, a key's allowed models, the model list clients see, the dry run and the speed test all treat it like a listed model. The summary at the top of the list counts the models the upstream lists apart from those added by hand. Hovering over a manual model shows **Remove**, which takes effect at once, with **Undo** in the message that follows. For an upstream that lists no models, the same dialog replaces the manual model list, and the models added there are its whole list. +An upstream's models are the ones it lists plus any added by hand. Upstreams often leave out models they serve: an account backend hides newer models, a relay lists only some. **Add models…** in an upstream's model list adds such models by their exact IDs, several at a time. An ID cannot contain `*` or `?` or have spaces around it, is at most 256 characters, and is added once; the dialog points out an ID that breaks these rules, or one the upstream already lists, before saving. A model added by hand is marked **Manual**, and routing and failover, aliases, routing rules that name a model, a key's allowed models, the model list clients see, the dry run and the speed test all treat it like a listed model. The summary at the top of the list counts the models the upstream lists apart from those added by hand. Hovering over a manual model shows **Remove**, which takes effect at once, with **Undo** in the message that follows. The **Models** section of the **New upstream** and **Edit upstream** dialogs ends with the same input, for an upstream with or without a list, so models can also be added while it is created or edited. They are saved with the rest of the upstream, and with **Selected models** a model added there is enabled at once. For an upstream that lists no models, the models added by hand are its whole list. The Aliases tab gives a model the name clients use for it. An alias lists the names the same model has on different upstreams, such as `claude-sonnet-5` on Anthropic and `us.anthropic.claude-sonnet-5-v1:0` on Bedrock; every upstream that offers one of them serves the alias under its own name, and they back each other up. Clients see aliases in their model lists, and answers carry the name the client asked for, while the request log shows the model each upstream was sent. When the same Claude model has different names on the official API, Bedrock, Vertex or OpenRouter, the tab and the new-alias dialog suggest grouping them. An alias can also be created from an upstream's model list. The dialog shows which upstream receives which name and warns when the name would take over another upstream's model of the same name. A key whose model scope allows an upstream model can also use the aliases that list it. @@ -53,7 +55,7 @@ A relay or vendor can hand out an import link, `thinkwatch://import?…` or its ## Routing and failover -Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream, a group or specified models, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Specified models name an upstream and one of its models, with backups tried in order, and send the model as written; to have one client use one model under another model's name, a rule on that client's key does it without changing what the name means for other clients. A model condition written for an upstream model also matches the aliases that list it. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. In a round-robin group each member has a weight from 1 to 100, its long-run share of requests, and **Distribute by** can multiply the weights by each upstream's speed (time to first token), its reliability (recent success rate) or both; conversations in progress stay on their upstream. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. With **Move to the next upstream when the start times out** on in Settings › Failover, a streamed answer that has sent nothing within the wait also moves on, when another upstream can take it; the last upstream always waits. A map at the top of the page traces every key through its route and groups to the upstreams. +Each key follows a route, and keys without one follow the default route. A route is a list of rules evaluated in order. A rule matches on the model, the key, the client's API format, input tokens, `max_tokens`, the number of tools, images, extended thinking, streaming, prompt caching or the kind of auxiliary request; it then forwards the request to an upstream, a group or specified models, or refuses it, and can rewrite the model, `max_tokens` or extended thinking. Specified models name an upstream and one of its models, with backups tried in order, and send the model as written; to have one client use one model under another model's name, a rule on that client's key does it without changing what the name means for other clients. A model condition written for an upstream model also matches the aliases that list it. Setting `max_tokens` caps the length of answers: the upstream stops at that point by itself, without an error. A group puts several upstreams behind one name and decides the order in which they are tried: as listed, manually selected, in turn, lowest latency first or lowest cost first. In a round-robin group each member has a weight from 1 to 100, its long-run share of requests, and **Distribute by** can multiply the weights by each upstream's speed (time to first token), its reliability (recent success rate) or both; conversations in progress stay on their upstream. When an attempt fails, the request moves on to the next upstream, and by default a session stays on one upstream so that its prompt cache keeps hitting. An upstream that sends no content for the **No-response timeout** in Settings › Failover, 300 seconds by default, is given up on; keep-alive pings do not count as content. Before anything has reached the client, the request moves to the next upstream, and when none is left the client receives a timeout error; once the answer has started, it ends with an error. A map at the top of the page traces every key through its route and groups to the upstreams. Auxiliary requests that clients send on their own (health checks, warm-ups, titles, topic detection and input suggestions) can be answered locally at no cost, or forwarded. Forwarded ones go through the routing rules like any other request, and a rule can send a given kind to a lower-cost upstream. @@ -87,7 +89,7 @@ Plugins are short JavaScript files that change requests before they go to an ups ## Settings -Settings has seven sections. Connection lists the local core and the saved remote cores, described in [Connecting to a remote core](/docs/lite/remote-core). General sets the language, the appearance, what the menu bar item shows on macOS (or that it is hidden), launch at login, whether notices arrive as system notifications, in the app only or not at all, and shows hidden guidance hints again. Listening sets who can reach the gateway (this machine only, the local network of a chosen interface, or every interface), its port and the allowed address ranges. Failover sets how long a failing upstream is paused before requests go to the next one: after how many consecutive failures, for how long, and separate pauses for an insufficient balance, a used-up quota and rate limits. It also sets how long to wait for a streamed answer to start, whether a stream that has not started by then moves to the next upstream (off by default), and how long in total a request may wait for a free slot on a full upstream or for a key's per-minute or per-hour limit, 30 seconds by default. Log retention sets how long request payloads and request records are kept, and a size cap for payloads. About shows the version, checks for updates and produces a diagnostics bundle with keys and addresses masked. Uninstall restores every connected client and removes the autostart entry, and is meant to be run before the app is deleted. +Settings has seven sections. Connection lists the local core and the saved remote cores, described in [Connecting to a remote core](/docs/lite/remote-core). General sets the language, the appearance, what the menu bar item shows on macOS (or that it is hidden), launch at login, whether notices arrive as system notifications, in the app only or not at all, and shows hidden guidance hints again. Listening sets who can reach the gateway (this machine only, the local network of a chosen interface, or every interface), its port and the allowed address ranges. Failover sets how long a failing upstream is paused before requests go to the next one: after how many consecutive failures, for how long, and separate pauses for an insufficient balance, a used-up quota and rate limits. It also sets the no-response timeout, after which an upstream that has sent no content is given up on (300 seconds by default, from 30 to 3,600), and how long in total a request may wait for a free slot on a full upstream or for a key's per-minute or per-hour limit, 30 seconds by default. Log retention sets how long request payloads and request records are kept, and a size cap for payloads. About shows the version, checks for updates and produces a diagnostics bundle with keys and addresses masked. Uninstall restores every connected client and removes the autostart entry, and is meant to be run before the app is deleted. ## Menu bar, system tray and notifications diff --git a/src/content/docs-lite/zh-CN/codex-other-models.md b/src/content/docs-lite/zh-CN/codex-other-models.md index 5da30d3..909f1d4 100644 --- a/src/content/docs-lite/zh-CN/codex-other-models.md +++ b/src/content/docs-lite/zh-CN/codex-other-models.md @@ -10,7 +10,7 @@ ThinkWatch Lite 把 Codex 接到本机的网关上,网关接收 OpenAI Respons ## 步骤 -1. 在「上游」页点击「新建上游」,选择「服务类型」:Claude 选「Anthropic」,Gemini 选「Google Gemini」,接口地址和接口协议会自动填入。中转站选「自定义」,「接口地址」填它的 Base URL(不含 `/chat/completions` 等接口路径),「接口协议」选「OpenAI Chat Completions」;例如 GLM 填 `https://api.z.ai/api/paas/v4`,GLM Coding Plan 填 `…/api/coding/paas/v4`。以自带版本号结尾的地址(如 `/v4`、火山方舟的 `/api/v3`),自 ThinkWatch Lite 2026.10.5 起按原样使用。填写「API 密钥」,点击「检测连接」,再点击「下一步」。中转站不提供模型列表时,在「手动清单」中每行填写一个模型 ID。再点击「下一步」,然后点击「创建」。 +1. 在「上游」页点击「新建上游」,选择「服务类型」:Claude 选「Anthropic」,Gemini 选「Google Gemini」,接口地址和接口协议会自动填入。中转站选「自定义」,「接口地址」填它的 Base URL(不含 `/chat/completions` 等接口路径),「接口协议」选「OpenAI Chat Completions」;例如 GLM 填 `https://api.z.ai/api/paas/v4`,GLM Coding Plan 填 `…/api/coding/paas/v4`。以自带版本号结尾的地址(如 `/v4`、火山方舟的 `/api/v3`),自 ThinkWatch Lite 2026.10.5 起按原样使用。填写「API 密钥」,点击「检测连接」,再点击「下一步」。中转站不提供模型列表或没有列全时,在「模型」一节末尾的输入框中逐个填写缺少的模型 ID 并回车。再点击「下一步」,然后点击「创建」。 2. 在「客户端」页 Codex 一行点击「接管…」。对话框列出对 `~/.codex/config.toml` 的修改: | 字段 | 写入 | @@ -31,8 +31,9 @@ ThinkWatch Lite 把 Codex 接到本机的网关上,网关接收 OpenAI Respons - **Codex 的内置模型表。**Codex 把自家模型的元数据(上下文窗口等)编在程序里。它不认识的模型,例如 Claude 或 Gemini 的模型,按兜底元数据运行,上下文窗口为 272,000 token,并提示「Model metadata for `<模型>` not found. Defaulting to fallback metadata; this can degrade performance and cause issues.」。模型的上下文窗口更小时,可在 `config.toml` 中用 `model_context_window` 设定 Codex 采用的窗口。网关的模型列表不会出现在 Codex 的模型选择中,因为 Codex 要求的是它自己格式的模型目录。 - **格式转换。**请求和流式回答双向转换。「流量」页把这类请求标为「已转换」,目标格式无法承载的字段被丢弃,并在请求详情中列出。网页搜索这类服务端工具只能由所属的服务商执行,转换时被丢弃。Codex 的 `wire_api` 只支持 `responses`,只提供 Chat Completions 的中转正是靠这一转换才能使用。 +- **工具、历史与压缩。**自 ThinkWatch Lite 2026.10.11 起,Codex 声明的所有工具都会发给上游,包括写在输入中而不在 `tools` 里的工具,工具调用按 Codex 给出的名称返回。历史中的本地 shell 调用、工具搜索和推理强度的变更同样会转换。长会话的上下文压缩可以正常进行:由上游写出摘要,Codex 把它作为压缩结果保存,并在之后的请求中带回。OpenAI 加密的压缩结果无法由其他上游读取,带有这类内容的请求会被拒绝,并说明原因。Codex 在对话中途加入的指令留在原位,系统提示词因此每轮保持不变,提示缓存持续命中。 - **凭据。**`requires_openai_auth = false` 使 Codex 用自己的网关密钥连接网关,不发送 OpenAI 密钥或 ChatGPT 令牌。应用内登录的 ChatGPT 账号仍可同时作为 OpenAI 模型的上游:每个请求都交给模型列表中有所请求模型的上游。 - **会话。**接管前后的会话在 Codex 中分开显示。运行 `codex resume <会话 ID> -c model_provider=thinkwatch` 可以通过网关继续之前的会话。「还原…」之后,接管期间的会话仍可打开,此时直连 OpenAI。 -- **费用。**改写过模型名的请求按实际发出的模型计价。 +- **费用。**改写过模型名的请求按实际发出的模型计价。Codex 不标注提示缓存断点,因此发往 Anthropic Messages 格式的上游,或 Bedrock 上 AWS 列为支持提示缓存的 Claude 模型(Claude 3.5 Sonnet v2、Claude 3.7 Sonnet 以及 Claude 4.5 起的模型)时,网关在工具、系统提示词和最后两轮用户消息处标注,每一轮都能读取上一轮写入的缓存。Bedrock 上更早的 Claude 模型(如 Sonnet 4)不标注;上游拒绝这些标注时,网关去掉标注重新发送一次。缓存写入与读取按价目表的缓存单价计费;Anthropic 的缓存写入比输入贵 25%,读取为输入的十分之一。 相关文档:[功能详解](/zh-CN/docs/lite/features/)、[让 Claude Code 使用 GLM、DeepSeek 或 Kimi](/zh-CN/docs/lite/claude-code-other-models/)、[安装与更新](/zh-CN/docs/lite/install/)。 diff --git a/src/content/docs-lite/zh-CN/failover-and-load-balancing.md b/src/content/docs-lite/zh-CN/failover-and-load-balancing.md index dd08cf0..a6801c2 100644 --- a/src/content/docs-lite/zh-CN/failover-and-load-balancing.md +++ b/src/content/docs-lite/zh-CN/failover-and-load-balancing.md @@ -4,7 +4,7 @@ ThinkWatch Lite 把每个中转站的每把密钥建成一个上游,再把多 ## 准备 -- 已[安装](/zh-CN/lite/#install) ThinkWatch Lite,并已在客户端页接管客户端。本文按 2026.10.10 版编写。 +- 已[安装](/zh-CN/lite/#install) ThinkWatch Lite,并已在客户端页接管客户端。本文按 2026.10.11 版编写。 - 各中转站的接口地址和 API 密钥。 ## 步骤 @@ -26,13 +26,13 @@ ThinkWatch Lite 把每个中转站的每把密钥建成一个上游,再把多 策略只决定顺序,所有成员都留作故障转移的后备。 -- **何时换到下一个上游。**在尚未向客户端发出任何内容时,上游出现以下情况,网关即改用下一个成员:无法连接或读不到它的凭据;返回 5xx、429、401、403、402 或 404;返回 400 或 422,且错误信息说明余额不足、额度用完或模型不可用;流式回答在第一段内容之前报错,例如过载。等待第一段内容的时长为「设置 › 故障转移」中的「等待回答开头」,默认 15 秒。其他 4xx,以及最后一个成员返回的 4xx(429 除外),原样返回客户端。 +- **何时换到下一个上游。**在尚未向客户端发出任何内容时,上游出现以下情况,网关即改用下一个成员:无法连接或读不到它的凭据;返回 5xx、429、401、403、402 或 404;返回 400 或 422,且错误信息说明余额不足、额度用完或模型不可用;流式回答在第一段内容之前报错,例如过载;在无响应超时内没有发出内容。其他 4xx,以及最后一个成员返回的 4xx(429 除外),原样返回客户端。 - **权重与分配依据。**「轮询」策略组中,成员的权重(1 到 100)就是它长期分到的请求比例:权重 7 与 3 时,每十个请求中七个交给前者,与另外三个交错分配。「按比例」只看权重;「按速度」再乘以该上游开始回答比成员中位数快多少(首 token 时间之比的平方,限定在 0.1 到 10 之间);「按稳定性」再乘以它 30 分钟内最近 50 次结果的成功率(取平方,最低 0.05);「按速度和稳定性」两者都乘。样本不足的上游按平均水平计算;因对话延续而留在某个上游的请求计入它的份额;暂停中或并发已满的成员本轮不参与。 -- **开头超时。**在「设置 › 故障转移」中打开「开头超时时转到下一个上游」后,流式回答等过「等待回答开头」仍没有内容时,只要此刻另有上游可以接下,网关就取消这次尝试,把请求交给下一个上游;最后一个上游照常等待,慢的上游也不会被暂停。被放弃的那次尝试的输入可能已被上游计费,请求的「路由」标签会标出。开启时等待时长至少为 5 秒;先思考再输出的模型宜设为 30 秒以上。 +- **无响应。**上游连续「设置 › 故障转移」中的「无响应超时」(默认 300 秒,可设 30 到 3,600 秒)没有发出内容时,网关放弃这个上游。只有内容才计入:文字、推理或工具调用,保活心跳不算。计时从请求发给这个上游时开始,等待空位的时间不计入,每收到一段内容重新计时。尚未向客户端发出内容时,请求转到下一个上游,没有响应的上游计一次失败,按暂停规则处理;没有可用的上游时,客户端收到其自身格式的超时错误。回答已经开始时,以客户端能识别的错误事件结束。非流式回答同样最多等待这么久。流式回答最多压住 15 秒,此后客户端先收到响应头和保活注释(Gemini 客户端只收到响应头),因此不会在网关仍可转到下一个上游时自行放弃。没有响应的上游,首 token 时间按等待的全部时长记录,「延迟最低」与「按速度」因此把它排到后面。请求的「路由」标签把这次尝试标为「无响应超时 · 300 秒」,并标出上游可能已计费的输入。长时间思考而不发出任何内容的模型,可能需要更长的超时。 - **并发上限。**设置了「并发上限」的上游同时最多接收这么多请求。上游已满时,留在它上面的对话等待空位,等不到再换下一个,其他请求直接交给下一个成员;所有成员都满时,请求等待最先空出的位置,最长为「设置 › 故障转移」中的「最多等待空位」(默认 30 秒),仍无空位则返回 429 并带 `Retry-After`;此前已有尝试发出时,返回那次尝试的错误。密钥的分钟、小时用量上限共用这段等待;流量页会标出每一次跳过和排队时间。 -- **暂停。**失败的上游按「设置 › 故障转移」暂停使用,默认值为:「连续失败」3 次后暂停 60 秒,此后每次加倍,最长 600 秒;「余额不足」暂停 30 分钟;「额度用完」暂停到重置时刻,未给出时暂停 60 分钟;「限流」按上游要求的等待时间暂停,最长 60 分钟。只有一个上游的规则不受暂停影响;所有成员都在暂停时,网关仍会逐个尝试。 +- **暂停。**失败的上游按「设置 › 故障转移」暂停使用,默认值为:「连续失败」3 次后暂停 60 秒,此后每次加倍,最长 600 秒;「余额不足」暂停 30 分钟;「额度用完」暂停到重置时刻,未给出时暂停 60 分钟;「限流」按上游要求的等待时间暂停,最长 60 分钟。只有一个上游的规则不受暂停影响;所有成员都在暂停时,网关仍会逐个尝试。在流量页手动中止的请求不计入上游的失败。 - **会话与提示缓存。**同一轮之内(客户端回传工具结果期间),请求沿用这一轮开头确定的规则和回答它的上游。跨轮时,如果上次回答在五分钟以内、且读写了至少 1,024 个缓存 token,对话继续使用该上游,因为换到别处要按全价重建缓存;否则,或者该上游正在暂停时,由策略重新排序,「轮询」正是在这时轮到下一个成员。请求「路由」标签中的「对话延续」一行说明请求是否因此留在原上游。 -- **模型名相同。**每个成员收到的都是客户端请求的模型名,或者规则改写之后的模型名。模型列表中没有该模型的成员会被跳过;没有模型列表的成员照常尝试,它返回 404 时请求换到下一个。成员能服务、却没有列进模型列表的模型,可以在上游页该成员的模型列表中选择「添加模型…」加上。 +- **模型名相同。**每个成员收到的都是客户端请求的模型名,或者规则改写之后的模型名。模型列表中没有该模型的成员会被跳过;没有模型列表的成员照常尝试,它返回 404 时请求换到下一个。成员能服务、却没有列进模型列表的模型,可以在上游页该成员的模型列表中选择「添加模型…」加上,也可以在它的编辑对话框的「模型」一节中添加。 - **回退到另一个模型。**添加一条规则:条件「选定上游」选择后备上游,「命中后」选择「继续匹配」,在「改写参数」的「模型改为」中填写它的模型。这样的规则在每次选定上游之后判断,故障转移之后同样适用。更换模型后已缓存的 prompt 不再命中,费用按发出的模型名计算。 相关文档:[切换中转站、上游或模型,无需重启 Claude Code 与 Codex](/zh-CN/docs/lite/switch-upstreams-without-restart/)、[功能详解](/zh-CN/docs/lite/features/#路由与故障转移)、[安装与更新](/zh-CN/docs/lite/install/)。 diff --git a/src/content/docs-lite/zh-CN/features.md b/src/content/docs-lite/zh-CN/features.md index fccaad7..0b83c67 100644 --- a/src/content/docs-lite/zh-CN/features.md +++ b/src/content/docs-lite/zh-CN/features.md @@ -14,11 +14,13 @@ ## 流量与会话 -流量页实时列出请求:状态、密钥、模型、上游、首 token 时间与总耗时(悬停时显示生成速度)、token 和费用,并标出格式转换、被脱敏的密钥、被内容过滤删除的文字,以及被拦截或可疑的工具调用。列表可以搜索(路径、密钥、上游、模型与错误信息),可以按密钥、上游和模型筛选,或只看失败、无法计价的请求。列表装有最近 2,000 条请求,搜索与筛选同时在全部请求记录中进行,可以继续向更早的记录搜索。打开「搜索内容」后,还会搜索每个请求新发送的内容(最后一轮用户消息,含工具结果)及其回答(含工具调用),范围以报文仍保留的请求为限;命中的片段显示在对应请求的下方。「会话」视图把同一段对话的请求归为若干轮次,给出每一轮的输入 token 与费用。会话的「对话」页按轮还原整段对话:每一轮新加入的消息和回答,包括文字、折叠的思考、工具调用及其结果,图片只显示类型和大小。系统提示词的变化和压缩上下文之后的重新开始会标出;报文已过保留期、过大未能完整保存或无法读取的轮次会注明原因。每一轮都可以打开对应的请求。 +流量页实时列出请求:状态、密钥、模型、上游、首 token 时间与总耗时(悬停时显示生成速度)、token 和费用,并标出格式转换、被脱敏的密钥、被内容过滤删除的文字,以及被拦截或可疑的工具调用。上游、模型和密钥的名称过长时截断显示,悬停可查看全名。列表可以搜索(路径、密钥、上游、模型与错误信息),可以按密钥、上游和模型筛选,或只看失败、无法计价的请求。列表装有最近 2,000 条请求,搜索与筛选同时在全部请求记录中进行,可以继续向更早的记录搜索。打开「搜索内容」后,还会搜索每个请求新发送的内容(最后一轮用户消息,含工具结果)及其回答(含工具调用),范围以报文仍保留的请求为限;命中的片段显示在对应请求的下方。「会话」视图把同一段对话的请求归为若干轮次,给出每一轮的输入 token 与费用。会话的「对话」页按轮还原整段对话:每一轮新加入的消息和回答,包括文字、折叠的思考、工具调用及其结果,图片只显示类型和大小。系统提示词的变化和压缩上下文之后的重新开始会标出;报文已过保留期、过大未能完整保存或无法读取的轮次会注明原因。每一轮都可以打开对应的请求。 打开一个请求可以查看时间线、路由(命中的规则、经过的策略组,以及每一次尝试的状态与耗时)、请求与响应正文、用量与费用。DeepSeek Harness 发出的请求还会显示所带会话日志的大小:这是客户端随每个请求附带的整段对话记录,发往 DeepSeek 以外的上游之前由网关去除。已结束的请求可以在预估费用后原样发送到另一个上游,两次的响应并排对照。 -尝试链还会标出因并发已满而跳过的上游、等待空位的排队时间,以及因开头超时而放弃的流式回答(「开头超时」),并注明该上游可能已计费的输入 token。因密钥的用量上限或上游并发均已满而未被接下的请求,在上游的位置显示「用量上限」或「并发已满」;以 Responses API 的 WebSocket 模式建立的连接按轮记为请求,每一轮单独统计用量与费用。 +尝试链还会标出因并发已满而跳过的上游、等待空位的排队时间、超过无响应超时仍未发出内容而被放弃的上游(默认值下显示为「无响应超时 · 300 秒」)及该上游可能已计费的输入 token,以及请求被手动中止时正在进行的尝试(「手动中止」)。因密钥的用量上限或上游并发均已满而未被接下的请求,在上游的位置显示「用量上限」或「并发已满」;以 Responses API 的 WebSocket 模式建立的连接按轮记为请求,每一轮单独统计用量与费用。 + +进行中的请求可以在它的操作菜单中选择「中止请求」停止。会话的操作菜单和会话详情顶部的「中止会话」经确认后停止该会话中所有进行中的请求。与上游的连接立即断开,客户端收到错误,请求以灰色标为「手动中止」,上游不会因此被暂停。Responses API 的 WebSocket 连接中的轮次不能中止。 ## 客户端接管 @@ -39,9 +41,9 @@ WSL 2 默认使用 NAT 网络,此时 Windows 上的网关无法从 WSL 内访 ## 上游 -上游是网关转发请求的目标:Anthropic、OpenAI、Google Gemini、DeepSeek 或任何兼容接口的 API 密钥,Amazon Bedrock(API 密钥、访问密钥或 AWS 配置文件),在应用内登录的 ChatGPT 账号或 Z.ai / BigModel 账号,OpenRouter 等中转服务,以及 Ollama 等本机模型。ChatGPT 账号显示订阅额度与重置时间;GLM Coding Plan 的上游(地址在 `api.z.ai` 或 `open.bigmodel.cn` 上,在应用内登录或手动填写密钥均可)同样显示:5 小时与每周额度,积分制套餐另外显示剩余积分(「剩余 1,976 / 2,000 积分」)。客户端与上游的 API 格式不同时,请求在 Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 与 Gemini 之间自动转换,无法转换的字段会在请求上逐一列出。上游可以经出站代理访问,也可以使用单独的价目表计价,别名、代理与价目表在同一页的标签中管理。可选的「并发上限」适用于限制并发的中转站或账号:上游已满时,留在它上面的对话等待空位,其他请求交给下一个上游;所有上游都满时,请求先等待,仍无空位则返回繁忙错误。在上游的模型列表中选择「规格…」,可以手动设置模型的上下文窗口、输出上限,以及是否支持推理与图片输入,适用于价目表中没有或数值有误的模型,手动设置的值优先于价目表。链路测速测量 DNS 解析以及 TCP、TLS、代理握手的耗时,不产生费用;推理测速测量首个 token 的时间,运行前先给出费用预估。 +上游是网关转发请求的目标:Anthropic、OpenAI、Google Gemini、DeepSeek 或任何兼容接口的 API 密钥,Amazon Bedrock(API 密钥、访问密钥或 AWS 配置文件),在应用内登录的 ChatGPT 账号或 Z.ai / BigModel 账号,OpenRouter 等中转服务,以及 Ollama 等本机模型。ChatGPT 账号显示订阅额度与重置时间;GLM Coding Plan 的上游(地址在 `api.z.ai` 或 `open.bigmodel.cn` 上,在应用内登录或手动填写密钥均可)同样显示:5 小时与每周额度,积分制套餐另外显示剩余积分(「剩余 1,976 / 2,000 积分」)。客户端与上游的 API 格式不同时,请求在 Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 与 Gemini 之间自动转换,无法转换的字段会在请求上逐一列出。Codex 等不标注提示缓存断点的客户端,请求转换后发往 Anthropic Messages 格式的上游,或 Bedrock 上 AWS 列为支持提示缓存的 Claude 模型(Claude 3.5 Sonnet v2、Claude 3.7 Sonnet 以及 Claude 4.5 起的模型)时,由网关标注断点,每一轮都能读取上一轮写入的缓存。上游拒绝这些标注时,网关去掉标注重新发送一次。上游可以经出站代理访问,也可以使用单独的价目表计价,别名、代理与价目表在同一页的标签中管理。可选的「并发上限」适用于限制并发的中转站或账号:上游已满时,留在它上面的对话等待空位,其他请求交给下一个上游;所有上游都满时,请求先等待,仍无空位则返回繁忙错误。在上游的模型列表中选择「规格…」,可以手动设置模型的上下文窗口、输出上限,以及是否支持推理与图片输入,适用于价目表中没有或数值有误的模型,手动设置的值优先于价目表。链路测速测量 DNS 解析以及 TCP、TLS、代理握手的耗时,不产生费用;推理测速测量首个 token 的时间,运行前先给出费用预估。 -一家上游的模型,是它自己列出的模型加上手动添加的模型。上游常常没有列全它能服务的模型:账号类的后端会隐藏较新的模型,中转站只列出一部分。在上游的模型列表中选择「添加模型…」,可以按确切的模型 ID 添加这类模型,一次可以添加多个。模型 ID 不能含 `*` 或 `?`,前后不能有空格,最多 256 个字符,且不能重复;不符合要求的,以及上游已经列出的,对话框在保存之前当场指出。手动添加的模型标有「手动」,路由与故障转移、别名、指定模型的路由规则、密钥的可用模型、客户端看到的模型列表、试算与推理测速都把它当作列出的模型。列表顶部的摘要把上游列出的模型与手动添加的模型分开计数。悬停在手动添加的模型上会出现「移除」,立即生效,随后的提示中可以「撤销」。上游不提供模型列表时,同一个对话框取代原来的手动清单,手动添加的模型就是它的全部模型。 +一家上游的模型,是它自己列出的模型加上手动添加的模型。上游常常没有列全它能服务的模型:账号类的后端会隐藏较新的模型,中转站只列出一部分。在上游的模型列表中选择「添加模型…」,可以按确切的模型 ID 添加这类模型,一次可以添加多个。模型 ID 不能含 `*` 或 `?`,前后不能有空格,最多 256 个字符,且不能重复;不符合要求的,以及上游已经列出的,对话框在保存之前当场指出。手动添加的模型标有「手动」,路由与故障转移、别名、指定模型的路由规则、密钥的可用模型、客户端看到的模型列表、试算与推理测速都把它当作列出的模型。列表顶部的摘要把上游列出的模型与手动添加的模型分开计数。悬停在手动添加的模型上会出现「移除」,立即生效,随后的提示中可以「撤销」。「新建上游」与「编辑上游」对话框的「模型」一节末尾有同样的输入框,无论上游是否提供模型列表,创建或编辑时都可以在这里添加模型,随上游的其他设置一起保存;启用范围为「指定模型」时,在这里添加的模型立即勾选启用。上游不提供模型列表时,手动添加的模型就是它的全部模型。 别名标签为模型设定客户端使用的名称。一个别名列出同一个模型在各家上游的名称,例如 Anthropic 上的 `claude-sonnet-5` 和 Bedrock 上的 `us.anthropic.claude-sonnet-5-v1:0`;提供其中任一名称的上游都能以自己的名称服务这个别名,并互为备用。客户端的模型列表里能看到别名,回答里的模型名写成客户端请求的名称,请求记录则保留每家上游实际收到的模型。同一个 Claude 模型在官方 API、Bedrock、Vertex 或 OpenRouter 上名称不同时,别名标签和新建别名对话框会建议合并。也可以在上游的模型列表里直接起别名。对话框列出每家上游将收到的名称,名称会接管另一家上游的同名模型时给出提示。密钥的可见模型允许某个上游模型时,列有它的别名也可以使用。 @@ -53,7 +55,7 @@ API 密钥和请求头的值可以写成 `${变量名}`,读取系统环境变 ## 路由与故障转移 -每把密钥使用一条路由,未指定的使用默认路由。路由由按顺序匹配的规则组成。规则的条件包括模型、密钥、客户端的 API 格式、输入 token 数、`max_tokens`、工具数量、图片、扩展思考、流式、提示缓存以及辅助请求的类型;命中后把请求交给某个上游、策略组或指定模型,或拒绝请求,也可以改写模型、`max_tokens` 或扩展思考。指定模型写明上游和它的某个模型,可按顺序列出备用,模型名原样发出;要让某个客户端把一个模型当另一个模型的名称使用,在这个客户端的密钥上写规则即可,不会改变这个名称对其他客户端的含义。条件写上游模型名时,也匹配列有它的别名。设置 `max_tokens` 可以限制回答的长度:上游到达上限时自行停止,不会报错。策略组把多个上游放在同一个名字下,并决定尝试的先后:按顺序、手动选择、轮询、延迟最低优先或费用最低优先。轮询策略组的每个成员有 1 到 100 的权重,即长期分到的请求比例;「分配依据」可以在权重之上再乘以各上游的速度(首 token 时间)、稳定性(最近的成功率)或两者,进行中的对话仍留在原上游。一次尝试失败时,请求转到下一个上游;同一会话默认保持在同一个上游上,以便提示缓存持续命中。在「设置 › 故障转移」中打开「开头超时时转到下一个上游」后,流式回答在等待时长内没有任何内容时,只要另有上游可以接下,请求也会转过去;最后一个上游照常等待。页面顶部的路由图显示每把密钥经过的路由、策略组和上游。 +每把密钥使用一条路由,未指定的使用默认路由。路由由按顺序匹配的规则组成。规则的条件包括模型、密钥、客户端的 API 格式、输入 token 数、`max_tokens`、工具数量、图片、扩展思考、流式、提示缓存以及辅助请求的类型;命中后把请求交给某个上游、策略组或指定模型,或拒绝请求,也可以改写模型、`max_tokens` 或扩展思考。指定模型写明上游和它的某个模型,可按顺序列出备用,模型名原样发出;要让某个客户端把一个模型当另一个模型的名称使用,在这个客户端的密钥上写规则即可,不会改变这个名称对其他客户端的含义。条件写上游模型名时,也匹配列有它的别名。设置 `max_tokens` 可以限制回答的长度:上游到达上限时自行停止,不会报错。策略组把多个上游放在同一个名字下,并决定尝试的先后:按顺序、手动选择、轮询、延迟最低优先或费用最低优先。轮询策略组的每个成员有 1 到 100 的权重,即长期分到的请求比例;「分配依据」可以在权重之上再乘以各上游的速度(首 token 时间)、稳定性(最近的成功率)或两者,进行中的对话仍留在原上游。一次尝试失败时,请求转到下一个上游;同一会话默认保持在同一个上游上,以便提示缓存持续命中。上游连续「设置 › 故障转移」中的「无响应超时」(默认 300 秒)没有发出内容时,网关放弃这个上游,保活心跳不算内容:尚未向客户端发出内容时,请求转到下一个上游,没有可用的上游时客户端收到超时错误;回答已经开始时,以错误结束。页面顶部的路由图显示每把密钥经过的路由、策略组和上游。 客户端自行发出的辅助请求(连通性检查、预热、生成标题、话题识别、输入建议)可以由网关在本地应答而不产生费用,也可以转发。转发的辅助请求和普通请求一样经过路由规则,规则可以按辅助请求的类型把它们分流到费用更低的上游。 @@ -87,7 +89,7 @@ MCP 页管理客户端从自己的配置文件中加载的内容,这些内容 ## 设置 -设置页分为七节。「连接」列出本机 core 和已保存的远程 core,详见[连接远程 core](/zh-CN/docs/lite/remote-core)。「通用」设置语言、外观、菜单栏显示的内容(仅 macOS,也可设为不显示)、开机启动、提醒以系统通知发送、仅在应用内显示还是关闭,并可让设为不再显示的引导提示重新显示。「网关监听」设置网关的访问范围(仅本机、所选网卡所在的局域网或所有网卡)、端口和放行网段。「故障转移」设置上游失败后暂停多久、请求交给下一个上游:连续失败几次后暂停、暂停多长,以及余额不足、额度用完和限流时各自的暂停时长;还设置流式回答等待开头的时长、开头超时时是否转到下一个上游(默认关闭),以及上游并发已满或密钥的分钟、小时上限用满时请求合计最多等待的时长(默认 30 秒)。「日志保留」分别设置请求报文与请求记录的保留天数,以及报文的空间上限。「关于」显示版本、检查更新,并可生成诊断包,其中的密钥与地址均已脱敏。「卸载」还原所有已接管的客户端并取消开机启动,应在删除应用之前执行。 +设置页分为七节。「连接」列出本机 core 和已保存的远程 core,详见[连接远程 core](/zh-CN/docs/lite/remote-core)。「通用」设置语言、外观、菜单栏显示的内容(仅 macOS,也可设为不显示)、开机启动、提醒以系统通知发送、仅在应用内显示还是关闭,并可让设为不再显示的引导提示重新显示。「网关监听」设置网关的访问范围(仅本机、所选网卡所在的局域网或所有网卡)、端口和放行网段。「故障转移」设置上游失败后暂停多久、请求交给下一个上游:连续失败几次后暂停、暂停多长,以及余额不足、额度用完和限流时各自的暂停时长;还设置无响应超时,即上游多久没有发出内容即被放弃(默认 300 秒,可设 30 到 3,600 秒),以及上游并发已满或密钥的分钟、小时上限用满时请求合计最多等待的时长(默认 30 秒)。「日志保留」分别设置请求报文与请求记录的保留天数,以及报文的空间上限。「关于」显示版本、检查更新,并可生成诊断包,其中的密钥与地址均已脱敏。「卸载」还原所有已接管的客户端并取消开机启动,应在删除应用之前执行。 ## 菜单栏、系统托盘与通知 diff --git a/src/content/docs/en/deployment-guide.md b/src/content/docs/en/deployment-guide.md index 323cd75..a7998f1 100644 --- a/src/content/docs/en/deployment-guide.md +++ b/src/content/docs/en/deployment-guide.md @@ -231,7 +231,7 @@ docker compose -f deploy/docker-compose.yml --env-file .env.production up -d To pin a release instead of `latest`, set `IMAGE_TAG` to its version, or to the SHA of a commit on `main`: ```bash -IMAGE_TAG=3.2.1 docker compose -f deploy/docker-compose.yml --env-file .env.production up -d +IMAGE_TAG=3.3.0 docker compose -f deploy/docker-compose.yml --env-file .env.production up -d ``` This starts: @@ -744,8 +744,8 @@ The server is stateless, so rolling updates work out of the box: ```bash # Update the image tag helm upgrade think-watch deploy/helm/think-watch \ - --set image.server.tag=3.2.1 \ - --set image.web.tag=3.2.1 \ + --set image.server.tag=3.3.0 \ + --set image.web.tag=3.3.0 \ --reuse-values ``` diff --git a/src/content/docs/zh-CN/deployment-guide.md b/src/content/docs/zh-CN/deployment-guide.md index 0c4f6a8..78e087f 100644 --- a/src/content/docs/zh-CN/deployment-guide.md +++ b/src/content/docs/zh-CN/deployment-guide.md @@ -231,7 +231,7 @@ docker compose -f deploy/docker-compose.yml --env-file .env.production up -d 如需固定某个版本而非 `latest`,将 `IMAGE_TAG` 设为该版本号,或 `main` 分支上某次提交的 SHA: ```bash -IMAGE_TAG=3.2.1 docker compose -f deploy/docker-compose.yml --env-file .env.production up -d +IMAGE_TAG=3.3.0 docker compose -f deploy/docker-compose.yml --env-file .env.production up -d ``` 这将启动: @@ -744,8 +744,8 @@ ThinkWatch 在 Web 控制台中内置了**配置指南**页面,位于 `/gatewa ```bash # Update the image tag helm upgrade think-watch deploy/helm/think-watch \ - --set image.server.tag=3.2.1 \ - --set image.web.tag=3.2.1 \ + --set image.server.tag=3.3.0 \ + --set image.web.tag=3.3.0 \ --reuse-values ``` diff --git a/src/data/core-docs/config.md b/src/data/core-docs/config.md index fae4f75..5775a5a 100644 --- a/src/data/core-docs/config.md +++ b/src/data/core-docs/config.md @@ -454,6 +454,26 @@ Identity fields that clients fill in themselves, such as Claude Code's `metadata.user_id`, are removed from the body. For an upstream that admits only certain clients, turn on `forward_client_identity`. +A request converted to Anthropic, or to Claude on Bedrock, marks where the +upstream may cache the prompt when the client marked nothing itself. Clients +in OpenAI or Gemini formats such as Codex cannot mark anything: those +providers cache a repeated prompt on their own, while Anthropic caches only +what is marked. The marks go at the end of the tools, at the end of the +system prompt and at the end of the last two user turns, at most four, each +kept for the default five minutes. The earlier of the two user marks is where +the previous request ended, so each turn reads back what the turn before +wrote and pays the cache price for it instead of the full input price. Cache +writes and reads are charged at the price table's cache prices. A request +that carries its own marks, as Claude Code's do, keeps exactly those, and a +request sent on in the upstream's own format is not changed. + +On Bedrock, marks are added only for the Claude models AWS lists as +supporting prompt caching: Claude 3.7 Sonnet, Claude 3.5 Sonnet v2, and every +Claude from version 4.5 on, including newer ones not yet listed. Older models, +such as Claude 3 Haiku, Sonnet 4 and Opus 4.1, get none. An upstream that +refuses the marks is sent the request once more without them, and that +upstream is not sent marks for that model again until the core restarts. + A ChatGPT account upstream (`protocol: chatgpt`) takes only the credential the desktop app obtains by signing in; it cannot be written by hand. Claude and Google subscription sign-ins are not supported; use an API key. @@ -1016,19 +1036,43 @@ Before the first content of a streamed answer reaches the client, an error the upstream sends in the stream moves the request to the next candidate, the same as an error status would. -An upstream can also be slow to start: it accepts the request and then sends -nothing for a long time. With `next_on_slow_start`, the request moves on to the -next candidate when no content has arrived `stream_start_wait_secs` after it -was sent. It is off by default, because models that think before they write -can take long to start; with it on, wait 30 seconds or more. The last -candidate always waits, and the upstream given up on is not set aside. A -candidate that is at its `max_concurrent` at that moment does not count as a -next one: the slow upstream keeps the request. +An upstream can also go quiet: it accepts the request and then sends nothing, +or stops partway through. After `idle_timeout_secs` (300 by default) without +content, the gateway stops waiting for it. The time counts from the moment the +request is sent to that upstream, so waiting for a free slot does not count, +and starts again with every piece of content: text, reasoning and tool calls +count; keep-alives (SSE comments, Anthropic's `ping`, empty chunks, Responses' +`response.in_progress`) do not, so an upstream that only keeps the connection +alive still runs out of time. A whole, non-streamed answer counts from sending +to the complete answer. + +- When no content has reached the client yet, the upstream counts as failed + (towards `failures_to_pause`, like a 5xx), the attempt appears with the + outcome `idle_timeout`, and the request moves to the next candidate. Until + then an upstream's answer is held back from the client, so the next upstream + starts it afresh. With no candidate left, the client gets a timeout error + (504) in its own format. The last candidate's stream is passed on as it + arrives, so once its response has started, a timeout there ends it with an + error event instead. +- A streamed answer is held for at most 15 seconds. If no content has + arrived by then, the client receives `200` and the streaming headers, so + that its own wait for headers does not run out, followed by an SSE comment + (`: keep-alive`) every 15 seconds; Gemini clients get no comments. The + upstream's events stay held until its first content, and failover continues + as before under the same `200`: the next upstream's stream starts cleanly, + and when no candidate is left, the stream ends with an error event in the + client's format instead of a 504. The comments do not count as content. +- When content has already reached the client, the request cannot move on + without repeating it: the answer ends with an error event in the client's + format, and the request is recorded as failed. + +The conversation then no longer stays on that upstream for the rest of its turn. +Models that think for a long time before they write anything need a longer +timeout. ```yaml failover: - stream_start_wait_secs: 30 - next_on_slow_start: true + idle_timeout_secs: 600 ``` When upstreams are at their `max_concurrent`, a request waits for a free slot @@ -1053,8 +1097,7 @@ failure instead. | `no_balance_pause_secs` | integer | `1800` | Seconds to set aside an upstream that reports an insufficient balance. | | `quota_pause_secs` | integer | `3600` | Seconds to set aside an upstream whose quota is used up when it does not say when the quota resets. When it does, the upstream is set aside until then. | | `rate_limit_max_pause_secs` | integer | `3600` | A rate-limited upstream is set aside for the time its `Retry-After` gives, at most this many seconds. Without `Retry-After` it counts as a failure without a stated reason. | -| `stream_start_wait_secs` | integer | `15` | Seconds to hold a streamed answer until its first content arrives. An error before then moves the request to the next upstream; after this long, what has arrived is passed on. From 1 to 120. | -| `next_on_slow_start` | bool | `false` | When a streamed answer still has no content `stream_start_wait_secs` after the request was sent, give up on that upstream and send the request to the next one. The last upstream always waits. The upstream given up on is not set aside. Needs `stream_start_wait_secs` of at least 5. | +| `idle_timeout_secs` | integer | `300` | Seconds an upstream may go without sending content before the gateway stops waiting for it. Counted from the moment the request is sent and started again by every piece of content: text, reasoning and tool calls count, keep-alives do not. A whole (non-streamed) answer counts from sending to the complete answer. Before any content has reached the client, the upstream counts as failed and the request moves to the next one; with none left, the client gets a timeout error. After content has reached the client, the answer ends with an error. From 30 to 3600. | | `slot_wait_secs` | integer | `30` | Seconds a request waits in all, counted once the key's own `max_concurrent` lets it in: for a key's `minute` or `hour` limit to free up, and for a free slot on upstreams at their `max_concurrent`. A key limit that does not free up in time refuses the request; without an upstream slot in time it goes to the next upstream, or, when every candidate is full, is answered with 429. `0`: never wait. From 0 to 300. | @@ -1168,11 +1211,11 @@ group shares out requests by the result in the same way as above. tenth. - `health`: upstreams that fail less get a larger share. It looks at the last 50 requests within the past 30 minutes. Server errors, rate limits, - used-up quota or balance, rejected credentials, timeouts and connection - errors count as failures; errors caused by the request itself do not, and - neither does a client that cancels, a switch away from a stream that is - slow to start, or an upstream skipped because it is at its - `max_concurrent`. An upstream that keeps failing keeps a twentieth of its + used-up quota or balance, rejected credentials, timeouts (including + `failover.idle_timeout_secs` before any content) and connection errors + count as failures; errors caused by the request itself do not, and neither + does a client that cancels, a request aborted by hand, or an upstream + skipped because it is at its `max_concurrent`. An upstream that keeps failing keeps a twentieth of its weight, so it still gets the occasional request and its recovery is noticed; one that fails outright is set aside by [`failover`](#cfg-failover) as before. @@ -1180,8 +1223,8 @@ group shares out requests by the result in the same way as above. Speed is measured on streamed answers only, from the moment the request is sent to that upstream, so waiting and upstreams that failed before it do not -count. An upstream given up on because its stream was slow to start -(`failover.next_on_slow_start`) counts as taking the whole wait. On a +count. An upstream given up on because it sent no content within +`failover.idle_timeout_secs` counts as taking the whole wait. On a Responses WebSocket connection, each `response.create` counts as one request for both speed and failures, its speed measured from the moment the upstream starts answering it. diff --git a/src/data/core-docs/config.zh-CN.md b/src/data/core-docs/config.zh-CN.md index 6e5f3e3..bb98524 100644 --- a/src/data/core-docs/config.zh-CN.md +++ b/src/data/core-docs/config.zh-CN.md @@ -333,6 +333,10 @@ providers: 每个请求只带请求本身和上游需要的请求头,客户端的其他信息一律不发:凭据和 `headers` 中写的请求头、ThinkWatch 自己的 `User-Agent`,以及客户端请求中该上游协议使用的请求头(Anthropic 为 `anthropic-*`,OpenAI 为 `Idempotency-Key` 和 `X-Client-Request-Id`,Gemini 没有)。客户端自动填写的身份字段(如 Claude Code 的 `metadata.user_id`)从请求体中去掉。只接受特定客户端的上游,打开 `forward_client_identity`。 +请求转换为 Anthropic 格式、或转给 Bedrock 上的 Claude 时,客户端自己没有标出缓存位置的,由网关标出可以缓存的位置。Codex 等使用 OpenAI 或 Gemini 格式的客户端无从标注:这两家自动缓存重复的提示,Anthropic 只缓存标出的部分。标注的位置是工具列表末尾、系统提示末尾和最后两条用户消息的末尾,最多四处,缓存时长为默认的五分钟。两条用户消息中靠前的那一处正是上一个请求结束的位置,因此每一轮都能读回上一轮写入的缓存,这部分按缓存价而不是全额输入价计费。缓存的写入和读取按价目表中的缓存单价计费。自带标注的请求(如 Claude Code 的)保持原样,按上游自身格式直接转发的请求不做改动。 + +Bedrock 上只为 AWS 列出支持提示缓存的 Claude 模型标注:Claude 3.7 Sonnet、Claude 3.5 Sonnet v2,以及 4.5 起的所有 Claude,包括尚未列出的更新版本。更早的模型(如 Claude 3 Haiku、Sonnet 4、Opus 4.1)不标注。上游拒绝这些标注时,网关去掉标注重新发送一次,此后在 core 重启之前不再为这家上游的该模型标注。 + ChatGPT 账号上游(`protocol: chatgpt`)只接受桌面应用登录得到的凭据,不能手写。不支持 Claude 和 Google 的订阅登录,请使用 API 密钥。 有的中转站和账号同时只接受几个请求,多出来的直接拒绝。`max_concurrent` 让网关守住这个数:请求发出时占用这家的一个位置,回答完整交给客户端、或者客户端断开时归还。这家满了的时候,为复用提示缓存而留在这家的对话等空位,别的请求直接换下一家。最多等多久由 `failover.slot_wait_secs` 决定。等待不算失败,这家不会因此停用。只计算 token 数的请求不占位置。Responses 的 WebSocket 连接上,每个 `response.create` 从发出起占一个位置,直到它的回答结束,密钥的 `max_concurrent` 也一样;空闲的连接不占位置。 @@ -785,16 +789,28 @@ security: 流式回答在第一段内容交给客户端之前,上游在流里报的错误和错误状态码一样,会把 请求换到下一个候选。 -上游也可能开头很慢:收下请求之后很久都不发内容。开启 `next_on_slow_start` 后,请求 -发出 `stream_start_wait_secs` 秒仍没有内容,就换到下一个候选。默认关闭,因为先思考 -再输出的模型本来就可能很久才开始;开启时建议等 30 秒以上。最后一个候选总是等下去, -被放弃的上游不会停用。到点那一刻并发数已满(`max_concurrent`)的候选不算下一个:请求 -留在慢的那一家。 +上游也可能不出声:收下请求之后一直不发内容,或者发到一半停住。连续 +`idle_timeout_secs` 秒(默认 300)没有内容,网关就不再等它。计时从请求发给这家上游的 +那一刻起,等空位的时间不算在内;每来一段内容重新计时:正文、推理、工具调用都算,心跳 +(SSE 注释、Anthropic 的 `ping`、空块、Responses 的 `response.in_progress`)不算,所以 +只保持连接、不出内容的上游同样会超时。整包(非流式)的回答从发出算到整份回来。 + +- 还没有内容交给客户端时,这家上游记一次失败(和 5xx 一样计入 `failures_to_pause`), + 尝试链上这一跳的结果是 `idle_timeout`,请求换到下一个候选。在此之前上游的回答不交给 + 客户端,下一家从头开始回答。没有候选了,客户端收到它自己格式的超时错误(504)。 + 最后一个候选的流是边收边交给客户端的,它的响应开始之后再超时,回答以一条错误事件收尾。 +- 流式回答最多暂存 15 秒。到时还没有内容,客户端先收到 `200` 和流式响应头,以免它自己 + 等响应头的时限先到,之后每 15 秒收到一行 SSE 注释(`: keep-alive`);Gemini 的客户端 + 不发注释。上游的事件照旧暂存到第一段内容,故障转移在同一个 `200` 之下照常进行:下一家 + 的流从头开始;没有候选了,流以客户端格式的错误事件结束,不再是 504。注释不算内容。 +- 已经有内容交给客户端时,换一家会重复已经发出的内容,所以不换:回答按客户端的格式 + 以一条错误事件收尾,请求记为失败。 + +之后这段对话在这一轮里不再留在这家上游。先思考很久才输出的模型,需要把超时调长。 ```yaml failover: - stream_start_wait_secs: 30 - next_on_slow_start: true + idle_timeout_secs: 600 ``` 上游的并发数满了(`max_concurrent`)时,一个请求等空位合计最多 `slot_wait_secs` @@ -815,8 +831,7 @@ failover: | `no_balance_pause_secs` | 整数 | `1800` | 上游报告余额不足时停用的秒数。 | | `quota_pause_secs` | 整数 | `3600` | 上游报告额度用完、但没有给出重置时间时停用的秒数。给出了重置时间的,停用到那一刻。 | | `rate_limit_max_pause_secs` | 整数 | `3600` | 被限流的上游按它给的 `Retry-After` 停用,最多这么多秒。没有 `Retry-After` 的按没有说明原因的失败计。 | -| `stream_start_wait_secs` | 整数 | `15` | 流式回答在第一段内容到达前最多暂存的秒数。在此之前上游报错,请求换到下一家;超过这个时间,已收到的部分照常交给客户端。取值 1 到 120。 | -| `next_on_slow_start` | 布尔 | `false` | 流式回答在请求发出 `stream_start_wait_secs` 秒后仍没有内容时,放弃这家上游,把请求交给下一家。最后一家总是等下去。被放弃的上游不会停用。开启时 `stream_start_wait_secs` 至少为 5。 | +| `idle_timeout_secs` | 整数 | `300` | 上游多少秒没有发出内容,网关就不再等它。从请求发出的那一刻算起,每来一段内容重新计时:正文、推理、工具调用都算,心跳不算。整包(非流式)的回答从发出算到整份回来。还没有内容交给客户端时,这家上游记一次失败,请求换到下一家;没有下一家了,客户端收到超时错误。已经有内容交给客户端的,回答以一条错误收尾。取值 30 到 3600。 | | `slot_wait_secs` | 整数 | `30` | 一个请求合计最多等的秒数,从过了密钥自己的 `max_concurrent` 时算起:等密钥的 `minute`、`hour` 上限空出名额,和等并发数满了(`max_concurrent`)的上游空出位置,都算在里面。密钥的上限到时空不出来就拒绝;等不到上游的空位就换下一家,候选全满时回 429。`0`:不等。取值 0 到 300。 | @@ -885,10 +900,10 @@ groups: - `weights`(默认):只按权重。 - `latency`:越快的上游分得越多。快慢看典型的从发出请求到回答第一段内容的时间,与 `url-test` 使用同一份测量。比组内居中者快一倍的上游,权重乘以四;最多乘以十,最少乘以十分之一。 -- `health`:越少失败的上游分得越多。依据是最近 30 分钟内的最近 50 次请求:服务器错误、限流、额度或余额用尽、凭据被拒、超时和连接失败算作失败;请求本身导致的错误不算,客户端取消、因开头太慢而换走、因并发数满了(`max_concurrent`)而跳过也不算。经常失败的上游至少保留权重的二十分之一,仍会偶尔分到请求,以便发现它已经恢复;完全失败的上游照旧由 [`failover`](#cfg-failover) 暂停。 +- `health`:越少失败的上游分得越多。依据是最近 30 分钟内的最近 50 次请求:服务器错误、限流、额度或余额用尽、凭据被拒、超时和连接失败算作失败;请求本身导致的错误不算,客户端取消、手动中止、因并发数满了(`max_concurrent`)而跳过也不算;超过 `failover.idle_timeout_secs` 仍没有内容算作超时。经常失败的上游至少保留权重的二十分之一,仍会偶尔分到请求,以便发现它已经恢复;完全失败的上游照旧由 [`failover`](#cfg-failover) 暂停。 - `latency-health`:两个系数相乘。 -快慢只在流式回答上测,从请求发给这家上游的那一刻算起:之前的等待、之前失败的上游都不算在内。因开头太慢而被放弃的上游(`failover.next_on_slow_start`),按等满的那段时间计。Responses 的 WebSocket 连接上,每个 `response.create` 在快慢和成败上都算一个请求,快慢从上游开始回答它的那一刻算起。 +快慢只在流式回答上测,从请求发给这家上游的那一刻算起:之前的等待、之前失败的上游都不算在内。因超过 `failover.idle_timeout_secs` 没有内容而被放弃的上游,按等满的那段时间计。Responses 的 WebSocket 连接上,每个 `response.create` 在快慢和成败上都算一个请求,快慢从上游开始回答它的那一刻算起。 测量还不够的上游按中等对待。与只按权重时一样,进行中的对话留在原来的上游,差额由新对话补齐。 diff --git a/src/data/core-docs/manifest.json b/src/data/core-docs/manifest.json index 3e81802..fd0064a 100644 --- a/src/data/core-docs/manifest.json +++ b/src/data/core-docs/manifest.json @@ -1,9 +1,9 @@ { "repository": "ThinkWatchProject/ThinkWatch-Core", - "ref": "v0.66.0", + "ref": "v0.67.1", "files": { - "docs/config.md": "91818dc4be92174ddcae4c41c8fdd5d514d80a0913dea84b4dcdc656380752f0", - "docs/config.zh-CN.md": "9550dd8d65a69b3096dc53337d8ee439f239ddc8185f4faa7b1171373876aad0", + "docs/config.md": "31d4ff32e28e94d2e53816147c51ef2818a9558265c147612ddfc39b0942beb8", + "docs/config.zh-CN.md": "d7ab9cde90fb6eb910b7b4261a8ccdc9bff1e70b97262860fecc1713dd285117", "docs/server.md": "5e1e9b901bef2b46d417aea057a1db24c78457d3c1938b8540ecb4766515b68f", "docs/server.zh-CN.md": "1f508e7c39b8ded28653773ca4d8701e6bdc2247bcd359c3a8fd00fd1401a6ff" }