Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions src/components/ImportLinkBuilder.astro
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ const t = zh
protocol: "接口协议",
auto: "自动识别(不写入链接)",
key: "API 密钥(可选)",
models: "手动模型清单(可选,逗号分隔)",
models: "手动添加的模型(可选,逗号分隔)",
appLink: "应用链接",
webLink: "网页链接",
copy: "复制",
Expand All @@ -35,7 +35,7 @@ const t = zh
url: "接口地址不符合要求:须为 https 地址(http 仅限 localhost、127.0.0.1、[::1]),不含账号、查询串、片段、空白与 % $ { } 等字符。",
protocol: "接口协议不受支持。",
key: "API 密钥只能包含字母、数字与 - _ . ~ + / = :,最长 512 个字符。",
models: "模型清单不符合要求:以逗号分隔,每项只含字母、数字与 - _ . : / @ +,最多 64 项。",
models: "手动添加的模型不符合要求:以逗号分隔,每项只含字母、数字与 - _ . : / @ +,最多 64 项。",
},
}
: {
Expand All @@ -46,7 +46,7 @@ const t = zh
protocol: "Protocol",
auto: "Auto-detect (left out of the link)",
key: "API key (optional)",
models: "Manual model list (optional, comma-separated)",
models: "Models added by hand (optional, comma-separated)",
appLink: "App link",
webLink: "Web link",
copy: "Copy",
Expand All @@ -63,7 +63,7 @@ const t = zh
url: "The base URL is not accepted: it must be an https address (http only for localhost, 127.0.0.1 and [::1]) without credentials, a query, a fragment, spaces or characters such as % $ { }.",
protocol: "The protocol is not supported.",
key: "The API key may contain only letters, digits and - _ . ~ + / = :, up to 512 characters.",
models: "The model list is not accepted: comma-separated, each entry only letters, digits and - _ . : / @ +, at most 64 entries.",
models: "The models added by hand are not accepted: comma-separated, each entry only letters, digits and - _ . : / @ +, at most 64 entries.",
},
};

Expand Down
8 changes: 4 additions & 4 deletions src/components/pages/ImportPage.astro
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ const t = zh
noKey: "未提供",
show: "显示",
hide: "隐藏",
models: "手动模型清单",
models: "手动添加的模型",
open: "在 ThinkWatch Lite 中打开",
notInstalled: "尚未安装 ThinkWatch Lite 时,先下载并安装,再回到此页面。",
invalidTitle: "此导入链接无效",
Expand All @@ -54,7 +54,7 @@ const t = zh
url: "接口地址不符合要求:须为 https 地址(http 仅限本机),且不含账号、查询串或片段。",
protocol: "接口协议不受支持。",
key: "API 密钥包含不支持的字符。",
models: "模型清单不符合要求。",
models: "手动添加的模型不符合要求。",
},
}
: {
Expand All @@ -73,7 +73,7 @@ const t = zh
noKey: "Not provided",
show: "Show",
hide: "Hide",
models: "Manual model list",
models: "Models added by hand",
open: "Open in ThinkWatch Lite",
notInstalled: "If ThinkWatch Lite is not installed yet, download and install it, then return to this page.",
invalidTitle: "This import link is not valid",
Expand All @@ -92,7 +92,7 @@ const t = zh
url: "The base URL is not accepted: it must be an https address (http only for this computer) without credentials, a query or a fragment.",
protocol: "The protocol is not supported.",
key: "The API key contains unsupported characters.",
models: "The model list is not accepted.",
models: "The models added by hand are not accepted.",
},
};

Expand Down
4 changes: 2 additions & 2 deletions src/content/docs-lite/en/failover-and-load-balancing.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ ThinkWatch Lite turns each relay key into an upstream and puts several upstreams

## Before you start

- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.6.
- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.10.
- The base URL and API keys of each relay.

## Steps
Expand Down Expand Up @@ -32,7 +32,7 @@ A strategy only sets the order; every member remains available for failover.
- **Concurrency limits.** An upstream with a **Concurrency limit** takes at most that many requests at once. When it is full, a conversation that stays on it waits for a free slot and then moves on, and other requests go straight to the next member; when every member is full, a request waits for the first free slot for up to **Wait for a free slot at most** in Settings › Failover, 30 seconds by default, and then receives a 429 with `Retry-After`, or the error of an earlier attempt when one was sent. The same wait covers a key's per-minute and per-hour usage limits, and the Traffic page shows each skip and the time queued.
- **Pauses.** A failing upstream is paused for as long as Settings › Failover sets: by default 60 seconds after 3 **Consecutive failures**, doubling up to 600; 30 minutes for **Insufficient balance**; until the reset time, or 60 minutes, for **Quota used up**; the wait a rate-limited upstream asks for, up to 60 minutes. A rule with a single upstream is never held back, and when every member is paused they are tried anyway.
- **Sessions and prompt cache.** Within a turn, while the client sends tool results back, requests keep the rule chosen at the start of the turn and the upstream that answered. Across turns, a conversation stays with the upstream that answered last if that answer read or wrote at least 1,024 cached tokens within the last five minutes; moving would rebuild the cache at full price. Otherwise, or when that upstream is paused, the strategy orders the members again, which is when **Round robin** moves on. The **Conversation** line on a request's **Routing** tab shows when a request stayed.
- **Same model name.** Every member is asked for the model the client sent, or the name a rule rewrote it to. A member whose model list lacks it is skipped; one without a list is tried, and its 404 moves the request on.
- **Same model name.** Every member is asked for the model the client sent, or the name a rule rewrote it to. A member whose model list lacks it is skipped; one without a list is tried, and its 404 moves the request on. A model that a member serves but leaves out of its list can be added with **Add models…** in the member's model list on the Upstreams page.
- **Falling back to another model.** Add a rule with the condition **Selected upstream** set to the backup upstream, **On match** set to **Continue matching**, and its model in **Change model to** under **Parameter rewrites**. Such a rule is evaluated for each upstream as it is tried, failover included. The changed model no longer hits the cached prompt, and the request is priced by the name sent.

Related: [Switch relays, upstreams or models without restarting Claude Code or Codex](/docs/lite/switch-upstreams-without-restart/), [Features](/docs/lite/features/#routing-and-failover), [Install and update](/docs/lite/install/).
8 changes: 5 additions & 3 deletions src/content/docs-lite/en/features.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,8 @@ Usage limits keep one client from using up a budget: each caps the key at a numb

Upstreams are the services requests are forwarded to: API keys for Anthropic, OpenAI, Google Gemini, DeepSeek or any compatible endpoint, Amazon Bedrock (with an API key, access keys or an AWS profile), a ChatGPT account or a Z.ai / BigModel account signed in from the app, relays such as OpenRouter, and local models such as Ollama. A ChatGPT account shows its usage limits and reset times. So does an upstream on a GLM Coding Plan, that is, one whose address is on `api.z.ai` or `open.bigmodel.cn`, whether it was signed in from the app or added with a key: its 5-hour and weekly limits and, on a plan billed in credits, the credits left (“1,976 / 2,000 credits left”). When a client and an upstream use different API formats, requests are converted between Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini, and the fields that cannot be carried over are listed on the request. Upstreams can be reached through an outbound proxy and priced with a price sheet of their own; aliases, proxies and price sheets have tabs on the same page. An optional **Concurrency limit** suits relays and accounts that allow only so many requests at once: when the upstream is full, a conversation that stays on it waits for a free slot, other requests go to the next upstream, and when every upstream is full a request waits and then receives a busy error. **Specs…** in an upstream's model list sets by hand a model's context window, maximum output, and whether it reasons and takes images, for a model the price table lacks or gets wrong; the values set there are used in place of the price table's. A connection test times the DNS lookup and the TCP, TLS and proxy handshakes without incurring any cost; an inference test measures the time to first token and estimates its cost before it runs.

An upstream's models are the ones it lists plus any added by hand. Upstreams often leave out models they serve: an account backend hides newer models, a relay lists only some. **Add models…** in an upstream's model list adds such models by their exact IDs, several at a time. An ID cannot contain `*` or `?` or have spaces around it, is at most 256 characters, and is added once; the dialog points out an ID that breaks these rules, or one the upstream already lists, before saving. A model added by hand is marked **Manual**, and routing and failover, aliases, routing rules that name a model, a key's allowed models, the model list clients see, the dry run and the speed test all treat it like a listed model. The summary at the top of the list counts the models the upstream lists apart from those added by hand. Hovering over a manual model shows **Remove**, which takes effect at once, with **Undo** in the message that follows. For an upstream that lists no models, the same dialog replaces the manual model list, and the models added there are its whole list.

The Aliases tab gives a model the name clients use for it. An alias lists the names the same model has on different upstreams, such as `claude-sonnet-5` on Anthropic and `us.anthropic.claude-sonnet-5-v1:0` on Bedrock; every upstream that offers one of them serves the alias under its own name, and they back each other up. Clients see aliases in their model lists, and answers carry the name the client asked for, while the request log shows the model each upstream was sent. When the same Claude model has different names on the official API, Bedrock, Vertex or OpenRouter, the tab and the new-alias dialog suggest grouping them. An alias can also be created from an upstream's model list. The dialog shows which upstream receives which name and warns when the name would take over another upstream's model of the same name. A key whose model scope allows an upstream model can also use the aliases that list it.

Each upstream is also compared over the last 7 days with the others serving the same model, and a deviation is marked next to its name in the list: **Model name differs** when answers name a model other than the one sent, **Input reported high** or **Input reported low** when the input tokens it reports, as a multiple of the gateway's own estimate, are well off the other upstreams', and **Low cache reads** when follow-up turns read a smaller share of their input from the prompt cache. Hovering over a mark shows the evidence and the sample sizes, and a deviation is marked only when both sides have enough samples.
Expand Down Expand Up @@ -77,19 +79,19 @@ The MCP page covers what clients load from their own configuration files, which
- **Skills and hooks:** the installed skills and configured hooks, with the client each belongs to; skills in the shared `~/.agents/skills` folder are listed as such.
- **Findings:** client configuration, skills, hooks, slash commands, subagents and project instruction files are scanned for hidden characters, prompt injection, dangerous commands and overly broad permissions, and each finding is graded high, medium or low. The scan only reports; it never changes a file.

The app watches these files while it runs, and a new finding raises a system notification.
The app watches these files while it runs, and a new finding raises a system notification. After a file is edited, only the findings that are new are reported; those already in it are not reported again, even when they moved to another line.

## Plugins

Plugins are short JavaScript files that change requests before they go to an upstream and answers before they reach the client, such as asking for answers in a chosen language or converting file paths in tool calls between WSL and Windows. They run after routing, in a sandbox inside core with no network, files or memory between requests, see placeholders instead of the keys in a request, and pass through the same protections afterwards. Two plugins ship with the app, both off until turned on. Each plugin is one file that also holds its scope, its behavior on errors and its settings. It is edited in one editor with a Settings tab, whose changes are written into the file, and a Code tab; adding a plugin opens the same editor. Routine changes are saved directly; for a plugin that may change tool calls, installing it, turning it on, changing its code and approving a changed file are confirmed in a system dialog. A plugin whose file changes outside the app stops running until the new version is approved. The page lists each plugin with its status, permissions, scope and statistics, and offers a trial run on a recent request and its log; requests changed by plugins are marked in Traffic. The API, permissions, limits and security model are described in [Plugins](/docs/lite/plugins).

## Settings

Settings has seven sections. Connection lists the local core and the saved remote cores, described in [Connecting to a remote core](/docs/lite/remote-core). General sets the language, the appearance, what the menu bar item shows on macOS, launch at login, whether notices arrive as system notifications, in the app only or not at all, and shows hidden guidance hints again. Listening sets who can reach the gateway (this machine only, the local network of a chosen interface, or every interface), its port and the allowed address ranges. Failover sets how long a failing upstream is paused before requests go to the next one: after how many consecutive failures, for how long, and separate pauses for an insufficient balance, a used-up quota and rate limits. It also sets how long to wait for a streamed answer to start, whether a stream that has not started by then moves to the next upstream (off by default), and how long in total a request may wait for a free slot on a full upstream or for a key's per-minute or per-hour limit, 30 seconds by default. Log retention sets how long request payloads and request records are kept, and a size cap for payloads. About shows the version, checks for updates and produces a diagnostics bundle with keys and addresses masked. Uninstall restores every connected client and removes the autostart entry, and is meant to be run before the app is deleted.
Settings has seven sections. Connection lists the local core and the saved remote cores, described in [Connecting to a remote core](/docs/lite/remote-core). General sets the language, the appearance, what the menu bar item shows on macOS (or that it is hidden), launch at login, whether notices arrive as system notifications, in the app only or not at all, and shows hidden guidance hints again. Listening sets who can reach the gateway (this machine only, the local network of a chosen interface, or every interface), its port and the allowed address ranges. Failover sets how long a failing upstream is paused before requests go to the next one: after how many consecutive failures, for how long, and separate pauses for an insufficient balance, a used-up quota and rate limits. It also sets how long to wait for a streamed answer to start, whether a stream that has not started by then moves to the next upstream (off by default), and how long in total a request may wait for a free slot on a full upstream or for a key's per-minute or per-hour limit, 30 seconds by default. Log retention sets how long request payloads and request records are kept, and a size cap for payloads. About shows the version, checks for updates and produces a diagnostics bundle with keys and addresses masked. Uninstall restores every connected client and removes the autostart entry, and is meant to be run before the app is deleted.

## Menu bar, system tray and notifications

On macOS the menu bar shows today's tokens above today's cost; the numbers turn orange when a subscription quota is nearly used up and red when it is, and Settings can reduce the item to the icon or to the numbers.
On macOS the menu bar shows today's tokens above today's cost; the numbers turn orange when a subscription quota is nearly used up and red when it is. **Menu bar** in Settings › General can reduce the item to the icon or to the numbers, or set it to **Hidden**, which takes it off the menu bar: the gateway keeps running in the background, after a launch at login as well, and opening ThinkWatch Lite again from Finder or Spotlight shows the main window.

Clicking it opens a native menu. **Open ThinkWatch Lite** always comes first, followed by unread notices and a **Today** block: today's tokens in large type, with requests, failures and cost on the line below and a small chart of tokens per hour beside them. The block's top line gives the gateway's state, the server's name when connected to a remote core, and the generation speed over the last minute. Below it, each quota window an upstream reports has a row with how much is used and when it resets (subscription accounts and GLM Coding Plan upstreams; a plan billed in credits shows the credits left under the bar), followed by the requests in progress. The actions come last: choosing the upstream of a manually selected group, copying the gateway address, which is shown beside the item, or the default key, installing a new version when one is available, settings, switching connections and checking for updates, all without opening the main window.

Expand Down
Loading
Loading