Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions src/content/docs-lite/en/codex-other-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ ThinkWatch Lite connects Codex to a gateway on the same computer that accepts th

## Steps

1. On the Upstreams page, choose **New upstream** and pick a **Service**: **Anthropic** for Claude, or **Google Gemini**; each fills in the address and protocol. For a relay, choose **Custom**, enter its **Base URL** without an endpoint path such as `/chat/completions`, and set **Protocol** to **OpenAI Chat Completions**; for GLM, for example, `https://api.z.ai/api/paas/v4`, or `…/api/coding/paas/v4` on a GLM Coding Plan. A base URL that ends with its own version, such as `/v4` or Volcengine Ark's `/api/v3`, is used as written from ThinkWatch Lite 2026.10.5. Enter the **API key**, choose **Check connection**, then **Next**. If the relay does not list its models, enter them one per line under **Manual list**. Choose **Next**, then **Create**.
1. On the Upstreams page, choose **New upstream** and pick a **Service**: **Anthropic** for Claude, or **Google Gemini**; each fills in the address and protocol. For a relay, choose **Custom**, enter its **Base URL** without an endpoint path such as `/chat/completions`, and set **Protocol** to **OpenAI Chat Completions**; for GLM, for example, `https://api.z.ai/api/paas/v4`, or `…/api/coding/paas/v4` on a GLM Coding Plan. A base URL that ends with its own version, such as `/v4` or Volcengine Ark's `/api/v3`, is used as written from ThinkWatch Lite 2026.10.5. Enter the **API key**, choose **Check connection**, then **Next**. If the relay lists no models, or leaves some out, type each missing model ID in the box at the end of **Models** and press Enter. Choose **Next**, then **Create**.
2. On the Clients page, choose **Connect…** on the Codex row. The dialog shows the change to `~/.codex/config.toml`:

| Field | Value |
Expand All @@ -31,8 +31,9 @@ ThinkWatch Lite connects Codex to a gateway on the same computer that accepts th

- **Codex's model table.** Codex carries metadata for its own models, such as the context window, inside the program. A model it does not know, such as a Claude or Gemini model, runs on fallback metadata with a 272,000-token context window, and Codex warns: "Model metadata for `<model>` not found. Defaulting to fallback metadata; this can degrade performance and cause issues." For a model with a smaller window, `model_context_window` in `config.toml` sets the window Codex assumes. The gateway's model list does not appear in Codex's model picker, as Codex expects a catalog in its own format.
- **Conversion.** Requests and streamed answers are converted in both directions. Traffic marks such requests **Converted**; fields the target format cannot carry are dropped and listed in the request details. Server-side tools such as web search run only at the provider they belong to and are dropped in conversion. Codex accepts only `responses` for `wire_api`, so this conversion is what makes a Chat Completions-only relay usable.
- **Tools, history and compaction.** From ThinkWatch Lite 2026.10.11, every tool Codex declares reaches the upstream, including those it lists in its input rather than in `tools`, and tool calls come back under the names Codex gave them. Local shell calls, tool search and reasoning-effort changes in the history are converted as well. Compacting a long session works: the upstream writes a summary, which Codex keeps as its compaction and sends back in later requests. A compaction encrypted by OpenAI cannot be read by another upstream, and a request that carries one is refused with an error saying so. Instructions Codex adds in the middle of a conversation stay where they are, so the system prompt stays the same from turn to turn and the prompt cache keeps hitting.
- **Credentials.** With `requires_openai_auth = false`, Codex authenticates to the gateway with its own key and sends no OpenAI key or ChatGPT token. A ChatGPT account signed in from the app can still serve the OpenAI models next to Claude or Gemini: each request goes to an upstream that lists the model asked for.
- **Sessions.** Codex lists sessions started before and after connecting separately. `codex resume <session ID> -c model_provider=thinkwatch` continues an earlier session through the gateway. After **Restore…**, sessions started while connected can still be opened, and go straight to OpenAI.
- **Cost.** A rewritten request is priced as the model actually sent.
- **Cost.** A rewritten request is priced as the model actually sent. Codex marks no prompt-cache breakpoints, so on a route to an upstream in Anthropic Messages format, or to a Claude model on Bedrock that AWS lists for prompt caching (Claude 3.5 Sonnet v2, Claude 3.7 Sonnet, and Claude 4.5 and later), the gateway marks the tools, the system prompt and the last two user turns, and each turn reads what the previous one cached. Older Claude models on Bedrock, such as Sonnet 4, get no marks, and an upstream that refuses them receives the request again without them. Cache writes and reads are priced at the price sheet's cache rates; at Anthropic, writes cost 25% more than input and reads a tenth of it.

Related: [Features](/docs/lite/features/), [Use Claude Code with GLM, DeepSeek or Kimi](/docs/lite/claude-code-other-models/), [Install and update](/docs/lite/install/).
10 changes: 5 additions & 5 deletions src/content/docs-lite/en/failover-and-load-balancing.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ ThinkWatch Lite turns each relay key into an upstream and puts several upstreams

## Before you start

- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.10.
- ThinkWatch Lite, [installed](/lite/#install), with a client connected on the Clients page. This guide follows version 2026.10.11.
- The base URL and API keys of each relay.

## Steps
Expand All @@ -26,13 +26,13 @@ ThinkWatch Lite turns each relay key into an upstream and puts several upstreams

A strategy only sets the order; every member remains available for failover.

- **When the next upstream is tried.** Before anything has reached the client, the gateway moves on when the upstream cannot be reached or its credential cannot be read; answers 5xx, 429, 401, 403, 402 or 404; answers 400 or 422 with an error about an insufficient balance, a used-up quota or an unavailable model; or, in a streamed answer, reports an error before the first content, such as an overload. The gateway waits for that first content for up to **Wait for the answer to start** in Settings › Failover, 15 seconds by default. Other 4xx responses, and the last member's 4xx other than 429, go back to the client unchanged.
- **When the next upstream is tried.** Before anything has reached the client, the gateway moves on when the upstream cannot be reached or its credential cannot be read; answers 5xx, 429, 401, 403, 402 or 404; answers 400 or 422 with an error about an insufficient balance, a used-up quota or an unavailable model; in a streamed answer, reports an error before the first content, such as an overload; or sends no content within the no-response timeout. Other 4xx responses, and the last member's 4xx other than 429, go back to the client unchanged.
- **Weights and distribution.** In a **Round robin** group, a member's weight, from 1 to 100, is its long-run share of requests: weights 7 and 3 send seven requests in ten to the first, interleaved with the other three. **By ratio** uses the weights alone; **By speed** multiplies each by how much faster than the members' median the upstream starts answering (time to first token, squared, kept between 0.1 and 10), **By reliability** by its success rate over its last 50 outcomes within 30 minutes (squared, at least 0.05), and **By speed and reliability** by both. An upstream with too few samples counts as average, requests that stay with an upstream for their conversation count toward its share, and paused or full members sit out.
- **Slow starts.** With **Move to the next upstream when the start times out** on in Settings › Failover, a streamed answer with no content after **Wait for the answer to start** is cancelled and the request goes to the next upstream, if one can take it at that moment; the last upstream always waits, and the slow one is not paused. The upstream may have billed the input of the abandoned attempt, which the request's **Routing** tab marks; with the switch on, the wait must be at least 5 seconds, and 30 or more suits models that think before they answer.
- **No response.** An upstream that sends no content for the **No-response timeout** in Settings › Failover, 300 seconds by default (30 to 3,600), is given up on. Only content counts: text, reasoning or a tool call, not keep-alive pings. The clock starts when the request is sent to that upstream, so time spent waiting for a free slot does not count, and starts again with every piece of content. Before anything has reached the client, the request moves to the next upstream and the silent one counts as a failure toward its pause; when none is left, the client receives a timeout error in its own format. Once the answer has started, it ends with an error event the client understands. A whole, non-streamed answer has the same time to arrive. A streamed answer is held back for at most 15 seconds; after that the client receives the response headers and keep-alive comments (Gemini clients get the headers only), so it does not give up while the gateway can still move to the next upstream. The silent upstream's time to first token is recorded as the full wait, which ranks it lower for **Lowest latency** and **By speed**. The request's **Routing** tab marks the attempt **No response · 300 s** with the input the upstream may have billed for it. A model that thinks for a long time without sending anything may need a longer timeout.
- **Concurrency limits.** An upstream with a **Concurrency limit** takes at most that many requests at once. When it is full, a conversation that stays on it waits for a free slot and then moves on, and other requests go straight to the next member; when every member is full, a request waits for the first free slot for up to **Wait for a free slot at most** in Settings › Failover, 30 seconds by default, and then receives a 429 with `Retry-After`, or the error of an earlier attempt when one was sent. The same wait covers a key's per-minute and per-hour usage limits, and the Traffic page shows each skip and the time queued.
- **Pauses.** A failing upstream is paused for as long as Settings › Failover sets: by default 60 seconds after 3 **Consecutive failures**, doubling up to 600; 30 minutes for **Insufficient balance**; until the reset time, or 60 minutes, for **Quota used up**; the wait a rate-limited upstream asks for, up to 60 minutes. A rule with a single upstream is never held back, and when every member is paused they are tried anyway.
- **Pauses.** A failing upstream is paused for as long as Settings › Failover sets: by default 60 seconds after 3 **Consecutive failures**, doubling up to 600; 30 minutes for **Insufficient balance**; until the reset time, or 60 minutes, for **Quota used up**; the wait a rate-limited upstream asks for, up to 60 minutes. A rule with a single upstream is never held back, and when every member is paused they are tried anyway. A request aborted on the Traffic page does not count against its upstream.
- **Sessions and prompt cache.** Within a turn, while the client sends tool results back, requests keep the rule chosen at the start of the turn and the upstream that answered. Across turns, a conversation stays with the upstream that answered last if that answer read or wrote at least 1,024 cached tokens within the last five minutes; moving would rebuild the cache at full price. Otherwise, or when that upstream is paused, the strategy orders the members again, which is when **Round robin** moves on. The **Conversation** line on a request's **Routing** tab shows when a request stayed.
- **Same model name.** Every member is asked for the model the client sent, or the name a rule rewrote it to. A member whose model list lacks it is skipped; one without a list is tried, and its 404 moves the request on. A model that a member serves but leaves out of its list can be added with **Add models…** in the member's model list on the Upstreams page.
- **Same model name.** Every member is asked for the model the client sent, or the name a rule rewrote it to. A member whose model list lacks it is skipped; one without a list is tried, and its 404 moves the request on. A model that a member serves but leaves out of its list can be added with **Add models…** in the member's model list on the Upstreams page, or in the **Models** section of its edit dialog.
- **Falling back to another model.** Add a rule with the condition **Selected upstream** set to the backup upstream, **On match** set to **Continue matching**, and its model in **Change model to** under **Parameter rewrites**. Such a rule is evaluated for each upstream as it is tried, failover included. The changed model no longer hits the cached prompt, and the request is priced by the name sent.

Related: [Switch relays, upstreams or models without restarting Claude Code or Codex](/docs/lite/switch-upstreams-without-restart/), [Features](/docs/lite/features/#routing-and-failover), [Install and update](/docs/lite/install/).
Loading
Loading