Skip to content

Idle timeout with automatic failover, and manual abort of requests and sessions (protocol 43) - #303

Merged
fylorn merged 2 commits into
mainfrom
feat/idle-timeout-abort
Oct 9, 2026
Merged

fylorn merged 2 commits into
mainfrom
feat/idle-timeout-abort

Conversation

@fylorn

@fylorn fylorn commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Why

An upstream can answer 200 and then send nothing, or stop partway through, and nothing ever gave up on it. There is deliberately no overall timeout (a six-minute answer must not be cut), and failover.next_on_slow_start was off by default, only covered the start of a stream, and always let the last candidate wait forever. Users also had no way to stop a stuck request from the app.

What

Idle timeout: failover.idle_timeout_secs

Default 300, from 30 to 3600.

  • The timer starts when the request is sent to that upstream, so time spent waiting for a slot does not count. Every piece of real content restarts it: text, reasoning (deltas or summaries), tool-call starts and arguments, refusals, finish reasons, closing usage, [DONE], and in-stream errors. Keep-alives do not restart it: SSE comments, ping/keepalive/heartbeat events, Anthropic message_start, Responses response.created / in_progress / queued and codex.*, Chat role-only or empty chunks, Gemini chunks with no parts and no finish reason, and Bedrock messageStart. Anything the gateway does not recognise counts as content. This is defined once, in tw_gateway::pulse, and the opening hold uses the same definition.
  • A whole (non-streamed) answer is timed from sending until its first non-whitespace body byte, so in practice until the complete answer arrives.
  • Before any content has reached the client: the upstream counts as a failure for cooldown. The attempt is recorded as idle_timeout, with gw.upstream.idle_timeout {upstream, secs} and any input it may already have billed. It gets a latency sample covering the full wait, and the conversation stops staying on that upstream for the rest of the turn. The request then moves to the next candidate. Every candidate except the last has its stream or whole answer held until the first content arrives. That way the next upstream starts the answer from the beginning. When no candidate is left before the headers have gone out, the client gets a 504 in its own format (x-thinkwatch-error: upstream; Anthropic timeout_error, Gemini DEADLINE_EXCEEDED).
  • Streams held at most 15 s, then 200 + keep-alives: waiting up to the idle timeout for headers would let clients' own header timeouts fire first. A streamed answer is therefore held for at most OPENING_HOLD (15 s, a constant, not configurable), so quick upstream errors still fail over and get a proper status code. After that, the client gets 200 and the streaming headers, then : keep-alive SSE comments every 15 s. Gemini clients get no comments: Google's Python SDK parses comment lines as JSON.
    • The rest of the pipeline owns everything it needs and moves into the response body. The upstream's events stay held until its first content.
    • An idle timeout, an in-stream error or an early close still fails over, and the next upstream's stream starts cleanly under the same 200.
    • If every candidate fails, or the last one answers with an error status, the stream ends with the client-format error event instead of a 504.
    • The keep-alive comments never touch the idle timer.
    • Whole answers and Gemini's JSON-array streams are held as before.
  • After content has reached the client, or on the last candidate: the last candidate's stream is passed on as it arrives. Once a response is on its way to the client, a timeout ends it with an error event in the client's format, reusing the existing stream-break path. The request is recorded as failed. The code is gw.upstream.idle_timeout when nothing had been said yet and gw.upstream.idle_timeout_mid_stream when something had.
  • For the last candidate, a success is recorded only when its first content arrives, so a silent last upstream also counts towards cooldown.
  • stream_start_wait_secs and next_on_slow_start are removed. If the hold ended earlier than the idle timeout, failing over before any content had arrived would no longer be possible. An old config that still contains either key fails to load with an unknown-field error, and the safe-mode repair offers to delete it (tested). config.slow_start_too_short and gw.slow_start are gone.

Manual abort

  • POST /request/{id}/abort and POST /sessions/{id}/abort return Aborted { requests }, the ids that were stopped. When nothing is running they return 404 with control.request_not_running or control.session_not_running.
  • Each request carries a switch, registered when it starts and dropped along with its ending. Throwing the switch:
    • drops the upstream call at once, while waiting for headers, during the opening hold, while waiting for a slot, or mid-stream;
    • ends the client's response with an error in its own format (499 before the answer starts, x-thinkwatch-error: aborted);
    • records the request as RequestFailed with the new FailureSource::Aborted and the code gw.request.aborted (tw_api::ABORTED);
    • marks a hop that was still in progress as the new attempt outcome aborted.
  • The upstream is not set aside. Upstream health counts aborts together with cancellations rather than as failures.
  • WebSocket turns are not registered, so trying to abort one returns a 404.

Contract

  • Protocol 43, with a note in the existing style. FailoverView replaces stream_start_wait_secs and next_on_slow_start with idle_timeout_secs. AttemptOutcome drops slow_start and adds idle_timeout and aborted. FailureSource adds aborted.
  • The store schema goes to 26, because slow_start disappears from stored attempt chains. Under the project's no-migration rule, existing request history is rebuilt.
  • In tw-dialect, a Gemini error body with status 499 is now CANCELLED.
  • docs/config.md and docs/config.zh-CN.md are updated, both the prose and the generated table. msg-codes.txt is regenerated, and the smoke script checks that both abort endpoints are registered and that the real binary sends its headers at 15 s on a two-candidate route whose first upstream is quiet for 17 s.

Tests

crates/tw-gateway/tests/idle_timeout.rs replaces slow_start.rs. A one-second window is injected through AppState::idle_tick. The tests cover:

  • an upstream that answers 200 and then says nothing, so the next candidate answers (after three timeouts the silent upstream is paused);
  • headers that never come;
  • no candidate left, which ends in a 504 rather than an infinite wait (or in an in-stream error after keep-alives once the 15 s hold has passed);
  • past the hold: headers and keep-alives go out, and failover carries on through idle timeouts and in-stream errors, with no content from the abandoned upstream; a last upstream's 400 is told in the stream; an abort after the headers ends the stream with the abort; Gemini clients get the early headers without comments; whole answers never answer early;
  • a silent last candidate, which ends with an error event;
  • pings alone in Anthropic, Responses and Chat, which still time out;
  • a whole answer that never comes or carries only whitespace;
  • a slow but steady stream that runs longer than the window and is not cut;
  • reasoning, which counts as content;
  • content followed by silence in all four client formats, each ending with that format's error event;
  • slot waits not counting towards the timeout;
  • abort before headers, mid-stream and per session.

There are also unit tests for pulse, abort, the opening hold, affinity, error status codes, config ranges and repair, the health tally, and the control endpoints.

cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, cargo test --workspace (3108 passed), cargo test -p tw-api --features ts, scripts/smoke.sh (79/79) and scripts/release_notes_test.py all pass locally.

🤖 Generated with Claude Code

fylorn and others added 2 commits October 9, 2026 15:11
…t of requests and sessions

An upstream could answer 200 and then send nothing, or stop halfway, and
nothing ever gave up on it: there is deliberately no overall timeout, and
`failover.next_on_slow_start` (off by default) only covered the start of a
stream and always let the last candidate wait forever.

`failover.idle_timeout_secs` (default 300, 30 to 3600) replaces it. The timer
starts when the request is sent upstream, so slot waits do not count, and
restarts on every piece of real content; keep-alives do not count, so an
upstream that only pings still runs out of time. What counts is defined per
dialect in one place (`tw_gateway::pulse`) and is shared with the opening
hold, so "first content" means the same thing in both.

- Before any content has reached the client the upstream counts as a
  failure (cooldown rules apply), the attempt is recorded as `idle_timeout`
  with the input it may have billed, the conversation no longer stays on it
  for the turn, and the request moves on. Streams and whole answers are
  held until their first content on every candidate but the last, so the
  next upstream starts afresh. With no candidate left the client gets a 504
  in its own format.
- After content has reached the client (or on the last candidate, whose
  stream is passed on as it arrives) the answer ends with an error event in
  the client's format and the request is recorded as failed.

`stream_start_wait_secs` and `next_on_slow_start` are removed: holding the
opening only up to a shorter window would make the before-content failover
impossible, so the hold now lasts until the idle timeout. Old configs that
still name them fail as unknown fields and the safe-mode repair offers to
delete them (tested).

Manual abort: `POST /request/{id}/abort` and `POST /sessions/{id}/abort`
throw a per-request switch registered when the request starts and dropped
with its ending. The upstream call is dropped at once, the client gets an
error in its format (499 before the answer started), the request ends as
`RequestFailed` with the new source `aborted` and code `gw.request.aborted`,
an in-flight hop is recorded as `aborted`, and the upstream is not set
aside. Upstream health counts aborts with cancellations. Requests no longer
running are a 404 with their own codes. WebSocket turns are not abortable.

Protocol 43. The store schema goes to 26 because `slow_start` disappears
from stored attempt chains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iling over in the body

Holding a streamed answer's headers until its first content (up to the
idle timeout, 300 s by default) let clients' own header timeouts fire
before the gateway's failover ever got a chance.

A streamed answer is now held for at most `OPENING_HOLD` (15 s, a
constant, not configurable): quick upstream errors still fail over and
can be answered with a proper status. Past that, the client gets `200`
and the streaming headers, then an SSE comment (`: keep-alive`) every
`KEEPALIVE_EVERY` (15 s); Gemini clients get none, because Google's
Python SDK parses comment lines as JSON. The rest of the pipeline
(trying candidates and relaying the answer) owns everything it needs,
so it simply moves into the response body and carries on: the
upstream's own events stay held until its first content, an idle
timeout, an in-stream error or an early close still moves the request
to the next candidate, whose stream starts cleanly under the same `200`.
When every candidate fails after the headers went out, the stream ends
with the client-format error event used mid-stream instead of a 504; a
last upstream's error answer is told the same way. The comments are
ours and never touch the idle timer. Whole answers and Gemini's JSON
array streams are held as before.

An ending dropped after its abort switch was thrown now reports the
abort rather than a client cancel, so a request aborted while the
pipeline is not watching the switch is still recorded as aborted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit fc8838e into main Oct 9, 2026
5 checks passed
@fylorn
fylorn deleted the feat/idle-timeout-abort branch October 9, 2026 08:27
@fylorn fylorn mentioned this pull request Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant