Repository navigation
Idle timeout with automatic failover, and manual abort of requests and sessions (protocol 43) - #303
Merged
Merged
Conversation
…t of requests and sessions
An upstream could answer 200 and then send nothing, or stop halfway, and
nothing ever gave up on it: there is deliberately no overall timeout, and
`failover.next_on_slow_start` (off by default) only covered the start of a
stream and always let the last candidate wait forever.
`failover.idle_timeout_secs` (default 300, 30 to 3600) replaces it. The timer
starts when the request is sent upstream, so slot waits do not count, and
restarts on every piece of real content; keep-alives do not count, so an
upstream that only pings still runs out of time. What counts is defined per
dialect in one place (`tw_gateway::pulse`) and is shared with the opening
hold, so "first content" means the same thing in both.
- Before any content has reached the client the upstream counts as a
failure (cooldown rules apply), the attempt is recorded as `idle_timeout`
with the input it may have billed, the conversation no longer stays on it
for the turn, and the request moves on. Streams and whole answers are
held until their first content on every candidate but the last, so the
next upstream starts afresh. With no candidate left the client gets a 504
in its own format.
- After content has reached the client (or on the last candidate, whose
stream is passed on as it arrives) the answer ends with an error event in
the client's format and the request is recorded as failed.
`stream_start_wait_secs` and `next_on_slow_start` are removed: holding the
opening only up to a shorter window would make the before-content failover
impossible, so the hold now lasts until the idle timeout. Old configs that
still name them fail as unknown fields and the safe-mode repair offers to
delete them (tested).
Manual abort: `POST /request/{id}/abort` and `POST /sessions/{id}/abort`
throw a per-request switch registered when the request starts and dropped
with its ending. The upstream call is dropped at once, the client gets an
error in its format (499 before the answer started), the request ends as
`RequestFailed` with the new source `aborted` and code `gw.request.aborted`,
an in-flight hop is recorded as `aborted`, and the upstream is not set
aside. Upstream health counts aborts with cancellations. Requests no longer
running are a 404 with their own codes. WebSocket turns are not abortable.
Protocol 43. The store schema goes to 26 because `slow_start` disappears
from stored attempt chains.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iling over in the body Holding a streamed answer's headers until its first content (up to the idle timeout, 300 s by default) let clients' own header timeouts fire before the gateway's failover ever got a chance. A streamed answer is now held for at most `OPENING_HOLD` (15 s, a constant, not configurable): quick upstream errors still fail over and can be answered with a proper status. Past that, the client gets `200` and the streaming headers, then an SSE comment (`: keep-alive`) every `KEEPALIVE_EVERY` (15 s); Gemini clients get none, because Google's Python SDK parses comment lines as JSON. The rest of the pipeline (trying candidates and relaying the answer) owns everything it needs, so it simply moves into the response body and carries on: the upstream's own events stay held until its first content, an idle timeout, an in-stream error or an early close still moves the request to the next candidate, whose stream starts cleanly under the same `200`. When every candidate fails after the headers went out, the stream ends with the client-format error event used mid-stream instead of a 504; a last upstream's error answer is told the same way. The comments are ours and never touch the idle timer. Whole answers and Gemini's JSON array streams are held as before. An ending dropped after its abort switch was thrown now reports the abort rather than a client cancel, so a request aborted while the pipeline is not watching the switch is still recorded as aborted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
An upstream can answer 200 and then send nothing, or stop partway through, and nothing ever gave up on it. There is deliberately no overall timeout (a six-minute answer must not be cut), and
failover.next_on_slow_startwas off by default, only covered the start of a stream, and always let the last candidate wait forever. Users also had no way to stop a stuck request from the app.What
Idle timeout:
failover.idle_timeout_secsDefault 300, from 30 to 3600.
[DONE], and in-stream errors. Keep-alives do not restart it: SSE comments,ping/keepalive/heartbeatevents, Anthropicmessage_start, Responsesresponse.created/in_progress/queuedandcodex.*, Chat role-only or empty chunks, Gemini chunks with no parts and no finish reason, and BedrockmessageStart. Anything the gateway does not recognise counts as content. This is defined once, intw_gateway::pulse, and the opening hold uses the same definition.idle_timeout, withgw.upstream.idle_timeout {upstream, secs}and any input it may already have billed. It gets a latency sample covering the full wait, and the conversation stops staying on that upstream for the rest of the turn. The request then moves to the next candidate. Every candidate except the last has its stream or whole answer held until the first content arrives. That way the next upstream starts the answer from the beginning. When no candidate is left before the headers have gone out, the client gets a 504 in its own format (x-thinkwatch-error: upstream; Anthropictimeout_error, GeminiDEADLINE_EXCEEDED).OPENING_HOLD(15 s, a constant, not configurable), so quick upstream errors still fail over and get a proper status code. After that, the client gets200and the streaming headers, then: keep-aliveSSE comments every 15 s. Gemini clients get no comments: Google's Python SDK parses comment lines as JSON.200.gw.upstream.idle_timeoutwhen nothing had been said yet andgw.upstream.idle_timeout_mid_streamwhen something had.stream_start_wait_secsandnext_on_slow_startare removed. If the hold ended earlier than the idle timeout, failing over before any content had arrived would no longer be possible. An old config that still contains either key fails to load with an unknown-field error, and the safe-mode repair offers to delete it (tested).config.slow_start_too_shortandgw.slow_startare gone.Manual abort
POST /request/{id}/abortandPOST /sessions/{id}/abortreturnAborted { requests }, the ids that were stopped. When nothing is running they return 404 withcontrol.request_not_runningorcontrol.session_not_running.x-thinkwatch-error: aborted);RequestFailedwith the newFailureSource::Abortedand the codegw.request.aborted(tw_api::ABORTED);aborted.Contract
FailoverViewreplacesstream_start_wait_secsandnext_on_slow_startwithidle_timeout_secs.AttemptOutcomedropsslow_startand addsidle_timeoutandaborted.FailureSourceaddsaborted.slow_startdisappears from stored attempt chains. Under the project's no-migration rule, existing request history is rebuilt.CANCELLED.docs/config.mdanddocs/config.zh-CN.mdare updated, both the prose and the generated table.msg-codes.txtis regenerated, and the smoke script checks that both abort endpoints are registered and that the real binary sends its headers at 15 s on a two-candidate route whose first upstream is quiet for 17 s.Tests
crates/tw-gateway/tests/idle_timeout.rsreplacesslow_start.rs. A one-second window is injected throughAppState::idle_tick. The tests cover:There are also unit tests for
pulse,abort, the opening hold, affinity, error status codes, config ranges and repair, the health tally, and the control endpoints.cargo fmt --check,cargo clippy --workspace --all-targets -D warnings,cargo test --workspace(3108 passed),cargo test -p tw-api --features ts,scripts/smoke.sh(79/79) andscripts/release_notes_test.pyall pass locally.🤖 Generated with Claude Code