fix: stop rewriting Responses history by default - #6
Merged
Merged
Conversation
Every /v1/responses request was trimmed pre-flight against a byte ceiling derived from max_prompt_tokens, rewriting old input items into placeholders before Copilot ever saw the payload. Two problems with that: The ratio was wrong. CHARS_PER_TOKEN_ESTIMATE = 3.5 came from context-manager.ts, where it measures *unescaped content length*; here it measures serialized JSON. Measured against multi-turn payloads built from this repo's own source, the real figure is 4.4 bytes per content token (3.7 per token of the JSON itself), so trimming engaged at roughly 77% of a model's real window — destroying context Copilot would have accepted. The rewrite broke prompt caching. dropOldInputItems replaces from the oldest item forward, which is precisely the stable prefix a cache is keyed on. In a simulated growing session the reusable prefix collapsed from 277 items to 1 on the turn trimming first engaged, and was re-poisoned each time the boundary advanced. Default behavior is now to forward the client's history verbatim and let Copilot enforce its own limits, so a rejection is upstream's real answer instead of a local guess. The ~5 MB Azure Front Door transport cliff is still enforced pre-flight — that one is a verified wire limit, not an estimate. Token-derived trimming remains available behind --responses-context-trim, with the ratio corrected to 4.3. Also repairs the reactive path it now depends on: - Context-overflow errors returned upstream's reported token counts instead of a local bytes/4 estimate. A genuine "300000 > 272000" was being rewritten as "16 + 1000 > 272000", and Claude Code sizes its compaction pass off that gap. The original upstream body is preserved in error.upstream_error. - isContextOverflow treated any 5xx on a >2 MB payload as context overflow, so a transient 503 made clients compact away history to work around what was really a Copilot outage. Now restricted to explicit overflow markers and plain-500 dead-zone hangs, with transient-outage phrasings excluded. Ratio, cache-prefix collapse, and tokenizer cost were measured independently by two models reaching the same conclusions. Fixes #5 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
voidsteed
force-pushed
the
fix/responses-context-trimming
branch
from
September 17, 2026 05:18
b828940 to
214704f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #5.
Every
/v1/responsesrequest was trimmed pre-flight against a byte ceiling derived frommax_prompt_tokens, rewriting old input items into placeholders before Copilot ever saw the payload. The reporter is right that this is the wrong default, and measuring it turned up two separate problems.The ratio was miscalibrated
CHARS_PER_TOKEN_ESTIMATE = 3.5was inherited fromcontext-manager.ts, where it measures unescaped content length. Here it measures serialized JSON — copying the constant silently changed its basis.Measured against multi-turn payloads built from this repo's own source (user messages,
function_callitems, largefunction_call_outputbodies, assistant replies):So trimming engaged at roughly 77% of a model's real window, destroying context Copilot would have accepted:
max_prompt_tokensThe rewrite broke prompt caching
dropOldInputItemsreplaces from the oldest item forward — precisely the stable prefix a cache is keyed on. Simulating a growing session against the real algorithm, measuring how many leading items stay byte-identical turn over turn:The reusable prefix collapses from 277 items to 1 on the turn trimming first engages, and is re-poisoned each time the boundary advances.
What changed
Default is now passthrough. Client history forwards verbatim; Copilot enforces its own token limits, so a rejection is upstream's real answer instead of a local guess. Only the ~5 MB Azure Front Door cliff is enforced pre-flight — that one is a verified wire limit, not an estimate.
Token-derived trimming stays available behind
--responses-context-trim, with the ratio corrected to 4.3, and warns at startup when enabled.The reactive path is repaired, since passthrough now depends on it being trustworthy:
bytes/4estimate instead of upstream's reported counts. A genuine300000 > 272000was being rewritten as16 + 1000 > 272000— and Claude Code sizes its compaction pass off that gap. Upstream's numbers are now used when present, with the original body preserved inerror.upstream_error.isContextOverflowtreated any 5xx on a >2 MB payload as context overflow, so a transient 503 made clients compact away history to work around what was really a Copilot outage. Now restricted to explicit overflow markers and plain-500 dead-zone hangs, with transient-outage phrasings excluded.Notes
dropOldInputItemsstill drops oldest-first when it runs. I tried reversing it to preserve the cacheable prefix and reverted that — dropping from the tail discards the active request, which is worse than the cache miss it avoids. The real mitigation is that it no longer runs by default; the comment there now says so.Four existing trimming tests needed fixture sizes scaled to still overflow the raised ceiling. They cover the same behavior but no longer exercise the default path —
tests/create-responses-context.test.tsdoes that, with 5 new tests covering default passthrough (asserts all 8 turns survive), the transport ceiling still holding, upstream count surfacing, the 503 case, and the opt-in flag.Caveat: neither 4.3 nor 4.4 is Copilot's actual prompt-token count — confirming that needs a live measurement against the real endpoint, which I haven't done. It matters less now that the ratio doesn't gate the default path.
bun run typecheckandbun testare green (363 tests); lint is clean on all changed files.Ratio, cache-prefix collapse, and tokenizer cost (~39ms at 1MB, vs. the "300-1000ms" claimed in the
context-manager.tscomment) were measured independently by two models reaching the same conclusions.🤖 Generated with Claude Code