Repository navigation
Mark prompt-cache breakpoints for Claude when the client marked none - #306
Merged
Merged
Conversation
Anthropic caches only what a request marks with `cache_control`, and Bedrock only what it marks with `cachePoint`. Clients in OpenAI or Gemini formats (Codex, Chat clients) cannot mark anything, because those providers cache a repeated prefix on their own. So a converted request to Claude was never cached, and every turn paid full price for the whole conversation. When the IR has no breakpoints, the Anthropic encoder (any Anthropic-format upstream) and the Bedrock encoder (Claude model ids only) now mark, with the default 5-minute TTL: - the end of the tools - the end of the system prompt - the end of the last user turn and of the one before it, at most four in total The earlier user mark is the previous request's last breakpoint, so each turn reads back what the previous one wrote even when a step adds more blocks than Anthropic's ~20-block lookback. Thinking blocks are never marked. Requests that carry their own breakpoints (Claude Code) keep exactly those, and same-format passthrough is untouched. Cost accounting needed no change: usage parsing already reads `cache_creation_input_tokens` / `cache_read_input_tokens` and Bedrock's `cacheWriteInputTokens` / `cacheReadInputTokens`, the sniffer reads the upstream's own bytes for converted requests, and the recorder charges cache writes and reads at the price table's cache rates. New tests cover both. Documented in docs/config.md and docs/config.zh-CN.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 9, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #305. Converted Codex requests to Claude carried no cache breakpoints, so Anthropic never cached them.
Why
Anthropic caches only what a request marks with
cache_control, and Bedrock only what it marks withcachePoint. Clients in OpenAI or Gemini formats can't mark anything: Codex, Chat clients and Gemini clients rely on those providers caching a repeated prefix automatically. The IR only carried breakpoints that an Anthropic or Bedrock client had set, so a request converted to Claude had none. Every turn of an agent loop paid full input price for the whole conversation.What changes
When the IR has no breakpoints, i.e. the client marked nothing:
cache_control, because Claude Code sends it to them)cache_control: {"type": "ephemeral"}on the last tool, the last system block, and the last block of the last two user messagesanthropic.claude-…,us.anthropic.claude-…){"cachePoint": {"type": "default"}}after the same four placescachePointreject the whole request, and an ARN doesn't say which model is behind itcachePoint) keep exactly those. Same-format passthrough stays byte-identical, since this lives only in the encoders.docs/config.mdanddocs/config.zh-CN.md, in theproviderssection next to the description of what a forwarded request carries.Cost accounting
No code change was needed; this verifies it end to end:
cache_creation_input_tokens/cache_read_input_tokens(includingcache_creation.ephemeral_1h_input_tokens) and Bedrock'scacheWriteInputTokens/cacheReadInputTokens.cost_of: input, output,cache_read, andcache_writeat the 5-minute or 1-hour write rate of the price sheet resolved for that upstream and model.cache_creation_input_tokens: 2000andcache_read_input_tokens: 9000producesRequestFinished.usage = {input: 30, cache_write: 2000, cache_read: 9000, output: 12}.Tests
tw-dialect/tests/auto_cache.rs(new), built on Codex-shaped Responses Lite requests:cachePoints; Nova, Llama and an ARN get none.tw-gateway/tests/conversion.rs: the cache-usage e2e test above, plus an Anthropic request without marks passed straight through byte for byte.tw-storerecorder: cache pricing.Checks:
cargo fmt --all --check,cargo clippy --workspace --all-targets -- -D warnings,cargo test --workspace --no-fail-fast(3147 passed, 0 failed, 8 ignored),scripts/smoke.sh(79 passed, 0 failed).🤖 Generated with Claude Code