diff --git a/CHANGELOG.md b/CHANGELOG.md index 6463c8c9..10122650 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,6 +11,106 @@ target. ## [Unreleased] +## [3.3.0] — 2026-10-09 + +Requests converted for Claude now use prompt caching, and Codex works +through routes that convert its requests to another format: its tools, +its history and its compactions arrive. A Responses request that only +OpenAI can read — one that continues a conversation kept on OpenAI's +servers, or carries a compaction OpenAI wrote — now goes to a route that +speaks Responses instead of being refused. The thinkwatch-core crates move +from v0.65.0 to v0.67.1. + +### Read before upgrading + +- **Converted requests to Claude are cached, and billed as cache traffic.** + A request converted for an Anthropic-format upstream, or for a Claude + model on Bedrock that AWS lists for prompt caching (Claude 3.5 Sonnet v2, + Claude 3.7 Sonnet, and Claude 4.5 and later), from a client that marks no + cache breakpoints — Codex, Chat and Gemini clients, and an Anthropic + client whose request goes to Bedrock without `cache_control` — now gets + them at the end of the tools, the system prompt and the last two user + turns, with the 5-minute lifetime. + - The upstream bills the prefix a request writes at its cache-write rate + (1.25× input at Anthropic) and what the next turns read back at its + cache-read rate (0.1×). ThinkWatch prices and counts them the same + way: cost, budgets and `tokens` rate limits weigh a cache write by the + model's `cache_write_weight` and a read by its `cache_read_weight`, + 1.25× and 0.1× its input weight when unset. Set them on the model to + match what your upstream charges. + - A conversation of several turns costs less. A one-off request with a + long prompt costs up to a quarter more on its input. + - Prompt token counts (`input_tokens` on the log rows, quotas) include + cached input in full, as before. + - Requests that mark their own breakpoints, such as Claude Code's, and + requests forwarded in their own format are unchanged. + - An upstream that refuses the added breakpoints — a `400` that names + `cache_control`, `cachePoint` or prompt caching, as an + Anthropic-compatible endpoint that does not know them answers — gets + the request once more without them, and later requests to it for that + model leave them out from the start. Each instance remembers this + until a provider or model is changed or it restarts. Breakpoints a + client marked itself are never taken out, and the probes that check a + route's API format go without breakpoints. +- **Rolling back with Codex sessions.** A Codex session that compacted + through a converted route on 3.3.0 carries a compaction that 3.2.1 + refuses, so after a rollback, or on a 3.2.1 instance during the rollout, + it cannot continue and needs a new session. + +Nothing else needs action: no setting, schema, API route or Redis key +changes. + +### Fixed + +- **Responses requests that only OpenAI can read.** A Responses request + that continues a conversation kept on OpenAI's servers + (`previous_response_id`, `conversation`, `prompt`, `background`), points + at a stored item (`item_reference`) or carries a compaction OpenAI wrote + was refused with `400`, even on a route that speaks Responses and would + have forwarded it as sent: every request was decoded for its usage + estimate, and one that could not be decoded was refused. Codex sends such + a compaction in every request after compacting a session through an + OpenAI upstream, so the session could not go on; and a WebSocket turn + naming a response other than the connection's last one, meant to go + upstream as sent, was refused too. These requests now go to the model's + routes that speak Responses, and only a model with none refuses them, + with the reason. Their input estimate counts the request's text. + +### Changed + +- thinkwatch-core crates (tw-bedrock, tw-breaker, tw-dialect, tw-guard) + v0.65.0 → v0.67.1. Only tw-dialect, the format conversion, changes: + - **Prompt caching for Claude** on converted requests (see Read before + upgrading). + - **Codex tools.** For some models Codex declares every tool in an + `additional_tools` input item and sends no top-level `tools`. Those + items were dropped, so a route converting to Anthropic, Chat, Gemini + or Bedrock sent no tools and Codex could not read files or run + commands. Every tool now arrives, and tool calls come back under the + names and kinds Codex declared. Local shell calls, tool search, + reasoning-effort changes and agent messages in the history are + converted instead of dropped. + - **Codex compaction on converted routes.** Codex compacts a long + session through a provider named `OpenAI` by asking for exactly one + compaction item, which a converted route could not give, so compacting + failed. The upstream now writes a handoff summary, Codex receives it as + its compaction, and later requests carry it back. The summary request + is billed like any other. The summary travels in the item's + `encrypted_content` (`tw1.c.…`) base64-encoded, not encrypted; with + `audit.body_redact_pii` on, captured bodies have it redacted like the + rest of the body. OpenAI's own compactions still cannot be read by an + upstream of another format, and are refused saying so. + - **System messages in mid-conversation stay in place.** A `system` or + `developer` message after the conversation has started was added to + the system prompt on a converted route. It now stays where it was + given: a developer message for a Responses upstream, and a user turn + wrapped in `` for Anthropic, Chat, Gemini and Bedrock + upstreams. The system prompt stays the same from one turn to the next, + which is what keeps the prompt cache working. + - **`verbosity`** (Chat) and **`text.verbosity`** (Responses) carry + across the conversion to GPT-5 and later models. Other models still + leave it out. + ## [3.2.1] — 2026-10-09 The Helm chart and the Compose file now pull the published images: they @@ -1469,7 +1569,8 @@ unreleased builds should: stop the gateway, run `db/schema.sql` against PostgreSQL, restart against this tag. The schema is idempotent end-to-end, so the apply is safe to repeat. -[Unreleased]: https://github.com/ThinkWatchProject/ThinkWatch/compare/v3.2.1...HEAD +[Unreleased]: https://github.com/ThinkWatchProject/ThinkWatch/compare/v3.3.0...HEAD +[3.3.0]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v3.3.0 [3.2.1]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v3.2.1 [3.2.0]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v3.2.0 [3.1.0]: https://github.com/ThinkWatchProject/ThinkWatch/releases/tag/v3.1.0 diff --git a/Cargo.lock b/Cargo.lock index bf2babfd..59bef983 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -4189,7 +4189,7 @@ dependencies = [ [[package]] name = "think-watch-auth" -version = "3.2.1" +version = "3.3.0" dependencies = [ "anyhow", "argon2", @@ -4219,7 +4219,7 @@ dependencies = [ [[package]] name = "think-watch-common" -version = "3.2.1" +version = "3.3.0" dependencies = [ "aes-gcm", "anyhow", @@ -4262,7 +4262,7 @@ dependencies = [ [[package]] name = "think-watch-gateway" -version = "3.2.1" +version = "3.3.0" dependencies = [ "anyhow", "arc-swap", @@ -4301,7 +4301,7 @@ dependencies = [ [[package]] name = "think-watch-mcp-gateway" -version = "3.2.1" +version = "3.3.0" dependencies = [ "anyhow", "arc-swap", @@ -4330,7 +4330,7 @@ dependencies = [ [[package]] name = "think-watch-server" -version = "3.2.1" +version = "3.3.0" dependencies = [ "anyhow", "arc-swap", @@ -4381,7 +4381,7 @@ dependencies = [ [[package]] name = "think-watch-test-support" -version = "3.2.1" +version = "3.3.0" dependencies = [ "anyhow", "async-stream", @@ -4837,8 +4837,8 @@ dependencies = [ [[package]] name = "tw-bedrock" -version = "0.65.0" -source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.65.0#2f1abbe9af75fd8caf3971b9bd954284f14d182a" +version = "0.67.1" +source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.67.1#fb32ae9607771e77526d501106b2642498d64482" dependencies = [ "aws-credential-types", "aws-sigv4", @@ -4852,16 +4852,16 @@ dependencies = [ [[package]] name = "tw-breaker" -version = "0.65.0" -source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.65.0#2f1abbe9af75fd8caf3971b9bd954284f14d182a" +version = "0.67.1" +source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.67.1#fb32ae9607771e77526d501106b2642498d64482" dependencies = [ "serde", ] [[package]] name = "tw-dialect" -version = "0.65.0" -source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.65.0#2f1abbe9af75fd8caf3971b9bd954284f14d182a" +version = "0.67.1" +source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.67.1#fb32ae9607771e77526d501106b2642498d64482" dependencies = [ "memchr", "serde", @@ -4870,8 +4870,8 @@ dependencies = [ [[package]] name = "tw-guard" -version = "0.65.0" -source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.65.0#2f1abbe9af75fd8caf3971b9bd954284f14d182a" +version = "0.67.1" +source = "git+https://github.com/ThinkWatchProject/ThinkWatch-Core.git?tag=v0.67.1#fb32ae9607771e77526d501106b2642498d64482" dependencies = [ "base64 0.22.1", "bytes", diff --git a/Cargo.toml b/Cargo.toml index 6214a8e9..eb1e48c9 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -10,7 +10,7 @@ members = [ ] [workspace.package] -version = "3.2.1" +version = "3.3.0" edition = "2024" # The MSRV: the newest `rust-version` among the locked dependencies # (sqlx 0.9 and the aws-smithy crates under aws-sigv4). Without it, @@ -70,10 +70,10 @@ opt-level = 3 # never re-exported through a local shim. And the reverse: something only # this side uses (the at-rest crypto, IMDSv2 credentials, the gateway error) # lives here, not in core. -tw-bedrock = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.65.0" } -tw-breaker = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.65.0" } -tw-dialect = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.65.0" } -tw-guard = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.65.0" } +tw-bedrock = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.67.1" } +tw-breaker = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.67.1" } +tw-dialect = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.67.1" } +tw-guard = { git = "https://github.com/ThinkWatchProject/ThinkWatch-Core.git", tag = "v0.67.1" } # Web framework axum = { version = "0.8", features = ["macros", "ws"] } diff --git a/crates/common/Cargo.toml b/crates/common/Cargo.toml index 4c113971..b196e8bc 100644 --- a/crates/common/Cargo.toml +++ b/crates/common/Cargo.toml @@ -7,6 +7,8 @@ edition.workspace = true tw-breaker = { workspace = true } tw-guard = { workspace = true } tw-bedrock = { workspace = true } +# Reads the compaction summaries a conversion carries (see `pii`). +tw-dialect = { workspace = true } axum = { workspace = true } sqlx = { workspace = true } fred = { workspace = true } @@ -46,6 +48,3 @@ aws-sigv4 = { workspace = true } aws-credential-types = { workspace = true } http_1x = { package = "http", version = "1" } async-trait = "0.1" - -[dev-dependencies] -tw-dialect = { workspace = true } diff --git a/crates/common/src/pii.rs b/crates/common/src/pii.rs index bac99f89..bd37fd0b 100644 --- a/crates/common/src/pii.rs +++ b/crates/common/src/pii.rs @@ -14,6 +14,7 @@ //! thinkwatch-core's (`tw_guard::redact`), the engine the desktop gateway //! redacts with. +use std::borrow::Cow; use std::sync::Arc; use tw_guard::policy::RedactPolicy; @@ -27,22 +28,70 @@ use tw_guard::redact::rules::RuleSet; /// JSON escapes are read, a custom rule's match stays inside one JSON /// string so the result is still JSON, base64 payloads are left alone, /// and so are the gateway's own placeholders (a captured answer still -/// carries them). +/// carries them). The one base64 payload that is read is a compaction +/// summary a conversion carries (see [`redact_carried_summaries`]). pub fn redact_blob(rules: &RuleSet, input: &str) -> String { if rules.is_empty() { return input.to_string(); } - let hits = tw_guard::redact::flow::hits(input, rules); - let mut out = input.to_string(); + let input = redact_carried_summaries(rules, input); + let hits = tw_guard::redact::flow::hits(&input, rules); + let mut out = input.into_owned(); for h in hits.iter().rev() { - out.replace_range( - h.bytes.clone(), - &format!("{{{{REDACTED_{}}}}}", h.rule.id()), - ); + out.replace_range(h.bytes.clone(), &redacted(&h.rule)); } out } +fn redacted(rule: &tw_guard::redact::rules::Rule) -> String { + format!("{{{{REDACTED_{}}}}}", rule.id()) +} + +/// The summaries in `input` that a conversion carries, redacted. +/// +/// A Codex compaction sent to an upstream of another format comes back as +/// the summary that upstream wrote — the conversation restated, with its +/// paths, commands and values — base64-encoded in the item's +/// `encrypted_content` (`tw1.c.…`, see `tw_dialect::compaction`), and +/// later requests carry it back. Unlike OpenAI's own compactions it is not +/// encrypted, but as base64 the search above passes over it. Each one is +/// read, redacted as plain text and written back the same way, so the +/// body keeps its shape. +fn redact_carried_summaries<'a>(rules: &RuleSet, input: &'a str) -> Cow<'a, str> { + const PREFIX: &str = tw_dialect::compaction::CARRIED_PREFIX; + if !input.contains(PREFIX) { + return Cow::Borrowed(input); + } + let mut out = String::with_capacity(input.len()); + let mut rest = input; + while let Some(at) = rest.find(PREFIX) { + let start = at + PREFIX.len(); + let end = rest[start..] + .find(|c: char| !(c.is_ascii_alphanumeric() || c == '-' || c == '_')) + .map_or(rest.len(), |n| start + n); + let carried = &rest[at..end]; + out.push_str(&rest[..at]); + match tw_dialect::compaction::read(carried) { + Some(summary) => { + let hits = tw_guard::redact::flow::hits_plain(&summary, rules); + if hits.is_empty() { + out.push_str(carried); + } else { + let mut text = summary; + for h in hits.iter().rev() { + text.replace_range(h.bytes.clone(), &redacted(&h.rule)); + } + out.push_str(&tw_dialect::compaction::carry(&text)); + } + } + None => out.push_str(carried), + } + rest = &rest[end..]; + } + out.push_str(rest); + Cow::Owned(out) +} + /// The at-rest redactor, hot-swapped with the outbound redaction policy. #[derive(Clone)] pub struct BlobRedactor { @@ -133,6 +182,31 @@ mod tests { assert_eq!(v["b"], "keep", "{out}"); } + #[test] + fn a_carried_compaction_summary_is_redacted_where_it_is() { + use tw_dialect::compaction::{carry, read}; + let rules = only(&[("EMAIL", r"[\w.]+@[\w.]+")]); + let carried = carry("Reply to alice@example.com about src/lib.rs."); + let body = format!( + r#"{{"input":[{{"type":"compaction","encrypted_content":"{carried}"}},{{"role":"user","content":"bob@example.com"}}]}}"# + ); + let out = redact_blob(&rules, &body); + let v: serde_json::Value = serde_json::from_str(&out).expect("still JSON"); + let summary = read(v["input"][0]["encrypted_content"].as_str().unwrap()); + assert_eq!( + summary.as_deref(), + Some("Reply to {{REDACTED_EMAIL}} about src/lib.rs."), + "{out}" + ); + assert_eq!(v["input"][1]["content"], "{{REDACTED_EMAIL}}", "{out}"); + + // Nothing found in it, or not one of ours: left as it was. + let clean = format!(r#"{{"encrypted_content":"{}"}}"#, carry("tests pass")); + assert_eq!(redact_blob(&rules, &clean), clean); + let theirs = r#"{"encrypted_content":"gAAAAABoQ2xpZW50"}"#; + assert_eq!(redact_blob(&rules, theirs), theirs); + } + #[test] fn the_policys_rules_are_the_ones_used_built_in_and_custom() { let key = "sk-ant-api03-AAAAAAAAAAAAAAAAAAAAAAAAAAAA"; diff --git a/crates/gateway/src/proxy/cache_marks.rs b/crates/gateway/src/proxy/cache_marks.rs new file mode 100644 index 00000000..e07e7b50 --- /dev/null +++ b/crates/gateway/src/proxy/cache_marks.rs @@ -0,0 +1,156 @@ +//! Prompt-cache breakpoints the conversion added, when an upstream refuses +//! them. +//! +//! A request converted for an Anthropic-format upstream, or for a Claude +//! model on Bedrock, from a caller that marked no breakpoints gets them +//! from the conversion layer (`tw_dialect`'s `auto_cache`). Not every +//! upstream takes them: an Anthropic-compatible endpoint that is not +//! Anthropic's may not know `cache_control`, and a Bedrock model AWS +//! stops listing refuses `cachePoint`. Either answers `400` and the whole +//! request fails. +//! +//! So the request goes to the same upstream once more without them (see +//! `generate::send`), and that upstream is remembered as refusing them for +//! that model: later requests leave them out before they are sent, instead +//! of being refused first every time. The upstream itself is fine, so it +//! is not failed over and its health is not touched. +//! +//! **Only breakpoints the conversion added are taken out.** Ones the caller +//! marked are its own decision; they go out as written, and a refusal of +//! them goes back to the caller like any other. The desktop gateway does +//! the same (`tw_gateway::cache_marks` in thinkwatch-core). + +use std::collections::HashMap; +use std::sync::Mutex; + +use crate::error::GatewayError; + +/// At most this many models remembered per upstream. A full table drops +/// the one noted first: forgetting costs one more refusal and resend. +const MAX: usize = 1024; + +/// The models one upstream refused the added breakpoints for. +/// +/// In memory only, per instance, and rebuilt with the router: whatever +/// is forgotten costs one more refusal and resend. +#[derive(Default)] +pub struct Refused { + /// Model → when it was noted, as a running count. + by: Mutex<(HashMap, u64)>, +} + +impl Refused { + /// This upstream refused the added breakpoints for `model`. + pub fn note(&self, model: &str) { + let mut guard = self.by.lock().unwrap_or_else(|p| p.into_inner()); + let (by, clock) = &mut *guard; + *clock += 1; + if by.len() >= MAX + && !by.contains_key(model) + && let Some(oldest) = by.iter().min_by_key(|(_, at)| **at).map(|(m, _)| m.clone()) + { + by.remove(&oldest); + } + by.insert(model.to_string(), *clock); + } + + /// Has this upstream refused the added breakpoints for `model`? + pub fn refused(&self, model: &str) -> bool { + let guard = self.by.lock().unwrap_or_else(|p| p.into_inner()); + guard.0.contains_key(model) + } +} + +/// Is this the upstream refusing prompt-cache breakpoints: a `400` whose +/// wording names `cache_control`, `cachePoint` or prompt caching? +/// +/// The wording, not the error's structure: compatible endpoints say +/// "messages.0.content.0.cache_control: Extra inputs are not permitted" or +/// "unknown field `cache_control`", Bedrock a `ValidationException` that +/// mentions `cachePoint` or prompt caching, each in a field of its own. +/// `label` is the upstream's name, which the error's message starts with. +pub fn is_refusal(err: &GatewayError, label: &str) -> bool { + let GatewayError::ProviderHttpError { + status: 400, + message, + } = err + else { + return false; + }; + let said = message + .strip_prefix(label) + .and_then(|m| m.strip_prefix(": ")) + .unwrap_or(message) + .to_ascii_lowercase(); + [ + "cache_control", + "cachepoint", + "cache_point", + "cache point", + "prompt caching", + "prompt_caching", + ] + .iter() + .any(|w| said.contains(w)) +} + +#[cfg(test)] +mod tests { + use super::*; + + fn http(status: u16, said: &str) -> GatewayError { + GatewayError::ProviderHttpError { + status, + message: format!("relay: {said}"), + } + } + + #[test] + fn what_upstreams_say_when_they_do_not_take_breakpoints() { + for said in [ + r#"{"type":"error","error":{"type":"invalid_request_error","message":"messages.0.content.0.cache_control: Extra inputs are not permitted"}}"#, + r#"{"error":{"message":"unknown field `cache_control`, expected one of `type`, `text`","code":"invalid_request"}}"#, + r#"{"message":"The model returned the following errors: cachePoint is not supported for this model."}"#, + r#"{"message":"Prompt caching is not supported for this model"}"#, + ] { + assert!(is_refusal(&http(400, said), "relay"), "{said}"); + } + for (status, said) in [ + ( + 400, + r#"{"type":"error","error":{"type":"invalid_request_error","message":"max_tokens: Field required"}}"#, + ), + ( + 400, + r#"{"message":"Malformed input request: #/messages/0/content: expected type: JSONArray"}"#, + ), + // Not a refusal of the request: the upstream failing. + (500, r#"{"message":"cache_control backend unavailable"}"#), + ] { + assert!(!is_refusal(&http(status, said), "relay"), "{said}"); + } + // The upstream's own name is not what it said. + assert!(!is_refusal( + &GatewayError::ProviderHttpError { + status: 400, + message: "prompt caching proxy: max_tokens: Field required".into(), + }, + "prompt caching proxy" + )); + } + + #[test] + fn the_memory_is_per_model_and_stays_bounded() { + let r = Refused::default(); + r.note("claude-opus-4-7"); + assert!(r.refused("claude-opus-4-7")); + assert!(!r.refused("claude-sonnet-4-5")); + for i in 0..MAX + 10 { + r.note(&format!("m{i}")); + } + assert_eq!(r.by.lock().unwrap().0.len(), MAX); + // The one noted first goes first. + assert!(!r.refused("claude-opus-4-7")); + assert!(r.refused(&format!("m{}", MAX + 9))); + } +} diff --git a/crates/gateway/src/proxy/generate.rs b/crates/gateway/src/proxy/generate.rs index 9e93eab6..e109ad81 100644 --- a/crates/gateway/src/proxy/generate.rs +++ b/crates/gateway/src/proxy/generate.rs @@ -259,6 +259,53 @@ pub(crate) struct Wire { /// Restores this hop's answer: the request's ledger, and any value /// only this hop carried. pub ledger: Ledger, + /// The body's prompt-cache breakpoints are ones the conversion added: + /// the caller marked none. An upstream that refuses them gets the + /// request again without them (see `proxy::cache_marks`). + pub auto_cache: bool, +} + +impl Wire { + /// Take out the cache breakpoints the conversion added. False when + /// there were none to take out. + fn drop_added_marks(&mut self) -> bool { + if !std::mem::take(&mut self.auto_cache) { + return false; + } + match tw_dialect::cache::strip_marks(self.dialect, &self.body) { + Some(body) => { + self.body = body; + true + } + None => false, + } + } + + async fn send_to( + &self, + upstream: &super::transport::Upstream, + call_ctx: &CallCtx, + ) -> Result { + upstream + .send( + self.body.clone(), + &self.path, + self.query.as_deref(), + self.dialect, + &self.headers, + call_ctx, + ) + .await + } + + /// This hop as `upstream` takes it: without the added breakpoints + /// when it refused them for `model` before. + fn for_upstream(mut self, upstream: &super::transport::Upstream, model: &str) -> Wire { + if self.auto_cache && upstream.cache_marks.refused(model) { + self.drop_added_marks(); + } + self + } } impl Outbound { @@ -300,6 +347,30 @@ impl Outbound { tw_dialect::params::cap_max_output_tokens(dialect, body, cap, official); } + /// What assembles an answer forwarded in the caller's own format, for + /// the cache and the audit row. + /// + /// It is made from the request alone. Decoding also notes how the + /// caller wants a converted answer written — for a Codex compaction, + /// as one item carrying the summary the upstream wrote — and a + /// forwarded answer is the upstream's own: assembled that way, OpenAI's + /// compaction would be recorded as a failed one. A request the + /// conversion layer cannot read at all can still be forwarded, and its + /// answer is assembled knowing only the model and whether it streams. + fn forwarded_session(&self, body: &Value, model: &str, target: &Target) -> Session { + let client = self.surface.dialect; + let request = + match tw_dialect::convert::decode(client, body, &self.path, internal_query(client)) { + Ok(decoded) => decoded.request, + Err(_) => tw_dialect::ir::Request { + model: model.to_string(), + stream: self.stream, + ..Default::default() + }, + }; + tw_dialect::convert::encode(&request, target).session + } + /// Address the request to `protocol`, naming `model` upstream. pub(crate) fn address( &self, @@ -357,7 +428,7 @@ impl Outbound { } } self.cap_output(client, &mut body, model, official); - let collect = decode(&body)?.encode(&target(client)).session; + let collect = self.forwarded_session(&body, model, &target(client)); let mut bytes = serde_json::to_vec(&body).unwrap_or_default(); // Reasoning signatures a conversion wrote earlier in this // conversation (`tw1.`-prefixed) were not issued by this @@ -376,6 +447,7 @@ impl Outbound { convert: None, collect, ledger, + auto_cache: false, }); } @@ -400,6 +472,10 @@ impl Outbound { // The conversion moved the placeholders along with the text; one it // assembled from two pieces is numbered here. let (body, ledger) = self.redaction.replace(body, &self.ledger); + // With none of the caller's own, any breakpoints are the ones the + // conversion added for Claude. + let auto_cache = decoded.request.cache.is_empty() + && tw_dialect::cache::may_have_marks(protocol.dialect(), &body); Ok(Wire { body, path: prepared.path, @@ -409,6 +485,7 @@ impl Outbound { convert: Some(prepared.session.clone()), collect: prepared.session, ledger, + auto_cache, }) } } @@ -416,9 +493,14 @@ impl Outbound { /// Send to `entry`, and if the upstream rejects the dialect this route /// is configured for, try its alternates and remember whichever answers. /// -/// A rejected dialect is known before a single byte of body arrives — -/// the status check happens inside `send` — so a stream needs no special -/// handling: nothing has reached the client yet when the retry happens. +/// An upstream that refuses the prompt-cache breakpoints the conversion +/// added gets the request once more without them, and is remembered as +/// refusing them for this model (see `proxy::cache_marks`). +/// +/// A rejected dialect or breakpoint is known before a single byte of body +/// arrives — the status check happens inside `send` — so a stream needs +/// no special handling: nothing has reached the client yet when the retry +/// happens. pub(crate) async fn send( entry: &RouteEntry, outbound: &Outbound, @@ -426,24 +508,32 @@ pub(crate) async fn send( db: &sqlx::PgPool, model: &str, ) -> Result<(reqwest::Response, Wire), GatewayError> { - let official = entry.upstream.is_official(); - let first = outbound.address(entry.protocol, model, official)?; - let result = entry - .upstream - .send( - first.body.clone(), - &first.path, - first.query.as_deref(), - first.dialect, - &first.headers, - call_ctx, - ) - .await; - let mut last = match result { + let upstream = &entry.upstream; + let official = upstream.is_official(); + let first = outbound + .address(entry.protocol, model, official)? + .for_upstream(upstream, model); + let mut last = match first.send_to(upstream, call_ctx).await { Ok(resp) => return Ok((resp, first)), Err(e) => e, }; + if first.auto_cache && super::cache_marks::is_refusal(&last, &upstream.label) { + let mut wire = first; + if wire.drop_added_marks() { + upstream.cache_marks.note(model); + tracing::info!( + provider = %entry.provider_name, + model, + "Upstream refused the cache breakpoints added on conversion — sending again without them" + ); + match wire.send_to(upstream, call_ctx).await { + Ok(resp) => return Ok((resp, wire)), + Err(e) => last = e, + } + } + } + if super::protocol_relearn::is_protocol_mismatch(&last) { for protocol in &entry.alternates { tracing::info!( @@ -453,19 +543,10 @@ pub(crate) async fn send( to = %protocol, "Upstream rejected the configured protocol — retrying with an alternate" ); - let wire = outbound.address(*protocol, model, official)?; - match entry - .upstream - .send( - wire.body.clone(), - &wire.path, - wire.query.as_deref(), - wire.dialect, - &wire.headers, - call_ctx, - ) - .await - { + let wire = outbound + .address(*protocol, model, official)? + .for_upstream(upstream, model); + match wire.send_to(upstream, call_ctx).await { Ok(resp) => { super::protocol_relearn::persist(db, entry.route_id, *protocol).await; return Ok((resp, wire)); @@ -662,10 +743,13 @@ async fn run( None => (body, raw), }; - // 4. Decode once, for the input estimate. + // 4. Decode once, for the input estimate and to know whether the + // request can be converted at all. One that cannot — it continues a + // conversation kept on OpenAI's servers (`previous_response_id`), or + // carries a compaction only OpenAI can read — can still be forwarded + // to an upstream of its own format, so that is where it goes (step 9). let decoded = - tw_dialect::convert::decode(surface.dialect, &raw, path, internal_query(surface.dialect)) - .map_err(|r| ctx.emit(GatewayError::TransformError(r.0)))?; + tw_dialect::convert::decode(surface.dialect, &raw, path, internal_query(surface.dialect)); // 5. Outbound redaction: the whole request is searched and what is // found numbered once, in the order the caller wrote it. @@ -822,8 +906,32 @@ async fn run( "No provider found for model: {mapped_model}" ))) })?; - - let input_estimate = crate::usage_estimate::request_tokens(&decoded.request); + // A request the conversion layer cannot read goes only to routes that + // forward it as sent. With none, the reason it cannot be converted is + // the answer, as it would be from any of the routes. + let forwarding: Vec; + let (routes, input_estimate) = match &decoded { + Ok(decoded) => ( + routes.as_slice(), + crate::usage_estimate::request_tokens(&decoded.request), + ), + Err(rejection) => { + forwarding = routes + .iter() + .filter(|r| r.protocol.dialect() == surface.dialect) + .cloned() + .collect(); + if forwarding.is_empty() { + return Err(ctx + .emit(GatewayError::TransformError(rejection.0.clone())) + .into()); + } + ( + forwarding.as_slice(), + crate::usage_estimate::raw_request_tokens(&outbound_body), + ) + } + }; let outbound = Outbound { surface, path: path.to_string(), @@ -1061,25 +1169,29 @@ mod tests { assert_eq!(gemini_target("/v1beta/models/:generateContent"), None); } - /// The body a Chat caller's request goes out with to a Chat upstream, - /// under a model cap of `cap`. - fn sent_to_chat(ask: Value, cap: u32, official: bool) -> Value { + fn outbound(surface: ClientSurface, path: &str, body: Value, cap: Option) -> Outbound { let redaction = Redaction::new(&tw_guard::policy::RedactPolicy { mode: tw_guard::policy::Mode::Off, ..Default::default() }); let (_, ledger) = redaction.look(b"{}"); - let outbound = Outbound { - surface: CHAT, - path: "/v1/chat/completions".into(), - body: ask, - stream: false, + Outbound { + surface, + path: path.into(), + stream: body["stream"] == true, + body, dialect_headers: Vec::new(), input_estimate: 0, redaction, ledger, - max_output_tokens: Some(cap), - }; + max_output_tokens: cap, + } + } + + /// The body a Chat caller's request goes out with to a Chat upstream, + /// under a model cap of `cap`. + fn sent_to_chat(ask: Value, cap: u32, official: bool) -> Value { + let outbound = outbound(CHAT, "/v1/chat/completions", ask, Some(cap)); let wire = outbound .address(UpstreamProtocol::OpenAiChat, "gpt-5", official) .unwrap_or_else(|e| panic!("{e:?}")); @@ -1102,6 +1214,112 @@ mod tests { assert!(sent.get("max_completion_tokens").is_none(), "{sent}"); } + /// A Responses request that continues a conversation kept on OpenAI's + /// servers, or carries a compaction only OpenAI can read, cannot be + /// converted — and is forwarded as sent to an upstream that speaks + /// Responses. + #[test] + fn a_request_the_conversion_cannot_read_still_goes_out_in_its_own_format() { + let ask = serde_json::json!({ + "model": "gpt-5.5", + "stream": true, + "previous_response_id": "resp_1", + "input": [ + {"type": "compaction", "encrypted_content": "gAAAAABo"}, + {"type": "message", "role": "user", "content": "next"} + ] + }); + let out = outbound(RESPONSES, "/v1/responses", ask.clone(), None); + let wire = out + .address(UpstreamProtocol::OpenAiResponses, "gpt-5.5", true) + .unwrap_or_else(|e| panic!("{e:?}")); + assert!(wire.convert.is_none()); + let sent: Value = serde_json::from_slice(&wire.body).unwrap(); + assert_eq!(sent, ask); + assert!(wire.collect.stream && wire.collect.model == "gpt-5.5"); + + match out.address(UpstreamProtocol::OpenAiChat, "gpt-5.5", true) { + Err(GatewayError::TransformError(_)) => {} + Err(e) => panic!("{e:?}"), + Ok(_) => panic!("converted a request that points at OpenAI's servers"), + } + } + + /// Codex's compaction forwarded to OpenAI comes back as OpenAI's own + /// compaction item. Assembled for the audit row as if converted, it + /// would read as a compaction that wrote no summary and failed. + #[test] + fn a_forwarded_compaction_is_assembled_as_the_upstreams_answer() { + let ask = serde_json::json!({ + "model": "gpt-5.5", + "stream": true, + "input": [ + {"type": "message", "role": "user", "content": "fix the test"}, + {"type": "compaction_trigger"} + ] + }); + let out = outbound(RESPONSES, "/v1/responses", ask.clone(), None); + let forwarded = out + .address(UpstreamProtocol::OpenAiResponses, "gpt-5.5", true) + .unwrap_or_else(|e| panic!("{e:?}")); + assert!(!forwarded.collect.is_compaction()); + let sent: Value = serde_json::from_slice(&forwarded.body).unwrap(); + assert_eq!(sent, ask); + + // Converted, the upstream's summary is what Codex gets back. + let converted = out + .address(UpstreamProtocol::OpenAiChat, "gpt-5.5", true) + .unwrap_or_else(|e| panic!("{e:?}")); + assert!(converted.convert.as_ref().unwrap().is_compaction()); + } + + /// Only breakpoints the conversion added may be taken out when an + /// upstream refuses them; the caller's own are its decision. + #[test] + fn only_breakpoints_the_conversion_added_count_as_added() { + let chat = outbound( + CHAT, + "/v1/chat/completions", + serde_json::json!({"model": "m", "messages": [{"role": "user", "content": "hi"}]}), + None, + ); + let mut wire = chat + .address( + UpstreamProtocol::AnthropicMessages, + "claude-sonnet-4-5", + true, + ) + .unwrap_or_else(|e| panic!("{e:?}")); + assert!(wire.auto_cache); + assert!(wire.drop_added_marks()); + let body = String::from_utf8(wire.body.clone()).unwrap(); + assert!( + !body.contains("cache_control") && body.contains("hi"), + "{body}" + ); + + let claude_code = outbound( + MESSAGES, + "/v1/messages", + serde_json::json!({ + "model": "m", "max_tokens": 16, + "system": [{"type": "text", "text": "rules", "cache_control": {"type": "ephemeral"}}], + "messages": [{"role": "user", "content": "hi"}] + }), + None, + ); + let mut wire = claude_code + .address( + UpstreamProtocol::BedrockNative, + "us.anthropic.claude-sonnet-4-5-20250929-v1:0", + true, + ) + .unwrap_or_else(|e| panic!("{e:?}")); + assert!(String::from_utf8_lossy(&wire.body).contains("cachePoint")); + assert!(!wire.auto_cache); + assert!(!wire.drop_added_marks()); + } + #[test] fn both_chat_limits_are_held_to_the_cap() { // An upstream that reads only `max_tokens` would otherwise go diff --git a/crates/gateway/src/proxy/mod.rs b/crates/gateway/src/proxy/mod.rs index 45088b48..6370147d 100644 --- a/crates/gateway/src/proxy/mod.rs +++ b/crates/gateway/src/proxy/mod.rs @@ -23,6 +23,7 @@ use think_watch_common::limits::weight; mod accounting; mod body_capture; +pub(crate) mod cache_marks; mod early_cancel; pub(crate) mod generate; mod headers; diff --git a/crates/gateway/src/proxy/transport.rs b/crates/gateway/src/proxy/transport.rs index 9e484cce..c9b5b8a6 100644 --- a/crates/gateway/src/proxy/transport.rs +++ b/crates/gateway/src/proxy/transport.rs @@ -65,6 +65,9 @@ pub struct Upstream { pub shape: Shape, /// Names the upstream in error messages. pub label: String, + /// The models this upstream refused the cache breakpoints a + /// conversion adds for (see `proxy::cache_marks`). + pub(crate) cache_marks: crate::proxy::cache_marks::Refused, } impl Upstream { @@ -75,6 +78,7 @@ impl Upstream { headers, shape, label: label.to_string(), + cache_marks: Default::default(), } } diff --git a/crates/gateway/src/usage_estimate.rs b/crates/gateway/src/usage_estimate.rs index 5b889710..8f6c669b 100644 --- a/crates/gateway/src/usage_estimate.rs +++ b/crates/gateway/src/usage_estimate.rs @@ -58,10 +58,25 @@ pub fn answer_tokens(body: &[u8]) -> u64 { let Ok(v) = serde_json::from_slice::(body) else { return to_tokens(body.len()); }; - to_tokens(text_len(&v)) + to_tokens(text_len(&v, &[])) } -fn text_len(v: &Value) -> usize { +/// The input of a request the conversion layer cannot read, in tokens. +/// +/// Such a request can still be forwarded in its own format — one that +/// continues a conversation kept on OpenAI's servers, say — but there is +/// no decoded request to measure. Its text is counted the way an answer's +/// is, and images and files are left out as [`request_tokens`] leaves +/// them out. +pub fn raw_request_tokens(v: &Value) -> u64 { + /// Images and files: a Chat or Responses `image_url` (often a data + /// URI), the base64 `data` of an Anthropic or Gemini part, and a + /// Responses `file_data`. + const MEDIA: &[&str] = &["image_url", "data", "file_data"]; + to_tokens(text_len(v, MEDIA)) +} + +fn text_len(v: &Value, skip: &[&str]) -> usize { const NOT_TEXT: &[&str] = &[ "id", "model", @@ -80,11 +95,11 @@ fn text_len(v: &Value) -> usize { ]; match v { Value::String(s) => s.len(), - Value::Array(a) => a.iter().map(text_len).sum(), + Value::Array(a) => a.iter().map(|v| text_len(v, skip)).sum(), Value::Object(o) => o .iter() - .filter(|(k, _)| !NOT_TEXT.contains(&k.as_str())) - .map(|(_, v)| text_len(v)) + .filter(|(k, _)| !NOT_TEXT.contains(&k.as_str()) && !skip.contains(&k.as_str())) + .map(|(_, v)| text_len(v, skip)) .sum(), _ => 0, } @@ -162,6 +177,24 @@ mod tests { assert_eq!(answer_tokens(body.to_string().as_bytes()), 2); } + #[test] + fn an_unreadable_request_counts_its_text_and_not_its_images() { + let body = json!({ + "model": "gpt-5.5", + "previous_response_id": "resp_0123456789", + "input": [ + {"type": "message", "role": "user", "content": [ + {"type": "input_text", "text": "12345678"}, + {"type": "input_image", "image_url": format!("data:image/png;base64,{}", "A".repeat(4000))} + ]}, + {"type": "compaction", "encrypted_content": "x".repeat(4000)} + ] + }); + // "resp_0123456789" and "12345678": 23 bytes. The model, the item + // types, the role, the image and the compaction are not text. + assert_eq!(raw_request_tokens(&body), 6); + } + #[test] fn nothing_reported_is_estimated_whole() { let answer = json!({"choices": [{"message": {"content": "x".repeat(40)}}]}).to_string(); diff --git a/crates/server/src/protocol_probe.rs b/crates/server/src/protocol_probe.rs index 7d2d0212..6b0171ca 100644 --- a/crates/server/src/protocol_probe.rs +++ b/crates/server/src/protocol_probe.rs @@ -150,13 +150,19 @@ pub(crate) async fn clear_for_provider(db: &sqlx::PgPool, provider_id: Uuid) -> /// converts live traffic. A hand-written probe body drifts from what /// forwarding actually sends, and then "the probe passed, forwarding /// fails" has nothing to go on. +/// +/// Without the prompt-cache breakpoints the conversion adds for Claude, +/// though: a probe is far too short to be cached, and an upstream that +/// does not take them would refuse the probe and have the model recorded +/// as unavailable for good. Live traffic recovers from that refusal on +/// its own (see the gateway's `proxy::cache_marks`). fn probe_request( upstream_model: &str, protocol: UpstreamProtocol, official: bool, ) -> tw_dialect::convert::Prepared { use tw_dialect::ir::{Message, Part, Request, Role, Target}; - tw_dialect::convert::encode( + let mut prepared = tw_dialect::convert::encode( &Request { model: upstream_model.to_string(), messages: vec![Message { @@ -171,7 +177,11 @@ fn probe_request( official, default_max_tokens: 1, }, - ) + ); + if let Some(body) = tw_dialect::cache::strip_marks(protocol.dialect(), &prepared.body) { + prepared.body = body; + } + prepared } /// Does this failure tell us anything about the model, or only about @@ -408,6 +418,27 @@ pub(crate) async fn resolve( mod tests { use super::*; + /// The conversion marks Claude-bound requests for caching; a probe goes + /// out without the marks, so an upstream that does not take them + /// cannot make a model look unavailable. + #[test] + fn a_probe_carries_no_cache_breakpoints() { + for (protocol, model) in [ + (UpstreamProtocol::AnthropicMessages, "claude-sonnet-4-5"), + ( + UpstreamProtocol::BedrockNative, + "us.anthropic.claude-sonnet-4-5-20250929-v1:0", + ), + ] { + let body = String::from_utf8(probe_request(model, protocol, true).body).unwrap(); + assert!( + !body.contains("cache_control") && !body.contains("cachePoint"), + "{body}" + ); + assert!(body.contains("hi"), "{body}"); + } + } + #[test] fn transient_failures_never_become_a_permanent_verdict() { // Verdicts don't expire, so recording "unavailable" off an diff --git a/crates/test-support/tests/gateway_proxy.rs b/crates/test-support/tests/gateway_proxy.rs index 38684f91..f8207697 100644 --- a/crates/test-support/tests/gateway_proxy.rs +++ b/crates/test-support/tests/gateway_proxy.rs @@ -517,3 +517,151 @@ async fn a_passthrough_request_leaves_the_gateways_own_signatures_behind() { assert_eq!(sent["messages"][1]["content"][0]["text"], "answer one"); assert_eq!(sent["messages"][2]["content"][0]["signature"], "EqQBCkgIBx"); } + +/// A Responses request the gateway cannot convert — it continues a +/// conversation kept on OpenAI's servers and carries a compaction only +/// OpenAI can read, as Codex sends after compacting against OpenAI — is +/// still forwarded as sent to a route that speaks Responses. Only when +/// every route would have to convert it is it refused, saying why. +#[ignore = "integration test — run via `make test-it`"] +#[tokio::test] +async fn a_request_only_openai_can_read_is_forwarded_to_a_responses_route() { + use wiremock::matchers::{method, path}; + + let app = TestApp::spawn().await; + let upstream = MockProvider::openai_chat_ok("resp-model").await; + upstream + .mount( + wiremock::Mock::given(method("POST")) + .and(path("/v1/responses")) + .respond_with(MockProvider::json(json!({ + "id": "resp_2", + "object": "response", + "created_at": 1_700_000_000_i64, + "model": "resp-model", + "status": "completed", + "output": [{ + "type": "message", "id": "msg_1", "role": "assistant", + "status": "completed", + "content": [{"type": "output_text", "text": "continued", + "annotations": []}] + }], + "usage": {"input_tokens": 9, "output_tokens": 2, "total_tokens": 11} + }))), + ) + .await; + let api_key = seed_provider_and_key(&app, &upstream.uri(), "openai", "resp-model", None).await; + let ask = json!({ + "model": "resp-model", + "previous_response_id": "resp_1", + "input": [ + {"type": "compaction", "encrypted_content": "gAAAAABo"}, + {"type": "message", "role": "user", "content": "go on"} + ] + }); + let gw = app.gateway_client(); + gw.set_bearer(&api_key); + + // The route converts to Chat Completions, an OpenAI provider's + // default: refused with the reason, and nothing sent. + let resp = gw.post("/v1/responses", ask.clone()).await.unwrap(); + resp.assert_status(400); + let body: Value = resp.json().unwrap(); + assert!( + body["error"]["message"] + .as_str() + .is_some_and(|m| m.contains("previous_response_id")), + "{body}" + ); + assert!(upstream.received_requests().await.is_empty()); + + // A route that speaks Responses: forwarded as sent. + sqlx::query( + "UPDATE model_routes SET upstream_protocol = 'openai_responses' WHERE model_id = $1", + ) + .bind("resp-model") + .execute(&app.db) + .await + .unwrap(); + app.rebuild_gateway_router().await; + let resp = gw.post("/v1/responses", ask).await.unwrap(); + resp.assert_ok(); + let body: Value = resp.json().unwrap(); + assert_eq!( + body["output"][0]["content"][0]["text"], "continued", + "{body}" + ); + + let sent = upstream.received_requests().await; + assert_eq!(sent.len(), 1); + assert_eq!(sent[0].url.path(), "/v1/responses"); + let sent: Value = serde_json::from_slice(&sent[0].body).unwrap(); + assert_eq!(sent["previous_response_id"], "resp_1", "{sent}"); + assert_eq!(sent["input"][0]["encrypted_content"], "gAAAAABo", "{sent}"); +} + +/// A request converted for an Anthropic-format upstream gets prompt-cache +/// breakpoints when the caller marked none. A compatible upstream that +/// does not take `cache_control` refuses the whole request; it gets it +/// again without them, and later requests to it leave them out from the +/// start. +#[ignore = "integration test — run via `make test-it`"] +#[tokio::test] +async fn an_upstream_that_refuses_added_cache_breakpoints_gets_the_request_without_them() { + use wiremock::matchers::{body_string_contains, method, path}; + + let app = TestApp::spawn().await; + let upstream = MockProvider::anthropic_messages_ok("claude-relay").await; + upstream + .mount( + wiremock::Mock::given(method("POST")) + .and(path("/v1/messages")) + .and(body_string_contains("cache_control")) + .respond_with(wiremock::ResponseTemplate::new(400).set_body_json(json!({ + "type": "error", + "error": { + "type": "invalid_request_error", + "message": "messages.0.content.0.cache_control: Extra inputs are not permitted" + } + }))) + .with_priority(1), + ) + .await; + let api_key = + seed_provider_and_key(&app, &upstream.uri(), "anthropic", "claude-relay", None).await; + let gw = app.gateway_client(); + gw.set_bearer(&api_key); + let ask = json!({ + "model": "claude-relay", + "messages": [ + {"role": "system", "content": "Be brief."}, + {"role": "user", "content": "hi"} + ] + }); + + // A Chat caller, so the request is converted and marked. + let resp = gw.post("/v1/chat/completions", ask.clone()).await.unwrap(); + resp.assert_ok(); + let body: Value = resp.json().unwrap(); + assert_eq!(body["choices"][0]["message"]["content"], "hi", "{body}"); + let sent = upstream.received_requests().await; + assert_eq!(sent.len(), 2, "refused, then sent again"); + let marked = String::from_utf8_lossy(&sent[0].body).into_owned(); + assert!(marked.contains("cache_control"), "{marked}"); + let plain: Value = serde_json::from_slice(&sent[1].body).unwrap(); + assert!(!plain.to_string().contains("cache_control"), "{plain}"); + assert_eq!(plain["system"][0]["text"], "Be brief.", "{plain}"); + assert_eq!(plain["messages"][0]["content"][0]["text"], "hi", "{plain}"); + + // Remembered: the next request goes out without them at once. (Not + // the same request, which the response cache would answer.) + let mut next = ask; + next["messages"][1]["content"] = json!("and again"); + gw.post("/v1/chat/completions", next) + .await + .unwrap() + .assert_ok(); + let sent = upstream.received_requests().await; + assert_eq!(sent.len(), 3); + assert!(!String::from_utf8_lossy(&sent[2].body).contains("cache_control")); +} diff --git a/deploy/helm/think-watch/Chart.yaml b/deploy/helm/think-watch/Chart.yaml index bd03e7f6..c24e00b4 100644 --- a/deploy/helm/think-watch/Chart.yaml +++ b/deploy/helm/think-watch/Chart.yaml @@ -2,8 +2,8 @@ apiVersion: v2 name: think-watch description: Enterprise AI API Gateway & MCP Management Platform type: application -version: 3.2.1 -appVersion: "3.2.1" +version: 3.3.0 +appVersion: "3.3.0" keywords: - ai - gateway diff --git a/deploy/helm/think-watch/README.md b/deploy/helm/think-watch/README.md index a7b4e519..ca9e6ec1 100644 --- a/deploy/helm/think-watch/README.md +++ b/deploy/helm/think-watch/README.md @@ -39,7 +39,7 @@ make helm-deploy HELM_VALUES=deploy/helm/think-watch/values-production.yaml.exam # Or one-off flags helm upgrade --install thinkwatch deploy/helm/think-watch \ --namespace thinkwatch --create-namespace \ - --set image.server.tag=3.2.1 \ + --set image.server.tag=3.3.0 \ --set ingress.enabled=true \ --set ingress.gateway.host=api.example.com ``` diff --git a/deploy/helm/think-watch/values-production.yaml.example b/deploy/helm/think-watch/values-production.yaml.example index 5ee59c33..308c319c 100644 --- a/deploy/helm/think-watch/values-production.yaml.example +++ b/deploy/helm/think-watch/values-production.yaml.example @@ -10,7 +10,7 @@ replicaCount: 3 image: server: - tag: "" # "" → .Chart.AppVersion; set e.g. "3.2.1" to pin (no "v") + tag: "" # "" → .Chart.AppVersion; set e.g. "3.3.0" to pin (no "v") web: tag: "" diff --git a/web/package.json b/web/package.json index 00eec0b7..22f90a06 100644 --- a/web/package.json +++ b/web/package.json @@ -1,7 +1,7 @@ { "name": "web", "private": true, - "version": "3.2.1", + "version": "3.3.0", "type": "module", "packageManager": "pnpm@12.9.1", "scripts": {