Repository navigation
fix(cost): price Codex per request at OpenAI's current and dated rates, and count each rollout's tokens once - #10
Conversation
…s, and count each rollout's tokens once Codex cost on the desktop was computed from each day's token totals with a price table that predates most of the models people now run: - gpt-5.5 used a gpt-5.4 placeholder ($2.50 / $15 per 1M), written before its prices were published. It is $5 / $30. - gpt-6-astra and the gpt-5.6 Sol / Terra / Luna / Cyber models had no row, so they showed no cost at all. The table now follows the one bundled with steipete/CodexBar at 25bba9b7 (MIT, notice in pricing.rs), which cites OpenAI's pricing pages. Two of its rules only make sense per request, so Codex cost is now summed per request while parsing (packed slot 3), the way Claude's already is: - Long context: a request with more than 272K input tokens is billed at the model's long-context rates for the whole request. A busy day's total crosses 272K when no request does, so pricing the day would bill it all at the higher rate. - Dated rates: Sol was repriced on 2026-08-21 and Terra / Luna on 2026-07-30. A request is billed at the rate in force at its timestamp. Aliases OpenAI routes to a priced model (gpt-5.6 -> Sol, gpt-reserve -> Luna, the Daybreak names) resolve to it. gpt-5.5-codex follows gpt-5.5, as every -codex row follows its base model; gpt-5.5-mini / -nano keep their placeholder rates, since neither has a published price. Two counting guards: - The scan reads sessions/ and archived_sessions/ (and each WSL distro's) and told files apart only by path, so a rollout present in both would be counted twice. Files with the same rollout id (session_meta payload.id) whose token events overlap in time are now one rollout, and the most complete copy counts. The id alone is not enough: a rollout can be continued in a second file under the same id with no event in common, and both halves are real. Sub-agent rollouts have their own id and are not touched (their session_id is the parent's, which is why it is not used). - A cumulative counter that dropped lowered the baseline, so when it jumped back up the gap was counted again. Each component now counts growth above the highest snapshot seen, or growth since the previous snapshot while below it. A flip-flopping counter cannot double; a restarted counter still counts (a pure high-water mark would drop everything after a restart, and a restart is the shape real logs show). The counter also advances on events outside the scan window, so a rollout that began before the window no longer lands its earlier usage on the window's first day. The Codex cache carries a per-provider rules version, so it is rebuilt once under the new rules while Claude's cache is kept. Tests: pricing (every new row, the 272K boundary on both sides, each repricing boundary to the millisecond, aliases, a table consistency check), the counter (restart, interleaving, repeats, last-only events, resuming with and without the saved peak, each compared against the rule it replaces), copy detection (identical copies, stale copy, continuation, sub-agents), and end-to-end scans for per-request pricing, dated rates, the window edge, archive copies, warm-cache reuse and incremental parsing across a drop. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three pricing tests used 1M-token requests to make the arithmetic round. 1M input is a long-context request, so the code correctly billed it at the long-context rate and the tests failed (gpt-5.5 read $14.50 instead of $8; Sol $10 instead of $5). They now use 100K requests, and the Sol test also checks today's long-context rate next to the earlier one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pro long-context tier OpenAI released gpt-6-sol, gpt-6-luna and gpt-6.1-sol after the CodexBar commit the table follows, and lists a long-context tier ($60 / $270 above 272K) for gpt-5.4-pro and gpt-5.5-pro that the upstream table does not carry. All rates checked against developers.openai.com/api/docs/pricing (2026-09-30); the macOS app's table has the same rows. gpt-6.1-sol's cached rate is 5% of input rather than 10%, and has its own test case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Negative controlsEach guard was broken on purpose on a throwaway branch taken from Branch 1 (six mutations; each has at least one failure only it explains):
Branch 2 (four mutations):
The rows added after |
…he Mac's own cases Review of this PR found the desktop and the macOS app (1.56, same release) would report different Codex numbers from the same logs. This makes them agree, and makes a future disagreement fail a test. - Counter: the baseline only rises, and a snapshot below it in any component is skipped (the Mac's rule 1, CodexBar's watermark). The restart-aware rule this PR had counted requests after a counter restart that the Mac and CodexBar do not; one rule for both apps wins over the more generous one. - A file whose first event carries a counter from elsewhere (a fork, or a continuation file under the same rollout id) counts only its own request: total minus last becomes the baseline (the Mac's rule 3, all files). It counted the carried total a second time, and priced it as one request, so above 272K at long-context rates. - The Mac's shared fixture (codex-accounting-cases.json, copied unchanged from cli-pulse-private 89156f66) runs in the Rust tests three ways: full, warm and incremental. 10 of 11 cases must match exactly; the sub-agent copied-history case is a listed difference (sub-agent rules are not ported) and is asserted both ways. - Aliases (gpt-5.6, gpt-reserve, Daybreak) are resolved only to find the rates. The stored and uploaded model name stays the one Codex logged, as on the Mac; renaming it would have left the already-uploaded rows beside new ones and doubled those days on iPhone and web. - The Codex cache records a digest of the price table (rows, dated rates, aliases, 272K threshold) and is rebuilt when it changes, so a table edit cannot leave cached requests at an old price or at $0.00. - A last line without a newline is read only when it already parses; otherwise it waits, and the saved offset stays at its first byte (Codex and Claude). The line used to be skipped for good. - Test: a copy in archived_sessions is still left out after the live file grows and is parsed incrementally. - The "copies" log line is at debug level; it repeated on every scan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… call files copies only when nested The Mac's accounting PR (cli-pulse-private#625) moved on while this was in review (4c3dc377), and the shared fixture moved with it: 14 cases now. Two of its changes are the Mac's answer to shapes the desktop also gets wrong, so the desktop follows them and stays in step. - Copied history. A sub-agent or fork rollout can begin with its parent's history, token events included, stamped when Codex copies them, so after the file's own session_meta. What marks them is the line number: the session_meta names the first line of the file's own history (subagent_history_start_ordinal) and every line carries its ordinal. A child's event numbered before that start, or stamped before its own session_meta, is the parent's: not counted, and it does not move the counter or the file's event span. Only the file's first session_meta is its own. The child marker is kept in the cache, so an incremental parse applies the rule too. Real sub-agent logs have this shape, so the desktop was counting the parent's copied usage a second time. - A child's first event without last_token_usage is taken as carried in full, as on the Mac. - Copies. Two files of one rollout are copies only when one's event span lies within the other's (a copy is the whole file or an earlier state of it). Any overlap at all dropped a file whose other events appear nowhere else. Most-complete-first now breaks ties on final input + output, as the Mac. - Fixture re-copied from 4c3dc377; the desktop now matches all 14 cases, so the list of known differences is empty. - Token sums saturate instead of overflowing on a corrupt count. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review follow-upTwo commits on top of Counting rule (the desktop and the Mac disagreed). Both apps now use one rule, the Mac's. The baseline only rises, and a snapshot below it is skipped. The restart-aware rule this PR had is gone. A restart inside a file is now undercounted the same way on the Mac and in CodexBar. The shared case Carried counters. On a file's first event with The Mac's shared fixture runs in the Rust tests.
The desktop now matches all 14 cases. The test's list of deliberate differences is empty; any entry added to it has to be asserted both ways. Aliases. Aliases are resolved only to look up the price. The stored and uploaded model name stays the one Codex logged, as on the Mac, so rows already uploaded under that name are not left beside a renamed one. The integration assertion is flipped. Price table in the cache check. The Codex cache records a digest of the rows, the dated rates, the aliases and the 272K threshold. It is rebuilt whenever that digest changes, so no version bump has to be remembered. Tests check that the digest changes with each of those parts, and that a cache priced with another table, or with none, is rebuilt. Half-written last line. A last line without a newline is read only if it already parses. Otherwise it waits, and the saved offset stays at the line's first byte. This covers Codex and Claude, which share the loop. Tests cover both, plus a complete last line that has no newline. Long context from Copy detection after a resume. A new integration test grows the live file between two warm scans while a copy sits in Differences written down. The PR body now has a "Known differences from the Mac" section: unknown models, date-folder selection, unparseable timestamps, last-only events, and the long-context input. The "copies" log line is now at debug level. Negative controlsThrowaway branches taken from these commits ran
When #625 changes its fixture again, copy it into |
…own request decides its tier The Mac's accounting (cli-pulse-private#625) changed twice more before it merged (f18939f2), and the shared fixture grew from 14 cases to 20. The desktop follows, so the two count the same logs the same way again. - Copied history has two shapes. Codex's migration of older sub-agent rollouts moves subagent_history_start_ordinal to the end of the file and drops the copied session_meta lines, so every line, the sub-agent's own work included, is numbered before the boundary. The previous revision took all of it for the parent's copy. Now the lines before the boundary are copied only once an ancestor's session_meta numbered before it shows they are. Without one, the events before the parent's first inter-agent message are its replayed last requests and are dropped; the rest counts. Until a marker comes, those events are held; if none comes before a line past the boundary or the end of the parse, they count, in order. The answer is kept in the cache (CodexCopiedPrefix), so an incremental parse goes on applying it. - Only a file's first line can be its own session_meta. When it is something else the identity stays unknown, so a parent's copied session_meta is never taken for the file's own. - Lines over 32 KB are not decoded; their first 4 KB is enough to recognise a copied session_meta or an inter-agent message, and the line number is read from the first 512 bytes. As on the Mac. - Cost: the request an event reports (last_token_usage) decides the 272K tier, not the tokens counted. Growth of the counter beyond that request (a request that wrote no event of its own) is billed at standard rates. The Mac's codexEventCostUSD, field by field. - A dated spelling of an alias (gpt-5.6-2026-08-01) is billed as its model, and the price fingerprint records how a fixed list of names resolves, as the Mac's does. - Codex rules version 2 (1 was this PR's earlier revision, never released). - The fixture is the Mac's file byte for byte, pinned by commit and SHA-256; scripts/sync-codex-accounting-cases.sh copies a new one. On this Mac's own logs (aggregates only), the desktop's real-log scan and the Mac's replica agree on every day and model in three 31-day windows: 09-03.. 10-03, 09-01..10-01 and 08-21..09-20. Under the previous revision the last window lost 146M input tokens of migrated sub-agents' own work. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… copy Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…is not a date, not a panic parse_day_key_local fell back to the first ten bytes of a timestamp it could not parse, by slicing the string. A byte 10 inside a multi-byte character panicked the scan. It now takes the prefix only when it is whole. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brought up to the Mac's final rules (#625 as merged,
|
| Mac | desktop | same? |
|---|---|---|
observeSessionMeta from the first line only (≤ 1 MB); observeUnreadableFirstLine otherwise |
first line only, same limit | yes, except the desktop also reads a non-session_meta first line as an ordinary line (listed in the body) |
later session_meta → observeCopiedSessionMeta(codexLineOrdinal), never decoded |
same, line number from the first 512 bytes | yes |
inter_agent_communication_metadata → observeInterAgentMessage, only while awaiting a marker |
same | yes |
| lines over 32 KB: 4 KB head for markers, never decoded | same | yes |
receive: hold events before the boundary; at a line past it, held ones first; an event without a number is judged at once |
CodexResume::receive |
yes |
finish at the end of every read, noMarker only if something was held |
same | yes |
count: ordinal rule only after ancestorMetadata; time rule for any child |
is_copied_history |
yes |
| rules 1 and 3 | CodexCounter (unchanged) |
yes, except the last-only baseline (listed) |
CodexCopiedPrefix persisted in the file state |
CodexChildMeta.copied_prefix |
yes |
codexEventCostUSD: request decides the tier; growth beyond it at standard rates, field by field |
codex_event_cost_usd |
yes |
aliasTarget with a dated spelling |
codex_price_key |
yes (fixed here) |
| fingerprint covers name resolution | CODEX_FINGERPRINT_MODEL_NAMES |
yes (added here) |
| rate rows, dated rates, aliases | table | all 29 shared rows equal; 3 desktop-only rows as before |
Found and fixed while reviewing: the dated alias spelling was unpriced on the desktop, and parse_day_key_local sliced a timestamp at byte 10, which panics inside a multi-byte character (e711015).
An independent review (Gemini 3.1 Pro, read-only, given the diff, the desktop files and the Mac files): no findings, APPROVE.
Negative controls
Throwaway branches with several mutations each ran cargo test --all --no-fail-fast on ubuntu (runs 37119515801, 37119516848, 37119518036), and locally with the same result. Mutations that shared a failing test were also run one at a time locally, so each has a test that fails for it alone. All caught; the branches are deleted.
| mutation | caught by |
|---|---|
ordinal marks a copy without an ancestor's session_meta (the previous revision) |
copied_history_is_told_by_line_number_once_marked_or_by_time; shared cases migrated_subagent_counts_its_own_work, …_skips_the_replayed_parent_tail, …_without_an_inter_agent_message_counts_in_full (all scans) |
| inter-agent message not recognised | shared case migrated_subagent_skips_the_replayed_parent_tail (all scans) |
| the event past the boundary counted before the held ones | held_events_count_in_order_when_no_marker_comes |
| a marker's line number never read | the_line_number_is_read_without_decoding_the_line |
| decided copied part not kept for the next parse | the_copied_part_decided_in_one_parse_holds_in_the_next; shared case copied_history_marked_by_ordinal (incremental) |
| held events dropped at the end of a parse | shared case migrated_subagent_without_an_inter_agent_message_counts_in_full (all scans) |
first session_meta anywhere is the file's own |
only_the_first_line_can_be_the_files_own_session_meta |
| long lines skipped before their head is looked at | a_copied_session_meta_too_long_to_decode_still_marks_the_copy |
| tier decided by the counted tokens | codex_long_context_tier_is_decided_by_the_request_the_event_reports |
| growth beyond the request takes the request's tier | an_events_own_request_decides_its_long_context_tier, codex_growth_beyond_the_request_is_billed_at_standard_rates |
| dated alias not resolved | a_dated_alias_is_billed_as_its_model_and_keeps_its_name |
| fingerprint without name resolution | pricing_fingerprint_records_how_names_resolve |
| fixture edited in place | the_shared_cases_are_the_macs_file_unchanged; shared case migrated_subagent_counts_its_own_work |
Real logs
See "Checked against the Mac on real logs" in the body: three windows, every day and model equal.
🤖 Generated with Claude Code
The Mac's 1.55 has been cut, so this no longer waits for it. It now follows the Mac's final 1.56 Codex rules (cli-pulse-private#625 as merged).
Codex cost on the desktop came from a price table that predates most of the models people now run, and it was computed from each day's token totals.
gpt-5.5was priced at agpt-5.4placeholder, half its published rate.gpt-6-astraand thegpt-5.6Sol, Terra, Luna and Cyber models had no rates, so they showed no cost at all. This PR brings the table up to date and prices Codex per request. It also changes how Codex tokens are counted so that the desktop and the macOS app report the same numbers from the same logs, and so the same tokens cannot be counted twice.What changed
Prices (
src-tauri/src/pricing.rs)25bba9b7(MIT; the full notice is in the file header). That table cites OpenAI's pricing pages.gpt-5.5gpt-5.4gpt-5.4-pro,gpt-5.5-progpt-6-astragpt-5.6-solgpt-5.6-terragpt-5.6-lunagpt-5.6-cyber,gpt-5.5-cybergpt-6-solgpt-6-lunagpt-6.1-solgpt-6-sol,gpt-6-luna,gpt-6.1-soland the-prolong-context tier are not in CodexBar's table at that commit. They were released after it, or it does not carry them. They come from OpenAI's pricing page, checked 2026-09-30 and rechecked in review on 2026-10-01, and match the macOS app's table.gpt-5.6→ Sol,gpt-reserve→ Luna,gpt-daybreak-blue-latest→ Sol,gpt-daybreak-red-latest→ Cyber, and a dated spelling of any of them (gpt-5.6-2026-08-01), as on the Mac. Only the price lookup resolves them. The model is stored, uploaded and shown under the name Codex logged, as on the Mac. Renaming it would have left the daily rows already uploaded under the old name beside new ones under the target's name, because the upload only ever upserts, and those days would have counted twice on iPhone and the web.gpt-5.5-codexis not in OpenAI's list. It now costs the same asgpt-5.5, because every-codexrow in the table costs the same as its base model.gpt-5.5-miniandgpt-5.5-nanokeep their placeholder rates, since neither has a published price.None), which is different fromgpt-5.3-codex-sparkat a known $0.Per-request cost (
scanner.rs,cache.rs)last_token_usagedecides the tier. When the growth is more than that request (a turn aborted after a request that wrote no event of its own), the part beyond it is billed at standard rates: nothing shows how large those requests were. This is the Mac'scodexEventCostUSD, field by field.$0.00for its cached requests as if that were its price.Counting Codex tokens the way the Mac does
Codex logs a cumulative counter on each
token_countevent, usually with that request's ownlast_token_usage. The desktop now counts them with the same rules as the macOS app's 1.56 accounting (cli-pulse-private#625):The baseline only rises. A snapshot below it in any component is skipped and leaves it where it is. Before, the baseline followed every drop, so when the counter jumped back up the gap was counted again, once per flip for a file whose counter alternates between two series.
A counter carried over is not counted again. When a file's first event reports a total larger than its own request, the difference was counted elsewhere: by the rollout a fork continues, or by an earlier file of the same rollout. It becomes the baseline. This applies to every file. A fresh counter's first total equals its first request, so nothing changes for it.
A sub-agent's copy of its parent's history is not counted again, and its own work is. A sub-agent or fork rollout can begin with its parent's history copied in, token events included. Codex stamps the copied lines when it writes them, after the file's own
session_meta, so their times say nothing. Thesession_metanames the first line of the file's own history (subagent_history_start_ordinal), and every line carries itsordinal. What that boundary means depends on which of two shapes the file has:session_meta, numbered before the boundary. Then every event numbered before the boundary is the parent's.session_metalines and moves the boundary to the end of the file, so every line, the sub-agent's own work included, is numbered before it. The boundary marks nothing there. The events before the parent's first inter-agent message (inter_agent_communication_metadata) are its last requests, replayed, and are not counted; the rest is the sub-agent's own.An event stamped before the file's own
session_metais also the parent's (logs without ordinals). No copied event is counted or moves the counter. The answer, once known, is kept in the cache, so an incremental parse goes on applying it.Only a file's first line can be its own
session_meta. When the first line is something else, the identity stays unknown, so a parent's copiedsession_metais never taken for the file's own (which would make the file look like a copy of its parent).Lines over 32 KB are not decoded. Their first 4 KB is enough to recognise a copied
session_metaor an inter-agent message, and the line number is read from the first 512 bytes. Token events and turn contexts are far shorter (on this Mac's logs the longest are 887 B and 10 KB). Same limits as the Mac.Cost: a counter that restarts inside a file is counted only once its total passes the old high, so the requests after a restart are undercounted. The Mac and CodexBar make the same choice. A restart-aware rule would count them, but it would also make the two apps disagree, and it cannot tell a restart from a replayed snapshot. The shared case
counter_restart_counts_only_above_the_old_highpins the choice for both apps.The counter also advances on events outside the scan window, so a rollout that began before the window no longer puts its earlier usage on the window's first day.
The Mac's shared fixture (
codex-accounting-cases.json, the Mac's file byte for byte as of cli-pulse-privatef18939f2(#625), 20 cases) runs in the Rust tests as a full scan, a warm scan and an incremental scan. The desktop matches every case. The test keeps a list of deliberate differences, which is empty; an entry there has to be asserted both ways. A test pins the file to the Mac's commit and SHA-256, so it cannot be edited here by accident;scripts/sync-codex-accounting-cases.sh <commit>copies a new one and prints the values to pin.A rollout in two places is counted once
sessions/andarchived_sessions/, plus each WSL distro's copies of both. Until now it told files apart only by path, so a rollout present in two places was counted twice. That happens with a copy instead of a move, a restored or synced folder, or a WSL home linked to the Windows one.session_meta.payload.id) are copies when one file's token events lie within the other's time span, since a copy is the whole file or an earlier state of it. The most complete file counts. Completeness means most events, then the larger final input + output, then path order. This is the Mac's rule.session_idis the parent's, which is why the rule does not use it.A line caught half-written is read once it is complete
Cache
CostUsageCache::for_provider, or the next load would throw it away. A warm-scan test covers this.Known differences from the Mac
Left as they are, on purpose or because they predate this PR:
session_metais read like any other line on the desktop; the Mac skips it. The file's identity stays unknown on both. Codex always writes thesession_metafirst.last_token_usagemove the desktop's baseline up, so a later cumulative total does not count them again. The Mac leaves its baseline alone there. The two differ only for a file that mixes both kinds of event.Not changed
sync_nowuploads the recomputed daily rows for the scan window as usual. Model names are unchanged, so no existing uploaded row is orphaned. No schema or RPC change.gpt-5.5,gpt-6-astraandgpt-5.6users that is higher than before, which can trip a budget alert that did not fire until now.Tests
pricing.rs:-protier (no cached rate, so cached input is billed at the long-context input rate) andgpt-6.1-sol's 5% cached rate;gpt-5.5-codex=gpt-5.5;pricing.rs(new): an event's own request decides its tier, growth beyond it is billed at standard rates, field by field; a dated alias is billed as its model and keeps its name; the fingerprint records how names resolve.cache.rs: a Codex cache from older rules (including the unreleased rules version 1), or priced with another table (or none), is rebuilt; one under the current rules and prices is kept; Claude's is unaffected.scanner.rsunit tests: the counter on monotone, repeated, interleaved, down-and-back, restarted, below-baseline, negative, carried-over, child-without-lastand last-only snapshots. Copied history: by line number only once an ancestor'ssession_metamarks it, or by time; an ancestor'ssession_metaor an inter-agent message drops the held events, and a marker numbered past the boundary is not one; held events count in log order when no marker comes (at a line past the boundary or at the end of a parse); a copiedsession_metatoo long to decode still marks the copy; only the first line can be the file's ownsession_meta; the decided answer survives a resume; the line number is read without decoding the line. The shape tests also assert what the pre-1.56 rule and the restart-aware alternative would give, so each fixture tells the rules apart. Resuming from the saved baseline equals a full pass, and resuming without it does not. Copy detection covers identical copies, a stale copy, a continuation, a partial overlap, and different ids overlapping.tests/scanner_integration.rs, end-to-end scans:gpt-5.6alias and reported under that name;sessions/andarchived_sessions/counted once, including the origin split, and still once after the live file grows between scans;session_idall count;codex_scan_real_logs_when_asked, scans a real Codex home and writes per-day, per-model numbers (no paths) for comparison with the Mac.Checked against the Mac on real logs
The opt-in scan and the Mac's replica (
scripts/codex_accounting_replica.pyin cli-pulse-private at26702809, which agreed with the Mac app on every day of the 1.56 acceptance run) were run back to back on one machine's Codex logs, in three 31-day windows. Only aggregates were compared:The desktop's Codex cost for the 09-01 .. 10-01 window is the same, to the cent, as the Mac's 1.56 scanner gave for it on 10-01.
Under this PR's previous revision, which took every line numbered before a sub-agent's boundary for copied history, the last window would have lost more than a third of its input tokens: the migrated sub-agents' own work.
🤖 Generated with Claude Code