Skip to content

fix(cost): price Codex per request at OpenAI's current and dated rates, and count each rollout's tokens once - #10

Merged
JasonYeYuhe merged 8 commits into
mainfrom
p156-desktop-pricing
Oct 3, 2026
Merged

JasonYeYuhe merged 8 commits into
mainfrom
p156-desktop-pricing

Conversation

@JasonYeYuhe

@JasonYeYuhe JasonYeYuhe commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

The Mac's 1.55 has been cut, so this no longer waits for it. It now follows the Mac's final 1.56 Codex rules (cli-pulse-private#625 as merged).

Codex cost on the desktop came from a price table that predates most of the models people now run, and it was computed from each day's token totals. gpt-5.5 was priced at a gpt-5.4 placeholder, half its published rate. gpt-6-astra and the gpt-5.6 Sol, Terra, Luna and Cyber models had no rates, so they showed no cost at all. This PR brings the table up to date and prices Codex per request. It also changes how Codex tokens are counted so that the desktop and the macOS app report the same numbers from the same logs, and so the same tokens cannot be counted twice.

What changed

Prices (src-tauri/src/pricing.rs)

  • The Codex table follows the one bundled with steipete/CodexBar at 25bba9b7 (MIT; the full notice is in the file header). That table cites OpenAI's pricing pages.
model before (in / out per 1M) now long context (> 272K input)
gpt-5.5 $2.50 / $15 (placeholder) $5 / $30 $10 / $45
gpt-5.4 $2.50 / $15 unchanged $5 / $22.50 (new tier)
gpt-5.4-pro, gpt-5.5-pro $30 / $180 unchanged $60 / $270 (new tier)
gpt-6-astra none $10 / $50 $20 / $75
gpt-5.6-sol none $4 / $20 (before 2026-08-21: $5 / $30) $8 / $30 (before: $10 / $45)
gpt-5.6-terra none $2 / $12 (before 2026-07-30: $2.50 / $15) $4 / $18 (before: $5 / $22.50)
gpt-5.6-luna none $0.20 / $1.20 (before 2026-07-30: $1 / $6) $0.40 / $1.80 (before: $2 / $9)
gpt-5.6-cyber, gpt-5.5-cyber none $12.50 / $75 no tier
gpt-6-sol none $2 / $10 $4 / $15
gpt-6-luna none $0.10 / $0.50 $0.20 / $0.75
gpt-6.1-sol none $2 / $10 (cached $0.10, 5% of input) $4 / $15
  • gpt-6-sol, gpt-6-luna, gpt-6.1-sol and the -pro long-context tier are not in CodexBar's table at that commit. They were released after it, or it does not carry them. They come from OpenAI's pricing page, checked 2026-09-30 and rechecked in review on 2026-10-01, and match the macOS app's table.
  • Cached input is billed at each row's cached rate, as before. Cache-write rates are left out: Codex CLI logs report no cache writes.
  • Names OpenAI routes to a priced model are billed at that model's rates: gpt-5.6 → Sol, gpt-reserve → Luna, gpt-daybreak-blue-latest → Sol, gpt-daybreak-red-latest → Cyber, and a dated spelling of any of them (gpt-5.6-2026-08-01), as on the Mac. Only the price lookup resolves them. The model is stored, uploaded and shown under the name Codex logged, as on the Mac. Renaming it would have left the daily rows already uploaded under the old name beside new ones under the target's name, because the upload only ever upserts, and those days would have counted twice on iPhone and the web.
  • gpt-5.5-codex is not in OpenAI's list. It now costs the same as gpt-5.5, because every -codex row in the table costs the same as its base model. gpt-5.5-mini and gpt-5.5-nano keep their placeholder rates, since neither has a published price.
  • Unknown models still have no cost (None), which is different from gpt-5.3-codex-spark at a known $0.

Per-request cost (scanner.rs, cache.rs)

  • Two of these rules depend on the single request:
    • Long context. One request above 272K input tokens is billed at the long-context rates for the whole request. A busy day's total crosses 272K even when no single request does, so pricing the day would bill all of it at the higher rate.
    • Dated rates. A request is billed at the rate in force at its timestamp.
  • So Codex cost is now summed per request while parsing, into packed slot 3. Claude's cost already worked this way (slot 4).
  • The request is the one the event reports. What an event counts is the growth of a cumulative counter, which is not a request: for a file's first event that carries a counter over, it is less than the request, and after a counter restart it can be a small part of a long-context one. So the event's own last_token_usage decides the tier. When the growth is more than that request (a turn aborted after a request that wrote no event of its own), the part beyond it is billed at standard rates: nothing shows how large those requests were. This is the Mac's codexEventCostUSD, field by field.
  • Because cost is fixed when a file is parsed, the Codex cache now records a digest of the price table: every row, the dated rates, the aliases, the 272K threshold, and which row each of a fixed list of model names resolves to (the Mac's fingerprint names). A cache priced with another table is rebuilt on load, Codex only. Editing the table therefore reaches requests already cached. Without this, a model that gains a price would keep showing $0.00 for its cached requests as if that were its price.

Counting Codex tokens the way the Mac does

Codex logs a cumulative counter on each token_count event, usually with that request's own last_token_usage. The desktop now counts them with the same rules as the macOS app's 1.56 accounting (cli-pulse-private#625):

  • The baseline only rises. A snapshot below it in any component is skipped and leaves it where it is. Before, the baseline followed every drop, so when the counter jumped back up the gap was counted again, once per flip for a file whose counter alternates between two series.

  • A counter carried over is not counted again. When a file's first event reports a total larger than its own request, the difference was counted elsewhere: by the rollout a fork continues, or by an earlier file of the same rollout. It becomes the baseline. This applies to every file. A fresh counter's first total equals its first request, so nothing changes for it.

  • A sub-agent's copy of its parent's history is not counted again, and its own work is. A sub-agent or fork rollout can begin with its parent's history copied in, token events included. Codex stamps the copied lines when it writes them, after the file's own session_meta, so their times say nothing. The session_meta names the first line of the file's own history (subagent_history_start_ordinal), and every line carries its ordinal. What that boundary means depends on which of two shapes the file has:

    • A current child rollout copies its ancestor's history in with the ancestor's session_meta, numbered before the boundary. Then every event numbered before the boundary is the parent's.
    • Codex's migration of older sub-agent rollouts drops the copied session_meta lines and moves the boundary to the end of the file, so every line, the sub-agent's own work included, is numbered before it. The boundary marks nothing there. The events before the parent's first inter-agent message (inter_agent_communication_metadata) are its last requests, replayed, and are not counted; the rest is the sub-agent's own.
    • Until one of the two markers is seen, events numbered before the boundary are held. If neither comes before a line past the boundary or the end of the parse, they count, in log order.

    An event stamped before the file's own session_meta is also the parent's (logs without ordinals). No copied event is counted or moves the counter. The answer, once known, is kept in the cache, so an incremental parse goes on applying it.

  • Only a file's first line can be its own session_meta. When the first line is something else, the identity stays unknown, so a parent's copied session_meta is never taken for the file's own (which would make the file look like a copy of its parent).

  • Lines over 32 KB are not decoded. Their first 4 KB is enough to recognise a copied session_meta or an inter-agent message, and the line number is read from the first 512 bytes. Token events and turn contexts are far shorter (on this Mac's logs the longest are 887 B and 10 KB). Same limits as the Mac.

  • Cost: a counter that restarts inside a file is counted only once its total passes the old high, so the requests after a restart are undercounted. The Mac and CodexBar make the same choice. A restart-aware rule would count them, but it would also make the two apps disagree, and it cannot tell a restart from a replayed snapshot. The shared case counter_restart_counts_only_above_the_old_high pins the choice for both apps.

  • The counter also advances on events outside the scan window, so a rollout that began before the window no longer puts its earlier usage on the window's first day.

  • The Mac's shared fixture (codex-accounting-cases.json, the Mac's file byte for byte as of cli-pulse-private f18939f2 (#625), 20 cases) runs in the Rust tests as a full scan, a warm scan and an incremental scan. The desktop matches every case. The test keeps a list of deliberate differences, which is empty; an entry there has to be asserted both ways. A test pins the file to the Mac's commit and SHA-256, so it cannot be edited here by accident; scripts/sync-codex-accounting-cases.sh <commit> copies a new one and prints the values to pin.

A rollout in two places is counted once

  • The scan reads sessions/ and archived_sessions/, plus each WSL distro's copies of both. Until now it told files apart only by path, so a rollout present in two places was counted twice. That happens with a copy instead of a move, a restored or synced folder, or a WSL home linked to the Windows one.
  • Two files that share a rollout id (session_meta.payload.id) are copies when one file's token events lie within the other's time span, since a copy is the whole file or an earlier state of it. The most complete file counts. Completeness means most events, then the larger final input + output, then path order. This is the Mac's rule.
  • The id alone is not enough, and neither is an overlap in time. A rollout can be continued in a second file under the same id. Two files whose spans overlap without one containing the other each hold events the other lacks. All of these files are counted, and a counter a later file carries over is counted once (see above).
  • Sub-agent rollouts are not affected: each has its own rollout id. Their session_id is the parent's, which is why the rule does not use it.
  • The day totals and the native/WSL split both leave the copy out, so they still add up. This still holds after the live file grows and is parsed incrementally.

A line caught half-written is read once it is complete

  • A scan that ran while Codex or Claude Code was writing a line took the part it saw as a line of its own. It did not parse, and the saved offset moved past it, so the next scan started in the middle of the line and that request was lost for good.
  • Now a last line without a newline is read only when it already parses, which means it is a finished line written without a trailing newline. Otherwise it is left, the saved offset stays at its first byte, and the next scan reads the whole line. This is the same rule as the Mac and CodexBar #2168. It applies to Codex and Claude, which read their logs through the same loop.

Cache

  • Each provider's cache records a rules version, and the Codex cache also records the price digest. The Codex cache is rebuilt once after updating, and again whenever the price table changes. The Codex rules version is 2; 1 was this PR's earlier revision, never released. Claude's cache is kept, because nothing on the Claude side changes what a finished file contributes.
  • A fresh cache has to be created through CostUsageCache::for_provider, or the next load would throw it away. A warm-scan test covers this.

Known differences from the Mac

Left as they are, on purpose or because they predate this PR:

  • Unknown models have no cost on the desktop. The Mac prices them from the nearest known version and marks them "≈" (cli-pulse-private#623).
  • File selection. The desktop reads date folders from the start of the window to today. The Mac reads from one day earlier to one day later. This predates the PR.
  • A timestamp that does not parse still places the line by the date it starts with on the desktop; the Mac skips the line. This predates the PR.
  • A first line that is not a session_meta is read like any other line on the desktop; the Mac skips it. The file's identity stays unknown on both. Codex always writes the session_meta first.
  • Events with only last_token_usage move the desktop's baseline up, so a later cumulative total does not count them again. The Mac leaves its baseline alone there. The two differ only for a file that mixes both kinds of event.

Not changed

  • CodexBar's full fork and sub-agent accounting (history bases, interleaved lineages, compact sub-agents) is not ported, on either platform. Both use the smaller set of rules above.
  • API Fast (priority) multipliers are not applied; Standard rates are used.
  • No UI strings changed, so the locale catalogues are untouched.
  • Cost values change for Codex users, and sync_now uploads the recomputed daily rows for the scan window as usual. Model names are unchanged, so no existing uploaded row is orphaned. No schema or RPC change.
  • Budget alerts and the cost forecast read the same daily entries, so they see the corrected Codex cost too. For gpt-5.5, gpt-6-astra and gpt-5.6 users that is higher than before, which can trip a budget alert that did not fire until now.

Tests

  • pricing.rs:
    • every new row, including the -pro tier (no cached rate, so cached input is billed at the long-context input rate) and gpt-6.1-sol's 5% cached rate;
    • both sides of the 272K line (exactly 272,000 is standard);
    • each repricing boundary to the millisecond (the constants are checked against their UTC dates);
    • aliases: billed as their target, dated rates included, but kept under the logged name; every alias names a priced row;
    • gpt-5.5-codex = gpt-5.5;
    • priced-at-zero vs unknown;
    • a consistency check over every row: cached ≤ input ≤ output, and the long-context tier is never cheaper;
    • the price digest is stable, and changes with a new row, a rate, a removed cached rate or tier, a repricing date, an earlier rate, an alias and the threshold.
  • pricing.rs (new): an event's own request decides its tier, growth beyond it is billed at standard rates, field by field; a dated alias is billed as its model and keeps its name; the fingerprint records how names resolve.
  • cache.rs: a Codex cache from older rules (including the unreleased rules version 1), or priced with another table (or none), is rebuilt; one under the current rules and prices is kept; Claude's is unaffected.
  • scanner.rs unit tests: the counter on monotone, repeated, interleaved, down-and-back, restarted, below-baseline, negative, carried-over, child-without-last and last-only snapshots. Copied history: by line number only once an ancestor's session_meta marks it, or by time; an ancestor's session_meta or an inter-agent message drops the held events, and a marker numbered past the boundary is not one; held events count in log order when no marker comes (at a line past the boundary or at the end of a parse); a copied session_meta too long to decode still marks the copy; only the first line can be the file's own session_meta; the decided answer survives a resume; the line number is read without decoding the line. The shape tests also assert what the pre-1.56 rule and the restart-aware alternative would give, so each fixture tells the rules apart. Resuming from the saved baseline equals a full pass, and resuming without it does not. Copy detection covers identical copies, a stale copy, a continuation, a partial overlap, and different ids overlapping.
  • tests/scanner_integration.rs, end-to-end scans:
    • the Mac's 20 shared cases, each run full, warm and incremental (every file half written, scanned, then completed), and the fixture pinned to the Mac's file;
    • two 200K requests priced standard, not as one 400K day, and a 300K request at long-context rates;
    • 100K counted of a 300K request after a counter restart is billed at long-context rates; growth beyond the reported request is billed at standard rates;
    • Sol across its repricing, logged under the gpt-5.6 alias and reported under that name;
    • unknown model → no cost, Spark → $0;
    • interleaved and restarted counters; usage before the window stays out of it;
    • a rollout in both sessions/ and archived_sessions/ counted once, including the origin split, and still once after the live file grows between scans;
    • a stale copy loses to the full file; a continuation counts both files; sub-agents sharing the parent's session_id all count;
    • a half-written last line (Codex and Claude) is counted once complete, and a complete last line without a newline is counted;
    • a warm scan reuses the Codex cache; an incremental scan across a drop matches a full rescan.
  • An opt-in test, codex_scan_real_logs_when_asked, scans a real Codex home and writes per-day, per-model numbers (no paths) for comparison with the Mac.
  • Negative controls: see the comments below.

Checked against the Mac on real logs

The opt-in scan and the Mac's replica (scripts/codex_accounting_replica.py in cli-pulse-private at 26702809, which agreed with the Mac app on every day of the 1.56 acceptance run) were run back to back on one machine's Codex logs, in three 31-day windows. Only aggregates were compared:

window days day × model cells cells that differ
09-03 .. 10-03 18 19 0
09-01 .. 10-01 19 20 0
08-21 .. 09-20 26 28 0

The desktop's Codex cost for the 09-01 .. 10-01 window is the same, to the cent, as the Mac's 1.56 scanner gave for it on 10-01.

Under this PR's previous revision, which took every line numbered before a sub-agent's boundary for copied history, the last window would have lost more than a third of its input tokens: the migrated sub-agents' own work.

🤖 Generated with Claude Code

…s, and count each rollout's tokens once

Codex cost on the desktop was computed from each day's token totals with a
price table that predates most of the models people now run:

- gpt-5.5 used a gpt-5.4 placeholder ($2.50 / $15 per 1M), written before
  its prices were published. It is $5 / $30.
- gpt-6-astra and the gpt-5.6 Sol / Terra / Luna / Cyber models had no row,
  so they showed no cost at all.

The table now follows the one bundled with steipete/CodexBar at 25bba9b7
(MIT, notice in pricing.rs), which cites OpenAI's pricing pages. Two of its
rules only make sense per request, so Codex cost is now summed per request
while parsing (packed slot 3), the way Claude's already is:

- Long context: a request with more than 272K input tokens is billed at
  the model's long-context rates for the whole request. A busy day's total
  crosses 272K when no request does, so pricing the day would bill it all
  at the higher rate.
- Dated rates: Sol was repriced on 2026-08-21 and Terra / Luna on
  2026-07-30. A request is billed at the rate in force at its timestamp.

Aliases OpenAI routes to a priced model (gpt-5.6 -> Sol, gpt-reserve ->
Luna, the Daybreak names) resolve to it. gpt-5.5-codex follows gpt-5.5, as
every -codex row follows its base model; gpt-5.5-mini / -nano keep their
placeholder rates, since neither has a published price.

Two counting guards:

- The scan reads sessions/ and archived_sessions/ (and each WSL distro's)
  and told files apart only by path, so a rollout present in both would be
  counted twice. Files with the same rollout id (session_meta payload.id)
  whose token events overlap in time are now one rollout, and the most
  complete copy counts. The id alone is not enough: a rollout can be
  continued in a second file under the same id with no event in common, and
  both halves are real. Sub-agent rollouts have their own id and are not
  touched (their session_id is the parent's, which is why it is not used).
- A cumulative counter that dropped lowered the baseline, so when it jumped
  back up the gap was counted again. Each component now counts growth above
  the highest snapshot seen, or growth since the previous snapshot while
  below it. A flip-flopping counter cannot double; a restarted counter still
  counts (a pure high-water mark would drop everything after a restart, and
  a restart is the shape real logs show). The counter also advances on
  events outside the scan window, so a rollout that began before the window
  no longer lands its earlier usage on the window's first day.

The Codex cache carries a per-provider rules version, so it is rebuilt once
under the new rules while Claude's cache is kept.

Tests: pricing (every new row, the 272K boundary on both sides, each
repricing boundary to the millisecond, aliases, a table consistency check),
the counter (restart, interleaving, repeats, last-only events, resuming with
and without the saved peak, each compared against the rule it replaces),
copy detection (identical copies, stale copy, continuation, sub-agents), and
end-to-end scans for per-request pricing, dated rates, the window edge,
archive copies, warm-cache reuse and incremental parsing across a drop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings September 30, 2026 14:34

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

JasonYeYuhe and others added 2 commits September 30, 2026 23:38
Three pricing tests used 1M-token requests to make the arithmetic round.
1M input is a long-context request, so the code correctly billed it at the
long-context rate and the tests failed (gpt-5.5 read $14.50 instead of $8;
Sol $10 instead of $5). They now use 100K requests, and the Sol test also
checks today's long-context rate next to the earlier one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pro long-context tier

OpenAI released gpt-6-sol, gpt-6-luna and gpt-6.1-sol after the CodexBar
commit the table follows, and lists a long-context tier ($60 / $270 above
272K) for gpt-5.4-pro and gpt-5.5-pro that the upstream table does not
carry. All rates checked against developers.openai.com/api/docs/pricing
(2026-09-30); the macOS app's table has the same rows. gpt-6.1-sol's cached
rate is 5% of input rather than 10%, and has its own test case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JasonYeYuhe

Copy link
Copy Markdown
Collaborator Author

Negative controls

Each guard was broken on purpose on a throwaway branch taken from 13b494c, and CI ran with cargo test --all --no-fail-fast (ubuntu-latest; the other three Rust jobs failed the same way). Every mutation compiled clean under clippy -D warnings, so each failure below is a test catching the change. The branches have been deleted.

Branch 1 (six mutations; each has at least one failure only it explains):

mutation tests that failed
Codex cost from the day's total at today's rates (the old way) codex_long_context_tier_is_decided_per_request_not_per_day, codex_rates_follow_the_request_time
copies detected but not left out of the totals / origin split codex_rollout_in_both_sessions_and_archive_is_counted_once, codex_stale_copy_loses_to_the_full_rollout
old counter rule (baseline follows every drop) counter_interleaved_series_never_recount_the_gap, counter_restart_passing_the_old_peak_counts_from_the_peak, counter_resumes_from_saved_state_like_a_full_pass, codex_counter_jumping_between_two_series_is_not_recounted, codex_incremental_scan_matches_a_full_rescan_across_a_counter_drop
out-of-window events skipped before the counter sees them codex_usage_before_the_window_is_not_put_on_its_first_day
fresh caches not stamped with the rules version codex_cache_from_older_rules_is_rebuilt_but_claude_is_kept, codex_warm_scan_reuses_the_cache
dated rates ignored sol_is_billed_at_the_rate_of_the_request_day, repricing_boundaries_are_exact, codex_rates_follow_the_request_time

Branch 2 (four mutations):

mutation tests that failed
same rollout id is enough (overlap in time ignored) a_continuation_under_the_same_id_is_not_a_copy, codex_rollout_continued_in_a_second_file_counts_both_halves
rollouts grouped by session_id instead of their own id codex_sub_agent_rollouts_are_all_counted
pure high-water mark counter_restart_keeps_counting, counter_restart_passing_the_old_peak_counts_from_the_peak, codex_counter_restart_keeps_counting
peak not saved in the cache codex_incremental_scan_matches_a_full_rescan_across_a_counter_drop (the high-water mark alone gives the expected 1,100 here, so this failure is the missing peak)

The rows added after 13b494c (gpt-6-sol, gpt-6-luna, gpt-6.1-sol, the -pro tier) are covered by their own cases in new_models_are_priced and long_context_switches_the_whole_request_above_272k_input. The consistency check over every row also runs on them.

JasonYeYuhe and others added 2 commits October 1, 2026 00:29
…he Mac's own cases

Review of this PR found the desktop and the macOS app (1.56, same release)
would report different Codex numbers from the same logs. This makes them
agree, and makes a future disagreement fail a test.

- Counter: the baseline only rises, and a snapshot below it in any component
  is skipped (the Mac's rule 1, CodexBar's watermark). The restart-aware rule
  this PR had counted requests after a counter restart that the Mac and
  CodexBar do not; one rule for both apps wins over the more generous one.
- A file whose first event carries a counter from elsewhere (a fork, or a
  continuation file under the same rollout id) counts only its own request:
  total minus last becomes the baseline (the Mac's rule 3, all files). It
  counted the carried total a second time, and priced it as one request, so
  above 272K at long-context rates.
- The Mac's shared fixture (codex-accounting-cases.json, copied unchanged from
  cli-pulse-private 89156f66) runs in the Rust tests three ways: full, warm
  and incremental. 10 of 11 cases must match exactly; the sub-agent
  copied-history case is a listed difference (sub-agent rules are not ported)
  and is asserted both ways.
- Aliases (gpt-5.6, gpt-reserve, Daybreak) are resolved only to find the
  rates. The stored and uploaded model name stays the one Codex logged, as on
  the Mac; renaming it would have left the already-uploaded rows beside new
  ones and doubled those days on iPhone and web.
- The Codex cache records a digest of the price table (rows, dated rates,
  aliases, 272K threshold) and is rebuilt when it changes, so a table edit
  cannot leave cached requests at an old price or at $0.00.
- A last line without a newline is read only when it already parses;
  otherwise it waits, and the saved offset stays at its first byte (Codex and
  Claude). The line used to be skipped for good.
- Test: a copy in archived_sessions is still left out after the live file
  grows and is parsed incrementally.
- The "copies" log line is at debug level; it repeated on every scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… call files copies only when nested

The Mac's accounting PR (cli-pulse-private#625) moved on while this was in
review (4c3dc377), and the shared fixture moved with it: 14 cases now. Two of
its changes are the Mac's answer to shapes the desktop also gets wrong, so
the desktop follows them and stays in step.

- Copied history. A sub-agent or fork rollout can begin with its parent's
  history, token events included, stamped when Codex copies them, so after
  the file's own session_meta. What marks them is the line number: the
  session_meta names the first line of the file's own history
  (subagent_history_start_ordinal) and every line carries its ordinal. A
  child's event numbered before that start, or stamped before its own
  session_meta, is the parent's: not counted, and it does not move the
  counter or the file's event span. Only the file's first session_meta is
  its own. The child marker is kept in the cache, so an incremental parse
  applies the rule too. Real sub-agent logs have this shape, so the desktop
  was counting the parent's copied usage a second time.
- A child's first event without last_token_usage is taken as carried in
  full, as on the Mac.
- Copies. Two files of one rollout are copies only when one's event span lies
  within the other's (a copy is the whole file or an earlier state of it).
  Any overlap at all dropped a file whose other events appear nowhere else.
  Most-complete-first now breaks ties on final input + output, as the Mac.
- Fixture re-copied from 4c3dc377; the desktop now matches all 14 cases, so
  the list of known differences is empty.
- Token sums saturate instead of overflowing on a corrupt count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JasonYeYuhe

Copy link
Copy Markdown
Collaborator Author

Review follow-up

Two commits on top of c483a7b: ff0dd58 and b473c9d. Still on hold until 1.55 is cut. CI run 36738782661 on b473c9d: 13/13 green.

Counting rule (the desktop and the Mac disagreed). Both apps now use one rule, the Mac's. The baseline only rises, and a snapshot below it is skipped. The restart-aware rule this PR had is gone. A restart inside a file is now undercounted the same way on the Mac and in CodexBar. The shared case counter_restart_counts_only_above_the_old_high pins that on both sides, so changing it later has to be one decision for both apps. I picked the Mac's rule, not the desktop's, so that the desktop, cli-pulse-private#625 and CodexBar agree without reopening a second PR.

Carried counters. On a file's first event with total > last, total − last becomes the baseline, for every file (the Mac's rule 3). A child's first total without last is treated as carried, as on the Mac.

The Mac's shared fixture runs in the Rust tests. codex-accounting-cases.json is copied unchanged, and every case runs as a full scan, a warm scan and an incremental scan. While this was in review, #625 moved to 4c3dc377 and its fixture grew to 14 cases. That commit finds a sub-agent's copied history by line number and treats two files as copies only when one's span is nested in the other's. b473c9d ports both:

  • Real sub-agent logs start with the parent's history copied in, token events included, and stamped after the file's own session_meta. The desktop counted that history a second time for every sub-agent. The fix is the Mac's rule 2 (subagent_history_start_ordinal / ordinal, plus the time rule). It is kept in the cache so an incremental parse applies it too.
  • Copies: a partial overlap no longer drops a file whose other events appear nowhere else.

The desktop now matches all 14 cases. The test's list of deliberate differences is empty; any entry added to it has to be asserted both ways.

Aliases. Aliases are resolved only to look up the price. The stored and uploaded model name stays the one Codex logged, as on the Mac, so rows already uploaded under that name are not left beside a renamed one. The integration assertion is flipped.

Price table in the cache check. The Codex cache records a digest of the rows, the dated rates, the aliases and the 272K threshold. It is rebuilt whenever that digest changes, so no version bump has to be remembered. Tests check that the digest changes with each of those parts, and that a cache priced with another table, or with none, is rebuilt.

Half-written last line. A last line without a newline is read only if it already parses. Otherwise it waits, and the saved offset stays at the line's first byte. This covers Codex and Claude, which share the loop. Tests cover both, plus a complete last line that has no newline.

Long context from last_token_usage: not done. The Mac and CodexBar decide the tier from the tokens an event counts. After the fixes above, that amount is one request's usage or less in every shape the shared cases cover, and deciding it from last would make the desktop price differently from the Mac. This is listed in the PR body.

Copy detection after a resume. A new integration test grows the live file between two warm scans while a copy sits in archived_sessions/. The fixture's incremental mode also covers byte_identical_copy and archived_prefix_of_live_file.

Differences written down. The PR body now has a "Known differences from the Mac" section: unknown models, date-folder selection, unparseable timestamps, last-only events, and the long-context input. The "copies" log line is now at debug level.

Negative controls

Throwaway branches taken from these commits ran cargo test --all --no-fail-fast on ubuntu. Every mutation is caught. The branches have been deleted.

mutation failed
restart-aware counter (the rule this PR had) shared case counter_restart_… (all three scans); counter_restart_counts_only_above_the_old_high, counter_skips_a_snapshot_below_the_baseline_in_any_component, codex_counter_restart_counts_only_above_the_old_high
baseline follows every drop (pre-1.56) shared cases counter_goes_down_and_back, counter_restart_…; interleaved / down-and-back / resume unit tests; codex_counter_jumping_between_two_series_is_not_recounted, codex_incremental_scan_matches_a_full_rescan_across_a_counter_drop
carried counter counted shared cases fork_with_inherited_counter, continuation_that_carries_its_counter_over; counter_first_event_carrying_a_counter_counts_only_its_own_request
child first total without last counted counter_child_first_total_without_last_is_taken_as_carried
ordinal rule off shared case copied_history_marked_by_ordinal (all three scans); copied_history_is_told_by_line_number_or_by_time
child marker not restored on resume shared cases copied_history_before_meta, copied_history_marked_by_ordinal (incremental scan only); only_the_first_session_meta_is_the_files_own
any overlap makes a copy shared case partial_overlap_both_count (all three scans); a_partial_overlap_is_not_a_copy
resume forgets rollout id / span / count shared case byte_identical_copy (incremental scan); codex_copy_is_still_recognised_after_the_live_file_grows
alias renames the stored model codex_aliases_keep_the_logged_name, codex_rates_follow_the_request_time
cache check ignores prices codex_cache_priced_with_another_table_is_rebuilt
aliases left out of the digest pricing_fingerprint_changes_with_every_part_of_the_table
half-written last line consumed codex_half_written_last_line_is_counted_once_it_is_complete, claude_half_written_last_line_is_counted_once_it_is_complete

When #625 changes its fixture again, copy it into src-tauri/tests/fixtures/ so both apps keep running the same cases.

JasonYeYuhe and others added 3 commits October 3, 2026 20:22
…own request decides its tier

The Mac's accounting (cli-pulse-private#625) changed twice more before it
merged (f18939f2), and the shared fixture grew from 14 cases to 20. The
desktop follows, so the two count the same logs the same way again.

- Copied history has two shapes. Codex's migration of older sub-agent
  rollouts moves subagent_history_start_ordinal to the end of the file and
  drops the copied session_meta lines, so every line, the sub-agent's own
  work included, is numbered before the boundary. The previous revision took
  all of it for the parent's copy. Now the lines before the boundary are
  copied only once an ancestor's session_meta numbered before it shows they
  are. Without one, the events before the parent's first inter-agent message
  are its replayed last requests and are dropped; the rest counts. Until a
  marker comes, those events are held; if none comes before a line past the
  boundary or the end of the parse, they count, in order. The answer is kept
  in the cache (CodexCopiedPrefix), so an incremental parse goes on applying
  it.
- Only a file's first line can be its own session_meta. When it is something
  else the identity stays unknown, so a parent's copied session_meta is never
  taken for the file's own.
- Lines over 32 KB are not decoded; their first 4 KB is enough to recognise
  a copied session_meta or an inter-agent message, and the line number is
  read from the first 512 bytes. As on the Mac.
- Cost: the request an event reports (last_token_usage) decides the 272K
  tier, not the tokens counted. Growth of the counter beyond that request
  (a request that wrote no event of its own) is billed at standard rates.
  The Mac's codexEventCostUSD, field by field.
- A dated spelling of an alias (gpt-5.6-2026-08-01) is billed as its model,
  and the price fingerprint records how a fixed list of names resolves, as
  the Mac's does.
- Codex rules version 2 (1 was this PR's earlier revision, never released).
- The fixture is the Mac's file byte for byte, pinned by commit and SHA-256;
  scripts/sync-codex-accounting-cases.sh copies a new one.

On this Mac's own logs (aggregates only), the desktop's real-log scan and the
Mac's replica agree on every day and model in three 31-day windows: 09-03..
10-03, 09-01..10-01 and 08-21..09-20. Under the previous revision the last
window lost 146M input tokens of migrated sub-agents' own work.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… copy

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…is not a date, not a panic

parse_day_key_local fell back to the first ten bytes of a timestamp it could
not parse, by slicing the string. A byte 10 inside a multi-byte character
panicked the scan. It now takes the prefix only when it is whole.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JasonYeYuhe

Copy link
Copy Markdown
Collaborator Author

Brought up to the Mac's final rules (#625 as merged, f18939f2)

Three commits on top of b473c9d: a013032, dcc7773 and e711015. The shared fixture is now the Mac's file byte for byte (20 cases), and the desktop matches all 20 in full, warm and incremental scans. The list of deliberate differences stays empty.

Self-review against the Mac code

Each rule was read side by side with CodexTokenAccounting.swift, parseCodexFile/scanJsonl in CostUsageScanner.swift and CodexPricingTable.swift at 26702809.

Mac desktop same?
observeSessionMeta from the first line only (≤ 1 MB); observeUnreadableFirstLine otherwise first line only, same limit yes, except the desktop also reads a non-session_meta first line as an ordinary line (listed in the body)
later session_meta → observeCopiedSessionMeta(codexLineOrdinal), never decoded same, line number from the first 512 bytes yes
inter_agent_communication_metadata → observeInterAgentMessage, only while awaiting a marker same yes
lines over 32 KB: 4 KB head for markers, never decoded same yes
receive: hold events before the boundary; at a line past it, held ones first; an event without a number is judged at once CodexResume::receive yes
finish at the end of every read, noMarker only if something was held same yes
count: ordinal rule only after ancestorMetadata; time rule for any child is_copied_history yes
rules 1 and 3 CodexCounter (unchanged) yes, except the last-only baseline (listed)
CodexCopiedPrefix persisted in the file state CodexChildMeta.copied_prefix yes
codexEventCostUSD: request decides the tier; growth beyond it at standard rates, field by field codex_event_cost_usd yes
aliasTarget with a dated spelling codex_price_key yes (fixed here)
fingerprint covers name resolution CODEX_FINGERPRINT_MODEL_NAMES yes (added here)
rate rows, dated rates, aliases table all 29 shared rows equal; 3 desktop-only rows as before

Found and fixed while reviewing: the dated alias spelling was unpriced on the desktop, and parse_day_key_local sliced a timestamp at byte 10, which panics inside a multi-byte character (e711015).

An independent review (Gemini 3.1 Pro, read-only, given the diff, the desktop files and the Mac files): no findings, APPROVE.

Negative controls

Throwaway branches with several mutations each ran cargo test --all --no-fail-fast on ubuntu (runs 37119515801, 37119516848, 37119518036), and locally with the same result. Mutations that shared a failing test were also run one at a time locally, so each has a test that fails for it alone. All caught; the branches are deleted.

mutation caught by
ordinal marks a copy without an ancestor's session_meta (the previous revision) copied_history_is_told_by_line_number_once_marked_or_by_time; shared cases migrated_subagent_counts_its_own_work, …_skips_the_replayed_parent_tail, …_without_an_inter_agent_message_counts_in_full (all scans)
inter-agent message not recognised shared case migrated_subagent_skips_the_replayed_parent_tail (all scans)
the event past the boundary counted before the held ones held_events_count_in_order_when_no_marker_comes
a marker's line number never read the_line_number_is_read_without_decoding_the_line
decided copied part not kept for the next parse the_copied_part_decided_in_one_parse_holds_in_the_next; shared case copied_history_marked_by_ordinal (incremental)
held events dropped at the end of a parse shared case migrated_subagent_without_an_inter_agent_message_counts_in_full (all scans)
first session_meta anywhere is the file's own only_the_first_line_can_be_the_files_own_session_meta
long lines skipped before their head is looked at a_copied_session_meta_too_long_to_decode_still_marks_the_copy
tier decided by the counted tokens codex_long_context_tier_is_decided_by_the_request_the_event_reports
growth beyond the request takes the request's tier an_events_own_request_decides_its_long_context_tier, codex_growth_beyond_the_request_is_billed_at_standard_rates
dated alias not resolved a_dated_alias_is_billed_as_its_model_and_keeps_its_name
fingerprint without name resolution pricing_fingerprint_records_how_names_resolve
fixture edited in place the_shared_cases_are_the_macs_file_unchanged; shared case migrated_subagent_counts_its_own_work

Real logs

See "Checked against the Mac on real logs" in the body: three windows, every day and model equal.

🤖 Generated with Claude Code

@JasonYeYuhe
JasonYeYuhe merged commit 50b2d14 into main Oct 3, 2026
13 checks passed
@JasonYeYuhe
JasonYeYuhe deleted the p156-desktop-pricing branch October 3, 2026 11:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants