Skip to content

feat(server): stats request limits, session override, cost guard and the demand scan (rows-first p36a) - #1033

Closed
paddymul wants to merge 35 commits into
adr-003-stats-tiers-and-size-policyfrom
feat/rowsfirst-p36a-limits-override-cost-guard
Closed

paddymul wants to merge 35 commits into
adr-003-stats-tiers-and-size-policyfrom
feat/rowsfirst-p36a-limits-override-cost-guard

Conversation

@paddymul

@paddymul paddymul commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #1031 (feat/rowsfirst-p34-scalar-tier-units), which carries #1021, #1022, #1024, #1026, #1028, #1029 and the stats policy of #1019 and #1023. This PR is based on main so the repo's Checks workflow runs on it (its pull_request trigger uses branches: "*", which does not match a base branch containing a slash). The diff therefore includes the commits of those PRs until they merge. The commits of this phase are d4477c46 (failing tests) and 9af4042c (implementation); read only those (git diff origin/feat/rowsfirst-p34-scalar-tier-units...HEAD).

Problem

A client that advertised stats_ondemand can be told a session has no stats (not_computed) and what it could ask for (requestable), but it cannot ask. A stats_request carries no tier, so the only runs a server can start are the ones the policy target names, and a request for anything else is not_requestable. Three things the plan puts on the server are missing:

  • The ceiling and requestable are advertised but not enforced on a request, so a client that sends a tier the policy did not offer has nothing stopping it.
  • Nothing remembers that a user asked for more than the policy gave. A forced tier would be lost at the next dataflow-field change, and a run that was too slow to continue automatically would be restarted by the next stats_gen bump or /reload_expr.
  • A color map reads histogram_bins of its val_column. When the policy computes no stats, those columns get no bins, and nothing finds out which columns they are.

Phase and plan references

Rows-first p36a, from buckaroo2-reports/plans/: plan 3 (03-no-summary-stats-for-large-files.md) section 6 "Phase 6a", with section 3.2 (the cost guard, elapsed_ms), section 3.3 (the demand-driven minimum) and section 3.4 (the control's server side: force, the ceiling answer, the session override). Plan 1 sections 3 and 4.0, plan 2 sections 3 and 4.1.

Approach

  • Request fields. For a client that advertised both bits, on a session that has a policy, stats_request reads tier (scalar or full), force (a boolean) and, together with a tier, columns (the grid's a, b, c names) that scope the run to those columns. Without a tier, columns is only the ordering hint it was. A value the server cannot read is stats_aborted with reason bad_request. Every other client, and a session with no policy, ignores all three fields.
  • Judging a request (_gate_request), in this order. The ceiling: resolve_stats_policy is called with the requested tier as the host's, over the cells asked for (all of them, or rows times the named columns), so the ceiling is the same function load uses; a tier it lowers is answered stats_update {final: true, status: "not_computed", reason: "ceiling"} with no payload and nothing run. Then, for a whole-table request, the tier must be the target or above it with force (below the target, or above it without force, is not_requestable, as is a schema target with no tier). Then the cost pause: a request without force on a paused session is answered with reason cost. A column-scoped request needs no force, since it is how the demand columns are asked for.
  • The override. A force above the target stores stats_override and the generation it was made in on the session. effective_stats_policy(session) is the policy in force: the one resolved at load, or the one the override tier resolves to for the same entry (never lower than the target, still under the ceiling, thresholds read again). It applies to an unfiltered state and to the generation it was forced in. "Filtered" is a quick command (a search) on the operations, which is the difference between the dataflow's filt and clean chains; a cleaning method or a post-processor does not make a state filtered. begin_stats_generation, stats_meta and serves_scalar_tier read the effective policy, so a forced full makes the session pending for full again after a dataflow-field change, and a client that connects later is told so.
  • The cost guard. run_units can report how long each unit took (from the run's own timer). An automatic request, which is a request without force, that ran a unit over STATS_COST_BUDGET_S and whose reply leaves units to run sets cost_paused and the status not_computed for cost. A unit cannot be cut short, so the unit that went over has run; the guard stops the next one. A request that finishes the run pauses nothing, and a forced request never pauses. A force clears the pause, and the run continues where it stopped (ran is kept). The pause is on the session, so begin_stats_generation reads it (initial_stats_status(..., cost_paused)) and it survives a stats_gen bump and /reload_expr. A client without both bits is not held by it: it is told pending and gets the stats when it pulls, as before.
  • When the controls are forgotten. reset_stats_controls clears the override and the pause on /load, /load_compare, a /load_expr of another expression, and a /load_expr or /reload_expr whose body names a stats_tier. A rebuild of the same expression, or a reload that names none, keeps them.
  • The demand scan. stats_policy.demand_columns(display_args, pairs) collects the val_column of every color_map_config whose color_rule is color_map, in any display, and returns the rewritten names in the table's order. It reads the built config, so it covers rules in the column overrides and rules a klass adds at style time, and it sends no query. A val_column may be the rewritten name or an original one (a rewritten name wins), and one that names no column is dropped. A categorical rule, a tooltip and a color_map with no val_column create no demand. refresh_demand_columns writes the result to session.stats_policy["demand_columns"] at load and with every snapshot refresh (a dataflow-field change, /reload_expr), for a schema target only, since a higher target computes those columns anyway; df_meta.stats reports it to an ondemand client through the existing _policy_fields.
  • Scoped runs. A column-scoped request is served by _serve_run from start_stat_run(tier=..., columns=...), the run p34 added: its key is (stats_gen, scope, tier, columns), two requests for the same columns in any order share it, it never assigns, and a client that asks later reads its fragments with no query. Its final reply says what the session still is (status, and reason when there is one). The scalar run of p34 is served by the same function and now says the same on its final reply.

What changes

  • buckaroo/server/session.py: stats_override, stats_override_gen and cost_paused on SessionState; initial_stats_status(..., cost_paused); effective_stats_policy, reset_stats_controls, restore_stats_status; begin_stats_generation and stats_meta read the effective policy; session_dataflow moved here from stats_wire (still importable from there).
  • buckaroo/server/stats_wire.py: STATS_COST_BUDGET_S, REQUEST_TIERS, TierRequest, _tier_request, _gate_request, _apply_force, _pause_if_slow, _refused; _serve_scalar is now _serve_run (tier and columns); run_units(timings=); refresh_demand_columns, called from refresh_session_snapshot; serves_scalar_tier reads the effective policy.
  • buckaroo/server/stats_policy.py: demand_columns; the module docstring no longer says nothing imports it.
  • buckaroo/server/handlers.py: the resets above, same_expression at /load_expr, the demand scan after the snapshot is stored.
  • No change to websocket_handler.py, dataflow.py or any client.

Tests

Two commits: d4477c46 (failing tests, pushed alone) and 9af4042c (implementation and guard tests). Tests are in the existing files of their kind.

  • test_stats_policy.py (pure, every Python and Windows): TestRequestsOverTheCeiling, TestMalformedTierRequests, TestRequestableTiers, TestCostPausedStatus, TestEffectiveStatsPolicy, TestDemandColumns. They cover the ceiling answer for six combinations of host tier, limits and request, each unreadable field, which tiers a whole-table request may name, the status a paused session starts a generation with and the order of the checks, what a forced tier raises the policy to and where it applies, and the demand scan on every shape of config.
  • test_load_expr.py (xorq, over HTTP and WebSocket): TestForcedRequestsWire, TestCostGuardWire, TestCostGuardUnits, TestDemandScanWire, TestTierFieldsAndOlderClients. A forced scalar and a forced full run for real; the ceiling answer runs no query; the override across a field change, a search, a reload and a new expression; the pause across a generation, a reconnect and a reload; a unit under the budget, a forced unit and a request that finishes the run do not pause; elapsed_ms on updates and refusals; the demand columns of a klass at style time, of an override, and none for categorical rules, with zero stat queries at load; a demand request runs the named columns only, assigns nothing and is replayed to a second client with no query; the ceiling judged per named column.
  • Added with the implementation because they pass on the earlier code: nine tests that pin what does not change (a stats_update-only client has the fields ignored and a host-named schema target still refuses it, a session with no policy ignores them, an explicit full session serves plain requests as before, a client without both bits is not held by a pause, a session that computes the columns anyway has no demand, categorical rules alone have none, a session nobody forced stays not_computed through a field change, a scoped request on a complete session is answered from the dataflow).

The elapsed_ms test has bounds of 50 ms in d4477c46: a request that ran a 50 ms query takes at least that long, and a refusal takes less. 9af4042c widens both sides (a 300 ms query, bounds of 250 ms), because an upper bound of 50 ms on a refusal can fail on a loaded machine (the load average here reached 35 during the full suite). The test fails on the earlier code for the same reason in both versions: the scoped request is refused with not_requestable.

On d4477c46 all nine Python / Test jobs failed (3.11 to 3.14, Max Versions 3.11 to 3.14, Windows), with all 28 checks completed; the other 18 succeeded (Docs, Lint, Typecheck, the JS job, the wheel build and the Playwright jobs among them) and deploy was skipped. CI logs were not read (the REST log API is rate-limited for this account), so the reasons rest on the local run of the same commit: 107 of the new tests fail, each on a missing session attribute or function, on a reply with no tier, final or elapsed_ms where a stats_update is expected, or on a request refused with not_requestable or stats_aborted where an answer is expected, and the other 1940 unit tests pass.

CI on 9af4042c: all 28 checks completed, 27 succeeded (one of them the Read the Docs status) and deploy was skipped. That includes all nine Python / Test jobs, Python / Lint, Python / Typecheck, Docs / Build + Check Links, the JS job, the wheel build and the Playwright jobs.

Local checks on 9af4042c:

  • The full unit suite (pytest ./tests/unit -m "not slow") gives 2056 passed and 5 skipped: the 1940 of feat(stats): xorq scalar tier units, served as fragments and never assigned (rows-first p34) #1031, the 107 failing tests, and the 9 guard tests. The same 2056 and 5 in an environment resolved the way the Max Versions job does it (pandas 3.0.6, polars 1.44.2, xorq 0.4.5, numpy 2.5.3), in a copy of the tree with the lockfile removed and uv sync --resolution=highest. That copy left out packages/, so one test that reads a committed JS fixture (test_xorq_window_fixture.py) failed there until the fixture file was copied in; it then passed.
  • Forty-five deliberate regressions of the implementation (each check of _gate_request removed, the pause never set, set on a final reply, set on a forced request, never cleared by force, the override not recorded, ignoring the filtered scope or lasting only one generation, the reset removed from each of its triggers, the demand scan not refreshed, not limited to a schema target, including categorical rules, letting an original name win, the fields applying to every client or to a session with no policy, a scoped request served as the whole run, and others). On the first set of tests 40 were killed and five survived. One survivor was an equivalent mutation of a redundant check, which is gone now (simplifying it also exposed that an unhashable val_column raised TypeError in the scan, which is fixed and tested). The other four led to two new tests (a paused schema target has nothing to request; /load and /load_compare forget the controls) and two tighter ones (a forced unit is checked right after its reply, not after a later forced request has cleared the pause; a session with no policy ignores unreadable fields too). With those, all the others are killed.
  • Real server processes (python -m buckaroo.server --no-browser, started and stopped by me, killed by listener PID, ports free afterwards), on a 60,000-row, three-column xorq build with a deferred auto session. On port 8941 with BUCKAROO_STATS_FULL_AUTO_ROWS=1000, BUCKAROO_STATS_SCALAR_AUTO_CELLS=100000 and BUCKAROO_STATS_CEILING_FULL_ROWS=50000 and a color_map override naming price: the first frame of a client with both bits was not_computed, reason size, tier_target schema, requestable ["scalar"] and demand_columns ["a"]. A request with no tier was not_requestable; the demand request (tier: "scalar", columns: ["a"]) was a stats_update with the scalar rows of column a only in 16.3 ms; a bad tier was bad_request; scalar without force was not_requestable; full with force and a scoped full over three columns were refused with reason ceiling in 0.0 and 0.1 ms; scalar with force ran for all three columns (24.3 ms) and a client that connected afterwards was told tier_target: "scalar". A stats_update-only client was told pending and a no-caps client got the complete stats. On port 8942 with the cost budget set to 0 and one unit per request by a small launcher script (a 10 s unit cannot be produced on this table): the first automatic request returned a partial update and paused the session; a client that connected then was told not_computed, reason cost, tier_target full, while a stats_update-only client was told pending; the next automatic request was refused with reason cost; a force request continued the run where it stopped and the last one was final.

Why default behaviour is unchanged

  • Every new request behaviour is reached only through tier_fields_apply: a client that sent both stats_update and stats_ondemand, on a session that has a policy. No released client sends stats_ondemand, and a session has a policy only when it is deferred or a host named scalar or schema.
  • A request without the new fields takes the path it took before, unless the session is paused or the thresholds in the environment have changed since load (the ceiling is judged again at request time). The cost guard acts only for the clients the fields apply to, and nothing sets stats_override or cost_paused without them, so begin_stats_generation, stats_meta and serves_scalar_tier return what they returned.
  • df_meta.stats changes only for an ondemand client of a schema target whose config has a color_map rule (demand_columns), of a forced session (the effective target) and of a paused one.
  • One status rule changed: a session with a policy whose target is full is pending whatever its delivery. That case could not occur before this PR, since a policy exists only for a dataflow built at the schema tier and a host naming full there is the explicit-full case feat(server): stats policy wire contract and handler plumbing (rows-first p33) #1029 pins.
  • The existing tests pass unchanged. The only edit to an existing test file's text is the module docstring of test_stats_policy.py.

Deviations from the plan

  • STATS_COST_BUDGET_S is 10 s, a module constant and not an environment setting. The plan gives no number; this is the tallyman client's base timeout and is marked provisional beside the constant with the measurements it can be compared with.
  • The guard measures the slowest unit of a request and acts after the reply, not before the unit: units cannot be cancelled (plan 3 section 3.4), so the first unit over the budget has already run and the next one is what the pause prevents. A request that leaves nothing to run, and any forced request, pause nothing. A pause is cleared by one force, and a later automatic request can pause again.
  • "The unfiltered scope" (plan 3 question 3) is read as: no quick command on the operations. The plan leaves the answer open ("proposed: yes for the unfiltered scope only"); a cleaning method is on the clean chain, which survives a filter flip, so it does not end the override. A force made in a filtered state serves that state and does not reach another filtered one.
  • A whole-table force above the target sets the override; a force that only continues the target (a "Continue" after a pause) sets nothing. A force with no tier on a schema target is not_requestable and clears nothing.
  • The ceiling is judged at request time with the thresholds read again, over the cells of the request, with the count load took (estimate). Named columns are judged by rows times their number, so a ceiling that refuses the whole table can still allow a few columns.
  • A refusal is a stats_update with tier set to the tier asked for (the target when none was named), status: "not_computed" always, and no payload or df_display_args.
  • The final reply of a run that does not assign (the scalar run of p34 included) carries the session's status and reason. p34's scalar final reply carried neither; no existing test pinned that.
  • df_meta.stats reports tier_target, estimate and the other policy fields for an explicit full session while it is paused, which it does not otherwise (feat(server): stats policy wire contract and handler plumbing (rows-first p33) #1029). A client that reads reason: "cost" needs to know which tier a force continues.
  • The demand scan runs at load and with every snapshot refresh, and writes demand_columns to the session policy (the p33 interface note). It is empty unless the target is schema. Nothing on the server requests the demand columns: they are a scoped stats_request, which the client sends (phase 5), under the same ceiling and pause.
  • reset_stats_controls and its triggers are mine; the plan says only that the override and the pause survive a generation and a reload.
  • session_dataflow moved from stats_wire to session, because the effective policy needs it and stats_wire imports session.

Not in this PR

  • Client rendering of the control, demand_columns and the cost state (phase 5); sort and search guards (phase 7); the config upgrade after stats complete (phase 6b); default flips (phase 9).
  • Calibrating STATS_COST_BUDGET_S.
  • A bound on the scoped runs a session holds. Each distinct column group keeps a StatRun (an accumulator and its fragments, no data) until the generation changes; a client that asks for many different groups grows that dict.
  • Carrying filtered stats on the wire (unchanged from p34: a filtered state's stats_update holds schema rows only).

🤖 Generated with Claude Code

paddymul and others added 30 commits October 3, 2026 22:51
…the _handle_widget_change split (rows-first s1)

A default-tier XorqBuckarooWidget and BuckarooWidget publish df_data_dict,
then df_display_args, then the rest of the widget_args_tuple observers on a
search change, and merged_sd carries the full stat set. These pass on main
and pin the behaviour the split must keep.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…(rows-first p31)

New tests/unit/server/test_stats_policy.py for buckaroo/server/stats_policy.py,
which does not exist yet:

- resolve_stats_policy: a table of (backend, source_kind, rows, cols, host_tier)
  against (tier_target, auto_request, requestable, reason), including the
  boundaries of each threshold and the three tallyman entries over 70 s.
- The ceiling holds for host_tier="full", for a reload that re-resolves
  against a larger row count, and for a force request, and every tier a result
  lists as requestable is one a force would be granted.
- Probes add no data query: the parquet footer probe reads a small share of a
  counting file object and decodes no row group, the dtype probe never collects
  a LazyFrame or executes a xorq expression, and a known xorq count is an input.
- route_polars_entry returns "xorq" above R and "eager" at or below it.
- Threshold overrides from BUCKAROO_* environment variables.

Nothing imports the module yet. The tests error at fixture setup on the missing
module until the implementation lands.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… p31)

New buckaroo/server/stats_policy.py, pure and imported by nothing yet:

- resolve_stats_policy(backend, source_kind, rows, cols, bytes, host_tier,
  limits) returns tier_target, auto_request, requestable, reason and estimate
  over the tiers schema < scalar < full. A ceiling is computed inside the
  function, so every caller gets the lower of the requested and ceiling tiers
  with reason "ceiling". A host tier lowers freely and raises only to the
  ceiling. Eager pandas and polars resolve to full.
- route_polars_entry(rows, cols) returns "xorq" above R rows and "eager" at or
  below it.
- probe_dtypes reads a schema without collecting or executing anything, and
  probe_parquet_rows reads a parquet footer's num_rows without decoding data.
  A xorq count is an input to the policy, never computed by it.
- The thresholds are provisional constants gathered in StatsLimits, each with
  a BUCKAROO_* environment override read on every call.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…-handler fields (rows-first s1)

A schema-tier XorqServerDataflow should match full stats on pinned_rows,
data_key, summary_stats_key and (except the stats-derived minWidth)
column_config, issue no data query besides the cached count, and keep
init_sd hints and sorted windows working. A pending state must write
nothing under a full-tier cache key, and a later full assignment must reach
merged_sd for the raw, clean and filt scopes. assemble_merged_sd must equal
merged_sd, and _handle_widget_change must be built from separately callable
all_stats and display-args builders. /load_expr and /reload_expr accept
stats_tier and stats_delivery, replay them on reload, and keep them out of
the warm short-circuit's has_config tuple.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… s1)

Add a dataflow-level stats_tier ("full" default, "schema"). The xorq schema
tier builds identity and typing for every column from the expression's
schema, with no data query beyond the cached row count, so a dataflow
constructs in milliseconds rather than the stats' hundreds.

The tier is part of _scope_cache_key, so a schema entry is never read as a
full one, and _populate_sd_cache stores summary_sd under the filt key only
if it was computed for the current frame, klass list and tier. add_analysis
no longer builds DFStatsClass outside the hook when the tier is not full.
The merged_sd observer body is extracted as the pure assemble_merged_sd,
and _handle_widget_change is split into _build_df_data_dict and
_build_df_display_args.

/load_expr and /reload_expr accept stats_tier and stats_delivery, stored on
the session beside dataflow_kwargs and replayed on reload. They stay out of
the has_config tuple; the warm short-circuit compares the stored pair.
stats_delivery="deferred" builds the schema-tier dataflow and publishes it.
Defaults (full, inline) leave behaviour unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…esolve_stats_policy (rows-first p31)

bytes and source_kind were documented as validated but are not: bytes=-1,
'lots', 1.5 and object() are accepted, a np.int64 is echoed unchanged so
json.dumps of the result fails, a host tier passed positionally lands in the
bytes slot, and source_kind=None or 5 is accepted.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ation test (rows-first s1)

The Max Versions jobs resolve pandas 3, which reports a string column's
dtype as 'str' where pandas 2 says 'object'. The characterization test
asserted 'object'. Verified in a Max Versions environment (pandas 3.0.6,
polars 1.44.2, xorq 0.4.5): the unit suite passes.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…rows-first p31)

bytes goes through the same integer check as rows and cols, so a negative,
float or non-numeric value raises, a numpy integer is echoed as a plain int
and the result stays JSON-serialisable, and a host tier passed positionally
into the bytes slot raises instead of being echoed. source_kind must be a
str; its vocabulary stays open for the callers that will read it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…g on a skipped column (rows-first s1)

A column in skip_stat_columns gets only name, dtype and length from the
full-tier pipeline, so its _type comes from init_sd. The schema tier layers
the schema-derived _type and is_* keys over it, so an int64 column that
init_sd types as float merges as integer and renders with zero fraction
digits instead of the float displayer init_sd asked for.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ows-first s1)

_get_schema_sd never read skip_stat_columns, so a skipped column's
schema-derived _type and is_* keys overrode init_sd's _type once merged. The
full tier gives a skipped column only name, dtype and length. The schema tier
now does the same, so init_sd's _type decides the displayer at both tiers.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…er (rows-first s2)

ServerDataflow and PolarsServerDataflow with stats_tier="schema" should publish
the display state full stats give (column_config without stats-derived keys,
pinned_rows including a host-supplied one, data_key, summary_stats_key) with no
stat computed on the data, still apply init_sd, serve sorted windows, take a
later full assignment into merged_sd for every scope, and assemble to the same
sd as merged_sd. Both backends run through the same parametrized class, plus a
pandas test that pins how an object column is typed from its dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
ServerDataflow and PolarsServerDataflow now implement the _get_schema_sd hook,
so stats_tier="schema" builds a dataflow that types every column from its dtype
and runs no stat on the data. schema_sd (stat_pipeline.py) builds the sd the way
process_df shapes it: an empty frame gives {}, a skipped column keeps only its
names. pandas applies the existing typing_stats to a zero-row slice and derives
_type through the _type stat; polars factors pl_dtype_typing out of
pl_typing_stats and feeds it the dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…s (rows-first p31b)

The boundary tables now run on the module defaults and are written against
the phase-0 proposals: full auto up to 12M rows and 520M cells, scalar auto
up to 1.0B cells, full refused above 25M rows or 1.0B cells (force
included), a scalar ceiling of 4.0B cells, and polars routed to xorq above
8M rows. Each threshold has an equal, a one-below and a one-above row, in
rows and in cells. New tests pin each default to its literal value and
cover the two new environment overrides, BUCKAROO_STATS_FULL_AUTO_CELLS and
BUCKAROO_STATS_CEILING_FULL_CELLS.

They fail on the current constants (10M rows, 500M cells, 50M-row ceiling,
no scalar ceiling, R of 10M) and on the missing full_auto_cells and
ceiling_full_cells fields.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tats_request (rows-first s3)

On a deferred /load_expr session a stats_request {stats_gen, scope} should
return a stats_update with the matching stats_gen whose inline wide payload
equals the all_stats an inline session sends, a stale stats_gen should get
stats_aborted and run no query, and /load_expr and /reload_expr should bump the
generation. df_meta.stats should be injected on every frame and survive a
dataflow-field change, which returns the session to the schema tier. With a
caps client and a legacy client on one session, the legacy client should keep
getting complete messages through the websocket broadcast, the /load_expr,
/reload_expr, /load and /load_compare pushes and the highlight overlay, while
the caps client gets a stats-free frame and then pulls a stats_update. A spy
telemetry sink should see a stats.request span.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…measurements (rows-first p31b)

Full auto now needs at most 12M rows and 520M cells (was 10M rows), scalar
auto goes to a 1.0B cell budget (was 500M), `full` is refused above 25M rows
or 1.0B cells including for a forced request (was 50M rows), scalar gets a
4.0B cell ceiling by default (was unset; an extrapolation with no
measurement behind it), and polars routes to xorq above 8M rows (was 10M;
the figure assumes pre_limit False). Each DEFAULT_* constant carries the
measurement it comes from.

The two new bounds are full_auto_cells and ceiling_full_cells, appended to
StatsLimits, with BUCKAROO_STATS_FULL_AUTO_CELLS and
BUCKAROO_STATS_CEILING_FULL_CELLS as overrides under the same parsing rules.
_size_tier and _ceiling_tier read them; signatures and the result shape are
unchanged.

Three existing tests that hard-coded the old boundaries are updated: the
numpy-integer case (11M rows is now full), the reload case (a 40M-row entry
is now refused full) and the empty-environment case for the scalar ceiling
(an empty value now means the 4.0B default, not no ceiling).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…eration edge cases (rows-first s3)

Four more cases for the stats wire format, kept in their own commit so each is
seen failing on CI before the implementation lands. Returning to a state whose
stats were completed once is answered from summary_stats_cache with no query.
Completing the stats keeps the session's component_config on the refreshed
display config. A warm /load_expr, which rebuilds nothing, leaves stats_gen
alone. A stats_request on a session with no data is answered with
stats_aborted.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ed sessions (rows-first s3)

A client that advertises ?caps=stats_update gets a stats-free initial_state on a
deferred /load_expr session (df_meta.stats.status "pending") and pulls the stats
with stats_request {stats_gen, scope}. The reply is a stats_update carrying the
dataflow's all_stats as an inline wide envelope, or stats_aborted when the
generation is stale. The request is the whole run: one synchronous call that
computes the full stats, writes the full-tier summary_stats_cache entry, assigns
summary_sd and refreshes the session snapshot through one helper.

stats_gen is a server-owned counter bumped by every load handler and by a state
change that touches a dataflow field, which also returns a deferred session to
the schema tier. df_meta.stats is injected by build_state_message from the
session, since the dataflow rebuilds df_meta wholesale. Every send site goes
through build_state_message_for, so a client without the capability still gets
complete messages (its missing stats run synchronously first); broadcast_state
replaces the five copies of the send loop and sends to capable clients first.
The stats_request branch binds the session's telemetry sink and emits a
stats.request span.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…s (rows-first s4)

Both pass on the current code. process_df is written out as the
process_column loop it is, and each xorq histogram query's snapshot-cache
key is compared with the key of a query built in the test, so the refactor
into resumable units that follows can be checked against something that does
not go through the units.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tore (rows-first s4)

Each stats class gets plan(state) and run(unit, acc): the fragments of the
planned units, assembled through assemble_merged_sd, must equal the full-stats
merged_sd on pandas, polars and xorq with init_sd, cleaning and overrides; a
column group returns only its columns; skip_stat_columns columns get no unit;
the xorq batch can be split by column chunk behind a default-off flag and is
refused for anything but a plain parquet scan; the histogram cache keys stay
put. A StatRun held on the session, keyed by (stats_gen, scope), keeps an
append-only fragment list and accumulator, runs a unit only when asked, is
dropped when stats_gen changes, and each connection keeps its own cursor.

test_the_batch_is_one_query_by_default passes on the current code: it pins the
default the flag must not change.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
StatPipeline and XorqStatPipeline get plan(state), new_accumulator(state) and
run(unit, acc); process_df and process_table are those three in order, so the
inline path and a caller running units one at a time compute the same stats.
pandas and polars plan one unit per non-skipped column. xorq plans the scalar
batch (which also yields histogram_bins), then one histogram query per column,
visible columns first. The batch can be cut into column chunks sized by cells
when a host sets stat_chunk_cells, and only for a plain parquet scan: any other
source falls back to the single batch and logs why.

A StatRun, kept on the session by (stats_gen, scope), holds the planned units,
an append-only fragment list and the accumulator. It has no thread, timer or
callback, begin_stats_generation drops it, and each WebSocket connection keeps
its own StatCursor into the list. CustomizableDataflow.build_stats is the one
place a stats class is built, with run=False for plan/run callers.

Two of the new tests are corrected here: the default-path test now ignores the
PERVERSE_DF self-check units, and the xorq comparisons round floats to nine
digits because a parallel aggregate adds its partial sums in varying order.
The new-module imports in the tests move to module level.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… once (rows-first s4)

A column group, the priority columns and StatRun's prefer hint accept
original or rewritten (a, b, c) names and read each name in both. When an
original name equals another column's rewritten name, a group returns the
columns of both, priority ranks both first, and prefer picks the unit of
whichever comes first in the frame.

On a frame whose columns are c, b, a (rewritten a, b, c):

- StatState(df, columns=('c',)) plans the units of c and a.
- priority=('a',) plans c before a.
- StatRun.next_unit(prefer=('a',)) picks the unit of c.

The new tests also pin the namespace argument the fix adds, so a client that
holds the rewritten names can ask for exactly those: StatState(namespace=)
and StatRun.next_unit/run_next(namespace=), with "any" (the default),
"original" and "rewritten".

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…rst s4)

A column group, StatState.priority and StatRun's prefer hint matched a name
against both the original and the rewritten (a, b, c) column names. When an
original name equals another column's rewritten name, one name picked two
columns: a group returned columns that were not asked for, and prefer ran the
unit of whichever column came first in the frame.

resolve_names reads each name once. In "any", the default and what the code
did for names that do not collide, a name that is an original column name
picks that column and only a name that is not one is read as a rewritten
name. StatState(namespace=) and StatRun.next_unit/run_next(namespace=) also
take "original" and "rewritten", so a caller that holds the client's
rewritten names says so and gets exactly those columns. Priority and prefer
names are resolved against every column of the frame, not only the columns
in the group, so a name picks the same column whatever group is asked for.

prioritized now takes the state, since it needs the frame's columns and the
namespace to read the priority names.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t cursors and the final assignment (rows-first s5)

Covers the request branch that runs units for a time budget (at least one, one
cold unit per budget), the per-handler cursors over a shared StatRun, the
final assignment that frees the run and carries the rebuilt df_display_args,
a legacy client finishing a run in progress, a concurrent infinite_request
between units, the end-to-end order of first initial_state, rows and stats,
/reload_expr mid-run, and the stats.unit and firstpull.stats_total spans.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nt cursors and the final assignment (rows-first s5)

A stats_request with incremental: true runs the units of the session's StatRun
for about STATS_BUDGET_S (at least one) and answers with the fragments the
asking client has not seen, read through the handler's own StatCursor. The
request that runs the last unit does the final assignment in complete_stats:
the full-tier summary_stats_cache entry, summary_sd, the session snapshot and
the status in one step, the run freed, and the final reply carries the
rebuilt df_display_args when its digest differs from the one the client holds.
Without the field a request is still the whole run, now finishing the run's
remaining units when a client has made progress. A legacy client's complete
frame does the same.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…handlers (rows-first p33)

stats_tier takes auto, full, scalar or schema on /load_expr and /reload_expr
and is stored with the pair. The policy resolves after the schema-tier
dataflow and the count exist, is stored on the session, is reported in
df_meta.stats (tier_target, reason, auto_request, requestable, estimate,
omitted_keys, approx_keys, demand_columns, each with its documented default)
and is applied at WebSocket open only for a client that sends
?caps=stats_update,stats_ondemand. The version-skew cases (no caps,
stats_update only, both bits, an old server) and /load, which keeps
resolving to full, are covered.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…and the compare reset under a stats policy (rows-first p33)

Cases found untested after the first tests commit, each one a regression the
first set did not catch: a stats_update client pulling units from a policy
session gets full-tier updates and the final reply carries the rebuilt
df_display_args, a scalar tier named with inline delivery builds a schema
dataflow, and /load_compare clears the policy a session held.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…irst p33)

/load_expr and /reload_expr accept stats_tier auto, full, scalar or schema,
stored with the pair and kept out of has_config, so the warm short-circuit
holds. When the dataflow is built at the schema tier the handler resolves the
policy (resolve_stats_policy) from the count it already has, stores it on the
session and starts the stats generation from it: a target below full is
not_computed with the policy's reason.

df_meta.stats reports the policy as tier_target and estimate, plus
auto_request, requestable, omitted_keys, approx_keys and demand_columns where
they differ from their documented defaults. It is applied at WebSocket open
only for a client that sends ?caps=stats_update,stats_ondemand. A client with
stats_update only is told the session is pending and pulls the stats, and a
client with no caps gets them at connect, as for any deferred session. A tier
the host named (scalar, schema) reaches every client as before. A session on
an explicit stats_tier full within the ceiling sends the message it always has.

/load keeps resolving to full: it does not read the field and clears a policy
left on the session.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… p34)

A xorq stat run at the scalar tier is the scalar class of the batch without
approx_median and distinct_count, and no histogram query. Its histogram_bins
come from the batch's min and max. The run behaves like a filtered one:
fragments go to the clients and nothing is assigned to the dataflow, the
session snapshot or summary_stats_cache, so scalar stats are never served as
the complete ones.

Tests cover the plan, the queries a scalar run issues, the fragment union
against the full stats, the bins, the chunked batch, the StatRun key and
assignment rule, which clients are served the tier, and the stats_update
messages of an ondemand client on a scalar target.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…d generations (rows-first p34)

A deliberate regression of the implementation showed four behaviours the first
set of tests does not cover: a stat that reads a key the scalar tier leaves out
is left out with it, a stat that reads distinct_count runs on the unknown, a
scalar request emits the request and unit spans with the scalar tier and no
completion span, and a dataflow-field change drops the scalar run and starts
one over the new state.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
paddymul and others added 2 commits October 4, 2026 06:35
…signed (rows-first p34)

StatState takes a tier. A xorq run at the scalar tier is the batch alone, in
the same column chunks, without the aggregates behind median and
distinct_count and with no histogram unit. Its histogram_bins come out of the
batch's min and max, so color_map needs no histogram query. The accumulator
withholds median, distinct_count, distinct_per and histogram, so no fragment
and no assembled sd carries a key the tier did not compute. The pandas and
polars pipelines refuse a non-full state.

StatRun takes a state, and reports its tier and whether it assigns (only the
full tier over every column does). The full run keeps its (stats_gen, scope)
key; every other run has a key of its own, so it never takes the full run's
place. start_stat_run takes a tier and a column group. _run_summary refuses a
run that does not assign, so a scalar or column-scoped sd is never written to
summary_stats_cache as the complete one.

A stats_request from a client that advertised stats_update and
stats_ondemand, on a session whose policy target is scalar and has nothing
computed, runs the scalar units and answers stats_update with tier scalar,
final and remaining. The run behaves like a filtered one: the fragments go to
the clients that ask (each through its own cursor) and nothing is assigned to
the dataflow, the session snapshot, summary_stats_cache or the status. Other
clients are served as before.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…demand scan (rows-first p36a)

A stats_request from a client that advertised both capability bits can now
say tier, force and columns. The tests pin what the server does with them:
the ceiling refuses a tier for the cells asked for and runs nothing,
requestable and the target decide which tier a whole-table request may name,
force raises the target through a session override that survives a
dataflow-field change for the unfiltered scope only, an automatic unit over the
budget pauses the session (not_computed for cost) until a force request, and
the pause survives a generation and /reload_expr. The demand scan reads the
color_map rules of the built config (overrides and rules a klass adds at style
time) with no query, and a scoped request computes those columns only.

Every test here fails on the earlier code. The guard tests that pass there (a
client without both bits is served as before, a session with no policy ignores
the fields) go in with the implementation.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

📦 TestPyPI package published

pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37240210873

or with uv:

uv pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37240210873

MCP server for Claude Code

claude mcp add buckaroo-table -- uvx --from "buckaroo[mcp]==0.15.9.dev37240210873" --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo-table

📖 Docs preview

🎨 Storybook preview

…the demand scan (rows-first p36a)

A stats_request from a client that advertised both capability bits reads
tier, force and, with a tier, columns. The server judges it: the ceiling for
the cells asked for (resolve_stats_policy with the tier as the host's) answers
a refusal and runs nothing, a whole-table tier must be the target or above it
with force, and a paused session refuses a request that is not forced.

A force above the target stores stats_override on the session, and
effective_stats_policy raises the target for the unfiltered scope and for the
generation it was forced in. begin_stats_generation, stats_meta and
serves_scalar_tier read the effective policy. An automatic request that ran a
unit over STATS_COST_BUDGET_S and left units to run sets cost_paused, which
begin_stats_generation reads, so the pause survives a stats_gen bump and
/reload_expr until a force request. The override and the pause are forgotten
by /load, /load_compare, another expression and a body that names a tier.

stats_policy.demand_columns scans the built display config for color_map rules
and names the columns they read; the snapshot refresh writes them to
stats_policy["demand_columns"] for a schema target, and a scoped request
computes those columns only, with a run that is never assigned. The final
reply of a run that does not assign now carries the session's status.

Guard tests that pass on the earlier code go in with this commit, and the
elapsed_ms test has wider margins on both sides.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…n it continues after a cost pause (rows-first p36a)

A client that continues a cost-paused session (force with no tier, or a tier
equal to the target) never has its display_args_hash recorded, so the final
stats_update carries no df_display_args and the client keeps the config its
not_computed frame carried. Two clients are in that state: one that connected
while the session was paused, and one that changed state (a search) while it
was. Both complete the session with force and expect the rebuilt config in the
final reply. The client that connected pending before the pause is covered by
the existing tests.

Both tests fail on the earlier code: the final reply has no df_display_args.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…sed run (rows-first p36a)

_apply_force recorded the client's display_args_hash only when a force raised
the target above the policy's. A continue after a cost pause names no tier, or
the target, so the digest was never recorded: build_state_message_for had
already cleared it for the paused not_computed frame, and the final
stats_update left out df_display_args. A client that connected while the
session was paused, or changed state while it was, kept the schema-tier config
its frame carried after the session completed.

The digest is now recorded for any whole-table force that runs the full tier,
by raising the target or by continuing to it. A force for the scalar tier or
for named columns still records none, since neither run ends in an
assignment. A client that already holds a digest keeps it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@paddymul

paddymul commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Review pass on this PR: one finding, confirmed and fixed.

Finding (medium): a Continue after a cost pause never records the client's display digest. _apply_force set client.display_args_hash only inside the branch for a tier above the target. A Continue names no tier, or the tier the session is already headed for, so the digest was not recorded. build_state_message_for had already cleared it for the paused not_computed frame. The final stats_update then carried no df_display_args, and the client kept the schema-tier config after the session completed. Two clients were in that state: one that connected while the session was paused, and one that changed state (a search) while it was. The client that connected pending before the pause was not affected, which is why the existing tests passed.

I reproduced both cases before changing anything. In each, the final reply had no df_display_args, the session was complete, and the frame's config differed from the complete one.

How it was addressed

  • 6e85bd0 adds two tests to TestCostGuardWire: a client that connects while paused and continues (with force alone, and with tier: "full"), and a client that sends a search while paused and then continues. Each expects the final reply to carry the rebuilt config (with the client's highlight, in the search case). Locally they failed on the missing df_display_args and the rest of the suite passed. On CI, Python / Test (3.11), (3.12), (3.13) and the three matching Max Versions jobs failed on that commit, where the previous head had passed them. The 3.14 jobs passed because xorq is gated to Python below 3.14 and this test module is skipped without it.
  • 909ed59 changes _apply_force to record the digest for any whole-table force that runs the full tier, whether it raises the target or continues to it, and keeps the guard that a client already holding a digest keeps it. The override condition is the same as before. A forced scalar request or one scoped to columns still records nothing, since neither run ends in an assignment. The docstring now says this.

Every check on 909ed59 is complete and green (the deploy job is skipped), and the local unit suite is 2058 passed. No existing test was changed.

Not covered: the digest is taken when the client sends force. A client that was told not_computed for cost and then sends a request without force, after another client's force has cleared the pause and run the session to completion, would still get no df_display_args. The finding did not include that case and I did not change it here.

@paddymul

paddymul commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Closing. The cost guard pauses between units and the demand columns are served as scoped StatRuns, both pull machinery that the revised D2 of ADR-002 (#1043, 7943412) drops.

These still apply to an explicit stats_request under the push and are the parts to carry forward: _gate_request (the ceiling and requestable checks), the override (effective_stats_policy, reset_stats_controls) and the demand_columns scan. The branch is kept for them. #1034 is stacked on this branch and needs a rebase.

@paddymul paddymul closed this Oct 6, 2026
paddymul added a commit that referenced this pull request Oct 6, 2026
…count memo (rows-first p37)

stats_policy gains GuardLimits (BUCKAROO_SORT_DISABLE_ROWS, default 25M rows;
BUCKAROO_SEARCH_DISABLE_ROWS, default none) and resolve_source_guards, apart from
the stats thresholds. A session a host opened with a stats policy resolves them
at /load_expr and /reload_expr from the count load took. Above the sort
threshold the xorq dataflow finishes its display config with disable_sorting, so
no klass or override leaves a sortable column in a display infinite_request
serves, and a sorted infinite_request is refused with error_code sort_disabled
before any query runs. df_meta carries sort and search when one is disabled.

handle_infinite_request_xorq holds the searched expression per (base expression,
term), so the second window of a search is a hit in _expr_count's cache and
issues no count.

Rebased onto #1029 without the closed #1031 and #1033: the guard resets
no longer call reset_stats_controls, session_dataflow stays in stats_wire,
and the wire tests share a _LimitsWire base ported from #1033 without its
unit and cost helpers.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
testpypi — 909ed59c Deployed Oct 4, 2026 by paddymul via Publish to TestPyPI #1712
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant