Skip to content

feat(server): stats policy wire contract and handler plumbing (rows-first p33) - #1029

Draft
paddymul wants to merge 25 commits into
adr-003-stats-tiers-and-size-policyfrom
feat/rowsfirst-p33-policy-wire-contract
Draft

paddymul wants to merge 25 commits into
adr-003-stats-tiers-and-size-policyfrom
feat/rowsfirst-p33-policy-wire-contract

Conversation

@paddymul

@paddymul paddymul commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #1028 (feat/rowsfirst-s5-stats-request-units), which is stacked on #1026, #1024, #1022 and #1021, and it carries the stats policy of #1019 and #1023 through a merge commit. This PR is based on main so the repo's Checks workflow runs on it (its pull_request trigger uses branches: "*", which does not match a base branch containing a slash). The diff therefore includes the commits of #1021, #1022, #1024, #1026 and #1028, and the commits of #1019 and #1023 (the buckaroo/server/stats_policy.py module and its tests, brought in by the merge commit e552b408), until they merge; they drop out of this diff then. The commits of this phase are 751ea324 and 769b5c73 (failing tests) and c691825f (implementation); read only those three (git diff origin/feat/rowsfirst-s5-stats-request-units...HEAD, minus the merge).

Problem

resolve_stats_policy (#1019, #1023) decides a stats tier from an entry's size, and nothing calls it. The server has one stats policy pair, stats_tier (full or schema) and stats_delivery, and one status in df_meta.stats (status, tier, gen, reason). A host cannot say auto, so a server default that protects hosts that send nothing has nowhere to go, and a client has no way to learn why a session has no stats, what it may ask for, or whether to ask on its own.

Phase and plan references

Rows-first p33, from buckaroo2-reports/plans/: plan 3 (03-no-summary-stats-for-large-files.md) section 6 "Phase 3" with sections 3.2 (the first message and the client) and 3.6 (request field and server threshold), plan 1 sections 3 and 4.0, plan 2 sections 3 and 4.1.

Approach

  • Request field. /load_expr and /reload_expr take stats_tier of auto, full, scalar or schema (default full until plan 3 phase 9), stored on the session with stats_delivery. The pair is compared with the session's stored pair for the warm short-circuit and is not in has_config, as in feat(xorq): stats tier machinery and the xorq schema tier (rows-first s1) #1021. An omitted field keeps the session's value. /load does not read it: the eager backends resolve to full, and /load and /load_compare clear any policy a session held.
  • Resolution. When the dataflow is built at the schema tier (deferred delivery, or a tier named below full), LoadExprHandler calls resolve_stats_policy("xorq", "xorq_build", rows, cols, host_tier=stats_tier) after the dataflow and get_xorq_metadata (which holds the count) exist. No data query is added. The result is stored on session.stats_policy and the status restarts from it (initial_stats_status(stats_tier, stats_delivery, policy)): below full is not_computed with the policy's reason, full is pending as before. /reload_expr resolves again, with the stored host_tier (or the one in its body) and the count the load took, which session.metadata holds (the expression is reused, so there is no new count). A dataflow-field change keeps the stored policy.
  • Inline sessions. A dataflow built with its stats inline has run them before a policy could be resolved, so auto with inline delivery is full: no policy, the message it always was. auto takes effect with stats_delivery: "deferred", as plan 2 phase 4 and plan 3 phase 9 assume.
  • The wire. df_meta.stats keeps status, tier (reached so far), gen and reason, and a session that has a policy to report adds tier_target, estimate, and auto_request, requestable, omitted_keys, approx_keys, demand_columns only when they differ from their defaults (auto_request true, requestable ["full"], the three key lists empty). reason is one of size, host, cost, ceiling (cost is the cost guard's, phase 6a). Fields are added only while the status is pending or not_computed; a completed session reports {status: "complete", tier: "full", gen}. A session on an explicit stats_tier: "full" within the ceiling reports nothing new, so its messages are the ones sent before this PR. session.stats_with_defaults(df_meta) is the one place the defaults are written down; it also reads a message with no df_meta.stats (an old server) as complete.
  • Capabilities. ?caps=stats_update,stats_ondemand: the second bit counts only together with the first (a client that merges stats_update). The policy is applied at open for such a client only. For a target the server chose below full (a size rule, or the ceiling on a host that asked for auto or full), a client with stats_update only is told pending with no policy fields and pulls the stats as it does from any deferred session (stats_request runs them), and a client with no caps gets a complete initial_state, the stats run synchronously at connect, as for a deferred session today. A tier the host named below full (scalar, schema) is the host's choice and reaches every client as feat(server): stats_request, stats_update and df_meta.stats on deferred sessions (rows-first s3) #1024 already sends it (not_computed, reason host, no stats). stats_wire.stats_to_pull(session, client) decides whether a client's view has stats left to compute, and session.stats_meta(session, ondemand=...) builds the view, so every send site (build_state_message_for, broadcast_state, the highlight overlay, the open) gets it without a per-site change. broadcast_state already sends to clients with stats_update before the others, which keeps an ondemand client's frame ahead of a legacy client's synchronous completion.

What changes

  • buckaroo/server/session.py: STATS_TIER_REQUESTS, STATS_REASONS, STATS_FIELD_DEFAULTS, SessionState.stats_policy, dataflow_stats_tier, initial_stats_status(..., policy), begin_stats_generation, stats_meta(session, ondemand), build_state_message(..., ondemand), stats_deferred_by_policy, stats_with_defaults.
  • buckaroo/server/stats_wire.py: STATS_ONDEMAND_CAP, client_has_ondemand, stats_to_pull, resolve_session_policy; build_state_message_for, complete_stats and the stats_request handler read the view. A partial stats_update now says tier: "full" literally, where it read session.stats_tier, which can be auto.
  • buckaroo/server/handlers.py: stats_tier validation, the resolution in LoadExprHandler and ReloadExprHandler, the reset in LoadHandler and LoadCompareHandler.
  • No change to websocket_handler.py, dataflow.py or the policy module.

Tests

Three commits: 751ea324 and 769b5c73 (failing tests, pushed alone) and c691825f (implementation). The second tests commit covers four cases a deliberate regression of the implementation showed the first set did not catch.

  • test_stats_policy.py (pure, runs on every Python and on Windows): stats_with_defaults and the documented defaults, initial_stats_status over the pair and the policy, dataflow_stats_tier, stats_meta for eight resolved cases (auto within the thresholds, scalar and schema by size, full over the ceiling, scalar and schema named by the host, scalar over the scalar ceiling) and the three client views of them, a completed session, the generation restart, the pass-through of omitted_keys, approx_keys and demand_columns, resolve_session_policy, and client_has_ondemand and stats_to_pull for each cap set.
  • test_load_expr.py, TestStatsPolicyWire (xorq, over HTTP and WebSocket): every stats_tier value accepted and stored on /load_expr and /reload_expr; auto with inline delivery is full; a scalar tier with inline delivery builds a schema dataflow; the policy resolves after the schema-tier dataflow and the count, with no data stat query; it is stored on the session; the fields each case sends an ondemand client; a warm re-POST with an unchanged tier takes the short-circuit and a changed one rebuilds and resolves again; /reload_expr resolves again for a new tier, for new limits and against the stored count; the policy survives a dataflow-field change; the version-skew cases (no caps, stats_ondemand alone, stats_update only with a whole and an incremental pull, both bits with no stat query and not_requestable, the three kinds of client on one session, the ondemand client's frame sent before a legacy client's completion, a tier named by the host, an old server's message); an explicit stats_tier: "full" changes no frame.
  • test_server.py, TestLoad, and test_load_compare.py: /load ignores stats_tier and stats_delivery, resolves to full and clears a policy left on the session, and so does /load_compare.
  • Added with the implementation because they pass on the earlier code: an explicit full within the ceiling reports no policy (test_an_explicit_full_within_the_ceiling_reports_no_policy, and the same check in test_an_explicit_full_changes_no_frame). One test from the first commit was removed there, test_the_policy_is_what_the_pure_function_resolved: it compared a value the test had just assigned with the function that produced it, and could not fail.

The size limits in the wire tests are the BUCKAROO_STATS_* environment overrides of #1019, which resolve_stats_policy reads on every call, so a five-by-three table lands in each tier.

On 751ea324 all nine Python / Test jobs failed (3.11 to 3.14, Max Versions 3.11 to 3.14, Windows), with all 28 checks completed. On 769b5c73 all nine failed again, 28 checks completed. CI logs were not read (the REST log API is rate-limited for this account), so the reasons rest on the local run of the same commits: on 769b5c73, 86 of the 93 new tests fail (each on a missing attribute, an HTTP 400 for the new stats_tier values, or an assertion on a frame) and 7 pass (five rows of the dataflow_stats_tier table that the old function already satisfies, one row of the completed-session table, and the test removed later). The full unit suite on 769b5c73 gives 86 failed and 1800 passed, and every failure is one of the new tests.

CI on c691825f: all 28 checks completed, 27 succeeded (one of them the Read the Docs status) and deploy was skipped. That includes all nine Python / Test jobs, Python / Lint, Python / Typecheck, the JS job, the wheel build and the Playwright jobs.

Local checks on c691825f: the full unit suite (pytest ./tests/unit -m "not slow") gives 1887 passed and 5 skipped, and the same 1887 and 5 in an environment resolved with the newest versions the Max Versions job uses (pandas 3.0.6, polars 1.44.2, xorq 0.4.5, numpy 2.5.3). Twenty-seven deliberate regressions of the implementation (the policy always reported, never reported, the legacy view never pending, the ondemand check ignored or reduced to one bit, complete_stats or the request path counting only pending, a policy not stored by /load_expr or /reload_expr or not cleared by /load or /load_compare, a reload that counts again, the wrong host tier, the old tier set, the display digest recorded only when pending, a partial update that names the host's tier, and others) each make at least one test fail. A real server process (python -m buckaroo.server --no-browser on a free port, BUCKAROO_STATS_FULL_AUTO_ROWS=1000, a 50,000-row xorq build, stats_tier: "auto" and stats_delivery: "deferred") answered a client with both bits not_computed, reason size, tier_target scalar with a stats_request refused as not_requestable; a stats_update-only client pending and its request answered; and a client with no caps complete.

Why default behaviour is unchanged

  • The field defaults to full, and a session on the default pair (or an explicit full within the ceiling) sends the same initial_state as before: df_meta has no stats key for an inline session and the same {status, tier, gen} for a deferred one. test_an_explicit_full_changes_no_frame compares the frames of two sessions with and without the field for inline and deferred delivery and for each cap set, and checks the deferred frame is exactly {status, tier, gen}.
  • auto and scalar are new values. schema and full keep the behaviour of feat(xorq): stats tier machinery and the xorq schema tier (rows-first s1) #1021 and feat(server): stats_request, stats_update and df_meta.stats on deferred sessions (rows-first s3) #1024; a schema session now also carries policy fields for a client with both bits, which no client sends yet.
  • A client that sends no stats_ondemand bit sees the status quo: a complete state, or pending and a pull. Tallyman (buckaroo-js-core 0.15.8) sends neither bit.
  • Existing tests are unchanged.

Deviations from the plan

  • The plan says a message with no tier_target means the target equals tier. That is wrong for a pending message (tier schema, headed for full), so the default here is full unless the status is not_computed, where it is tier. A session that reports a policy always sends tier_target.
  • scalar is accepted and resolved, but no scalar-tier units exist (phase 4), so the dataflow is built at the schema tier and the status is not_computed for a scalar target. requestable and auto_request are the policy function's output, so a scalar target reports auto_request true and a schema target requestable ["scalar", "full"], though the server serves neither tier yet: a stats_request from an ondemand client on a not_computed session is still refused with not_requestable. Serving a tier and enforcing requestable, the ceiling and force belong to phases 4 and 6a.
  • A client with no caps, or with stats_update only, gets the full stats for a target the server chose below full, ceiling included. That is plan 3's version-skew rule (a client without stats_ondemand gets today's behaviour), and it means an older client can still run the 215 s case it can already run today. The ceiling protects clients that advertise stats_ondemand, and, once phase 6a lands, their force requests.
  • stats_ondemand without stats_update is treated as no caps. The plan defines the second bit as an addition to the first.
  • A tier the host named (scalar, schema) is applied to every client, where the plan's skew table speaks of "a resolved-schema entry". feat(server): stats_request, stats_update and df_meta.stats on deferred sessions (rows-first s3) #1024's tests pin that a schema session is not_computed for a stats_update client and gets the schema tier on a legacy connection, and a host that names a tier has chosen it.
  • auto with inline delivery is full (see Approach); the plan says the policy resolves only on a deferred session.
  • The policy is resolved at load and reload and kept across a dataflow-field change, not resolved again on a filtered count.
  • No stats.policy span (the detail file's phase 3 lists one; plan 3 section 6 does not).
  • The policy is stored as the dict resolve_stats_policy returns, with the three key lists added by later phases as optional keys (omitted_keys, approx_keys, demand_columns); stats_meta sends them when present.

Not in this PR

  • Scalar-tier units, the client, limits and force, the cost guard and the demand scan (the producers of demand_columns, approx_keys, omitted_keys and cost_paused), and the default flips. The three key lists are carried by the session policy and sent when present; nothing fills them yet.
  • /load taking stats_tier (L-track), and footer wiring.

🤖 Generated with Claude Code

paddymul and others added 17 commits October 3, 2026 22:51
…the _handle_widget_change split (rows-first s1)

A default-tier XorqBuckarooWidget and BuckarooWidget publish df_data_dict,
then df_display_args, then the rest of the widget_args_tuple observers on a
search change, and merged_sd carries the full stat set. These pass on main
and pin the behaviour the split must keep.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…(rows-first p31)

New tests/unit/server/test_stats_policy.py for buckaroo/server/stats_policy.py,
which does not exist yet:

- resolve_stats_policy: a table of (backend, source_kind, rows, cols, host_tier)
  against (tier_target, auto_request, requestable, reason), including the
  boundaries of each threshold and the three tallyman entries over 70 s.
- The ceiling holds for host_tier="full", for a reload that re-resolves
  against a larger row count, and for a force request, and every tier a result
  lists as requestable is one a force would be granted.
- Probes add no data query: the parquet footer probe reads a small share of a
  counting file object and decodes no row group, the dtype probe never collects
  a LazyFrame or executes a xorq expression, and a known xorq count is an input.
- route_polars_entry returns "xorq" above R and "eager" at or below it.
- Threshold overrides from BUCKAROO_* environment variables.

Nothing imports the module yet. The tests error at fixture setup on the missing
module until the implementation lands.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… p31)

New buckaroo/server/stats_policy.py, pure and imported by nothing yet:

- resolve_stats_policy(backend, source_kind, rows, cols, bytes, host_tier,
  limits) returns tier_target, auto_request, requestable, reason and estimate
  over the tiers schema < scalar < full. A ceiling is computed inside the
  function, so every caller gets the lower of the requested and ceiling tiers
  with reason "ceiling". A host tier lowers freely and raises only to the
  ceiling. Eager pandas and polars resolve to full.
- route_polars_entry(rows, cols) returns "xorq" above R rows and "eager" at or
  below it.
- probe_dtypes reads a schema without collecting or executing anything, and
  probe_parquet_rows reads a parquet footer's num_rows without decoding data.
  A xorq count is an input to the policy, never computed by it.
- The thresholds are provisional constants gathered in StatsLimits, each with
  a BUCKAROO_* environment override read on every call.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…-handler fields (rows-first s1)

A schema-tier XorqServerDataflow should match full stats on pinned_rows,
data_key, summary_stats_key and (except the stats-derived minWidth)
column_config, issue no data query besides the cached count, and keep
init_sd hints and sorted windows working. A pending state must write
nothing under a full-tier cache key, and a later full assignment must reach
merged_sd for the raw, clean and filt scopes. assemble_merged_sd must equal
merged_sd, and _handle_widget_change must be built from separately callable
all_stats and display-args builders. /load_expr and /reload_expr accept
stats_tier and stats_delivery, replay them on reload, and keep them out of
the warm short-circuit's has_config tuple.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… s1)

Add a dataflow-level stats_tier ("full" default, "schema"). The xorq schema
tier builds identity and typing for every column from the expression's
schema, with no data query beyond the cached row count, so a dataflow
constructs in milliseconds rather than the stats' hundreds.

The tier is part of _scope_cache_key, so a schema entry is never read as a
full one, and _populate_sd_cache stores summary_sd under the filt key only
if it was computed for the current frame, klass list and tier. add_analysis
no longer builds DFStatsClass outside the hook when the tier is not full.
The merged_sd observer body is extracted as the pure assemble_merged_sd,
and _handle_widget_change is split into _build_df_data_dict and
_build_df_display_args.

/load_expr and /reload_expr accept stats_tier and stats_delivery, stored on
the session beside dataflow_kwargs and replayed on reload. They stay out of
the has_config tuple; the warm short-circuit compares the stored pair.
stats_delivery="deferred" builds the schema-tier dataflow and publishes it.
Defaults (full, inline) leave behaviour unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…esolve_stats_policy (rows-first p31)

bytes and source_kind were documented as validated but are not: bytes=-1,
'lots', 1.5 and object() are accepted, a np.int64 is echoed unchanged so
json.dumps of the result fails, a host tier passed positionally lands in the
bytes slot, and source_kind=None or 5 is accepted.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ation test (rows-first s1)

The Max Versions jobs resolve pandas 3, which reports a string column's
dtype as 'str' where pandas 2 says 'object'. The characterization test
asserted 'object'. Verified in a Max Versions environment (pandas 3.0.6,
polars 1.44.2, xorq 0.4.5): the unit suite passes.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…rows-first p31)

bytes goes through the same integer check as rows and cols, so a negative,
float or non-numeric value raises, a numpy integer is echoed as a plain int
and the result stays JSON-serialisable, and a host tier passed positionally
into the bytes slot raises instead of being echoed. source_kind must be a
str; its vocabulary stays open for the callers that will read it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…g on a skipped column (rows-first s1)

A column in skip_stat_columns gets only name, dtype and length from the
full-tier pipeline, so its _type comes from init_sd. The schema tier layers
the schema-derived _type and is_* keys over it, so an int64 column that
init_sd types as float merges as integer and renders with zero fraction
digits instead of the float displayer init_sd asked for.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ows-first s1)

_get_schema_sd never read skip_stat_columns, so a skipped column's
schema-derived _type and is_* keys overrode init_sd's _type once merged. The
full tier gives a skipped column only name, dtype and length. The schema tier
now does the same, so init_sd's _type decides the displayer at both tiers.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…er (rows-first s2)

ServerDataflow and PolarsServerDataflow with stats_tier="schema" should publish
the display state full stats give (column_config without stats-derived keys,
pinned_rows including a host-supplied one, data_key, summary_stats_key) with no
stat computed on the data, still apply init_sd, serve sorted windows, take a
later full assignment into merged_sd for every scope, and assemble to the same
sd as merged_sd. Both backends run through the same parametrized class, plus a
pandas test that pins how an object column is typed from its dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
ServerDataflow and PolarsServerDataflow now implement the _get_schema_sd hook,
so stats_tier="schema" builds a dataflow that types every column from its dtype
and runs no stat on the data. schema_sd (stat_pipeline.py) builds the sd the way
process_df shapes it: an empty frame gives {}, a skipped column keeps only its
names. pandas applies the existing typing_stats to a zero-row slice and derives
_type through the _type stat; polars factors pl_dtype_typing out of
pl_typing_stats and feeds it the dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…s (rows-first p31b)

The boundary tables now run on the module defaults and are written against
the phase-0 proposals: full auto up to 12M rows and 520M cells, scalar auto
up to 1.0B cells, full refused above 25M rows or 1.0B cells (force
included), a scalar ceiling of 4.0B cells, and polars routed to xorq above
8M rows. Each threshold has an equal, a one-below and a one-above row, in
rows and in cells. New tests pin each default to its literal value and
cover the two new environment overrides, BUCKAROO_STATS_FULL_AUTO_CELLS and
BUCKAROO_STATS_CEILING_FULL_CELLS.

They fail on the current constants (10M rows, 500M cells, 50M-row ceiling,
no scalar ceiling, R of 10M) and on the missing full_auto_cells and
ceiling_full_cells fields.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tats_request (rows-first s3)

On a deferred /load_expr session a stats_request {stats_gen, scope} should
return a stats_update with the matching stats_gen whose inline wide payload
equals the all_stats an inline session sends, a stale stats_gen should get
stats_aborted and run no query, and /load_expr and /reload_expr should bump the
generation. df_meta.stats should be injected on every frame and survive a
dataflow-field change, which returns the session to the schema tier. With a
caps client and a legacy client on one session, the legacy client should keep
getting complete messages through the websocket broadcast, the /load_expr,
/reload_expr, /load and /load_compare pushes and the highlight overlay, while
the caps client gets a stats-free frame and then pulls a stats_update. A spy
telemetry sink should see a stats.request span.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…measurements (rows-first p31b)

Full auto now needs at most 12M rows and 520M cells (was 10M rows), scalar
auto goes to a 1.0B cell budget (was 500M), `full` is refused above 25M rows
or 1.0B cells including for a forced request (was 50M rows), scalar gets a
4.0B cell ceiling by default (was unset; an extrapolation with no
measurement behind it), and polars routes to xorq above 8M rows (was 10M;
the figure assumes pre_limit False). Each DEFAULT_* constant carries the
measurement it comes from.

The two new bounds are full_auto_cells and ceiling_full_cells, appended to
StatsLimits, with BUCKAROO_STATS_FULL_AUTO_CELLS and
BUCKAROO_STATS_CEILING_FULL_CELLS as overrides under the same parsing rules.
_size_tier and _ceiling_tier read them; signatures and the result shape are
unchanged.

Three existing tests that hard-coded the old boundaries are updated: the
numpy-integer case (11M rows is now full), the reload case (a 40M-row entry
is now refused full) and the empty-environment case for the scalar ceiling
(an empty value now means the 4.0B default, not no ceiling).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…eration edge cases (rows-first s3)

Four more cases for the stats wire format, kept in their own commit so each is
seen failing on CI before the implementation lands. Returning to a state whose
stats were completed once is answered from summary_stats_cache with no query.
Completing the stats keeps the session's component_config on the refreshed
display config. A warm /load_expr, which rebuilds nothing, leaves stats_gen
alone. A stats_request on a session with no data is answered with
stats_aborted.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ed sessions (rows-first s3)

A client that advertises ?caps=stats_update gets a stats-free initial_state on a
deferred /load_expr session (df_meta.stats.status "pending") and pulls the stats
with stats_request {stats_gen, scope}. The reply is a stats_update carrying the
dataflow's all_stats as an inline wide envelope, or stats_aborted when the
generation is stale. The request is the whole run: one synchronous call that
computes the full stats, writes the full-tier summary_stats_cache entry, assigns
summary_sd and refreshes the session snapshot through one helper.

stats_gen is a server-owned counter bumped by every load handler and by a state
change that touches a dataflow field, which also returns a deferred session to
the schema tier. df_meta.stats is injected by build_state_message from the
session, since the dataflow rebuilds df_meta wholesale. Every send site goes
through build_state_message_for, so a client without the capability still gets
complete messages (its missing stats run synchronously first); broadcast_state
replaces the five copies of the send loop and sends to capable clients first.
The stats_request branch binds the session's telemetry sink and emits a
stats.request span.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

📦 TestPyPI package published

pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37493398972

or with uv:

uv pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37493398972

MCP server for Claude Code

claude mcp add buckaroo-table -- uvx --from "buckaroo[mcp]==0.15.9.dev37493398972" --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo-table

📖 Docs preview

🎨 Storybook preview

…al_state is dropped (#998) (#1014)

* test(server): failing tests for state_seq / reply_seq on buckaroo_state_change (#998)

Server: a buckaroo_state_change carrying state_seq gets reply_seq back on
the initial_state sent to the originating client (dataflow broadcast and
highlight overlay); another client's broadcast copy carries none.

Client: WebSocketModel attaches an incrementing state_seq to each change
it sends and drops an initial_state whose reply_seq is older than the
latest one sent; replies without reply_seq still apply.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(server): sequence token on buckaroo_state_change so a stale initial_state is dropped (#998)

WebSocketModel attaches an incrementing state_seq to each
buckaroo_state_change it sends. The handler echoes it as reply_seq on
the initial_state sent back to the originating client, on the dataflow
broadcast and on the highlight overlay; other clients' broadcast
copies, the /load push and a fresh connection carry none. The client
drops an initial_state whose reply_seq is older than the latest
state_seq it sent, so the reply for "alle" arriving after "allen" was
sent no longer reverts buckaroo_state, purges the grid and re-sends
the search.

reply_seq is optional on the wire; a client that sends no state_seq
gets today's behaviour, so PROTOCOL_VERSION stays at 1.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(server): failing test for a state_seq reply on a no-op buckaroo_state_change (#998)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* fix(server): answer a numbered no-op buckaroo_state_change so the client's newest seq gets a reply (#998)

The client drops a reply whose reply_seq is older than its latest state_seq. A
dataflow change followed at once by a show_commands/df_display/sampled toggle
got no reply for the toggle, so the dataflow reply was dropped and the new
display config never arrived. The server now sends a state ack carrying the
change's own buckaroo_state for any numbered change that touches no dataflow
field and no search term. Clients that send no state_seq see no new messages.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* test(server): failing tests for the error reply, the overlay's buckaroo_state and the live highlight (#998)

- A numbered change that fails gets only an error frame. The client has
  already dropped the previous reply as stale, so it is left on the data
  from before both changes.
- The highlight overlay carries reply_seq but builds buckaroo_state from
  the session's, which puts back show_commands and df_display.
- A dataflow broadcast drops each client's live-search highlight.
- Clearing the live term also strips the committed search's highlight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(server): answer a failed numbered change, echo the change's state in the overlay, keep each client's highlight (#998)

The highlight overlay and the no-op ack become one per-client reply,
_send_client_state. Every reply it sends carries the change's own
buckaroo_state, the session's current data and this client's live-search
highlight.

- A numbered change that fails gets that reply after the error frame,
  carrying reply_seq. The client has already dropped the reply to its
  previous change, so it was left on the data from before both. The
  reply keeps the failed change's buckaroo_state: the session's would
  revert a search box holding the failed term, and the box would send it
  again.
- The overlay no longer builds buckaroo_state from the session's, which
  put back show_commands and df_display.
- Each client's copy of a dataflow broadcast carries its own live-search
  highlight.
- An empty live term leaves df_display_args as the session has it, so the
  committed quick_command_args.search highlight stays.
- The reply no longer skips an empty df_display_args, so every numbered
  change that reaches the dataflow is answered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
paddymul and others added 7 commits October 6, 2026 11:57
…ts-tier-core

# Conflicts:
#	buckaroo/server/handlers.py
…at/rowsfirst-s3-stats-wire-server

Brings in #1004 (telemetry.arm_session) and #1014 (state_seq / reply_seq).

Conflict resolution:
- broadcast_state takes reply_to, reply_seq and highlight, so the
  state-change broadcast keeps #1014's reply_seq for the originating client
  and per-client live-search highlight while staying capable-first and
  completing stats for legacy clients.
- build_state_message_for passes reply_seq through.
- #1014's _send_client_state replaces _send_highlight_overlay and builds its
  message with build_state_message_for, so a legacy client's stats are
  completed before the highlight is applied.
- LoadCompareHandler keeps both arm_session and the stats reset.
- The overlay-order test calls _send_client_state; a new test covers a
  numbered change on a deferred session (reply_seq to the sender only).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…handlers (rows-first p33)

stats_tier takes auto, full, scalar or schema on /load_expr and /reload_expr
and is stored with the pair. The policy resolves after the schema-tier
dataflow and the count exist, is stored on the session, is reported in
df_meta.stats (tier_target, reason, auto_request, requestable, estimate,
omitted_keys, approx_keys, demand_columns, each with its documented default)
and is applied at WebSocket open only for a client that sends
?caps=stats_update,stats_ondemand. The version-skew cases (no caps,
stats_update only, both bits, an old server) and /load, which keeps
resolving to full, are covered.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…d the compare reset under a stats policy (rows-first p33)

Cases found untested after the first tests commit: a scalar tier named with
inline delivery builds a schema dataflow, and /load_compare clears the policy
a session held.

Rebased off the closed unit PRs (#1026, #1028): the assertions on unit tiers
and on df_display_args in the final reply are dropped with the code they
tested.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…irst p33)

/load_expr and /reload_expr accept stats_tier auto, full, scalar or schema,
stored with the pair and kept out of has_config, so the warm short-circuit
holds. When the dataflow is built at the schema tier the handler resolves the
policy (resolve_stats_policy) from the count it already has, stores it on the
session and starts the stats generation from it: a target below full is
not_computed with the policy's reason.

df_meta.stats reports the policy as tier_target and estimate, plus
auto_request, requestable, omitted_keys, approx_keys and demand_columns where
they differ from their documented defaults. It is applied at WebSocket open
only for a client that sends ?caps=stats_update,stats_ondemand. A client with
stats_update only is told the session is pending and pulls the stats, and a
client with no caps gets them at connect, as for any deferred session. A tier
the host named (scalar, schema) reaches every client as before. A session on
an explicit stats_tier full within the ceiling sends the message it always has.

/load keeps resolving to full: it does not read the field and clears a policy
left on the session.

Rebased onto #1024 without the unit PRs (#1026, #1028): stats_request is the
whole run of #1024, so handle_stats_request now takes the client to decide
stats_to_pull, and the incremental and display-config parts are gone.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@paddymul
paddymul force-pushed the feat/rowsfirst-p33-policy-wire-contract branch from c691825 to ffd7862 Compare October 6, 2026 16:08
paddymul added a commit that referenced this pull request Oct 6, 2026
…count memo (rows-first p37)

stats_policy gains GuardLimits (BUCKAROO_SORT_DISABLE_ROWS, default 25M rows;
BUCKAROO_SEARCH_DISABLE_ROWS, default none) and resolve_source_guards, apart from
the stats thresholds. A session a host opened with a stats policy resolves them
at /load_expr and /reload_expr from the count load took. Above the sort
threshold the xorq dataflow finishes its display config with disable_sorting, so
no klass or override leaves a sortable column in a display infinite_request
serves, and a sorted infinite_request is refused with error_code sort_disabled
before any query runs. df_meta carries sort and search when one is disabled.

handle_infinite_request_xorq holds the searched expression per (base expression,
term), so the second window of a search is a hit in _expr_count's cache and
issues no count.

Rebased onto #1029 without the closed #1031 and #1033: the guard resets
no longer call reset_stats_controls, session_dataflow stays in stats_wire,
and the wire tests share a _LimitsWire base ported from #1033 without its
unit and cost helpers.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
testpypi — ffd7862e Deployed Oct 6, 2026 by paddymul via Publish to TestPyPI #1751
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant