Skip to content

feat(server): sort and search guards for huge sources and a searched count memo (rows-first p37) - #1034

Draft
paddymul wants to merge 28 commits into
adr-003-stats-tiers-and-size-policyfrom
feat/rowsfirst-p37-huge-source-guards
Draft

paddymul wants to merge 28 commits into
adr-003-stats-tiers-and-size-policyfrom
feat/rowsfirst-p37-huge-source-guards

Conversation

@paddymul

@paddymul paddymul commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Stack

Stacked on #1033 (feat/rowsfirst-p36a-limits-override-cost-guard), which carries #1021, #1022, #1024, #1026, #1028, #1029, #1031 and the stats policy of #1019 and #1023. This PR is based on main so the repo's Checks workflow runs on it (its pull_request trigger uses branches: "*", which does not match a base branch containing a slash). The diff therefore includes the commits of those PRs until they merge.

The commits of this phase are the three after 909ed59c, the head of #1033 (git diff origin/feat/rowsfirst-p36a-limits-override-cost-guard...HEAD):

  • c7fb1ff9 and d9be4e90, failing tests, each pushed alone.
  • 0181386e, the implementation.

Problem

A session whose source is very large can skip its summary stats (#1029, #1033), but nothing keeps the grid from running the two requests that stay expensive without them:

  • A sorted window on a xorq source of tens of millions of rows blocks the loop for the whole sort (261-287 ms warm at 10.8M rows x 43 columns, 326-337 ms at 3M CSV rows, from the phase-0 measurements; the 42M to 78M entries are extrapolations). The grid offers a sort on every column, and a klass or a host override can turn it back on.
  • A searched window pays a fresh count() per request. search_expr builds a new expression on every call and _expr_count caches by expression object, so the filtered expression always misses (504-1063 ms for a "NY" search on 10.8M rows, each window).

Phase and plan references

Rows-first p37, from buckaroo2-reports/plans/: plan 3 (03-no-summary-stats-for-large-files.md) section 6 "Phase 7" with section 3.5, and plan 2 (02-rows-first-xorq-and-lazy-polars.md) section 4.3 (the filtered count memo). Plan 1 sections 3 and 4.0 and plan 2 sections 3 and 4.1 for the session and wire contract it extends. The detail file's section 4.7 has the measurements.

Approach

  • Thresholds (stats_policy.py). GuardLimits(sort_disable_rows, search_disable_rows) with from_env(), and resolve_source_guards(backend, source_kind, rows, cols, limits=None), which returns {"sort": "enabled" | "disabled", "search": "enabled" | "disabled"}. They are separate from StatsLimits, which is unchanged: a sort scales with rows, the stats tiers with cells. BUCKAROO_SORT_DISABLE_ROWS (provisional default 25,000,000 rows) and BUCKAROO_SEARCH_DISABLE_ROWS (default none) are read on every call, like the stats ones. Eager backends are always enabled. 25M is also the ceiling on full stats (the midpoint of the telemetry gap between 11.8M and 42.3M rows). It is a guess marked provisional, with the extrapolated costs written beside it.
  • Who is guarded. Only a session a host opened with a stats policy: resolve_session_guards(dataflow_tier, rows, cols) returns None unless the dataflow was built at the schema tier, as resolve_session_policy does. A session built the way every host builds one today, with its stats inline, has no guard. The count is the one load took (metadata["rows"]), so the guard describes the source, not a search over it, and does not flip while a user types.
  • The final pass. XorqServerDataflow.set_sort_enabled(False) makes _build_df_display_args finish with stats_policy.disable_sorting(display_args). That sets ag_grid_specs.sortable false on every column and every left (index) column of each display whose data_key is main, the display infinite_request serves, and keeps the column's other specs. The summary display is a client-side grid and keeps its sort. The pass runs after the klasses and after column_config_overrides are merged, so neither can turn sorting back on. The flag lives on the dataflow, so every rebuild keeps it (a field change, completed stats, /reload_expr).
  • The refusal. session.sort_refusal(session, payload_args) returns infinite_resp {key, length: 0, error_info, error_code: "sort_disabled"} for a window that has a sort on a session whose sort is off. DataStreamHandler._dispatch asks it first, so both the first and the second_request window are judged and no query runs. No parquet frame follows.
  • The flags. session.source_guards holds the resolved flags. build_state_message adds df_meta.sort and df_meta.search when one is "disabled", to every client (the server refuses the sort whoever asks). An enabled flag is left out, as the stats fields are, so a session under the thresholds sends the message it always has. session.guards_with_defaults(df_meta) reads them with the default filled in. search is advice: there is no server switch for search (plan 3 question 7), so a window with a term is served either way, and a client that wants to hide the search cell reads the flag.
  • The count memo. handle_infinite_request_xorq holds the searched expression per (base expression, term) in a weak-keyed map, with a 16-term LRU per base. The second request for the same term gets the same expression object, so _expr_count's own cache serves the count, and a count that failed is still not cached. A state change builds a new base expression, which is a new key.

What changes

  • buckaroo/server/stats_policy.py: DEFAULT_SORT_DISABLE_ROWS, DEFAULT_SEARCH_DISABLE_ROWS, GuardLimits, resolve_source_guards, disable_sorting.
  • buckaroo/server/session.py: SessionState.source_guards, SOURCE_GUARD_DEFAULTS, guards_with_defaults, sort_refusal; build_state_message adds the flags.
  • buckaroo/server/stats_wire.py: resolve_session_guards, apply_source_guards.
  • buckaroo/server/xorq_loading.py: XorqServerDataflow.sort_enabled, set_sort_enabled and the _build_df_display_args override; the searched-expression memo in handle_infinite_request_xorq. search_expr is imported at module top instead of inside the function.
  • buckaroo/server/handlers.py: /load_expr and /reload_expr resolve and apply the guards; /load and /load_compare clear them.
  • buckaroo/server/websocket_handler.py: _dispatch asks sort_refusal first.
  • No change to dataflow.py, the widgets or any client.

Tests

120 new tests, in the existing files of their kind.

  • test_stats_policy.py (pure, runs on every Python and on Windows): TestGuardLimits, TestResolveSourceGuards (a boundary table in rows, the two thresholds apart from each other and from the stats ones, eager backends, bad inputs), TestGuardFlagsInDfMeta, TestSortRefusal, TestResolveSessionGuards, TestDisableSorting.
  • test_load_expr.py (xorq): TestSortGuardDataflow, TestSortGuardWire (every column across a klass and an override, the refusal with no query and no stray frame, second_request, the flags for each kind of client, a reload resolving the guards again in both directions, /load and /load_compare ending the guard), TestSearchedCountMemo and TestSearchedWindowCountsWire (one CountStar for a repeated searched window, for other windows of the same search and for two clients; the bound, the LRU order, a failed count not remembered, and the base expression not kept alive).

The tests came in two tests-only commits, as in #1029 and #1031. c7fb1ff9 has the first set. A deliberate-regression pass over it left cases no test pinned (the memo's bound and order, /load_compare, non-dict column entries), and d9be4e90 adds those and corrects one wire test that expected five rows after first_three, which leaves three. Guard tests that already pass on the earlier code (an inline session is not guarded, flags that are enabled change nothing, the stats policy ignores the guard thresholds, the memo does not hold the base expression alive) went in with the implementation, because they cannot be seen failing.

Verification

  • CI on c7fb1ff9 (tests only): all nine Python / Test jobs failed (3.11-3.14, Max Versions 3.11-3.14, Windows), 28 checks completed.
  • CI on d9be4e90 (tests only): all nine Python / Test jobs failed again, 28 checks completed, deploy skipped. Locally, against the earlier source, the tests of the first commit and the six new or corrected ones of the second fail on assertions, not on collection.
  • Full unit suite on 0181386e: 2178 passed, 5 skipped. The same 2178 passed and 5 skipped in a newest-versions environment (pandas 3.0.6, polars 1.44.2, xorq 0.4.5, numpy 2.5.3).
  • 44 deliberate regressions of the implementation (a threshold compared with >=, each flag reading the other's threshold, a pass that skips the left columns or also changes the summary display, a refusal that ignores the flag or the sort key, guards applied after the snapshot, /load keeping the guard, a memo that never hits, is unbounded, evicts the newest, is not LRU, or holds the base alive) each fail at least one test. One survived the first set (the memo not refreshing a term on a hit) and is killed by d9be4e90.
  • A real server on ports 8951 (this branch) and 8952 (the earlier source), both started and stopped by me, on a 2M-row, 4-column parquet entry opened with stats_tier auto, stats_delivery deferred and both thresholds at 1M rows:
    • this branch: df_meta carried sort and search as disabled; every column of the main display had sortable false and the summary display none; a sorted window got sort_disabled in 0.3 ms with no frame, for a client with the stats capabilities and for one with none; an unsorted window was served; searched windows took 53 ms, then 8.3, 8.1 and 7.7 ms.
    • earlier source: no flags, no sortable setting, the sorted window was served (19.6 ms on this table), and every searched window took 46-49 ms.
  • CI on 0181386e: all 28 checks completed, 27 success and deploy skipped, including all nine Python / Test jobs, Python / Lint and Python / Typecheck.

Why default behaviour is unchanged

  • A session is guarded only if the host opened it with a stats policy, and only above a threshold. Inline sessions, /load and /load_compare get no guard and the same messages (test_a_session_built_the_old_way_is_not_guarded).
  • An enabled flag is not sent, so df_meta is the same for every session under the thresholds (test_enabled_flags_leave_the_message_as_it_was).
  • The memo changes the number of backend counts. A response has the same rows and the same length.
  • The existing tests pass unchanged. The only edits to existing test text are the module docstring of test_stats_policy.py and three import lines in test_load_expr.py (Expr for counting queries, gc and weakref).

What the client does on sort_disabled

Nothing new; the client is not changed. getKeySmartRowCache sees error_info on the infinite_resp, calls addErrorResponse, which fails the waiting block, and passes error_info to its error callback. In BuckarooView, the server-view client, that callback is a console.error; in BuckarooWidgetInfinite the respError state that would hold the text is commented out and the grid gets error_info="". So the refusal is logged, the block fails, and the grid keeps its sort model. With sortable: false on every column the header offers no sort, so a refusal reaches only a client that holds an older config. Reverting the sort model on error_code, and showing the message, is client work for a later phase.

Deviations from the plan

  • The final pass is in XorqServerDataflow._build_df_display_args, not in style_columns. merge_column_config applies the host overrides after style_columns returns, so a pass inside it would be bypassed by {"qty": {"ag_grid_specs": {"sortable": true}}}. A test pins that case.
  • The pass covers displays served by infinite_request (data_key == "main"), not the client-side summary grid.
  • Guards apply only to sessions opened with a stats policy, a rule the plan does not state. It keeps this phase opt-in.
  • search is a flag with a threshold that defaults to none and that nothing enforces; the plan has no server switch for search.
  • The memo holds the searched expression per (base, term) rather than the integer count, so the failure rule of _expr_count (a failed count is not cached) stays in one place.
  • The guard judges the source's row count, not the count of a search over it, so a sort that was allowed stays allowed while a user searches.
  • A refusal has error_info as well as error_code, because the existing client path acts on error_info and ignores the code.
  • search_expr is imported at the top of xorq_loading.py, where it was imported inside handle_infinite_request_xorq, to follow the repo's rule on imports.
  • Two tests-only commits instead of one (see Tests).

Not in this PR

  • The position cache of the L-track, the client revert of the sort model, default flips.
  • A search switch on the server. The flag is advice.
  • A sort guard that follows a searched count.

🤖 Generated with Claude Code

paddymul and others added 17 commits October 3, 2026 22:51
…the _handle_widget_change split (rows-first s1)

A default-tier XorqBuckarooWidget and BuckarooWidget publish df_data_dict,
then df_display_args, then the rest of the widget_args_tuple observers on a
search change, and merged_sd carries the full stat set. These pass on main
and pin the behaviour the split must keep.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…(rows-first p31)

New tests/unit/server/test_stats_policy.py for buckaroo/server/stats_policy.py,
which does not exist yet:

- resolve_stats_policy: a table of (backend, source_kind, rows, cols, host_tier)
  against (tier_target, auto_request, requestable, reason), including the
  boundaries of each threshold and the three tallyman entries over 70 s.
- The ceiling holds for host_tier="full", for a reload that re-resolves
  against a larger row count, and for a force request, and every tier a result
  lists as requestable is one a force would be granted.
- Probes add no data query: the parquet footer probe reads a small share of a
  counting file object and decodes no row group, the dtype probe never collects
  a LazyFrame or executes a xorq expression, and a known xorq count is an input.
- route_polars_entry returns "xorq" above R and "eager" at or below it.
- Threshold overrides from BUCKAROO_* environment variables.

Nothing imports the module yet. The tests error at fixture setup on the missing
module until the implementation lands.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… p31)

New buckaroo/server/stats_policy.py, pure and imported by nothing yet:

- resolve_stats_policy(backend, source_kind, rows, cols, bytes, host_tier,
  limits) returns tier_target, auto_request, requestable, reason and estimate
  over the tiers schema < scalar < full. A ceiling is computed inside the
  function, so every caller gets the lower of the requested and ceiling tiers
  with reason "ceiling". A host tier lowers freely and raises only to the
  ceiling. Eager pandas and polars resolve to full.
- route_polars_entry(rows, cols) returns "xorq" above R rows and "eager" at or
  below it.
- probe_dtypes reads a schema without collecting or executing anything, and
  probe_parquet_rows reads a parquet footer's num_rows without decoding data.
  A xorq count is an input to the policy, never computed by it.
- The thresholds are provisional constants gathered in StatsLimits, each with
  a BUCKAROO_* environment override read on every call.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…-handler fields (rows-first s1)

A schema-tier XorqServerDataflow should match full stats on pinned_rows,
data_key, summary_stats_key and (except the stats-derived minWidth)
column_config, issue no data query besides the cached count, and keep
init_sd hints and sorted windows working. A pending state must write
nothing under a full-tier cache key, and a later full assignment must reach
merged_sd for the raw, clean and filt scopes. assemble_merged_sd must equal
merged_sd, and _handle_widget_change must be built from separately callable
all_stats and display-args builders. /load_expr and /reload_expr accept
stats_tier and stats_delivery, replay them on reload, and keep them out of
the warm short-circuit's has_config tuple.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… s1)

Add a dataflow-level stats_tier ("full" default, "schema"). The xorq schema
tier builds identity and typing for every column from the expression's
schema, with no data query beyond the cached row count, so a dataflow
constructs in milliseconds rather than the stats' hundreds.

The tier is part of _scope_cache_key, so a schema entry is never read as a
full one, and _populate_sd_cache stores summary_sd under the filt key only
if it was computed for the current frame, klass list and tier. add_analysis
no longer builds DFStatsClass outside the hook when the tier is not full.
The merged_sd observer body is extracted as the pure assemble_merged_sd,
and _handle_widget_change is split into _build_df_data_dict and
_build_df_display_args.

/load_expr and /reload_expr accept stats_tier and stats_delivery, stored on
the session beside dataflow_kwargs and replayed on reload. They stay out of
the has_config tuple; the warm short-circuit compares the stored pair.
stats_delivery="deferred" builds the schema-tier dataflow and publishes it.
Defaults (full, inline) leave behaviour unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…esolve_stats_policy (rows-first p31)

bytes and source_kind were documented as validated but are not: bytes=-1,
'lots', 1.5 and object() are accepted, a np.int64 is echoed unchanged so
json.dumps of the result fails, a host tier passed positionally lands in the
bytes slot, and source_kind=None or 5 is accepted.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ation test (rows-first s1)

The Max Versions jobs resolve pandas 3, which reports a string column's
dtype as 'str' where pandas 2 says 'object'. The characterization test
asserted 'object'. Verified in a Max Versions environment (pandas 3.0.6,
polars 1.44.2, xorq 0.4.5): the unit suite passes.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…rows-first p31)

bytes goes through the same integer check as rows and cols, so a negative,
float or non-numeric value raises, a numpy integer is echoed as a plain int
and the result stays JSON-serialisable, and a host tier passed positionally
into the bytes slot raises instead of being echoed. source_kind must be a
str; its vocabulary stays open for the callers that will read it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…g on a skipped column (rows-first s1)

A column in skip_stat_columns gets only name, dtype and length from the
full-tier pipeline, so its _type comes from init_sd. The schema tier layers
the schema-derived _type and is_* keys over it, so an int64 column that
init_sd types as float merges as integer and renders with zero fraction
digits instead of the float displayer init_sd asked for.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ows-first s1)

_get_schema_sd never read skip_stat_columns, so a skipped column's
schema-derived _type and is_* keys overrode init_sd's _type once merged. The
full tier gives a skipped column only name, dtype and length. The schema tier
now does the same, so init_sd's _type decides the displayer at both tiers.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…er (rows-first s2)

ServerDataflow and PolarsServerDataflow with stats_tier="schema" should publish
the display state full stats give (column_config without stats-derived keys,
pinned_rows including a host-supplied one, data_key, summary_stats_key) with no
stat computed on the data, still apply init_sd, serve sorted windows, take a
later full assignment into merged_sd for every scope, and assemble to the same
sd as merged_sd. Both backends run through the same parametrized class, plus a
pandas test that pins how an object column is typed from its dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
ServerDataflow and PolarsServerDataflow now implement the _get_schema_sd hook,
so stats_tier="schema" builds a dataflow that types every column from its dtype
and runs no stat on the data. schema_sd (stat_pipeline.py) builds the sd the way
process_df shapes it: an empty frame gives {}, a skipped column keeps only its
names. pandas applies the existing typing_stats to a zero-row slice and derives
_type through the _type stat; polars factors pl_dtype_typing out of
pl_typing_stats and feeds it the dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…s (rows-first p31b)

The boundary tables now run on the module defaults and are written against
the phase-0 proposals: full auto up to 12M rows and 520M cells, scalar auto
up to 1.0B cells, full refused above 25M rows or 1.0B cells (force
included), a scalar ceiling of 4.0B cells, and polars routed to xorq above
8M rows. Each threshold has an equal, a one-below and a one-above row, in
rows and in cells. New tests pin each default to its literal value and
cover the two new environment overrides, BUCKAROO_STATS_FULL_AUTO_CELLS and
BUCKAROO_STATS_CEILING_FULL_CELLS.

They fail on the current constants (10M rows, 500M cells, 50M-row ceiling,
no scalar ceiling, R of 10M) and on the missing full_auto_cells and
ceiling_full_cells fields.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tats_request (rows-first s3)

On a deferred /load_expr session a stats_request {stats_gen, scope} should
return a stats_update with the matching stats_gen whose inline wide payload
equals the all_stats an inline session sends, a stale stats_gen should get
stats_aborted and run no query, and /load_expr and /reload_expr should bump the
generation. df_meta.stats should be injected on every frame and survive a
dataflow-field change, which returns the session to the schema tier. With a
caps client and a legacy client on one session, the legacy client should keep
getting complete messages through the websocket broadcast, the /load_expr,
/reload_expr, /load and /load_compare pushes and the highlight overlay, while
the caps client gets a stats-free frame and then pulls a stats_update. A spy
telemetry sink should see a stats.request span.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…measurements (rows-first p31b)

Full auto now needs at most 12M rows and 520M cells (was 10M rows), scalar
auto goes to a 1.0B cell budget (was 500M), `full` is refused above 25M rows
or 1.0B cells including for a forced request (was 50M rows), scalar gets a
4.0B cell ceiling by default (was unset; an extrapolation with no
measurement behind it), and polars routes to xorq above 8M rows (was 10M;
the figure assumes pre_limit False). Each DEFAULT_* constant carries the
measurement it comes from.

The two new bounds are full_auto_cells and ceiling_full_cells, appended to
StatsLimits, with BUCKAROO_STATS_FULL_AUTO_CELLS and
BUCKAROO_STATS_CEILING_FULL_CELLS as overrides under the same parsing rules.
_size_tier and _ceiling_tier read them; signatures and the result shape are
unchanged.

Three existing tests that hard-coded the old boundaries are updated: the
numpy-integer case (11M rows is now full), the reload case (a 40M-row entry
is now refused full) and the empty-environment case for the scalar ceiling
(an empty value now means the 4.0B default, not no ceiling).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…eration edge cases (rows-first s3)

Four more cases for the stats wire format, kept in their own commit so each is
seen failing on CI before the implementation lands. Returning to a state whose
stats were completed once is answered from summary_stats_cache with no query.
Completing the stats keeps the session's component_config on the refreshed
display config. A warm /load_expr, which rebuilds nothing, leaves stats_gen
alone. A stats_request on a session with no data is answered with
stats_aborted.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ed sessions (rows-first s3)

A client that advertises ?caps=stats_update gets a stats-free initial_state on a
deferred /load_expr session (df_meta.stats.status "pending") and pulls the stats
with stats_request {stats_gen, scope}. The reply is a stats_update carrying the
dataflow's all_stats as an inline wide envelope, or stats_aborted when the
generation is stale. The request is the whole run: one synchronous call that
computes the full stats, writes the full-tier summary_stats_cache entry, assigns
summary_sd and refreshes the session snapshot through one helper.

stats_gen is a server-owned counter bumped by every load handler and by a state
change that touches a dataflow field, which also returns a deferred session to
the schema tier. df_meta.stats is injected by build_state_message from the
session, since the dataflow rebuilds df_meta wholesale. Every send site goes
through build_state_message_for, so a client without the capability still gets
complete messages (its missing stats run synchronously first); broadcast_state
replaces the five copies of the send loop and sends to capable clients first.
The stats_request branch binds the session's telemetry sink and emits a
stats.request span.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

📦 TestPyPI package published

pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37494092809

or with uv:

uv pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37494092809

MCP server for Claude Code

claude mcp add buckaroo-table -- uvx --from "buckaroo[mcp]==0.15.9.dev37494092809" --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo-table

📖 Docs preview

🎨 Storybook preview

…al_state is dropped (#998) (#1014)

* test(server): failing tests for state_seq / reply_seq on buckaroo_state_change (#998)

Server: a buckaroo_state_change carrying state_seq gets reply_seq back on
the initial_state sent to the originating client (dataflow broadcast and
highlight overlay); another client's broadcast copy carries none.

Client: WebSocketModel attaches an incrementing state_seq to each change
it sends and drops an initial_state whose reply_seq is older than the
latest one sent; replies without reply_seq still apply.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(server): sequence token on buckaroo_state_change so a stale initial_state is dropped (#998)

WebSocketModel attaches an incrementing state_seq to each
buckaroo_state_change it sends. The handler echoes it as reply_seq on
the initial_state sent back to the originating client, on the dataflow
broadcast and on the highlight overlay; other clients' broadcast
copies, the /load push and a fresh connection carry none. The client
drops an initial_state whose reply_seq is older than the latest
state_seq it sent, so the reply for "alle" arriving after "allen" was
sent no longer reverts buckaroo_state, purges the grid and re-sends
the search.

reply_seq is optional on the wire; a client that sends no state_seq
gets today's behaviour, so PROTOCOL_VERSION stays at 1.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(server): failing test for a state_seq reply on a no-op buckaroo_state_change (#998)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* fix(server): answer a numbered no-op buckaroo_state_change so the client's newest seq gets a reply (#998)

The client drops a reply whose reply_seq is older than its latest state_seq. A
dataflow change followed at once by a show_commands/df_display/sampled toggle
got no reply for the toggle, so the dataflow reply was dropped and the new
display config never arrived. The server now sends a state ack carrying the
change's own buckaroo_state for any numbered change that touches no dataflow
field and no search term. Clients that send no state_seq see no new messages.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* test(server): failing tests for the error reply, the overlay's buckaroo_state and the live highlight (#998)

- A numbered change that fails gets only an error frame. The client has
  already dropped the previous reply as stale, so it is left on the data
  from before both changes.
- The highlight overlay carries reply_seq but builds buckaroo_state from
  the session's, which puts back show_commands and df_display.
- A dataflow broadcast drops each client's live-search highlight.
- Clearing the live term also strips the committed search's highlight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(server): answer a failed numbered change, echo the change's state in the overlay, keep each client's highlight (#998)

The highlight overlay and the no-op ack become one per-client reply,
_send_client_state. Every reply it sends carries the change's own
buckaroo_state, the session's current data and this client's live-search
highlight.

- A numbered change that fails gets that reply after the error frame,
  carrying reply_seq. The client has already dropped the reply to its
  previous change, so it was left on the data from before both. The
  reply keeps the failed change's buckaroo_state: the session's would
  revert a search box holding the failed term, and the box would send it
  again.
- The overlay no longer builds buckaroo_state from the session's, which
  put back show_commands and df_display.
- Each client's copy of a dataflow broadcast carries its own live-search
  highlight.
- An empty live term leaves df_display_args as the session has it, so the
  committed quick_command_args.search highlight stays.
- The reply no longer skips an empty df_display_args, so every numbered
  change that reaches the dataflow is answered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
paddymul and others added 4 commits October 6, 2026 11:57
…ts-tier-core

# Conflicts:
#	buckaroo/server/handlers.py
…at/rowsfirst-s3-stats-wire-server

Brings in #1004 (telemetry.arm_session) and #1014 (state_seq / reply_seq).

Conflict resolution:
- broadcast_state takes reply_to, reply_seq and highlight, so the
  state-change broadcast keeps #1014's reply_seq for the originating client
  and per-client live-search highlight while staying capable-first and
  completing stats for legacy clients.
- build_state_message_for passes reply_seq through.
- #1014's _send_client_state replaces _send_highlight_overlay and builds its
  message with build_state_message_for, so a legacy client's stats are
  completed before the highlight is applied.
- LoadCompareHandler keeps both arm_session and the stats reset.
- The overlay-order test calls _send_client_state; a new test covers a
  numbered change on a deferred session (reply_seq to the sender only).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
paddymul and others added 6 commits October 6, 2026 12:05
…handlers (rows-first p33)

stats_tier takes auto, full, scalar or schema on /load_expr and /reload_expr
and is stored with the pair. The policy resolves after the schema-tier
dataflow and the count exist, is stored on the session, is reported in
df_meta.stats (tier_target, reason, auto_request, requestable, estimate,
omitted_keys, approx_keys, demand_columns, each with its documented default)
and is applied at WebSocket open only for a client that sends
?caps=stats_update,stats_ondemand. The version-skew cases (no caps,
stats_update only, both bits, an old server) and /load, which keeps
resolving to full, are covered.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…d the compare reset under a stats policy (rows-first p33)

Cases found untested after the first tests commit: a scalar tier named with
inline delivery builds a schema dataflow, and /load_compare clears the policy
a session held.

Rebased off the closed unit PRs (#1026, #1028): the assertions on unit tiers
and on df_display_args in the final reply are dropped with the code they
tested.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…irst p33)

/load_expr and /reload_expr accept stats_tier auto, full, scalar or schema,
stored with the pair and kept out of has_config, so the warm short-circuit
holds. When the dataflow is built at the schema tier the handler resolves the
policy (resolve_stats_policy) from the count it already has, stores it on the
session and starts the stats generation from it: a target below full is
not_computed with the policy's reason.

df_meta.stats reports the policy as tier_target and estimate, plus
auto_request, requestable, omitted_keys, approx_keys and demand_columns where
they differ from their documented defaults. It is applied at WebSocket open
only for a client that sends ?caps=stats_update,stats_ondemand. A client with
stats_update only is told the session is pending and pulls the stats, and a
client with no caps gets them at connect, as for any deferred session. A tier
the host named (scalar, schema) reaches every client as before. A session on
an explicit stats_tier full within the ceiling sends the message it always has.

/load keeps resolving to full: it does not read the field and clears a policy
left on the session.

Rebased onto #1024 without the unit PRs (#1026, #1028): stats_request is the
whole run of #1024, so handle_stats_request now takes the client to decide
stats_to_pull, and the incremental and display-config parts are gone.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…arched count memo (rows-first p37)

Above the sort threshold every column of the grid served by infinite_request
has ag_grid_specs sortable false, including a column a klass or a host
override styles, and a sorted infinite_request gets error_code sort_disabled.
df_meta carries sort and search flags. A repeated searched xorq window issues
one count. The thresholds come from the stats policy module.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…count memo (rows-first p37)

A regression check on the first tests commit left cases that no test pinned.
A searched expression a base holds is bounded and the least recently asked
term goes first, a /load_compare ends the guard like /load does, and the pass
that turns sorting off leaves a column entry that is not a dict as it is and
replaces ag_grid_specs that is not a dict. The wire test of a state change
reads the three rows first_three leaves, not five.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…count memo (rows-first p37)

stats_policy gains GuardLimits (BUCKAROO_SORT_DISABLE_ROWS, default 25M rows;
BUCKAROO_SEARCH_DISABLE_ROWS, default none) and resolve_source_guards, apart from
the stats thresholds. A session a host opened with a stats policy resolves them
at /load_expr and /reload_expr from the count load took. Above the sort
threshold the xorq dataflow finishes its display config with disable_sorting, so
no klass or override leaves a sortable column in a display infinite_request
serves, and a sorted infinite_request is refused with error_code sort_disabled
before any query runs. df_meta carries sort and search when one is disabled.

handle_infinite_request_xorq holds the searched expression per (base expression,
term), so the second window of a search is a hit in _expr_count's cache and
issues no count.

Rebased onto #1029 without the closed #1031 and #1033: the guard resets
no longer call reset_stats_controls, session_dataflow stays in stats_wire,
and the wire tests share a _LimitsWire base ported from #1033 without its
unit and cost helpers.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@paddymul
paddymul force-pushed the feat/rowsfirst-p37-huge-source-guards branch from 0181386 to d07a7f2 Compare October 6, 2026 16:14

This branch was successfully deployed

1 active deployment
testpypi — d07a7f22 Deployed Oct 6, 2026 by paddymul via Publish to TestPyPI #1752
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant