Skip to content

feat(server): pandas and polars schema stats tier (rows-first s2) - #1022

Draft
paddymul wants to merge 10 commits into
adr-002-rows-first-stats-deliveryfrom
feat/rowsfirst-s2-schema-tier-pandas-polars
Draft

paddymul wants to merge 10 commits into
adr-002-rows-first-stats-deliveryfrom
feat/rowsfirst-s2-schema-tier-pandas-polars

Conversation

@paddymul

@paddymul paddymul commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #1021 (feat/rowsfirst-s1-stats-tier-core). This PR is based on main so the repo's Checks workflow runs on it (its pull_request trigger uses branches: "*", which does not match a base branch containing a slash). The diff therefore includes #1021's commits until that PR merges; the commits of this phase are 9cac279a (failing tests) and 334b7367 (implementation).

Problem

#1021 adds a dataflow-level stats_tier (full | schema) and implements the schema tier for xorq. ServerDataflow (pandas) and PolarsServerDataflow accept stats_tier="schema" but raise NotImplementedError, because the base _get_schema_sd hook has no override for them. Plan 1 lists the schema tier for these two dataflows as unrun: "Pandas and polars are unrun [I]" (section 3, constraint 1) and "schema-tier equivalence on pandas and polars" under "What was not verified" (section 12). This PR builds it and runs that comparison.

Phase and plan references

Rows-first s2, from buckaroo2-reports/plans/: plan 1 (01-rows-first-stats-separate-plumbing.md) section 9 "Phase 1", the pandas and polars halves, with sections 3 (constraint 1) and 4.0; plan 2 (02-rows-first-xorq-and-lazy-polars.md) sections 3 and 4.1 for the cache and final-assignment constraints; plan 3 (03-no-summary-stats-for-large-files.md) for why /load is untouched: eager backends resolve to full, and /load takes the field only on the lazy track.

Approach

  • A new schema_sd(df, column_typing, skip_columns=None) in stat_pipeline.py, next to StatPipeline.process_df, builds the sd the pipeline would give for the keys a column's schema determines. Each entry has orig_col_name, rewritten_col_name, length and whatever column_typing(series) returns. It follows process_df where that shapes the result: a frame with no rows gives {}, and a column in skip_columns gets only its two names, so its typing comes from init_sd.
  • pandas: schema_stats(ser) in pd_stats_v2.py runs the existing typing_stats on a zero-row slice of the column (ser.iloc[:0]), drops memory_usage (it measures the data), and derives _type through the existing _type stat (with_type(flags)). No value is read.
  • polars: pl_dtype_typing(dt) is pl_typing_stats's body without memory_usage, factored out so it needs a dtype and not a series. pl_schema_stats(ser) feeds it ser.dtype and with_type. pl_typing_stats now calls it and returns the same dict as before.
  • ServerDataflow._get_schema_sd and PolarsServerDataflow._get_schema_sd call schema_sd with those typing functions and the dataflow's skip_stat_columns.
  • The tier machinery from feat(xorq): stats tier machinery and the xorq schema tier (rows-first s1) #1021 is unchanged: the tier is in _scope_cache_key, add_analysis skips DFStatsClass off the full tier, and a later full assignment reaches merged_sd through the final-assignment protocol it documents.
  • /load does not take stats_tier, so there is no handler change. The tier is reachable through the constructors, as it is through XorqServerDataflow in feat(xorq): stats tier machinery and the xorq schema tier (rows-first s1) #1021 before its handler fields.

What changes

  • buckaroo/pluggable_analysis_framework/stat_pipeline.py: schema_sd.
  • buckaroo/customizations/pd_stats_v2.py: with_type, schema_stats.
  • buckaroo/customizations/pl_stats_v2.py: pl_dtype_typing, pl_schema_stats; pl_typing_stats uses pl_dtype_typing.
  • buckaroo/server/data_loading.py, buckaroo/server/data_loading_polars.py: _get_schema_sd on ServerDataflow and PolarsServerDataflow.

Tests

Two commits: failing tests (9cac279), then the fix (334b736). All new tests are in tests/unit/server/test_data_loading_polars.py, a class TestStatsTierSchema parametrized over the two backends so each assertion runs on pandas and polars, plus one pandas-only test. The file already holds the polars loader tests and pins parity with the pandas path; there is no unit test module for the pandas server dataflow, and splitting one class across two files would hide that the two backends are held to the same assertions.

  • test_matches_full_stats_display_state: column_config without minWidth, pinned_rows, data_key and summary_stats_key equal full stats for every display.
  • test_host_supplied_pinned_rows_match_full_stats: a host pinned_rows that names a stat the schema tier does not compute (mean) is the same in both tiers.
  • test_min_width_is_the_stats_derived_difference: minWidth is the one key that differs, for a float with a 1e9 maximum.
  • test_typing_matches_full_stats_for_every_dtype: on a frame with one column per dtype kind (pandas: float, int, nullable int, bool, string, datetime, tz-aware datetime, timedelta, category; polars: float, int32, bool, string, date, datetime, duration, time, decimal, binary, categorical) every schema key equals the full-tier value.
  • test_publishes_no_data_derived_stat: no data stat in merged_sd, and the all_stats wire payload (decoded) holds only dtype. The same check on the full tier sees histogram_bins, so it can see them.
  • test_computes_no_stat_on_the_data: a spy on StatPipeline.process_df sees no call on the loaded frame through construction, a search and add_analysis. It ignores PERVERSE_DF, which the pipeline's DAG self-check runs and which holds none of the loaded data. test_the_spy_sees_the_pipeline_at_the_full_tier is the control, and passes on the base branch by design.
  • test_init_sd_hints_and_overrides_still_apply, test_skipped_column_keeps_init_sd_typing, test_empty_frame_matches_full_stats.
  • test_sorted_infinite_request_works: a descending sort on the schema-tier dataflow returns the right rows through each backend's real infinite_request handler.
  • test_pending_state_writes_no_full_tier_cache_key, test_later_full_assignment_reaches_merged_sd_for_all_scopes, test_full_scope_sds_cached_first_are_used_without_the_pipeline: the cache and final-assignment protocol on pandas (raw, clean and filt scopes: a user op and a search) and polars (raw and clean are one scope, because the polars conf ships no command to run outside a search).
  • test_assembled_sd_equals_merged_sd_with_init_sd_and_a_user_op: assemble_merged_sd over the schema-tier scope sds equals merged_sd, with init_sd overrides and a filter layer (and a cleaning layer on pandas).
  • test_object_column_is_typed_as_a_string_at_the_schema_tier (pandas only): pins the object-dtype behaviour described under Deviations.

On the failing-tests commit all 29 new tests that exercise the schema tier fail locally with NotImplementedError: ... has no schema stats tier, and the two control tests pass. On CI all nine Python / Test jobs failed on that commit (3.11 to 3.14, Max Versions 3.11 to 3.14, Windows). CI's annotation only says the step exited 1, so the reason rests on the local run of the same commit. On the fix commit every job of the Checks run completed: 23 succeeded and Publish to TestPyPI was skipped, including all nine Python / Test jobs, Python / Lint, the JS job and the Playwright jobs.

Locally the full unit suite (pytest ./tests/unit -m "not slow") gives 1245 passed and 5 skipped on pandas 2.2.3 / polars 1.35.2, and the same on a Max Versions environment built the way CI builds it (pandas 3.0.6, polars 1.44.2, no lock file).

CI does not run on this PR by itself. checks.yml triggers on pull_request with branches: "*", which does not match a base branch whose name contains /, so a PR stacked on feat/rowsfirst-s1-stats-tier-core shows only the Read the Docs check (so does #991, based on fix/988-column-order; every PR based on main has 28). I ran Checks with workflow_dispatch on this branch for each commit: run 1673 on the tests commit and run 1675 on the fix commit (run 1674 is an accidental second dispatch on the tests commit). Their jobs are on the Actions tab and do not appear in this PR's checks list. The trigger is not changed here.

One test needed untied data. test_later_full_assignment_reaches_merged_sd_for_all_scopes recomputes full stats on an independently built dataflow and compares the result with a second full-tier dataflow's. On polars the two differ in mode whenever two values of a column have the same count, because pl.Series.mode() returns the tied values in no fixed order: 30 rebuilds of a 5-row frame with ties, compared with the first, gave a different mode for column a 24 times and none matched in all three columns, while 30 of 30 matched on a frame where every value of every column has a different count. That test uses the second kind of frame. The other tests compare only keys the schema tier owns or reuse one computed sd, so they are not affected.

Measurements

Construction of the dataflow (skip_main_serial=True, best of three, in process) on synthetic frames of floats and 500-word strings, and the size of the all_stats payload in df_data_dict (base64 of the wide parquet):

Frame Backend Full Schema all_stats, full to schema
52,814 x 27 pandas 225.5 ms 8.5 ms 226,884 B to 9,632 B
52,814 x 27 polars 80.7 ms 2.8 ms 226,892 B to 9,632 B
200,000 x 20 pandas 190.6 ms 7.5 ms 165,772 B to 7,204 B
200,000 x 20 polars 60.2 ms 2.5 ms 165,676 B to 7,204 B

Above one million cells both pipelines downsample to 50,000 rows, which is why the 200,000-row frame costs about what the 52,814-row one does. Eager pandas and polars stay at full and inline (plans 1 and 3), so these figures are what a schema tier could save on an eager backend, not a change to how /load runs.

Why default behaviour is unchanged

  • stats_tier defaults to full, and _get_summary_sd only reaches _get_schema_sd when it is schema; nothing sets schema outside the constructors.
  • pl_typing_stats returns the same dict in the same key order as before; pl_dtype_typing is its body.
  • schema_sd, schema_stats, with_type and pl_schema_stats have no caller on the full tier.
  • The existing tests pass unchanged (see Tests for the full-suite runs).

Deviations from the plan

  • The plan routes polars typing through pl_dtype_typing (plan 2 section 5.4), as if it existed. It exists only on the unmerged perf(server): hold a LazyFrame for backend=polars /load sessions (#993) #1011 branch (feat/993-lazy-polars-scan), so this PR adds it to pl_stats_v2.py with the same signature and body (a dtype in, the flags out, no memory_usage) so the two merge without a semantic conflict. Nothing here depends on perf(server): hold a LazyFrame for backend=polars /load sessions (#993) #1011.
  • "Existing typing stat functions applied to dtype only": typing_stats takes a series, so it runs on ser.iloc[:0], which carries the dtype and no values. The result drops memory_usage.
  • An object column cannot be typed from its dtype. pandas' is_string_dtype inspects the values of an object series and calls an empty one a string, so the schema tier types every object column as string. Full stats may say obj for a column of mixed values (the test pins both). String columns, the common case, agree. Typing them by reading values would put a data scan back into the tier. polars and pandas 3's str dtype have no such gap.
  • Equal column_config holds except ag_grid_specs.minWidth, as on xorq: the width of an integer or float column reads its min and max. Other klasses that read stats at style time will differ the same way.
  • length is len(processed_df). At the full tier it is the length of the stats frame, which both pipelines downsample to 50,000 rows once rows times columns exceeds FAST_SUMMARY_WHEN_GREATER. The schema tier reports the real row count there.
  • A column in skip_stat_columns gets only orig_col_name and rewritten_col_name, which is what the pandas and polars pipeline gives it, and not dtype and length as the xorq tier does (the xorq pipeline gives those).
  • create_dataflow and create_polars_dataflow, the helpers /load calls, do not take stats_tier; the scope says /load does not expose it.

Not in this PR

Any /load request field; wire messages and df_meta.stats; units; the xorq half (#1021); lazy polars. Selecting a cleaning_method still runs the cleaning analyses' own stat pass (PandasAutocleaning._run_cleaning) at the schema tier; it is a separate stat pipeline, not the summary stats this tier skips.

🤖 Generated with Claude Code

paddymul and others added 8 commits October 3, 2026 22:51
…the _handle_widget_change split (rows-first s1)

A default-tier XorqBuckarooWidget and BuckarooWidget publish df_data_dict,
then df_display_args, then the rest of the widget_args_tuple observers on a
search change, and merged_sd carries the full stat set. These pass on main
and pin the behaviour the split must keep.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…-handler fields (rows-first s1)

A schema-tier XorqServerDataflow should match full stats on pinned_rows,
data_key, summary_stats_key and (except the stats-derived minWidth)
column_config, issue no data query besides the cached count, and keep
init_sd hints and sorted windows working. A pending state must write
nothing under a full-tier cache key, and a later full assignment must reach
merged_sd for the raw, clean and filt scopes. assemble_merged_sd must equal
merged_sd, and _handle_widget_change must be built from separately callable
all_stats and display-args builders. /load_expr and /reload_expr accept
stats_tier and stats_delivery, replay them on reload, and keep them out of
the warm short-circuit's has_config tuple.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… s1)

Add a dataflow-level stats_tier ("full" default, "schema"). The xorq schema
tier builds identity and typing for every column from the expression's
schema, with no data query beyond the cached row count, so a dataflow
constructs in milliseconds rather than the stats' hundreds.

The tier is part of _scope_cache_key, so a schema entry is never read as a
full one, and _populate_sd_cache stores summary_sd under the filt key only
if it was computed for the current frame, klass list and tier. add_analysis
no longer builds DFStatsClass outside the hook when the tier is not full.
The merged_sd observer body is extracted as the pure assemble_merged_sd,
and _handle_widget_change is split into _build_df_data_dict and
_build_df_display_args.

/load_expr and /reload_expr accept stats_tier and stats_delivery, stored on
the session beside dataflow_kwargs and replayed on reload. They stay out of
the has_config tuple; the warm short-circuit compares the stored pair.
stats_delivery="deferred" builds the schema-tier dataflow and publishes it.
Defaults (full, inline) leave behaviour unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ation test (rows-first s1)

The Max Versions jobs resolve pandas 3, which reports a string column's
dtype as 'str' where pandas 2 says 'object'. The characterization test
asserted 'object'. Verified in a Max Versions environment (pandas 3.0.6,
polars 1.44.2, xorq 0.4.5): the unit suite passes.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…g on a skipped column (rows-first s1)

A column in skip_stat_columns gets only name, dtype and length from the
full-tier pipeline, so its _type comes from init_sd. The schema tier layers
the schema-derived _type and is_* keys over it, so an int64 column that
init_sd types as float merges as integer and renders with zero fraction
digits instead of the float displayer init_sd asked for.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ows-first s1)

_get_schema_sd never read skip_stat_columns, so a skipped column's
schema-derived _type and is_* keys overrode init_sd's _type once merged. The
full tier gives a skipped column only name, dtype and length. The schema tier
now does the same, so init_sd's _type decides the displayer at both tiers.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…er (rows-first s2)

ServerDataflow and PolarsServerDataflow with stats_tier="schema" should publish
the display state full stats give (column_config without stats-derived keys,
pinned_rows including a host-supplied one, data_key, summary_stats_key) with no
stat computed on the data, still apply init_sd, serve sorted windows, take a
later full assignment into merged_sd for every scope, and assemble to the same
sd as merged_sd. Both backends run through the same parametrized class, plus a
pandas test that pins how an object column is typed from its dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
ServerDataflow and PolarsServerDataflow now implement the _get_schema_sd hook,
so stats_tier="schema" builds a dataflow that types every column from its dtype
and runs no stat on the data. schema_sd (stat_pipeline.py) builds the sd the way
process_df shapes it: an empty frame gives {}, a skipped column keeps only its
names. pandas applies the existing typing_stats to a zero-row slice and derives
_type through the _type stat; polars factors pl_dtype_typing out of
pl_typing_stats and feeds it the dtype.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@paddymul
paddymul changed the base branch from feat/rowsfirst-s1-stats-tier-core to main October 4, 2026 05:06
@paddymul

paddymul commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Closing and reopening to start the Checks workflow now that the base is main.

@paddymul paddymul closed this Oct 4, 2026
@paddymul paddymul reopened this Oct 4, 2026
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

📦 TestPyPI package published

pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37178920905

or with uv:

uv pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37178920905

MCP server for Claude Code

claude mcp add buckaroo-table -- uvx --from "buckaroo[mcp]==0.15.9.dev37178920905" --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo-table

📖 Docs preview

🎨 Storybook preview

@paddymul
paddymul force-pushed the adr-002-rows-first-stats-delivery branch from c500780 to f61df3d Compare October 6, 2026 18:27

This branch was successfully deployed

1 active (outdated) deployment
testpypi — 334b7367 Deployed Oct 4, 2026 by paddymul via Publish to TestPyPI #1676
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant