Repository navigation
Conversation
…the _handle_widget_change split (rows-first s1) A default-tier XorqBuckarooWidget and BuckarooWidget publish df_data_dict, then df_display_args, then the rest of the widget_args_tuple observers on a search change, and merged_sd carries the full stat set. These pass on main and pin the behaviour the split must keep. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…-handler fields (rows-first s1) A schema-tier XorqServerDataflow should match full stats on pinned_rows, data_key, summary_stats_key and (except the stats-derived minWidth) column_config, issue no data query besides the cached count, and keep init_sd hints and sorted windows working. A pending state must write nothing under a full-tier cache key, and a later full assignment must reach merged_sd for the raw, clean and filt scopes. assemble_merged_sd must equal merged_sd, and _handle_widget_change must be built from separately callable all_stats and display-args builders. /load_expr and /reload_expr accept stats_tier and stats_delivery, replay them on reload, and keep them out of the warm short-circuit's has_config tuple. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… s1)
Add a dataflow-level stats_tier ("full" default, "schema"). The xorq schema
tier builds identity and typing for every column from the expression's
schema, with no data query beyond the cached row count, so a dataflow
constructs in milliseconds rather than the stats' hundreds.
The tier is part of _scope_cache_key, so a schema entry is never read as a
full one, and _populate_sd_cache stores summary_sd under the filt key only
if it was computed for the current frame, klass list and tier. add_analysis
no longer builds DFStatsClass outside the hook when the tier is not full.
The merged_sd observer body is extracted as the pure assemble_merged_sd,
and _handle_widget_change is split into _build_df_data_dict and
_build_df_display_args.
/load_expr and /reload_expr accept stats_tier and stats_delivery, stored on
the session beside dataflow_kwargs and replayed on reload. They stay out of
the has_config tuple; the warm short-circuit compares the stored pair.
stats_delivery="deferred" builds the schema-tier dataflow and publishes it.
Defaults (full, inline) leave behaviour unchanged.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ation test (rows-first s1) The Max Versions jobs resolve pandas 3, which reports a string column's dtype as 'str' where pandas 2 says 'object'. The characterization test asserted 'object'. Verified in a Max Versions environment (pandas 3.0.6, polars 1.44.2, xorq 0.4.5): the unit suite passes. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…g on a skipped column (rows-first s1) A column in skip_stat_columns gets only name, dtype and length from the full-tier pipeline, so its _type comes from init_sd. The schema tier layers the schema-derived _type and is_* keys over it, so an int64 column that init_sd types as float merges as integer and renders with zero fraction digits instead of the float displayer init_sd asked for. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ows-first s1) _get_schema_sd never read skip_stat_columns, so a skipped column's schema-derived _type and is_* keys overrode init_sd's _type once merged. The full tier gives a skipped column only name, dtype and length. The schema tier now does the same, so init_sd's _type decides the displayer at both tiers. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…er (rows-first s2) ServerDataflow and PolarsServerDataflow with stats_tier="schema" should publish the display state full stats give (column_config without stats-derived keys, pinned_rows including a host-supplied one, data_key, summary_stats_key) with no stat computed on the data, still apply init_sd, serve sorted windows, take a later full assignment into merged_sd for every scope, and assemble to the same sd as merged_sd. Both backends run through the same parametrized class, plus a pandas test that pins how an object column is typed from its dtype. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
ServerDataflow and PolarsServerDataflow now implement the _get_schema_sd hook,
so stats_tier="schema" builds a dataflow that types every column from its dtype
and runs no stat on the data. schema_sd (stat_pipeline.py) builds the sd the way
process_df shapes it: an empty frame gives {}, a skipped column keeps only its
names. pandas applies the existing typing_stats to a zero-row slice and derives
_type through the _type stat; polars factors pl_dtype_typing out of
pl_typing_stats and feeds it the dtype.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
paddymul
changed the base branch from
feat/rowsfirst-s1-stats-tier-core
to
main
October 4, 2026 05:06
Collaborator
Author
|
Closing and reopening to start the Checks workflow now that the base is main. |
Contributor
📦 TestPyPI package publishedpip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37178920905or with uv: uv pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.9.dev37178920905MCP server for Claude Codeclaude mcp add buckaroo-table -- uvx --from "buckaroo[mcp]==0.15.9.dev37178920905" --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo-table📖 Docs preview🎨 Storybook preview |
This was referenced Oct 4, 2026
Closed
Draft
paddymul
changed the base branch from
main
to
adr-002-rows-first-stats-delivery
October 6, 2026 14:40
…ts-tier-core # Conflicts: # buckaroo/server/handlers.py
…s2-schema-tier-pandas-polars
paddymul
force-pushed
the
adr-002-rows-first-stats-delivery
branch
from
October 6, 2026 18:27
c500780 to
f61df3d
Compare
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1021 (
feat/rowsfirst-s1-stats-tier-core). This PR is based onmainso the repo's Checks workflow runs on it (itspull_requesttrigger usesbranches: "*", which does not match a base branch containing a slash). The diff therefore includes #1021's commits until that PR merges; the commits of this phase are9cac279a(failing tests) and334b7367(implementation).Problem
#1021 adds a dataflow-level
stats_tier(full|schema) and implements the schema tier for xorq.ServerDataflow(pandas) andPolarsServerDataflowacceptstats_tier="schema"but raiseNotImplementedError, because the base_get_schema_sdhook has no override for them. Plan 1 lists the schema tier for these two dataflows as unrun: "Pandas and polars are unrun [I]" (section 3, constraint 1) and "schema-tier equivalence on pandas and polars" under "What was not verified" (section 12). This PR builds it and runs that comparison.Phase and plan references
Rows-first s2, from
buckaroo2-reports/plans/: plan 1 (01-rows-first-stats-separate-plumbing.md) section 9 "Phase 1", the pandas and polars halves, with sections 3 (constraint 1) and 4.0; plan 2 (02-rows-first-xorq-and-lazy-polars.md) sections 3 and 4.1 for the cache and final-assignment constraints; plan 3 (03-no-summary-stats-for-large-files.md) for why/loadis untouched: eager backends resolve tofull, and/loadtakes the field only on the lazy track.Approach
schema_sd(df, column_typing, skip_columns=None)instat_pipeline.py, next toStatPipeline.process_df, builds the sd the pipeline would give for the keys a column's schema determines. Each entry hasorig_col_name,rewritten_col_name,lengthand whatevercolumn_typing(series)returns. It followsprocess_dfwhere that shapes the result: a frame with no rows gives{}, and a column inskip_columnsgets only its two names, so its typing comes frominit_sd.schema_stats(ser)inpd_stats_v2.pyruns the existingtyping_statson a zero-row slice of the column (ser.iloc[:0]), dropsmemory_usage(it measures the data), and derives_typethrough the existing_typestat (with_type(flags)). No value is read.pl_dtype_typing(dt)ispl_typing_stats's body withoutmemory_usage, factored out so it needs a dtype and not a series.pl_schema_stats(ser)feeds itser.dtypeandwith_type.pl_typing_statsnow calls it and returns the same dict as before.ServerDataflow._get_schema_sdandPolarsServerDataflow._get_schema_sdcallschema_sdwith those typing functions and the dataflow'sskip_stat_columns._scope_cache_key,add_analysisskipsDFStatsClassoff the full tier, and a later full assignment reachesmerged_sdthrough the final-assignment protocol it documents./loaddoes not takestats_tier, so there is no handler change. The tier is reachable through the constructors, as it is throughXorqServerDataflowin feat(xorq): stats tier machinery and the xorq schema tier (rows-first s1) #1021 before its handler fields.What changes
buckaroo/pluggable_analysis_framework/stat_pipeline.py:schema_sd.buckaroo/customizations/pd_stats_v2.py:with_type,schema_stats.buckaroo/customizations/pl_stats_v2.py:pl_dtype_typing,pl_schema_stats;pl_typing_statsusespl_dtype_typing.buckaroo/server/data_loading.py,buckaroo/server/data_loading_polars.py:_get_schema_sdonServerDataflowandPolarsServerDataflow.Tests
Two commits: failing tests (9cac279), then the fix (334b736). All new tests are in
tests/unit/server/test_data_loading_polars.py, a classTestStatsTierSchemaparametrized over the two backends so each assertion runs on pandas and polars, plus one pandas-only test. The file already holds the polars loader tests and pins parity with the pandas path; there is no unit test module for the pandas server dataflow, and splitting one class across two files would hide that the two backends are held to the same assertions.test_matches_full_stats_display_state:column_configwithoutminWidth,pinned_rows,data_keyandsummary_stats_keyequal full stats for every display.test_host_supplied_pinned_rows_match_full_stats: a hostpinned_rowsthat names a stat the schema tier does not compute (mean) is the same in both tiers.test_min_width_is_the_stats_derived_difference:minWidthis the one key that differs, for a float with a 1e9 maximum.test_typing_matches_full_stats_for_every_dtype: on a frame with one column per dtype kind (pandas: float, int, nullable int, bool, string, datetime, tz-aware datetime, timedelta, category; polars: float, int32, bool, string, date, datetime, duration, time, decimal, binary, categorical) every schema key equals the full-tier value.test_publishes_no_data_derived_stat: no data stat inmerged_sd, and theall_statswire payload (decoded) holds onlydtype. The same check on the full tier seeshistogram_bins, so it can see them.test_computes_no_stat_on_the_data: a spy onStatPipeline.process_dfsees no call on the loaded frame through construction, a search andadd_analysis. It ignoresPERVERSE_DF, which the pipeline's DAG self-check runs and which holds none of the loaded data.test_the_spy_sees_the_pipeline_at_the_full_tieris the control, and passes on the base branch by design.test_init_sd_hints_and_overrides_still_apply,test_skipped_column_keeps_init_sd_typing,test_empty_frame_matches_full_stats.test_sorted_infinite_request_works: a descending sort on the schema-tier dataflow returns the right rows through each backend's realinfinite_requesthandler.test_pending_state_writes_no_full_tier_cache_key,test_later_full_assignment_reaches_merged_sd_for_all_scopes,test_full_scope_sds_cached_first_are_used_without_the_pipeline: the cache and final-assignment protocol on pandas (raw, clean and filt scopes: a user op and a search) and polars (raw and clean are one scope, because the polars conf ships no command to run outside a search).test_assembled_sd_equals_merged_sd_with_init_sd_and_a_user_op:assemble_merged_sdover the schema-tier scope sds equalsmerged_sd, withinit_sdoverrides and a filter layer (and a cleaning layer on pandas).test_object_column_is_typed_as_a_string_at_the_schema_tier(pandas only): pins the object-dtype behaviour described under Deviations.On the failing-tests commit all 29 new tests that exercise the schema tier fail locally with
NotImplementedError: ... has no schema stats tier, and the two control tests pass. On CI all ninePython / Testjobs failed on that commit (3.11 to 3.14, Max Versions 3.11 to 3.14, Windows). CI's annotation only says the step exited 1, so the reason rests on the local run of the same commit. On the fix commit every job of theChecksrun completed: 23 succeeded andPublish to TestPyPIwas skipped, including all ninePython / Testjobs,Python / Lint, the JS job and the Playwright jobs.Locally the full unit suite (
pytest ./tests/unit -m "not slow") gives 1245 passed and 5 skipped on pandas 2.2.3 / polars 1.35.2, and the same on a Max Versions environment built the way CI builds it (pandas 3.0.6, polars 1.44.2, no lock file).CI does not run on this PR by itself.
checks.ymltriggers onpull_requestwithbranches: "*", which does not match a base branch whose name contains/, so a PR stacked onfeat/rowsfirst-s1-stats-tier-coreshows only the Read the Docs check (so does #991, based onfix/988-column-order; every PR based onmainhas 28). I ranCheckswithworkflow_dispatchon this branch for each commit: run 1673 on the tests commit and run 1675 on the fix commit (run 1674 is an accidental second dispatch on the tests commit). Their jobs are on the Actions tab and do not appear in this PR's checks list. The trigger is not changed here.One test needed untied data.
test_later_full_assignment_reaches_merged_sd_for_all_scopesrecomputes full stats on an independently built dataflow and compares the result with a second full-tier dataflow's. On polars the two differ inmodewhenever two values of a column have the same count, becausepl.Series.mode()returns the tied values in no fixed order: 30 rebuilds of a 5-row frame with ties, compared with the first, gave a differentmodefor columna24 times and none matched in all three columns, while 30 of 30 matched on a frame where every value of every column has a different count. That test uses the second kind of frame. The other tests compare only keys the schema tier owns or reuse one computed sd, so they are not affected.Measurements
Construction of the dataflow (
skip_main_serial=True, best of three, in process) on synthetic frames of floats and 500-word strings, and the size of theall_statspayload indf_data_dict(base64 of the wide parquet):all_stats, full to schemaAbove one million cells both pipelines downsample to 50,000 rows, which is why the 200,000-row frame costs about what the 52,814-row one does. Eager pandas and polars stay at
fulland inline (plans 1 and 3), so these figures are what a schema tier could save on an eager backend, not a change to how/loadruns.Why default behaviour is unchanged
stats_tierdefaults tofull, and_get_summary_sdonly reaches_get_schema_sdwhen it isschema; nothing setsschemaoutside the constructors.pl_typing_statsreturns the same dict in the same key order as before;pl_dtype_typingis its body.schema_sd,schema_stats,with_typeandpl_schema_statshave no caller on the full tier.Deviations from the plan
pl_dtype_typing(plan 2 section 5.4), as if it existed. It exists only on the unmerged perf(server): hold a LazyFrame for backend=polars /load sessions (#993) #1011 branch (feat/993-lazy-polars-scan), so this PR adds it topl_stats_v2.pywith the same signature and body (a dtype in, the flags out, nomemory_usage) so the two merge without a semantic conflict. Nothing here depends on perf(server): hold a LazyFrame for backend=polars /load sessions (#993) #1011.typing_statstakes a series, so it runs onser.iloc[:0], which carries the dtype and no values. The result dropsmemory_usage.objectcolumn cannot be typed from its dtype. pandas'is_string_dtypeinspects the values of an object series and calls an empty one a string, so the schema tier types every object column asstring. Full stats may sayobjfor a column of mixed values (the test pins both). String columns, the common case, agree. Typing them by reading values would put a data scan back into the tier. polars and pandas 3'sstrdtype have no such gap.column_configholds exceptag_grid_specs.minWidth, as on xorq: the width of an integer or float column reads itsminandmax. Other klasses that read stats at style time will differ the same way.lengthislen(processed_df). At the full tier it is the length of the stats frame, which both pipelines downsample to 50,000 rows once rows times columns exceedsFAST_SUMMARY_WHEN_GREATER. The schema tier reports the real row count there.skip_stat_columnsgets onlyorig_col_nameandrewritten_col_name, which is what the pandas and polars pipeline gives it, and notdtypeandlengthas the xorq tier does (the xorq pipeline gives those).create_dataflowandcreate_polars_dataflow, the helpers/loadcalls, do not takestats_tier; the scope says/loaddoes not expose it.Not in this PR
Any
/loadrequest field; wire messages anddf_meta.stats; units; the xorq half (#1021); lazy polars. Selecting acleaning_methodstill runs the cleaning analyses' own stat pass (PandasAutocleaning._run_cleaning) at the schema tier; it is a separate stat pipeline, not the summary stats this tier skips.🤖 Generated with Claude Code