Skip to content

Performance on large tables: native CSV loading, dense working state, streamed output, memory layout; release 0.7.2 - #11

Merged
shigarov merged 12 commits into
mainfrom
perf/large-tables
Sep 8, 2026
Merged

shigarov merged 12 commits into
mainfrom
perf/large-tables

Conversation

@shigarov

@shigarov shigarov commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Performance-only release 0.7.2: RTL semantics, parity with jRegTab 0.7.1 and the CLI runner contract are unchanged; the public API is only extended.

What changed

  • Loading in one call: TableSyntax.from_csv(path), from_csv_text(text), from_rows(rows, *, num_cols=None) and a native parse_csv (exact port of RtlRunner.parseCsv); pyregtab.runner.parse_csv/load_table stay public and delegate to the core.
  • Interpretation without copies: the interpret binding borrows the table instead of cloning it and the semantics; Recordset is shared (Arc) by a transformation-free transform.
  • Working state: dense per-item vectors instead of hash maps, interned attribute names, one record arena, avp by attribute id, val/attr as overrides over the item strings; text shared between cells, items, state and records (Arc<str>).
  • Action templates shared by all anchors of an ActionSpec (the matcher used to clone the providers per matched cell).
  • Compact structures: 32-byte cells with formatting behind an Option<Box>, packed action anchors, u32 indices, flat recordset (one value vector for all records).
  • Streamed to_csv through a BufWriter (same bytes as before).
  • mimalloc as the allocator of the extension module (python feature only; building needs a C compiler, documented in README).

Results (comp-on-large-tables, medians of 5 runs)

case cells wall 0.7.1 → 0.7.2 jRegTab pandas peak RSS 0.7.1 → 0.7.2 jRegTab pandas
stack_test11 1 188 432 6.13 → 0.76 s 5.02 1.24 1656 → 335 MB 2179 125
explode_test22 200 005 1.00 → 0.20 s 1.31 0.59 280 → 87 MB 477 98
transpose_test21 338 528 0.76 → 0.21 s 1.10 0.52 219 → 77 MB 390 75

All 10 cases and 7 scaling points (up to 4.75 M cells: 2.9 s / 1.27 GB vs 19.3 s / 6.2 GB for jRegTab) are in plans/PERF_LARGE_TABLES.md and plans/PERF_MEMORY_LAYOUT.md.

Verification (after every commit)

  • pytest tests -q: 2013 passed; cargo test / cargo test --no-default-features: 49 + 49; clippy clean.
  • tools/differential.py against a jRegTab 0.7.1 dump: 750/750 variants identical.
  • tools/atbench_bytecmp.py (new): output.csv and exit codes byte-identical to the Java runner on all 244 regtab-eval-on-atbench solutions.

🤖 Generated with Claude Code

shigarov and others added 12 commits September 8, 2026 15:41
…e_csv

The CLI runner parsed CSV char by char in Python and set every cell through
a PyCell handle (1.10 s of the 1.19M-cell stack_test11 against 0.17 s in
Java). The RFC 4180 parser of RtlRunner.parseCsv is now a byte scanner in
the core (src/csv.rs), SyntaxCore::from_rows builds the grid in one pass,
and TableSyntax gains from_rows(rows, *, num_cols=None), from_csv(path)
(UTF-8, universal newlines as Path.read_text, so the runner's behaviour is
unchanged) and from_csv_text(text). pyregtab.runner.parse_csv/load_table
stay public and delegate to the core; read on stack_test11: 1.10 s -> 0.11 s.

tools/atbench_bytecmp.py compares the runner byte for byte with the Java
runner on all 244 regtab-eval-on-atbench solutions (Java outputs cached).
tests/test_runner.py pins the parser to the former pure-Python one on
boundary inputs and covers the runner contract and exit codes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…table

The Python binding of TableInterpreter.interpret cloned the whole SyntaxCore
and SemanticsCore to release the GIL and TablePattern.transform cloned the
whole recordset; interpret now borrows the table and PyRecordset shares an
Arc<RecordsetCore> (a transformation-free transform shares instead of
copying).

WorkingState stored val/attr/avp in IndexMap<ItemId, String> and rec in
IndexMap<usize, ...>: about ten hash lookups and four to five allocations
per record. They are now dense per-item vectors (ItemMap, AnchorMap with an
explicit insertion order, AnchorSet), attribute names are interned, and the
text of an item is an Arc<str> (util::Text) shared by the item, the working
state and the records. Providers append to one reusable buffer; schema
construction and record generation work on attribute ids.

stack_test11 (1.19M cells / records): transform 3.83 s -> 0.73 s (jRegTab
3.36 s), wall 6.13 s -> 2.03 s, peak RSS 1656 -> 1238 MB. examples/perf_runner
is a phase-timed runner over the pure-Rust core (interp::interpret_timed).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
RecordsetCore::write_csv writes record by record into any Write (quotes
doubled without an intermediate string); to_csv(path) uses a 1 MB
BufWriter<File> with the GIL released, to_csv() without a path still
returns the text. Same bytes as before (244/244 atbench solutions equal to
the Java runner). write on stack_test11: 0.17 s -> 0.05 s, at k=4 0.70 s ->
0.17 s, no outliers in 6 consecutive runs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… entries

The matcher instantiated a full ActionInst (providers with filter condition
and scope cloned from the spec) per matched cell: ~550 MB of the 1.2 GB of
stack_test11. Actions now hold an Arc<ActionTemplate> shared by all anchors
of the same ActionSpec (cached by spec address during the match; specs with
a constant-value context literal stay per instance, as in Java).

CellData keeps the text (an Arc<str> shared with the atomic item derived
from it), the text flags and an Option<Box<CellFormat>> materialized on the
first formatting write; getters return defaults. rec stores Records::One
for the common single-record anchor; parse_csv produces Text fields
directly.

stack_test11: peak RSS 1238 -> 523 MB (0.7.1: 1656 MB), match 0.70 -> 0.23
s, wall 1.15 s; at k=4 RSS 4944 -> 2067 MB.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A field without quotes or escapes is built from the input slice instead of
being copied through the accumulator first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Version 0.7.2 (performance only, semantics and parity with jRegTab 0.7.1
unchanged): README gains the bulk loading/streaming API in the runner
section and a "Performance on large tables" section with the measured
numbers; docs/api.md lists TableSyntax.from_rows/from_csv/from_csv_text,
parse_csv and the streaming to_csv; plans/PERF_LARGE_TABLES.md records the
before/after tables for the 10 cases and 7 scaling points.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Millions of small allocations (cell texts, items, records) pay less
per-object overhead and run faster than on the system heap: stack_test11
wall 1.11 s -> 0.93 s, peak RSS 523 -> 484 MB; at k=4 4.46 s -> 3.65 s,
2065 -> 1903 MB. Enabled by the `python` feature only, so the pure-Rust core
keeps the host's allocator. plans/PERF_MEMORY_LAYOUT.md opens the second
memory round.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…, lazy val/attr

Records live in one arena of the working state (RecEntry::One{start, len}
per anchor, ranges after JOIN) instead of a Vec per record; an avp stores
only the interned attribute id because its value is val(item) by
construction (all string operations precede AVP); val/attr hold overrides
over the item strings, so the initialization phase disappears. Action
groups are u32 indices and schema construction visits triples without
materializing them.

stack_test11: transform 0.457 s -> 0.373 s, peak RSS 484 -> 429 MB; at k=4
1903 -> 1680 MB.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
RecordsetCore stores the values of all records in one Vec<Option<Text>>
with the schema width as stride (record(i), records(), push, map_values)
instead of a RecordCore with its own Vec per record; transformations and
the Python bindings read slices. stack_test11: peak RSS 429 -> 395 MB,
transform 0.373 s -> 0.355 s.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
CellData drops its row/col (the caller knows the position: format(row,
col)) and keeps a u32 indent: 56 -> 32 bytes. ActionInst stores its anchor
as a packed u32 (ItemId::pack/unpack): 24 -> 16 bytes. RecEntry::Many points
into a side table of ranges: 24 -> 12 bytes. The item index and Range use
u32 indices, CellItem.index/span are u32 (88 -> 72 bytes).
CellPredicate::test takes the cell position explicitly.

stack_test11: peak RSS 395 -> 317 MB, wall 0.83 s -> 0.76 s; at k=4
1672 -> 1255 MB, 3.27 s -> 2.94 s.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@shigarov
shigarov merged commit 756005a into main Sep 8, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant