Skip to content

fix(in): admit the typed kernel unless both sides carry a null - #644

Merged
singaraiona merged 4 commits into
devfrom
fix/in-one-sided-nulls-593
Sep 29, 2026
Merged

singaraiona merged 4 commits into
devfrom
fix/in-one-sided-nulls-593

Conversation

@singaraiona

Copy link
Copy Markdown
Collaborator

Follow-up to #607, closing out #593.

Problem

#607 let in reach the typed membership kernel for text columns, but only when both operands are exactly null-free. That's stricter than the semantics need. The kernel (a null matches nothing) and the hashset fallback (null equals null) disagree only when a null row meets a null set element, and that requires a null on each side. So a single null row in the column sent the whole column to the per-row hashset probe.

Fix

Gate on !(has_nulls(set) && has_nulls(col)). The set is tested first: it's usually the small side, and when it has no nulls the column's null scan is skipped.

Numbers (351,393 rows, -c 8, release)

shape before after
SYM, one null row, (in col one) 4,820 us 50–60 us
I64, one null row, (in icol [0 1]) 5,730 us 80 us
SYM null-free, (in col one) 145 us 55 us
SYM null-free, needles [s, null] 4,620 us 145 us

Correctness

The hashset hashes i64 and f64 cells with different functions. Mixed int/float membership with a null on one side therefore missed matches: (in [1.0 0Nf 3.0] [1 3]) returned [true false false] and (in [1 0Nl 3] [1.0 3.0]) returned all false. Those shapes now take the kernel's double promotion and answer correctly.

A 23-case null edge matrix produces byte-identical output before and after, except for that corrected row.

Not fixed here: mixed int/float with nulls on both sides still reaches the hashset and still answers wrongly, e.g. (in [1 0Nl 3] [1.0 0Nf]) returns [false true false]. That's a hashset hashing bug, not a gate bug, so it's left for a separate fix.

Tests

test/rfl/collection/in.rfl: fixes the comment that claimed a needle-only null "must still reject", and pins one-sided nulls for I64/I32/I16/F64/DATE/STR/SYM, the INT32_MIN-sentinel needle, and the two formerly wrong mixed-type rows. make test: 3935/3935 pass, with no UBSan runtime errors.

#607 opened the typed membership kernel to text columns but kept it behind
"both operands exactly null-free".  That is stricter than the semantics
require: the kernel (null matches nothing) and the hashset fallback (null
equals null) disagree only on a null row against a null set element, which
needs a null on each side.  One null row in the column therefore sent the
whole column to the per-row hashset probe.

351,393-row column, -c 8, one null row, null-free needles:

  SYM (in col one)      4,820 us -> 60 us
  I64 (in icol [0 1])   5,730 us -> 80 us
  SYM null-free (in col one)  145 us -> 55 us   (set tested first: a
                                                 null-free set skips the
                                                 column's null scan)

It also fixes wrong answers: the hashset hashes i64 and f64 cells with
different functions, so mixed int/float membership with a null on one side
missed matches -- (in [1.0 0Nf 3.0] [1 3]) was [true false false].  Those
shapes now take the kernel's double promotion.  The both-sides-null
mixed-type case still reaches the hashset and is still wrong; it is a
hashset bug, not a gate one, and is left for a separate fix.

A 23-case null edge matrix is byte-identical before and after apart from
that corrected row.  The in.rfl comment claiming a needle-only null "must
still reject" is corrected, and one-sided nulls across every null-bearing
width are pinned.
@singaraiona

Copy link
Copy Markdown
Collaborator Author

The widened gate exposes a temporal-type mismatch: on this head, (in [2000.01.01 0Nd] [0]) returns [true false], and reversing the operands returns [true] (reproduced in the sandbox). The typed kernel compares raw date/integer values, whereas the fallback keeps those types distinct; null-versus-null is therefore not their only semantic difference. Please restrict admission to compatible equality semantics, or correct the kernel type checks, and add temporal/integer and distinct-temporal regressions with nulls on either side before merging.

@singaraiona

Copy link
Copy Markdown
Collaborator Author

The both-sides-null mixed int/float case this PR leaves unfixed is filed as #645. The same hashset bug also gives wrong find, except and union results on mixed int/float inputs, even without nulls.

@ser-vasilich

Copy link
Copy Markdown
Collaborator

Ran a differential review of the PR head against the pre-PR gate: 1056 small shapes plus 112 at 300k rows (12 column types × nulls {none, column, set, both} × set sizes {1, 3, 1000 distinct, with duplicates}, 21 cross-type pairs) and the edge cases (SYM id 0 vs "", STR "", INT32_MIN / INT16_MIN as real needles and as sentinels, NaN in both roles, -0.0, a FILE-domain SYM column from a splayed table against runtime needles, slices, atom needles, empty set / column, width mismatches, stale HAS_NULLS producers such as (as 'I32 [-2147483648])). Same-type results are byte-identical; the int/float one-sided-null rows change to the correct answers as described; the fused select where: and (count (where …)) agree with the bare in. ASan/UBSan clean.

Timings, release build, 4 cores, best of 5, ms, before → after:

shape 351k rows 3M rows
SYM column with 1 null × 3 needles 6.07 → 0.17 52.7 → 1.08
I64 column with 1 null × 3 needles 5.09 → 0.17 40.1 → 1.29
F64 column with 1 null × 3 needles 4.67 → 0.80 36.6 → 4.73
SYM clean × set with 1 null 5.44 → 0.24 45.9 → 2.77
I64 clean × set with 1 null 2.85 → 0.11 21.9 → 0.70
SYM clean × 100k needles with 1 null 11.7 → 0.96 92.3 → 3.49
STR column with 1 null × 3 needles 6.92 → 6.71 60.9 → 60.4
I64 clean × clean 3 needles 0.110 → 0.120 0.68 → 0.70
I64 column with 1 null × 1k needles 8.98 → 15.1 77.1 → 149
I64 column with 1 null × 10k needles 6.90 → 160 58.8 → 1263
I64 column with 1 null × 100k needles 9.21 → 1198 69.4 → 10901
I64 clean × 100k needles with 1 null 8.32 → 1223 60.1 → 10365
I64 clean × clean 100k needles (path unchanged by this PR) 1139 → 1120 9610 → 9453

Two things the description does not mention:

  1. Large numeric needle sets. The typed kernel probes the set linearly per row (exec_in_worker, no cap on the set size), so a nullable int/float column against a set of thousands of needles now runs O(rows × needles) where the hashset was O(rows + needles): the last four rows of the table. The linear probe was already the cost for null-free operands (last row: ~1.1 s at 351k, ~9.5 s at 3M on both binaries), so the PR widens an existing hole rather than opening one, and SYM is unaffected thanks to the verdict LUT. Hashing the numeric needles once inside the kernel (an open-addressing table over the int64 values / double bit patterns, used above ~16 needles) removes the cliff for every shape and keeps this gate as is; I prototyped that locally and the answers stay identical, so it can be a follow-up rather than part of this PR.
  2. Temporal vs integer with a null on one side. (in [2000.01.01 0Nd 2000.01.03] [0i 2i]) was [false false false] and is now [true false true] (same for TIME/TIMESTAMP vs I32/I64, both directions; (count (where (in d [0Ni 4i]))) on a 200k table went 0 → 2000). The hashset compared through the type-strict atom_eq, the kernel classifies both sides as int-family and compares raw values. The null-free case and the fused where: already answered like the kernel, so the bare in is now consistent with them, but (== 2000.01.01 0i) is a type error, so it is worth a sentence in the description.

Also noting that the both-sides-null int/float case is reachable through an ordinary pipeline: (count (where (in c [0Nf 2.0 999.0]))) on an I64 column with one null returns 1 instead of 400 (the fused where: returns 400) — the #646 class. Suites collection / query pass; of the new lines in in.rfl, :143 and :144 are the ones that fail against the old gate.

… the kernel

Review follow-ups on widening the `in` kernel gate.

Temporal types.  The typed kernel classified DATE/TIME/TIMESTAMP as plain
int-family and compared raw payloads, so (in [2000.01.01 0Nd] [0]) was
[true false] while find, except and the hashset `in` -- all atom_eq, which
never matches a DATE to an int or to a TIMESTAMP -- said no.  It also
compared a DATE's day count to a TIMESTAMP's nanoseconds:
(in [2000.01.02] [2000.01.01D00:00:00.000000001]) was [true].  A temporal
operand against a different type is now an empty probe, as SYM vs non-SYM
already was.  Bare `in`, the fused where: and the hashset agree whichever
side carries a null.

Large needle sets.  Past the SIMD small-set size the kernel scanned every
needle per row, O(rows x needles).  The widened gate sent nullable int/float
columns there, and null-free ones already went: a 351k-row column against
100k needles took ~1.2 s.  Sets larger than IN_SIMD_SET (8) now get an
open-addressing table over the live needles -- int64 values, or f64 bit
patterns with -0.0 folded to +0.0 and NaN never inserted -- built once per
call and shared by all workers.  An allocation failure keeps the scan.

351,393 rows, -c 8, release, dev before #644 -> this commit:

  I64 with a null   x 3 needles       6,500 us ->    50 us
  I64 with a null   x 1k             10,600 us -> 1,400 us
  I64 with a null   x 100k           10,000 us -> 2,500 us
  I64 null-free     x 100k        1,155,000 us -> 1,500 us
  F64 with a null   x 100k           10,500 us -> 2,500 us
  select where: in, null-free x 100k  1,155,500 us -> 2,000 us

in.rfl pins temporal vs int and distinct temporal types with nulls on
neither, either and both sides, the fused where:, and the 8/9-needle
boundary, -0.0, NaN, duplicates and narrow ints on the hash path.
make test 3952/3952; reverting exec.c fails in.rfl:169.
@singaraiona

Copy link
Copy Markdown
Collaborator Author

Both points are addressed in 0afbee3.

Temporal vs integer (blocking comment). I corrected the kernel type checks rather than restricting admission, so every path now agrees. The kernel treated DATE/TIME/TIMESTAMP as plain ints and compared raw payloads, while atom_eq never matches a DATE to an int or to a TIMESTAMP. atom_eq is what the hashset in, find and except all use. The raw compare was wrong even without nulls, and across temporal types it compared a DATE's day count with a TIMESTAMP's nanoseconds: (in [2000.01.02] [2000.01.01D00:00:00.000000001]) returned [true]. A temporal operand against a different type is now an empty probe, the way SYM vs non-SYM already was.

before after
(in [2000.01.01 2000.01.03] [0i 2i]) [true true] [false false]
(in [2000.01.01 0Nd] [0]) [true false] [false false]
(in [2000.01.02] [2000.01.01D00:00:00.000000001]) [true] [false]
select where: (in d [0i 2i]) 2 rows 0 rows

This changes the null-free answer as well as the widened gate's: the null-free case and the fused where: both answered like the raw compare before. Nothing in the suite depended on it (3952/3952). in.rfl now pins temporal vs int in both directions and distinct temporal types, with nulls on neither, either and both sides, and includes the fused where: and find. Reverting exec.c makes in.rfl:169 fail.

Large numeric needle sets. Agreed that the widened gate made this worse, so I fixed it here rather than as a follow-up, along the lines you prototyped. Sets larger than the SIMD small-set size (8) get an open-addressing table over the live needles. It stores int64 values, or f64 bit patterns with -0.0 folded to +0.0 and NaN never inserted. The table is built once per call and shared by all workers; if the allocation fails, the kernel keeps the linear scan. 351k rows, -c 8, pre-PR dev → now:

before after
I64 with a null × 1k needles 10.6 ms 1.4 ms
I64 with a null × 100k needles 10.0 ms 2.5 ms
I64 null-free × 100k needles 1,155 ms 1.5 ms
F64 with a null × 100k needles 10.5 ms 2.5 ms
select where: (in x n100k), null-free 1,155 ms 2.0 ms

I checked the hash path against find for 5 / 9 / 50 / 1k / 100k needles across I64/I32/F64 columns with nulls, and for duplicate needles. in.rfl pins the 8/9-needle boundary, -0.0, NaN, duplicates and narrow ints.

The both-sides-null int/float case you pointed out ((count (where (in c [0Nf 2.0 999.0]))) returning 1) is the #645 class, which #646 fixes.

@singaraiona
singaraiona merged commit 4a7276d into dev Sep 29, 2026
9 checks passed
@singaraiona
singaraiona deleted the fix/in-one-sided-nulls-593 branch September 29, 2026 17:58
singaraiona added a commit that referenced this pull request Sep 29, 2026
* fix(group): bound the tied groups the radix top-N selection keeps

The native top-N of the radix path kept every group at or beyond the
threshold.  Groups strictly beyond it number fewer than N, but the
groups AT the threshold can be nearly all of them: a count-per-group
top-10 over a near-unique key has a threshold of 1, every group ties
it, and the "kept superset" is the whole grouping, heap-sorted by
first row — a comparison sort over 100M pairs for a query whose answer
is ten rows.

The emitted prefix is the tied groups' first-seen order, so only the N
tied groups with the smallest first row can ever be taken.  The
selection now keeps the strictly-better groups plus exactly those N,
through a bounded max-heap, and sorts at most 2N entries.  The rows
emitted are the ones the full kept set produced.

Tests: the pinned kept count of the radix native top-N case becomes N;
a near-unique two-key grouping whose threshold every group ties, in
both directions and under a row selection, against the full grouping.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(lang): add a `while` special form

Rayfall had no way to stop iterating before the end of a sequence.  Every
iteration primitive — map, pmap, fold, fold-left/right, scan, scan-left/right,
prior — consumes its whole input, and there was no loop construct, so a
"repeat until done" loop had to be written as a fold over a fixed range whose
full length was paid on every call however early the work finished.  `return`
did not help: it exits the lambda, not the iteration, so the fold kept calling
the lambda for every remaining element.  Recursion was not an alternative
either, since there is no TCE and the stack tops out around 1-2k frames.

(while cond body...) evaluates cond, and while it is truthy evaluates each
body expression in order, then tests again.  It always returns null — a
statement form run for effect, and a never-taken loop has no last value to
report.  Zero body expressions is legal, so a condition with side effects can
be the whole loop; that is the shape a drain wants, where there is no sequence
to iterate over and the range was only ever scaffolding.

Implemented on both evaluators, which must agree:

  - ray_while_fn (eval.c), registered beside `if` and `do`.
  - an sf_while case in the bytecode compiler emitting JMPF over a backward
    JMP — the first backward branch the compiler produces.  The VM's op_jmp
    already checked for a pending interrupt when the displacement is negative
    (dormant until now), so a runaway loop is Ctrl-C-able for free.

Compiling the body inline, rather than letting it fall through to the generic
special-form path, is what keeps `return` working inside a loop: it reaches
the sf_return case and unwinds the lambda instead of degrading to the tree
walker's identity `return`.

Neither path pushes a scope around the body.  `let` binds only in the top
frame (env_bind_local), so a per-iteration frame would discard loop-carried
`let` state on the tree-walking path while the compiled path — whose `let`
writes a function-level slot — kept it, and the two evaluators would disagree
on the same source.  A caller wanting a fresh frame per pass writes
(while cond (do ...)), which composes.

patch_jump and emit_jump_back now share write_jump_offset; an out-of-range
displacement still sets c->error, dropping the lambda to the interpreter.

Measured on the drain shape from the issue (4-step drain under a 4096 safety
bound, release build, driver baseline subtracted): fold-left over (til 4096)
1420 us/batch, the nested 64x64 workaround 45.8 us/batch, `while` 3.8
us/batch — 374x over the original and 12x over the workaround.

Closes #588

* feat(lang): add `times` and `fold-while`

The two remaining early-termination forms requested in #588, alongside the
`while` that landed in #590.  Neither is an unblock — `while` already covers
the reporter's case — but both complete the vocabulary he asked for.

`times`
-------
(times n body...) runs the body exactly n times and returns null.  `do` is
already progn in this language, so the bounded loop could not reuse that name;
overloading `do` on an integer head was rejected as genuinely ambiguous —
(do 5) would have to mean either "loop five times over nothing" or "return 5",
and a computed first expression that happened to be an integer would silently
change meaning.

The count is evaluated ONCE on entry, so a body mutating whatever produced it
cannot change how many passes remain.  A count of zero or less runs zero times
rather than trapping; a non-integer is a type error.

Compiled as a counted loop over a hidden local slot, reusing the backward
branch added for `while`.  ray_times_norm_fn type-checks the count and clamps
a negative bound to zero on entry, which lets the per-pass test be a bare
truthiness check on the counter — 0 is falsy, so no comparison call is needed
per pass and a negative bound cannot run away.  That helper and the decrement
are pushed as constant-pool objects rather than resolved by name, so the
loop's own arithmetic is unnamable from source and cannot be swapped out by a
`(set - ...)` override.  The counter's slot is addressed by index and its sym
carries a space, so no source token can collide with it and nested `times`
counters stay apart.  As with `while`, the body compiles inline, so `return`
unwinds the enclosing lambda from inside the loop.

fold-while
----------
(fold-while pred f init xs) offers the accumulator to `pred` before each step
and stops on a falsy answer, yielding the accumulator as it stands.  The test
precedes the first element, so a predicate false at the start returns `init`
untouched.  The predicate takes the accumulator rather than the element: that
is the form that expresses "iterate until the running result says stop", which
is the early termination actually being asked for.

One deliberate divergence from ray_fold_fn: that routes its collection through
unbox_vec_arg -> to_boxed_list, boxing every element up front.  For a
primitive whose purpose is to stop early, paying for the tail it never reaches
is the cost being removed, so elements are pulled one at a time via
collection_elem.  A plain variadic builtin — no compiler work, since it
dispatches through the normal call path.

Measured, release builds: `times` 100 ns/pass against 180 ns for `while` plus
a manual counter over 1e6 passes.  `fold-while` stopping after three elements
of a 1e6-element vector costs 3.9 us against 369 ms for the `fold-left`
equivalent, which had to box and walk all million to discover it was done.

Tests: test/rfl/lang/times.rfl and test/rfl/collection/fold_while.rfl, each
behaviour asserted on both evaluator paths where applicable — including the
count validation and negative clamp on the compiled path, which runs through
entirely different code from the tree walker's check.

* perf(join): skip per-cell key null tests on provably null-free columns

ray_vec_is_null is out-of-line (no LTO) and the join called it once per key
column per row in hash_row_keys — on the build side, the probe side, and the
prefetch lookahead — plus twice per key column per hash-chain step in
join_keys_eq, across both the count and the fill pass.  A reported profile
put 13.28% of a service's samples there, on a join keyed by two SYM columns
that structurally never hold a null.

Prove once per join that no key column can hold a null and drop the call.
SYM/STR nulls are canonical empty payloads (id 0 / length 0) that HAS_NULLS
does not track, so text columns are proven by the chunked zero-scan from
#533 rather than by the flag; everything else reads the flag through slices
and takes a set bit at face value, so the proof stays O(n) and never
degrades into ray_vec_has_nulls' per-element walk.  Flag-readable columns
are settled first, so a nullable numeric key short-circuits before any text
column is scanned.  OP_CONST (atom) key slots are refused as unprovable.

Also skip the #458 null-run pre-scan under the proof.  It gates on
ray_vec_may_have_nulls, which is unconditionally true for SYM/STR, so a
SYM-keyed join ran a full per-key-per-build-row ray_vec_is_null scan before
the join proper on every execution; a null-free key set cannot contain an
all-null row, so the scan is dead.

Measured with bench/join_nullfree (release, 1:1 book join, 4M probe x 500K
build): two SYM keys 275 -> 249 ms (-9.5% median, -7.8% min); single I64 key
106 -> 99 ms (-6.3% median, -7.9% min).  A proof that fails on a text key
costs a partial scan (~8ms on a 4M-row SYM column) and gains nothing — only
nullable SYM/STR key columns pay it.

ray_join_force_null_checks forces the null-aware loops so the differential
tests and the perf gate can compare both paths in one binary;
ray_join_nullfree_keys counts the joins that took the fast path.

Closes #597

* fix(system): validate launcher and timeit integer arguments

* fix(cli): reject a flag as another flag's value, and unknown options

Every value-taking startup flag was gated on `i + 1 < argc` and then
consumed argv[++i] blindly, so one flag silently ate the next.  The shape
that was reported is the worst one for a service:

  rayforce -Q -p 5099 svc.rfl

`-Q` took `-p` as its value, the port was never bound, and nothing on
stdout or stderr said so.  Under a supervisor the process starts, the
script runs, the unit looks healthy, and every client gets connection
refused.  `-c` and `-t` had the identical hole.

Two neighbouring cases came out of the same code.  A trailing flag with
no value fell through to the positional-file arm and reported the
misleading `cannot open '-Q'`.  And that arm took ANY unrecognized token
as the script name, which the real positional then overwrote, so a
typo'd option was not merely ignored — it was silently swallowed:
`rayforce -x svc.rfl` ran the script as if `-x` had never been typed.

Every value-taking flag now goes through flag_value(), which refuses a
missing value and a value that is EXACTLY one of this program's flag
tokens (including the `--` app-args terminator), with a diagnostic and
exit 2 before the script is loaded.  A value that merely STARTS with '-'
stays legal: a password or a negative number is a legitimate value, and
refusing those would break working command lines for no gain.  An
unrecognized `-`-prefixed token is now an error too; a lone `-` is still
a positional, and tokens after `--` belong to the app and are never
validated.

`-Q, --querylog N` was a real flag, referenced twice in the docs, that
--help never listed beside its sibling `-t, --timeit N` — added, along
with `[-Q 0|1]` in the synopsis.

The suggestion to fail when `-p` was requested but nothing bound is
already in (#473, listen_fatal.rfl), and cannot catch this: the parser
never saw the eaten `-p`, so there was no request to check against.  The
guard on the value is what closes the class.

Also: the two pre-existing `-p` validation bail-outs now release the
runtime like every other error path, since these paths are exercised
under ASan by the new test.

Closes #600

* fix(exec): release the input table of window, sort and limit unconditionally

Every node the executor evaluates returns an owned reference — the
constant table node included, it retains its literal.  The OP_WINDOW,
OP_SORT and OP_HEAD cases released their input only when it differed
from the graph's table, taking an equal pointer for a borrowed one.  A
query's root is a constant node over that very table, so the reference
was never released: one input table per windowed, sorted or limited
query.  Invisible while a global kept the table alive; a whole table per
call when the input was built for that call, as a service that windowed
a freshly concatenated buffer on a timer found (#602).

The four cases now release the input on every path, as the join and
the plain head/tail cases already did.

Test: window, sorted and limited selects over a table built per call,
measured with the two-window bytes-allocated method of the other memory
probes, plus a shared input reused across fifty calls.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(lang): pin the refcount of a table across window, sort and limit queries

Twenty calls of each shape over one global table; the table's refcount
must be exactly what it was before, and the table still whole after.
Fails on the unfixed executor at the first window shape.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(exec): the reduction case owns its input too

The same borrowed-query-table guard sat under the reductions; no child
evaluates to the query table there today, so nothing leaked, but the
premise is the one the sort, window and limit cases just dropped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* perf(pool): wake only the workers a dispatch can keep busy

ray_pool_dispatch and ray_pool_dispatch_n signalled the whole pool on
every dispatch, however narrow the window.  The main thread participates
as worker 0, so at most n_tasks-1 helpers can ever claim anything; the
surplus threads woke, raced to an already-drained window, and went
straight back to the semaphore.  On a dispatch narrower than the machine
that is pure overhead — a partition-parallel step over three partitions
woke every core on the box to run two tasks.

Signal min(n_tasks-1, n_workers) instead.  Signalling FEWER is safe:
completion is governed by the `pending` spin-wait, never by the signal
count, so an unsignalled worker simply stays asleep, and under-signalling
cannot produce the surplus-signal problem between consecutive dispatches
that the spin-wait exists to avoid.

Measured (10M-row ClickBench, 43 queries x 3, splayed): sum of per-query
hot times 5690.7 ms -> 5667.3 ms (-0.4%, within noise), user CPU 55.1 s
-> 54.8 s.  So this is waste removal rather than a speedup on workloads
whose dispatches are already wider than the pool.

Not a fix for #599.  That report's shape — a 130k-row join probe, which
splits into 16 tasks against 8 workers — is unaffected, because the clamp
lands on the same worker count it already used (measured: 5.19 s -> 5.18 s
of CPU, unchanged).  The investigation there points at the pool defaulting
to logical rather than physical cores; that is a separate change with its
own benchmark trade and is filed on its own.

test_pool.c gains pool/dispatch_narrow, which is the case the existing
coverage missed: dispatch_n_small uses 4 tasks on a 2-worker pool, so the
clamp never binds.  The new test drives windows of 1..4 tasks on a
4-worker pool, then alternates narrow and wide windows 25 times over the
same pool, asserting exact task and element counts throughout — the two
risks the change introduces are a task nobody claims and signal
accounting that drifts across dispatches and starves a later wide window.

* build: Windows (MSYS2 CLANG64/MINGW64) toolchain in the Makefile

Detect Windows via $(OS) or uname (MSYS2's make hides $(OS)), link
Winsock and a statically linked winpthreads, build with 64-bit off_t
(_FILE_OFFSET_BITS=64) and an 8 MiB stack like a Linux main thread.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(core): WSAPoll event loop for Windows

Replace the IOCP stub with a readiness-based loop that mirrors the epoll
backend's dispatch order, so the selector state machine is identical on
every platform.  stdin (console or pipe) is not a socket, so RAY_SEL_STDIN
selectors are probed directly and the socket wait is sliced while one is
registered.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(net): Winsock errno mapping, exclusive bind and IPC on Windows

- mirror WSAGetLastError()/SO_ERROR into errno for send/recv/connect/
  accept/bind/listen: callers branch on EAGAIN/ECONNREFUSED;
- ipc_send_fn always overwrites errno, so a stale EAGAIN can no longer
  park a frame on a dead socket;
- SO_EXCLUSIVEADDRUSE instead of SO_REUSEADDR (which lets a bind steal
  a port that is in use);
- WSAStartup at load time; sock.h pulls in platform.h;
- verbose IPC capture uses a temp file that works without admin rights;
- the SIGURG out-of-band cancel stays POSIX-only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: Windows platform layer (file mapping, pools, crash report, paths)

- ray_vm_unmap_file only unmaps at a view's own base: UnmapViewOfFile
  drops the whole view for an interior pointer, which freed columns that
  carry a passenger index (munmap is a no-op there);
- ray_vm_alloc_aligned returns its own allocation base, so pools are
  really released by ray_vm_free;
- ray_vm_map_fd_ro maps for real (CSV reads always failed with io);
- crash report via SetUnhandledExceptionFilter;
- heap: file-backed spill stays POSIX-only (docs/architecture/memory.md);
- domain: realpath substitute; symfile paths compare case/separator-
  insensitively so one file never gets two domains;
- aof/csr/hnsw: platform handle for fsync, portable ray_mkdir.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: REPL, profiler and system builtins build and work on Windows

- term.h/profile.h include platform.h and a lean <windows.h>;
- term_write for the Windows console, errno.h and core count in the REPL;
- .sys.info reports page-size and total-mem on Windows too;
- KEY_READ no longer collides with <winreg.h>.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: run the suite on Windows

- runner: ';; @requires: posix' marks .rfl files whose fixtures or checks
  need a POSIX shell/filesystem; on Windows they are reported as SKIP;
- shell-free ray_test_rm_rf / ray_test_mkdir_p replace system("rm -rf")
  and "mkdir -p" in C tests; test.h maps the few POSIX helpers tests use;
- tests of POSIX-only behaviour (setrlimit, ENOTDIR, read-only dirs,
  AF_UNIX, file spill, Winsock send buffering) skip with the reason;
- fixture fixes that were latent on any platform: binary-mode CSV
  fixtures, a per-test AOF dir, journal closed before the crash rename.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: Windows build instructions and platform status

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: readable REPL prompt and CPU name in the banner on Windows

No console font Windows ships has U+2023 (checked Consolas, Cascadia Mono,
Lucida Console, Courier New) and the classic console does no font fallback,
so the prompt rendered as '?'.  Use the nearest filled triangle they all do
have, U+25BA, which is also three UTF-8 bytes.

The banner's CPU line said 'unknown': read ProcessorNameString instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* bench: build the micro-benchmarks on Windows, record a Windows/Linux check

alloc and agg_v2 read peak RSS through getrusage; use
GetProcessMemoryInfo there.  windows_vs_linux.md records the numbers the
port was checked against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* bench: record the recent perf benches on Windows vs Linux

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(format): render i64 at full width, not through 32-bit long

Interpolation (format/println), print and the pivot / column-name helpers
formatted an i64 with "%ld" and a (long) cast.  long is 64-bit on LP64 and
32-bit on Windows, so there 10^18 printed as -1486618624 and a pivot keyed
on values above 2^31 produced truncated column NAMES — a data defect, not
only a display one.  Use PRId64 throughout; regression test included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(crash): format the banner with snprintf instead of macro string splicing

The banner was spliced from string literals and the RAYFORCE_VERSION /
RAYFORCE_GIT_COMMIT macros inside #ifdef arms; without the -D values a
static analyser reads the literal followed by the bare macro name as two
adjacent tokens and reports a syntax error.  Format it with one snprintf
at install time (the handler itself still never formats), with the
macros defaulting to empty strings.  Output unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(regress): drop the 24 GB til from the i64 width probe

`til` is eager, so counting a three-billion-element range built a 24 GB
vector for no extra coverage — the neighbouring literals already cross
2^31, 2^32 and 2^63.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(mem): hand workers their freed blocks back at the end of every dispatch (#621)

* fix(mem): hand workers their freed blocks back at the end of every dispatch

A parallel operator's workers allocate per-task buffers from their own
heaps and the main thread frees them once the dispatch has completed, so
each block lands on the owning worker's foreign list. The owner drained
that list only when an allocation found its freelists empty, which a
warm worker with a partly cut pool rarely does: it kept splitting fresh
pool space instead, touching new pages every round while its own freed
blocks waited. Under a steady stream of such operators (a join against a
large table per incoming batch) the process grew by a full 32 MB pool
per worker before any block was reused, and only the idle decay, which
needs the process to sit quiet, drained it earlier.

The dispatcher now drains every registered heap's foreign list at the
end of each parallel region (ray_heap_reclaim_workers), under the same
conditions the idle decay already relies on: the flag is clear, so every
worker has finished its last task and neither allocates nor frees until
the next dispatch. The blocks go back to freelists only, no pages are
released, so the next round reuses them without faulting.

Tests: heap/reclaim_workers_drains_owner (a block owned by one heap and
freed from another leaves the owner's list on reclaim and is handed out
again at the same address, no new pool) and
pool/dispatch_reclaims_worker_blocks (blocks allocated by real workers
and freed by the main thread are back with their owners by the end of
the next dispatch).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(mem): reclaim only the pool's own worker heaps after a dispatch

The heap registry holds every thread that ever allocated, not only pool
workers: a server's poll thread or an embedding's own threads are there
too, and they may be allocating while a dispatch on another thread ends.
Draining such a heap's foreign list from the dispatcher coalesces into
freelists their owner is using at that moment. Only a pool worker parked
on its semaphore is known to touch nothing of its own until the next
dispatch, so each worker now publishes its heap in the pool
(worker_heaps, set after ray_heap_init, cleared before it exits) and the
dispatcher reclaims exactly those, plus its own list on its own thread.
ray_heap_reclaim_worker takes one heap; the registry walk is gone, which
also removes its per-dispatch scan of every registry slot.

The heap test now checks the block's bytes leave the owner's books only
on reclaim, instead of expecting the next allocation at the same
address; the pool test reads the worker heaps the pool publishes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(pool,heap): make the reclaim tests independent of worker start-up and adopted heaps

The pool test read a worker's heap slot right after ray_pool_create,
which returns before the workers have started, and passed vacuously when
the main thread took every task. It now waits for each worker to publish
its heap, records which worker allocated each block and repeats the
round until a worker really allocated something, then checks that the
next dispatch left no worker with a foreign list. The heap test drains
whatever an adopted heap already held before taking its baseline.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pool): allocate the worker heap slots without an _Atomic cast

cppcheck cannot build an AST for a cast to `_Atomic(void*)*` and fails
the static-analysis run on it. The cast is not needed in C: ray_sys_alloc
returns void*, which converts to the slot pointer type on its own, and
sizeof(*pool->worker_heaps) names the element size without spelling the
type.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(serde): reject trailing payload bytes (#622)

* fix(select): a projection sees the projections before it (#620)

* perf(like): one column kernel for the direct builtin and the executor, one row pass (#623)

* fix(system): report arity errors for extra arguments (#616)

* fix(in): admit the typed kernel for text columns, which it never reached (#607)

* perf: shared-node memo, STR descriptor views, fused top-k on nullable like, bulk derived key, parallel dense grouping (#626)

* v2.9.1 (#625) (#629)

* docs(release): merge master back into dev after each release

A merged release PR leaves a commit on master that dev lacks, so the next
release PR is blocked as out of date with the base, and after a squash it
also shows conflicts that aren't real. Add the back-merge as step 5.

* perf(select): whole-table count/min/max/sum/avg from chunk-zone metadata

A select of whole-table aggregates with no other clause scanned every
row, although a stored column already carries per-chunk min/max in its
chunk-zone index.

- The chunk-zone index of an integer or temporal column also records,
  per chunk, the sum of its non-null values (int64 with wraparound, as
  the engine sums) and their count, in a fourth child vector.  Indexes
  written before hold zero in that slot and load without it.
- `sum` of an integer column reads the per-chunk sums; `avg` does too
  when no partial sum can leave double's integer range, so the value is
  exactly the row-wise one (otherwise it scans as before).
- `(select {from: T a: (sum c) b: (min d) …})` with no where/by/take/
  sort, every output a count / min / max / sum / avg of a plain column
  the metadata answers, is built from the metadata in O(chunks).  Any
  other output leaves the query to the planner.

Test: store/splayed_zone_aggs (each answer, nulls and types equal to the
same query over an in-memory copy; float sums, a filter and an avg past
2^53 take the usual path and still agree; scalar sum/avg).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(index): mapped index regions — bounds-checked children, never released on drop

- A column's inline index region is mapped with the column file, and its
  child vectors are addressed by region-relative offsets.  The offsets
  were trusted as written.  A region re-saved by a binary that knows
  fewer child slots keeps a stale offset in the slot it did not write
  (the new chunk-zone aggregates slot), pointing at or past the end of
  the file; the free path then sized its munmap from it and unmapped a
  page it did not own.  ray_index_inline_map now takes the region size
  and admits only children that lie inside it; an out-of-bounds or
  malformed aggregates child is dropped, any other one leaves the column
  unindexed.
- Copies of a mapped column borrow its index without a reference.
  Dropping the index from such a copy (a delete or an in-place write on
  it) released the index anyway, freeing the loaded column's children
  in place: its min/max then read NULL and crashed.  A mapped index is
  now only detached from the vector; the mapping's owner unmaps it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(agg): zone min/max tell all-null from the non-null counts; empty tables keep the planner

- The zone min/max treated INT64_MAX / INT64_MIN as "no non-null value",
  so a column of int64 maxima answered null.  When the zone carries
  non-null counts they decide; min/max also require the zone's extrema
  to be present.
- The metadata select no longer answers an empty table (its count
  differed from the planner's there).

Tests: store/splayed_zone_aggs (edit through a copy of a loaded column
then aggregate the original; an all-INT64_MAX column; an empty filter).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(store): portable in-place sed in splayed_zone_aggs

BSD sed takes the argument after -i as the backup suffix, so the null
injection failed on macOS; use -i.bak and remove the backup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(system): reject arguments to gc (#615)

* fix(system): reject arguments to gc

* docs(memory): use zero-argument gc form

* fix(system): report gc arity errors

* fix(string): a STR view keeps only the bytes it points at (#631)

* fix(string): a STR view keeps only the bytes it points at

The descriptor views introduced for substr and `if` over STR columns pinned
more than they used.  An `if` whose two sides came from different pools —
including two columns of one table under a where: that kept a few rows —
copied BOTH whole parent pools into the result (the compacted branch
inputs share the full column pool, so a 10k-row result out of 500k rows
carried the two columns' 43 MB); the eager arm did the same for an
unfiltered `if`.  A substr with a per-row start or length retained the
parent pool even when every result fitted inline and nothing pointed
into it.

`if` now builds a pool of exactly the chosen rows' bytes whenever the
sides come from different pools, a pooled scalar is involved, or the rows
keep a small share (under an eighth) of one shared pool; one shared pool
whose rows keep most of it is still pointed into as before, so a
derived-key expression over one column stays copy-free.  A substr view
whose descriptors are all inline drops its pool reference; one that
points into the pool stays a view (the column is alive anyway, and a
sibling `if` over two such views can then pick either side without a
copy).

test/rfl/strop/str_view_pool_compact.rfl pins the retained bytes through
direct-bytes: an `if` under a where over two columns, over two substring
views, with a pooled scalar side, the eager `if` against the two-pool
sum, and an inline substr view outliving its table.

* test(string): cover the eager `if` over two STR pools

The case labelled eager in str_view_pool_compact went through the
selected path (a per-row condition column is always a row mask there),
so the eager arm's two-pool offset/copy had no regression coverage.

A scalar condition (a variable) is not a row mask: the selected path
declines and the eager fill runs over two different pools.  New cases
take it at 500k rows (parallel fill) and 1000 rows (below the parallel
threshold, serial fill), for both sides, check every row after all the
sources are released, and bound the retained bytes to one side's pool.
On the pre-fix tree the retention checks fail.  The old case is
relabelled as the per-row selected path it is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Anton Kundenko <singaraiona@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(log): reject arguments to purge (#638)

* perf(vec): avoid rewriting copied text nulls during concat (#641)

Co-authored-by: singaraiona <3473381+singaraiona@users.noreply.github.com>

* fix(hnsw): reject invalid persisted graph ids (#642)

* fix(log): reject arguments to roll (#639)

* perf(expr): 32 registers per fused expression; nulls produced mid-program, narrow scratch (#636)

* perf(expr): 32 registers / 96 instructions per fused expression

A filter of four conditions, one of them a `within`, needs more than 16
expression registers; the compiler then bailed and the whole predicate
was evaluated column at a time over every row, without the chunk-zone
morsel decisions.  A grouping under
`(and (== CounterID 62) (within EventDate [..]) (== IsRefresh 0) (== DontCountHits 0))`
on a 100M-row table took 71 ms warm / 540 ms cold, against 24 ms for
the same filter with three conditions.

The limits are now 32 registers and 96 instructions; the per-task
scratch is sized to the registers the expression uses instead of the
maximum.  The same query: 24 ms warm, 130 ms cold, same answer.

Test: expr/wide_predicate (four and six conditions with arithmetic and
`within`, a grouping under the four-condition filter, a wide computed
column — each against the conditions applied one at a time; these
shapes bailed on the register limit before).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(expr): nulls produced mid-program survive; narrow scratch widened

Two wrong answers in the fused expression compiler, both reachable with
more registers (the wider programs that now compile are exactly the
shapes that chain a null-producing op into further arithmetic):

- A null produced by an instruction — x/0, overflow to Inf, sqrt(<0),
  |INT64_MIN| — was only flagged when that instruction was the LAST one.
  `(abs (/ h b))`, `(as 'I64 (/ h b))`, `(+ (div h b) 1)` and any wider
  chain came back with the sentinel lanes but without HAS_NULLS, so sums
  and averages were poisoned (0Nf) or, after a non-null-aware i64 kernel,
  garbage.  Registers now carry the origin of their nullability: a
  column-derived flag keeps the conservative HAS_NULLS attr; an
  op-generated one makes every downstream kernel null-aware and the
  output is scanned for sentinels, whatever position the producer had.
  A pure-finite result still leaves HAS_NULLS unset.

- An I32/I16 narrowing cast or a comparison result used as an operand
  of an i64 op, a comparison or AND/OR was read as 8-byte lanes:
  `(- a (as 'I32 -1))` gave [15 -4294967275 30], `(+ (> a 15) 1)`
  gave [65793 3 3].  Such sources are widened to the lane type the
  kernel reads (the widening casts map the narrow null sentinels); a
  promotion that cannot be placed bails to the planner instead of
  running the mismatched kernel.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Anton Kundenko <singaraiona@gmail.com>

* perf(select): filtered positional take stops early; sort by an unprojected column (#634)

* perf(select): filtered positional take stops once the answer is known

`(select {… from: T where: <pred> take: K})` with no ordering evaluated
the predicate over the whole table, gathered every passing row and cut
K afterwards; `take: -K` did the same for the last rows.

The fused path now answers it: the table is walked in 64k-row chunks
from the end the answer comes from, one pool task per chunk in that
order, each worker keeping at most |K| passing row ids.  A worker that
holds |K| publishes the row beyond which nothing can be part of the
answer, and every later chunk returns at once.  The lists are merged by
row id and only the |K| rows are gathered.  `asc: key take: K` on a
column marked sorted (no nulls) takes the same path: its first K passing
rows are the K smallest, ties in table order.

Test: fused/fused_take_early_stop (first and last K over many chunks,
AND predicates, fewer matches than K, none, all columns, aliases, K
larger than one worker's share, the sorted-key form).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(select): sort by a source column the projection does not output

`(select {a: a from: T asc: c})` failed with `nyi` whenever the sort did
not go through the fused top-k: the sort runs over the projected table
and looked the key up there.  A sort key that names a source column no
output carries is now projected too and dropped from the result after
the sort (an output with the same name is used as before).

Test: query/sort_key_not_projected (no filter, a filter the fused path
does not take, with and without take, two keys, computed outputs, a key
that is also an output).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(select): hidden sort key vs a renamed output; alter set clears sorted

- An output that renames the key column (`{b: a … asc: a}`) is a bare
  scan the projection names by its source column; the key was added as
  a hidden column under the same name and the drop took the output with
  it.  Such a scan now counts as the key being present, and hidden keys
  are dropped by position (they are projected last).
- `alter … set` wrote into a vector marked sorted and kept the marker;
  the ascending take on a sorted key trusts it.  The marker is cleared
  on write.
- Without a sort, the positional take path admits only a literal K, so
  a `take:` expression is evaluated once, by the general path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(select): fused take/top-k leave alias-chained projections to the planner

Projections resolve left to right: in `{b: a z: b …}` the output z reads
the alias b (column a), not the source column b.  The fused positional
take and the fused top-k gather source columns by name, so z came back
as column b — `take: 2` gave [30 10] instead of [3 1], and the sorted
top-k had the same wrong answer before this branch.  When an output
names an alias bound by an earlier output (other than itself), the
fused paths now decline and the planner resolves the projections.

Tests: fused/fused_take_early_stop (positive and negative take, sorted
take with and without a filter, an identity alias that stays fused).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(select): sort keys bind to projected columns by position, name source columns

The sort of a projection resolved its keys by NAME over the projected
table, whose columns are named by source column for scans and `_e<c>`
for expressions.  Three wrong outcomes followed:

- a hidden key whose source column is called `_e0` collided with the
  first computed output: `{z: (+ a 10) asc: _e0}` sorted by z;
- `{b: a asc: b}` treated the output alias b as the key, found no column
  b in the projection (that scan is named a) and failed with `nyi`;
- hidden keys were capped at 16, so a 17-key sort failed with `nyi`.

A sort key now names a SOURCE column, like where: and by:.  In the
planner it binds to the projected column that scans it — an output, or
the hidden key appended for it — and exec_sort reads a key that is one
of the SELECT's columns by position, so no name lookup can pick another
column.  Output aliases are not consulted.  The hidden-key capacity is
one slot per sort key name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Anton Kundenko <singaraiona@gmail.com>

* perf(group): partitioned dense accumulate; derived-key grouping over the column's distinct values (#633)

* perf(group): partition the slots when the dense accumulate cannot afford one set per worker; pin the SYM vocabulary once per accumulate

The direct-array group path keeps one accumulator set per worker and bounds
that state by the row count.  Past the bound it ran ranged tasks over as
many sets as the budget allowed — one set, i.e. a serial scan of every
row, whenever the slots times the per-group state exceeded the rows, which
a wide key with a nullable min and a strlen average over a large table
does at once (5M slots, three aggregates: budget 1 out of 24 workers).

Instead the slots are partitioned: one parallel pass computes each row's
group and buckets its row id by slot span per worker (no shared writes),
then one task per span drains every worker's bucket for it into the single
accumulator set.  No two tasks touch a slot, nothing is merged, and the
state is one set however many workers run; the buckets cost four bytes a
row.  Row-range tasks stay for the small-slot case and as the fallback
when the buckets cannot be allocated.

Inside the accumulate, strlen over a SYM column pinned the FILE-domain
vocabulary snapshot and ran ray_vec_is_null per row, and the lexical
min/max compare pinned it per compare.  The snapshot is pinned once per
SYM aggregate column for the whole accumulate; null is the id-0 check.

test/rfl/group/dense_slot_split.rfl: 200k rows over 130k slots with a
nullable min/max/sum/avg/first/last, strlen over a SYM column with empty
cells, a two-key composite and a where: selection — each against the same
grouping forced onto the hash path by one far-away key.

* perf(group): decide a derived-key grouping over the column's distinct values when every aggregate reads that column

A grouping keyed by an expression over one SYM column C whose aggregates
are count(C), min(C), max(C), sum(strlen C) or avg(strlen C) is decided by
the distinct values of C: a group's count is the sum of its values' row
counts, its min the smallest of its values, its strlen sum the
length-weighted count.  It used to be evaluated per row — a spread key over
every row and an accumulate of every row's length and lexical compare —
while the key expression itself already ran once per distinct value.

derived_key_vocab_aggs now groups the rows by C once (a column key and a
count, the engine's cheapest grouping, under the select's where:), runs the
key once per distinct value (derived_key_str_chunks) and rewrites the
aggregates over the per-value table: count → sum of counts (null rows
count, as before), avg(strlen) → weighted length sum over the non-null
count, min/max → over the values (null skipped).  Sort and take clauses on
output aliases pass through; the outputs come back in the written order
under the computed key's usual name.  Any other aggregate, a runtime-domain
column or a small table keeps the row path.

Interning the key strings cached dotted segments for every new value — a
host or URL interned every one of its dot-separated parts as a symbol too,
which was most of the cold-run cost.  ray_sym_intern_batch_no_split interns
them as values; a value that later appears as an identifier gets its
segments on that intern.  The per-value length pass runs on the workers.

An `if` whose two STR branches are cheap and total (scans, constants,
substrings, case/trim/length, add/sub/mul, comparisons, and/or/not) takes
the eager descriptor arm instead of compacting the rows for each branch.

test/rfl/group/derived_key_vocab_aggs.rfl: the five aggregates with and
without where:, against the row path (forced by one aggregate over
another column), the same table in memory, the null group's avg and min,
sort/take on an alias, and the key name yielding to an alias.

* fix(group): keep FIRST/LAST groupings past the slot budget on the serial row-order scan

Several FIRST/LAST aggregates share one first_row/last_row per slot (a
documented limitation of da_accum_row), so the answer for a null-mixed
pair depends on the order the rows arrive in; the slot-partitioned drain
walks the workers' buckets, not the rows, and made `last` differ by
worker count.  Past the budget such a grouping stays on the single
accumulator in row order.

* fix(group): partitioned accumulate and vocabulary rewrite independent of core count

The slot-partitioned accumulate scattered rows by pool morsel into
per-worker buckets, so a slot received its rows in scheduler order and
float sums (and var/stddev) rounded differently from run to run.  The
scatter now runs nw tasks over contiguous row ranges in order, bucket
row = task, and the drain walks them in order: each slot gets its rows
in table order, the serial scan's accumulation.

The vocabulary rewrite took its group order from the per-value count
grouping, whose order depends on the core count.  The counts are now
put in domain-position order first (bitmap rank + scatter), so the
output order is the same at every core count.

Tests: dense_slot_split (1e16 + 1 - 1e16 + 1 per group sums to 1.0 in
table order), derived_key_vocab_aggs (unsorted output order pinned).
Both fail with their fix disabled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* perf(csv): parallel splayed load, batched symfile interning, parallel hash index build (#632)

* perf(csv): scan, append and index the splayed load in parallel

Streaming a CSV into a splayed table ran its per-chunk row scan, the
per-column writes and the index build on the calling thread; only the
field parse used the pool.

- Row offsets for a chunk come from the parallel quote-parity scanner
  over a byte window sized from the running average row length, widened
  until it holds the chunk (falls back to the serial limited scan).
- The chunk's columns are appended by one task per column; SYM cells
  go through a direct-mapped runtime-id to position cache in the writer,
  so only a value's first occurrence probes the symfile domain.
- ray_splay_build_indexes builds one column per task.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* perf(csv): intern chunk dictionaries straight into the symfile domain

Every distinct string of a splayed load was interned twice, serially:
once into the runtime symbol table by the chunk materialisation, then
again into the table's symfile domain by the column writer, one cell at
a time under the global domain spinlock — the two together were more
than half of the load time, and the runtime table grew by the whole
vocabulary for nothing.

- csv_materialize_rows takes a target domain: SYM chunk columns are
  built over the symfile domain, and the writer copies their positions.
- ray_sym_domain_intern_batch (domain.c): one batch per chunk over all
  SYM columns.  The existing vocabulary is probed in parallel against a
  snapshot of the reverse index (replaced bucket tables are retired,
  never freed), misses are deduplicated per hash partition, and the new
  atoms are built in parallel into per-partition arena regions
  (ray_arena_alloc_raw / ray_arena_str_at) before the count is
  published.  Index growth re-slots the old table instead of rehashing
  every atom; the symfile flush packs records into a 1 MB buffer.
- SYM fields are hashed by the parse tasks; with a domain target each
  SYM column's dedupe is split by hash partition across tasks.
- The window scan is preceded by a readahead hint and finished chunks
  are dropped from the mapping, so a long file does not stay resident.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* perf(index): build hash indexes partition-parallel

ray_index_attach_hash grouped the rows in one serial open-addressing
pass; on a high-cardinality numeric column of a large table it was the
longest single stretch of a splayed load, and running the columns as
pool tasks left the build itself single-threaded.

Numeric and SYM keys above 64k rows now build in partition-parallel
passes: key words and hash-partition buckets over row ranges, a
per-partition dedupe that marks each group's first row in a bitmap,
then keys, counts, the row scatter and the key table (CAS on the empty
slot) per partition.  A group's number is the rank of its first row
among the marked rows, so the layout is the serial walk's: groups in
first-occurrence order, rows ascending inside a group.  Every array a
pass writes is faulted in beforehand from all workers in slices — a
fresh mapping faulted at random from every worker serialised on the
page-table locks.  STR keys and small columns keep the serial walk.

ray_splay_build_indexes defers the columns that qualify for a hash
index out of the per-column tasks, builds them one after another with
the parallel path and writes all deferred columns together.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(index,domain): cast the CAS targets in place

cppcheck cannot parse a declaration of `_Atomic(T)*` type (AST broken
on the `=`), which failed static analysis; the two batch inserts now
cast at the compare-exchange, the form the rest of domain.c uses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(domain): retire the reverse index on external extend; batch intern tests

The batch intern probes a snapshot of a domain's reverse index outside
the lock, relying on replaced tables being retired rather than freed.
The external-growth path (a reader catching up with a symfile another
process appended to) still freed the table, so a probe running at that
moment could read freed memory.  It now retires it like the rebuild
does.

A probe that meets an entry it cannot compare (no atom and no raw bytes,
possible only if publishing the raw snapshot failed) no longer counts
it as a miss: the whole batch then resolves under the lock, so no
string is appended twice.

The splayed writer's per-column dispatch goes through
ray_pool_par_dispatch_ok like the other sites.  The batch-intern header
comment no longer claims appends in batch order: new strings get
positions grouped by hash partition.

Tests: domain/intern_batch (repeats across partitions, pre-interned
hits keep their positions, repeat batch appends nothing, "" is 0,
flush/reopen); csv_splayed_dedupe_overflow (140k distinct strings in
one SYM column overflow a partition's dictionary in debug builds and
take the row-by-row fallback); index/hash_large_parallel now creates
the pool so the parallel build is the one tested.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(csv): split every window as the whole file; symfile and index bytes independent of cores

A chunk's byte window chose the scanner's quote mode from its own bytes.
In a file with quotes elsewhere, a window without any took the
quote-free fast path, where a lone '\r' does not end a row: two rows
merged and a value was lost (the serial walk and .csv.read split them).
The window now scans in the file's quote mode.

The symfile no longer depends on the worker count.  The chunk batch is
laid out the way the cell-by-cell writer meets the strings — columns in
order, each column's strings by first occurrence (partition runs merged
by the first row, recorded in the dictionary entry) — and the batch
intern appends new strings in batch order.  The hash index key table is
filled in group order on the calling thread, as the serial build does,
so the persisted index is the same bytes on any core count.

A replaced reverse-index table is freed at once unless a batch intern is
probing a snapshot outside the lock (counted under the lock), so a
long-lived reader of a growing symfile keeps no extra tables.

Tests: splay/csv_symfile_order (positions follow the writer order over
three chunks, parallel batch and partitioned dedupe), and
splay/csv_quote_mode_per_file (quotes only in chunk 0, a lone '\r' in
chunk 2: same rows and values as .csv.read).  Both fail with their fix
disabled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Anton Kundenko <singaraiona@gmail.com>

* perf(select): fused top-k prunes chunks by zone extrema and visits the best chunks first (#643)

* perf(select): fused top-k skips chunks whose zone cannot beat the K-th key

The fused top-k evaluated the predicate over every row even when the
first sort key is a stored integer/temporal column whose chunk-zone
index shows most chunks cannot hold a top-K row (a table clustered by
another key, sorted within it, puts a time column's early values in a
few chunks).

Each worker whose heap is full publishes its worst first-key value; the
best of those bounds the final K-th key.  A chunk whose zone minimum
(asc) is above the bound, or maximum (desc) below it, is skipped before
the predicate runs — its pages are never read.  Chunks with nulls are
kept when nulls sort first.

On a 100M-row table ordering by a time column clustered under a
counter (52 of 1526 chunks can hold the top 10): cold 1205 ms -> 157 ms,
warm 30 ms -> 9 ms; with a second sort key 24 ms -> 5.5 ms.

Test: fused/topk_zone_prune (asc, desc, two keys, K near the match
count, all equal to an unindexed in-memory copy).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(select): publish the top-k pruning bound once per morsel

Publishing on every heap replacement made the bound a contended cache
line when rows arrive in the order being sought (a descending top-k over
a column stored ascending replaces the root on nearly every row, and
nothing is pruned there): 25-70% slower than without pruning.  The bound
is now published once per morsel, on its own cache line.  Those shapes
are back to the unpruned time; the prunable ones keep their gain.

Test: topk_zone_prune gains a nullable key with an all-null chunk,
asc/desc and two keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(select): fused top-k visits the most promising chunks first

With a chunk-zone index on the first sort key the K-th key only tightens
as fast as the physical order lets it: an `asc` over a descending column
scans everything before the bound becomes useful, and a key whose best
values sit in the middle prunes nothing until the workers get there.

Order the chunks by their zone extremum (smallest minimum for asc,
largest maximum for desc, null chunks first when nulls sort ahead) and
dispatch the workers over that virtual row space, mapping each slot back
to its chunk.  The K-heaps then fill from the best chunks, the bound is
tight after the first few, and the rest are skipped whatever the layout.
Results are unchanged: the heap keeps the K best rows with a source-row
tie break, so the visiting order cannot alter them.

Local A/B on a 4M-row splayed table (release, 8 cores): asc over a
descending key K=100 43 -> 2.7 ms, K=8000 212 -> 104 ms, with a filter
22 -> 0.2-1.7 ms; already-favourable and unindexed shapes unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(select): `if` over two STR columns under a sparse filter stays on the selected path

Two cheap STR branches route `if` to the eager arm, which picks
descriptors for every row of the table in one pass.  Under a `where:`
that keeps few rows that is a bad trade: the filtered result keeps the
whole intermediate — both parents' bytes — alive for its few rows
(`select {x: (if c S S2) where: (== st 1)}` over 500k rows held 20 MB
for 10k rows; the selected path holds nothing beyond them).

Take the eager arm only when the outer selection keeps more than a
quarter of the rows (the projection's pre-compaction threshold); a
sparse selection evaluates the branches over the selected rows only.
No filter and a filter keeping most rows are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* perf(select): fused top-k re-tests the pruning bound before every morsel

The chunk test ran once at chunk entry, so a bound that tightened while
the chunk was being scanned (another worker's heap filled, or this one's)
only took effect from the next chunk.  Re-test it before every morsel —
one relaxed load and a compare, no writes — and abandon the rest of the
chunk once it cannot beat the bound.

Local A/B on the 4M-row table (release, 8 cores): asc over a descending
key K=100 2.6-3.0 -> 0.4 ms, the same with a filter 0.8-1.4 -> 0.25 ms;
the other shapes unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(collection): hash int cells through f64 when a set meets a float probe (#646)

* fix(collection): hash int cells through f64 when a set meets a float probe

The row hashset behind find/except/union/sect/in hashed int cells with
ray_hash_i64 and float cells with ray_hash_f64.  Numeric equality crosses
the two (atom_eq compares any two numerics through f64, and hs_eq_rows
already defers to it), but equal values landed in different buckets, so
the probe never reached the compare.  A correct answer was a collision:

  (find [1.0 2.0 3.0] [3 1])    [0Nl 0]    -> [2 0]
  (except [1 2 3] [1.0 3.0])    [1 2 3]    -> [2]
  (sect [1 2 3] [2.0 3.0])      []         -> [2 3]
  (union [1 2] [2.0 3.0])       [1 2 2 3]  -> [1 2 3]
  (in [1 0Nl 3] [1.0 0Nf])      [false true false] -> [true true false]

The set gains a num_f64 mode in which int-family cells hash through f64.
The first probe of a new type goes to a cold helper; if it is the other
numeric class (or a list, whose numeric atoms already hash through f64),
the set rehashes once and keeps the mode.  A cached last-probe type keeps
the per-probe cost to one compare, so same-type callers are unchanged.
hs_eq_rows compares mixed numeric typed cells as f64 directly instead of
boxing two atoms per probe.  If the one-off rehash hits OOM, that probe
is answered by a full scan and the next probe retries.

1M x 500k, -c 8, release, 3 alternating runs:

  except/sect i64-i64, f64-f64   within ~1% of dev
  except i64-f64                 153 ms -> 80 ms   (and now correct)
  sect   i64-f64                 155 ms -> 80 ms   (and now correct)
  find   i64 hay / f64 needles   34.5 ms -> 0.1 ms (and now correct)

collection_branch_cov.rfl pinned the old miss -- its comment says the
probe "may miss" because the hashes differ -- and now expects 2.  New
mixed_numeric.rfl covers find/except/union/sect/in in both directions,
narrow ints, fractional misses, both-sides-null in, and sizes that grow
the table before and after the switch.

Closes #645

* fix(collection): skip the no-op rehash of a float-built set; pin ints past 2^53

Review follow-ups on #646.  Only int cells change hash in num_f64 mode, so
a float- or list-built set is already in the f64 layout: set the mode and
skip hashset_rehash unless the source is int-class (except with an i64
probe against a 500k f64 set, 79 -> 73 ms).  mixed_numeric.rfl gains the
float-set/int-probe order and a note, with cases, that ints past 2^53
compare through f64 -- consistent with == -- while same-type ints sharing
a double image stay distinct.

* test(agg): pin the LLC in dense_cache_bound so the route is runner-independent (#647)

agg_contract/dense_cache_bound failed intermittently on the macOS debug CI
job (got strategy 2, expected 1) across unrelated PRs.  The group route
bounds replicated task-local slabs by the probed last-level cache, and
switches to partition ownership when fewer than three slabs fit.  A macOS
CI VM reports a few MB of L2, which fits fewer than three of the test's
2.5 MB slabs, so the route correctly chose partitioning while the test
asserted the task-local strategy.

ray_cache_llc_bytes gains a DEBUG-only pin, ray_cache_llc_set_for_test,
following ray_ipc_set_auto_journal_eval_for_test.  The three platform
probes become static cache_llc_probe and one wrapper honours the pin.  The
test pins 32 MB (the size the route assumes when none is reported) around
the routed query and restores it before any assert can return.  It also
asserts the pinned bound actually bites (3..19 slabs).

Pinning 4 MB instead reproduces the CI failure exactly:
test_agg_contract.c:1759 got 2, expected 1.  make test 3953/3953.

* fix(in): admit the typed kernel unless both sides carry a null (#644)

* fix(in): admit the typed kernel unless both sides carry a null

#607 opened the typed membership kernel to text columns but kept it behind
"both operands exactly null-free".  That is stricter than the semantics
require: the kernel (null matches nothing) and the hashset fallback (null
equals null) disagree only on a null row against a null set element, which
needs a null on each side.  One null row in the column therefore sent the
whole column to the per-row hashset probe.

351,393-row column, -c 8, one null row, null-free needles:

  SYM (in col one)      4,820 us -> 60 us
  I64 (in icol [0 1])   5,730 us -> 80 us
  SYM null-free (in col one)  145 us -> 55 us   (set tested first: a
                                                 null-free set skips the
                                                 column's null scan)

It also fixes wrong answers: the hashset hashes i64 and f64 cells with
different functions, so mixed int/float membership with a null on one side
missed matches -- (in [1.0 0Nf 3.0] [1 3]) was [true false false].  Those
shapes now take the kernel's double promotion.  The both-sides-null
mixed-type case still reaches the hashset and is still wrong; it is a
hashset bug, not a gate one, and is left for a separate fix.

A 23-case null edge matrix is byte-identical before and after apart from
that corrected row.  The in.rfl comment claiming a needle-only null "must
still reject" is corrected, and one-sided nulls across every null-bearing
width are pinned.

* fix(in): temporal equals only its own type; hash large needle sets in the kernel

Review follow-ups on widening the `in` kernel gate.

Temporal types.  The typed kernel classified DATE/TIME/TIMESTAMP as plain
int-family and compared raw payloads, so (in [2000.01.01 0Nd] [0]) was
[true false] while find, except and the hashset `in` -- all atom_eq, which
never matches a DATE to an int or to a TIMESTAMP -- said no.  It also
compared a DATE's day count to a TIMESTAMP's nanoseconds:
(in [2000.01.02] [2000.01.01D00:00:00.000000001]) was [true].  A temporal
operand against a different type is now an empty probe, as SYM vs non-SYM
already was.  Bare `in`, the fused where: and the hashset agree whichever
side carries a null.

Large needle sets.  Past the SIMD small-set size the kernel scanned every
needle per row, O(rows x needles).  The widened gate sent nullable int/float
columns there, and null-free ones already went: a 351k-row column against
100k needles took ~1.2 s.  Sets larger than IN_SIMD_SET (8) now get an
open-addressing table over the live needles -- int64 values, or f64 bit
patterns with -0.0 folded to +0.0 and NaN never inserted -- built once per
call and shared by all workers.  An allocation failure keeps the scan.

351,393 rows, -c 8, release, dev before #644 -> this commit:

  I64 with a null   x 3 needles       6,500 us ->    50 us
  I64 with a null   x 1k             10,600 us -> 1,400 us
  I64 with a null   x 100k           10,000 us -> 2,500 us
  I64 null-free     x 100k        1,155,000 us -> 1,500 us
  F64 with a null   x 100k           10,500 us -> 2,500 us
  select where: in, null-free x 100k  1,155,500 us -> 2,000 us

in.rfl pins temporal vs int and distinct temporal types with nulls on
neither, either and both sides, the fused where:, and the 8/9-needle
boundary, -0.0, NaN, duplicates and narrow ints on the hash path.
make test 3952/3952; reverting exec.c fails in.rfl:169.

* fix(agg): integer mean divides the exact 128-bit sum; mapped index tail unmapped after drop (#648)

* fix(agg): integer mean divides the exact 128-bit sum

The mean of an integer column depended on where it was computed.  The
group engines divided a WRAPPED int64 sum — a whole-table
`select {a: (avg id)}` over signed 64-bit ids answered -5.59e10 for a
true mean of 2.53e18, and any per-group total past 2^63 was garbage —
while the vector path accumulated in double and lost low bits past
2^53, so its last bits moved with the accumulation order.  The
chunk-zone metadata could only answer a mean when every partial sum
stayed below 2^53.

Every integer mean now divides the exact 128-bit sum of its rows,
converted to double by one formula, so the vector reduction, the parted
column path, the columnar, hash-row, scatter and slice group engines,
the streaming engine, pivot and the metadata agree bit for bit:

- the chunk-zone aggregates keep the high word of each chunk's sum in a
  third block (older two-block indexes still load; without high words
  the metadata answers only when the extrema prove the total fits);
- the keyless reduction and the group accumulators carry a high word
  beside the wrapped sum, allocated only when an integer avg over an
  I64 / TIMESTAMP column (or a table of 2^31 rows or more) is present —
  narrower inputs cannot leave int64 and keep their plain sum;
- the streaming engine keeps its 16-byte state: a signed running sum
  plus the wrap count packed with the row count, exact while a group
  has fewer than 2^32 rows (larger tables stay on the legacy engines);
  narrow inputs use an exact int64 state;
- a bare I64 column sums in blocks whose mode (values below 2^53, below
  2^53 in magnitude, or general) is escalated once.

`sum` keeps its int64 wraparound contract.  The grouped combine over a
parted table (per-partition sums joined afterwards) is not changed here
and still divides the wrapped total.

Tests: group/avg_exact_i128 (exact means over values that overflow
int64 within every group, across the keyless, columnar, hash, scatter,
slice, streaming and pivot paths, nulls, narrow types, a splayed copy),
test_agg_engine avg_exact_i128_engines (v2 on and off),
store/splayed_zone_aggs (metadata means past 2^53 and past 2^63).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(store): a mapped column that drops its index still unmaps the index tail

A column loaded by mmap with an inline index is longer than its payload,
and ray_free sized the unmap from the attached index.  Once the loaded
column itself dropped that index — an in-place edit of its only
reference detaches a mapped index without releasing it — or when
ray_index_inline_map had discarded a child that is still in the file,
the unmap covered the payload only and the index tail stayed mapped
for the life of the process.

A mapped column that carries an index now registers its region in the
file-map registry under its own address with the true mapped length,
the way a string column's pool already does, and ray_free consults the
registry for every mapped column before falling back to the size …
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants