fix(agg): integer mean divides the exact 128-bit sum; mapped index tail unmapped after drop - #648
Conversation
The mean of an integer column depended on where it was computed. The
group engines divided a WRAPPED int64 sum — a whole-table
`select {a: (avg id)}` over signed 64-bit ids answered -5.59e10 for a
true mean of 2.53e18, and any per-group total past 2^63 was garbage —
while the vector path accumulated in double and lost low bits past
2^53, so its last bits moved with the accumulation order. The
chunk-zone metadata could only answer a mean when every partial sum
stayed below 2^53.
Every integer mean now divides the exact 128-bit sum of its rows,
converted to double by one formula, so the vector reduction, the parted
column path, the columnar, hash-row, scatter and slice group engines,
the streaming engine, pivot and the metadata agree bit for bit:
- the chunk-zone aggregates keep the high word of each chunk's sum in a
third block (older two-block indexes still load; without high words
the metadata answers only when the extrema prove the total fits);
- the keyless reduction and the group accumulators carry a high word
beside the wrapped sum, allocated only when an integer avg over an
I64 / TIMESTAMP column (or a table of 2^31 rows or more) is present —
narrower inputs cannot leave int64 and keep their plain sum;
- the streaming engine keeps its 16-byte state: a signed running sum
plus the wrap count packed with the row count, exact while a group
has fewer than 2^32 rows (larger tables stay on the legacy engines);
narrow inputs use an exact int64 state;
- a bare I64 column sums in blocks whose mode (values below 2^53, below
2^53 in magnitude, or general) is escalated once.
`sum` keeps its int64 wraparound contract. The grouped combine over a
parted table (per-partition sums joined afterwards) is not changed here
and still divides the wrapped total.
Tests: group/avg_exact_i128 (exact means over values that overflow
int64 within every group, across the keyless, columnar, hash, scatter,
slice, streaming and pivot paths, nulls, narrow types, a splayed copy),
test_agg_engine avg_exact_i128_engines (v2 on and off),
store/splayed_zone_aggs (metadata means past 2^53 and past 2^63).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ex tail A column loaded by mmap with an inline index is longer than its payload, and ray_free sized the unmap from the attached index. Once the loaded column itself dropped that index — an in-place edit of its only reference detaches a mapped index without releasing it — or when ray_index_inline_map had discarded a child that is still in the file, the unmap covered the payload only and the index tail stayed mapped for the life of the process. A mapped column that carries an index now registers its region in the file-map registry under its own address with the true mapped length, the way a string column's pool already does, and ray_free consults the registry for every mapped column before falling back to the size it can derive. Test: index/mapped_drop_unmaps_tail (the tail page must be unmapped after the sole reference drops its index and is freed). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The index region must reach a page of its own for the test to observe the short unmap; with 16 KiB pages (Apple silicon) the 4 KiB layout put the whole file in one page and msync on a misaligned address failed. The payload now ends 64 bytes short of the second page whatever the page size, and msync is issued on a page-aligned range. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Reviewed the diff at
Two findings: 1. TIME read as int64 in the slice-group fast branch (medium, predates this PR, but this PR routes the new path through it). if (contig && (t == RAY_I64 || t == RAY_TIME)) {reads the column as The same pattern is at I couldn't trigger it from the language: an unfiltered and a filtered 2. |
…or consults the registry The slice-group fast branches took RAY_TIME through the int64 path (`t == RAY_I64 || t == RAY_TIME`, and the F64 x int product's `tb == RAY_I64 || tb == RAY_TIME`), although a TIME cell is a 4-byte millisecond count: a contiguous chunk of n rows would read 2n int32 cells, pair neighbouring times into one value and run past the column. TIME and DATE now take the 4-byte branch with I32; the int64 branch serves I64 and TIMESTAMP. mapped_block_bytes, the read-only mirror of the size ray_free hands to the unmap, consulted the file-map registry for string columns only; since an indexed mapped column registers its region too, it consults the registry for every mapped column, as ray_free does. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Re-checked One nit, a leftover from the TIME change: the shared-stream pairing gate at if (it != RAY_I64 && it != RAY_TIME && it != RAY_I32) continue;After this commit, no typed branch in |
…roduct branches handle The F64 x int product pairing still admitted a TIME int side after TIME moved to the 4-byte branch, while no typed branch of sg_prod_range handles it — the generic branch zeroes the sibling sum under a comment that relies on the gate. Admit I64 and I32 only, in step with the branches (not reachable from the language: a temporal factor is rejected before a product forms). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(group): bound the tied groups the radix top-N selection keeps
The native top-N of the radix path kept every group at or beyond the
threshold. Groups strictly beyond it number fewer than N, but the
groups AT the threshold can be nearly all of them: a count-per-group
top-10 over a near-unique key has a threshold of 1, every group ties
it, and the "kept superset" is the whole grouping, heap-sorted by
first row — a comparison sort over 100M pairs for a query whose answer
is ten rows.
The emitted prefix is the tied groups' first-seen order, so only the N
tied groups with the smallest first row can ever be taken. The
selection now keeps the strictly-better groups plus exactly those N,
through a bounded max-heap, and sorts at most 2N entries. The rows
emitted are the ones the full kept set produced.
Tests: the pinned kept count of the radix native top-N case becomes N;
a near-unique two-key grouping whose threshold every group ties, in
both directions and under a row selection, against the full grouping.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(lang): add a `while` special form
Rayfall had no way to stop iterating before the end of a sequence. Every
iteration primitive — map, pmap, fold, fold-left/right, scan, scan-left/right,
prior — consumes its whole input, and there was no loop construct, so a
"repeat until done" loop had to be written as a fold over a fixed range whose
full length was paid on every call however early the work finished. `return`
did not help: it exits the lambda, not the iteration, so the fold kept calling
the lambda for every remaining element. Recursion was not an alternative
either, since there is no TCE and the stack tops out around 1-2k frames.
(while cond body...) evaluates cond, and while it is truthy evaluates each
body expression in order, then tests again. It always returns null — a
statement form run for effect, and a never-taken loop has no last value to
report. Zero body expressions is legal, so a condition with side effects can
be the whole loop; that is the shape a drain wants, where there is no sequence
to iterate over and the range was only ever scaffolding.
Implemented on both evaluators, which must agree:
- ray_while_fn (eval.c), registered beside `if` and `do`.
- an sf_while case in the bytecode compiler emitting JMPF over a backward
JMP — the first backward branch the compiler produces. The VM's op_jmp
already checked for a pending interrupt when the displacement is negative
(dormant until now), so a runaway loop is Ctrl-C-able for free.
Compiling the body inline, rather than letting it fall through to the generic
special-form path, is what keeps `return` working inside a loop: it reaches
the sf_return case and unwinds the lambda instead of degrading to the tree
walker's identity `return`.
Neither path pushes a scope around the body. `let` binds only in the top
frame (env_bind_local), so a per-iteration frame would discard loop-carried
`let` state on the tree-walking path while the compiled path — whose `let`
writes a function-level slot — kept it, and the two evaluators would disagree
on the same source. A caller wanting a fresh frame per pass writes
(while cond (do ...)), which composes.
patch_jump and emit_jump_back now share write_jump_offset; an out-of-range
displacement still sets c->error, dropping the lambda to the interpreter.
Measured on the drain shape from the issue (4-step drain under a 4096 safety
bound, release build, driver baseline subtracted): fold-left over (til 4096)
1420 us/batch, the nested 64x64 workaround 45.8 us/batch, `while` 3.8
us/batch — 374x over the original and 12x over the workaround.
Closes #588
* feat(lang): add `times` and `fold-while`
The two remaining early-termination forms requested in #588, alongside the
`while` that landed in #590. Neither is an unblock — `while` already covers
the reporter's case — but both complete the vocabulary he asked for.
`times`
-------
(times n body...) runs the body exactly n times and returns null. `do` is
already progn in this language, so the bounded loop could not reuse that name;
overloading `do` on an integer head was rejected as genuinely ambiguous —
(do 5) would have to mean either "loop five times over nothing" or "return 5",
and a computed first expression that happened to be an integer would silently
change meaning.
The count is evaluated ONCE on entry, so a body mutating whatever produced it
cannot change how many passes remain. A count of zero or less runs zero times
rather than trapping; a non-integer is a type error.
Compiled as a counted loop over a hidden local slot, reusing the backward
branch added for `while`. ray_times_norm_fn type-checks the count and clamps
a negative bound to zero on entry, which lets the per-pass test be a bare
truthiness check on the counter — 0 is falsy, so no comparison call is needed
per pass and a negative bound cannot run away. That helper and the decrement
are pushed as constant-pool objects rather than resolved by name, so the
loop's own arithmetic is unnamable from source and cannot be swapped out by a
`(set - ...)` override. The counter's slot is addressed by index and its sym
carries a space, so no source token can collide with it and nested `times`
counters stay apart. As with `while`, the body compiles inline, so `return`
unwinds the enclosing lambda from inside the loop.
fold-while
----------
(fold-while pred f init xs) offers the accumulator to `pred` before each step
and stops on a falsy answer, yielding the accumulator as it stands. The test
precedes the first element, so a predicate false at the start returns `init`
untouched. The predicate takes the accumulator rather than the element: that
is the form that expresses "iterate until the running result says stop", which
is the early termination actually being asked for.
One deliberate divergence from ray_fold_fn: that routes its collection through
unbox_vec_arg -> to_boxed_list, boxing every element up front. For a
primitive whose purpose is to stop early, paying for the tail it never reaches
is the cost being removed, so elements are pulled one at a time via
collection_elem. A plain variadic builtin — no compiler work, since it
dispatches through the normal call path.
Measured, release builds: `times` 100 ns/pass against 180 ns for `while` plus
a manual counter over 1e6 passes. `fold-while` stopping after three elements
of a 1e6-element vector costs 3.9 us against 369 ms for the `fold-left`
equivalent, which had to box and walk all million to discover it was done.
Tests: test/rfl/lang/times.rfl and test/rfl/collection/fold_while.rfl, each
behaviour asserted on both evaluator paths where applicable — including the
count validation and negative clamp on the compiled path, which runs through
entirely different code from the tree walker's check.
* perf(join): skip per-cell key null tests on provably null-free columns
ray_vec_is_null is out-of-line (no LTO) and the join called it once per key
column per row in hash_row_keys — on the build side, the probe side, and the
prefetch lookahead — plus twice per key column per hash-chain step in
join_keys_eq, across both the count and the fill pass. A reported profile
put 13.28% of a service's samples there, on a join keyed by two SYM columns
that structurally never hold a null.
Prove once per join that no key column can hold a null and drop the call.
SYM/STR nulls are canonical empty payloads (id 0 / length 0) that HAS_NULLS
does not track, so text columns are proven by the chunked zero-scan from
#533 rather than by the flag; everything else reads the flag through slices
and takes a set bit at face value, so the proof stays O(n) and never
degrades into ray_vec_has_nulls' per-element walk. Flag-readable columns
are settled first, so a nullable numeric key short-circuits before any text
column is scanned. OP_CONST (atom) key slots are refused as unprovable.
Also skip the #458 null-run pre-scan under the proof. It gates on
ray_vec_may_have_nulls, which is unconditionally true for SYM/STR, so a
SYM-keyed join ran a full per-key-per-build-row ray_vec_is_null scan before
the join proper on every execution; a null-free key set cannot contain an
all-null row, so the scan is dead.
Measured with bench/join_nullfree (release, 1:1 book join, 4M probe x 500K
build): two SYM keys 275 -> 249 ms (-9.5% median, -7.8% min); single I64 key
106 -> 99 ms (-6.3% median, -7.9% min). A proof that fails on a text key
costs a partial scan (~8ms on a 4M-row SYM column) and gains nothing — only
nullable SYM/STR key columns pay it.
ray_join_force_null_checks forces the null-aware loops so the differential
tests and the perf gate can compare both paths in one binary;
ray_join_nullfree_keys counts the joins that took the fast path.
Closes #597
* fix(system): validate launcher and timeit integer arguments
* fix(cli): reject a flag as another flag's value, and unknown options
Every value-taking startup flag was gated on `i + 1 < argc` and then
consumed argv[++i] blindly, so one flag silently ate the next. The shape
that was reported is the worst one for a service:
rayforce -Q -p 5099 svc.rfl
`-Q` took `-p` as its value, the port was never bound, and nothing on
stdout or stderr said so. Under a supervisor the process starts, the
script runs, the unit looks healthy, and every client gets connection
refused. `-c` and `-t` had the identical hole.
Two neighbouring cases came out of the same code. A trailing flag with
no value fell through to the positional-file arm and reported the
misleading `cannot open '-Q'`. And that arm took ANY unrecognized token
as the script name, which the real positional then overwrote, so a
typo'd option was not merely ignored — it was silently swallowed:
`rayforce -x svc.rfl` ran the script as if `-x` had never been typed.
Every value-taking flag now goes through flag_value(), which refuses a
missing value and a value that is EXACTLY one of this program's flag
tokens (including the `--` app-args terminator), with a diagnostic and
exit 2 before the script is loaded. A value that merely STARTS with '-'
stays legal: a password or a negative number is a legitimate value, and
refusing those would break working command lines for no gain. An
unrecognized `-`-prefixed token is now an error too; a lone `-` is still
a positional, and tokens after `--` belong to the app and are never
validated.
`-Q, --querylog N` was a real flag, referenced twice in the docs, that
--help never listed beside its sibling `-t, --timeit N` — added, along
with `[-Q 0|1]` in the synopsis.
The suggestion to fail when `-p` was requested but nothing bound is
already in (#473, listen_fatal.rfl), and cannot catch this: the parser
never saw the eaten `-p`, so there was no request to check against. The
guard on the value is what closes the class.
Also: the two pre-existing `-p` validation bail-outs now release the
runtime like every other error path, since these paths are exercised
under ASan by the new test.
Closes #600
* fix(exec): release the input table of window, sort and limit unconditionally
Every node the executor evaluates returns an owned reference — the
constant table node included, it retains its literal. The OP_WINDOW,
OP_SORT and OP_HEAD cases released their input only when it differed
from the graph's table, taking an equal pointer for a borrowed one. A
query's root is a constant node over that very table, so the reference
was never released: one input table per windowed, sorted or limited
query. Invisible while a global kept the table alive; a whole table per
call when the input was built for that call, as a service that windowed
a freshly concatenated buffer on a timer found (#602).
The four cases now release the input on every path, as the join and
the plain head/tail cases already did.
Test: window, sorted and limited selects over a table built per call,
measured with the two-window bytes-allocated method of the other memory
probes, plus a shared input reused across fifty calls.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(lang): pin the refcount of a table across window, sort and limit queries
Twenty calls of each shape over one global table; the table's refcount
must be exactly what it was before, and the table still whole after.
Fails on the unfixed executor at the first window shape.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(exec): the reduction case owns its input too
The same borrowed-query-table guard sat under the reductions; no child
evaluates to the query table there today, so nothing leaked, but the
premise is the one the sort, window and limit cases just dropped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* perf(pool): wake only the workers a dispatch can keep busy
ray_pool_dispatch and ray_pool_dispatch_n signalled the whole pool on
every dispatch, however narrow the window. The main thread participates
as worker 0, so at most n_tasks-1 helpers can ever claim anything; the
surplus threads woke, raced to an already-drained window, and went
straight back to the semaphore. On a dispatch narrower than the machine
that is pure overhead — a partition-parallel step over three partitions
woke every core on the box to run two tasks.
Signal min(n_tasks-1, n_workers) instead. Signalling FEWER is safe:
completion is governed by the `pending` spin-wait, never by the signal
count, so an unsignalled worker simply stays asleep, and under-signalling
cannot produce the surplus-signal problem between consecutive dispatches
that the spin-wait exists to avoid.
Measured (10M-row ClickBench, 43 queries x 3, splayed): sum of per-query
hot times 5690.7 ms -> 5667.3 ms (-0.4%, within noise), user CPU 55.1 s
-> 54.8 s. So this is waste removal rather than a speedup on workloads
whose dispatches are already wider than the pool.
Not a fix for #599. That report's shape — a 130k-row join probe, which
splits into 16 tasks against 8 workers — is unaffected, because the clamp
lands on the same worker count it already used (measured: 5.19 s -> 5.18 s
of CPU, unchanged). The investigation there points at the pool defaulting
to logical rather than physical cores; that is a separate change with its
own benchmark trade and is filed on its own.
test_pool.c gains pool/dispatch_narrow, which is the case the existing
coverage missed: dispatch_n_small uses 4 tasks on a 2-worker pool, so the
clamp never binds. The new test drives windows of 1..4 tasks on a
4-worker pool, then alternates narrow and wide windows 25 times over the
same pool, asserting exact task and element counts throughout — the two
risks the change introduces are a task nobody claims and signal
accounting that drifts across dispatches and starves a later wide window.
* build: Windows (MSYS2 CLANG64/MINGW64) toolchain in the Makefile
Detect Windows via $(OS) or uname (MSYS2's make hides $(OS)), link
Winsock and a statically linked winpthreads, build with 64-bit off_t
(_FILE_OFFSET_BITS=64) and an 8 MiB stack like a Linux main thread.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(core): WSAPoll event loop for Windows
Replace the IOCP stub with a readiness-based loop that mirrors the epoll
backend's dispatch order, so the selector state machine is identical on
every platform. stdin (console or pipe) is not a socket, so RAY_SEL_STDIN
selectors are probed directly and the socket wait is sliced while one is
registered.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(net): Winsock errno mapping, exclusive bind and IPC on Windows
- mirror WSAGetLastError()/SO_ERROR into errno for send/recv/connect/
accept/bind/listen: callers branch on EAGAIN/ECONNREFUSED;
- ipc_send_fn always overwrites errno, so a stale EAGAIN can no longer
park a frame on a dead socket;
- SO_EXCLUSIVEADDRUSE instead of SO_REUSEADDR (which lets a bind steal
a port that is in use);
- WSAStartup at load time; sock.h pulls in platform.h;
- verbose IPC capture uses a temp file that works without admin rights;
- the SIGURG out-of-band cancel stays POSIX-only.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: Windows platform layer (file mapping, pools, crash report, paths)
- ray_vm_unmap_file only unmaps at a view's own base: UnmapViewOfFile
drops the whole view for an interior pointer, which freed columns that
carry a passenger index (munmap is a no-op there);
- ray_vm_alloc_aligned returns its own allocation base, so pools are
really released by ray_vm_free;
- ray_vm_map_fd_ro maps for real (CSV reads always failed with io);
- crash report via SetUnhandledExceptionFilter;
- heap: file-backed spill stays POSIX-only (docs/architecture/memory.md);
- domain: realpath substitute; symfile paths compare case/separator-
insensitively so one file never gets two domains;
- aof/csr/hnsw: platform handle for fsync, portable ray_mkdir.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: REPL, profiler and system builtins build and work on Windows
- term.h/profile.h include platform.h and a lean <windows.h>;
- term_write for the Windows console, errno.h and core count in the REPL;
- .sys.info reports page-size and total-mem on Windows too;
- KEY_READ no longer collides with <winreg.h>.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: run the suite on Windows
- runner: ';; @requires: posix' marks .rfl files whose fixtures or checks
need a POSIX shell/filesystem; on Windows they are reported as SKIP;
- shell-free ray_test_rm_rf / ray_test_mkdir_p replace system("rm -rf")
and "mkdir -p" in C tests; test.h maps the few POSIX helpers tests use;
- tests of POSIX-only behaviour (setrlimit, ENOTDIR, read-only dirs,
AF_UNIX, file spill, Winsock send buffering) skip with the reason;
- fixture fixes that were latent on any platform: binary-mode CSV
fixtures, a per-test AOF dir, journal closed before the crash rename.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: Windows build instructions and platform status
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: readable REPL prompt and CPU name in the banner on Windows
No console font Windows ships has U+2023 (checked Consolas, Cascadia Mono,
Lucida Console, Courier New) and the classic console does no font fallback,
so the prompt rendered as '?'. Use the nearest filled triangle they all do
have, U+25BA, which is also three UTF-8 bytes.
The banner's CPU line said 'unknown': read ProcessorNameString instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* bench: build the micro-benchmarks on Windows, record a Windows/Linux check
alloc and agg_v2 read peak RSS through getrusage; use
GetProcessMemoryInfo there. windows_vs_linux.md records the numbers the
port was checked against.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* bench: record the recent perf benches on Windows vs Linux
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(format): render i64 at full width, not through 32-bit long
Interpolation (format/println), print and the pivot / column-name helpers
formatted an i64 with "%ld" and a (long) cast. long is 64-bit on LP64 and
32-bit on Windows, so there 10^18 printed as -1486618624 and a pivot keyed
on values above 2^31 produced truncated column NAMES — a data defect, not
only a display one. Use PRId64 throughout; regression test included.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(crash): format the banner with snprintf instead of macro string splicing
The banner was spliced from string literals and the RAYFORCE_VERSION /
RAYFORCE_GIT_COMMIT macros inside #ifdef arms; without the -D values a
static analyser reads the literal followed by the bare macro name as two
adjacent tokens and reports a syntax error. Format it with one snprintf
at install time (the handler itself still never formats), with the
macros defaulting to empty strings. Output unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(regress): drop the 24 GB til from the i64 width probe
`til` is eager, so counting a three-billion-element range built a 24 GB
vector for no extra coverage — the neighbouring literals already cross
2^31, 2^32 and 2^63.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(mem): hand workers their freed blocks back at the end of every dispatch (#621)
* fix(mem): hand workers their freed blocks back at the end of every dispatch
A parallel operator's workers allocate per-task buffers from their own
heaps and the main thread frees them once the dispatch has completed, so
each block lands on the owning worker's foreign list. The owner drained
that list only when an allocation found its freelists empty, which a
warm worker with a partly cut pool rarely does: it kept splitting fresh
pool space instead, touching new pages every round while its own freed
blocks waited. Under a steady stream of such operators (a join against a
large table per incoming batch) the process grew by a full 32 MB pool
per worker before any block was reused, and only the idle decay, which
needs the process to sit quiet, drained it earlier.
The dispatcher now drains every registered heap's foreign list at the
end of each parallel region (ray_heap_reclaim_workers), under the same
conditions the idle decay already relies on: the flag is clear, so every
worker has finished its last task and neither allocates nor frees until
the next dispatch. The blocks go back to freelists only, no pages are
released, so the next round reuses them without faulting.
Tests: heap/reclaim_workers_drains_owner (a block owned by one heap and
freed from another leaves the owner's list on reclaim and is handed out
again at the same address, no new pool) and
pool/dispatch_reclaims_worker_blocks (blocks allocated by real workers
and freed by the main thread are back with their owners by the end of
the next dispatch).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(mem): reclaim only the pool's own worker heaps after a dispatch
The heap registry holds every thread that ever allocated, not only pool
workers: a server's poll thread or an embedding's own threads are there
too, and they may be allocating while a dispatch on another thread ends.
Draining such a heap's foreign list from the dispatcher coalesces into
freelists their owner is using at that moment. Only a pool worker parked
on its semaphore is known to touch nothing of its own until the next
dispatch, so each worker now publishes its heap in the pool
(worker_heaps, set after ray_heap_init, cleared before it exits) and the
dispatcher reclaims exactly those, plus its own list on its own thread.
ray_heap_reclaim_worker takes one heap; the registry walk is gone, which
also removes its per-dispatch scan of every registry slot.
The heap test now checks the block's bytes leave the owner's books only
on reclaim, instead of expecting the next allocation at the same
address; the pool test reads the worker heaps the pool publishes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(pool,heap): make the reclaim tests independent of worker start-up and adopted heaps
The pool test read a worker's heap slot right after ray_pool_create,
which returns before the workers have started, and passed vacuously when
the main thread took every task. It now waits for each worker to publish
its heap, records which worker allocated each block and repeats the
round until a worker really allocated something, then checks that the
next dispatch left no worker with a foreign list. The heap test drains
whatever an adopted heap already held before taking its baseline.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(pool): allocate the worker heap slots without an _Atomic cast
cppcheck cannot build an AST for a cast to `_Atomic(void*)*` and fails
the static-analysis run on it. The cast is not needed in C: ray_sys_alloc
returns void*, which converts to the slot pointer type on its own, and
sizeof(*pool->worker_heaps) names the element size without spelling the
type.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(serde): reject trailing payload bytes (#622)
* fix(select): a projection sees the projections before it (#620)
* perf(like): one column kernel for the direct builtin and the executor, one row pass (#623)
* fix(system): report arity errors for extra arguments (#616)
* fix(in): admit the typed kernel for text columns, which it never reached (#607)
* perf: shared-node memo, STR descriptor views, fused top-k on nullable like, bulk derived key, parallel dense grouping (#626)
* v2.9.1 (#625) (#629)
* docs(release): merge master back into dev after each release
A merged release PR leaves a commit on master that dev lacks, so the next
release PR is blocked as out of date with the base, and after a squash it
also shows conflicts that aren't real. Add the back-merge as step 5.
* perf(select): whole-table count/min/max/sum/avg from chunk-zone metadata
A select of whole-table aggregates with no other clause scanned every
row, although a stored column already carries per-chunk min/max in its
chunk-zone index.
- The chunk-zone index of an integer or temporal column also records,
per chunk, the sum of its non-null values (int64 with wraparound, as
the engine sums) and their count, in a fourth child vector. Indexes
written before hold zero in that slot and load without it.
- `sum` of an integer column reads the per-chunk sums; `avg` does too
when no partial sum can leave double's integer range, so the value is
exactly the row-wise one (otherwise it scans as before).
- `(select {from: T a: (sum c) b: (min d) …})` with no where/by/take/
sort, every output a count / min / max / sum / avg of a plain column
the metadata answers, is built from the metadata in O(chunks). Any
other output leaves the query to the planner.
Test: store/splayed_zone_aggs (each answer, nulls and types equal to the
same query over an in-memory copy; float sums, a filter and an avg past
2^53 take the usual path and still agree; scalar sum/avg).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(index): mapped index regions — bounds-checked children, never released on drop
- A column's inline index region is mapped with the column file, and its
child vectors are addressed by region-relative offsets. The offsets
were trusted as written. A region re-saved by a binary that knows
fewer child slots keeps a stale offset in the slot it did not write
(the new chunk-zone aggregates slot), pointing at or past the end of
the file; the free path then sized its munmap from it and unmapped a
page it did not own. ray_index_inline_map now takes the region size
and admits only children that lie inside it; an out-of-bounds or
malformed aggregates child is dropped, any other one leaves the column
unindexed.
- Copies of a mapped column borrow its index without a reference.
Dropping the index from such a copy (a delete or an in-place write on
it) released the index anyway, freeing the loaded column's children
in place: its min/max then read NULL and crashed. A mapped index is
now only detached from the vector; the mapping's owner unmaps it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(agg): zone min/max tell all-null from the non-null counts; empty tables keep the planner
- The zone min/max treated INT64_MAX / INT64_MIN as "no non-null value",
so a column of int64 maxima answered null. When the zone carries
non-null counts they decide; min/max also require the zone's extrema
to be present.
- The metadata select no longer answers an empty table (its count
differed from the planner's there).
Tests: store/splayed_zone_aggs (edit through a copy of a loaded column
then aggregate the original; an all-INT64_MAX column; an empty filter).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(store): portable in-place sed in splayed_zone_aggs
BSD sed takes the argument after -i as the backup suffix, so the null
injection failed on macOS; use -i.bak and remove the backup.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(system): reject arguments to gc (#615)
* fix(system): reject arguments to gc
* docs(memory): use zero-argument gc form
* fix(system): report gc arity errors
* fix(string): a STR view keeps only the bytes it points at (#631)
* fix(string): a STR view keeps only the bytes it points at
The descriptor views introduced for substr and `if` over STR columns pinned
more than they used. An `if` whose two sides came from different pools —
including two columns of one table under a where: that kept a few rows —
copied BOTH whole parent pools into the result (the compacted branch
inputs share the full column pool, so a 10k-row result out of 500k rows
carried the two columns' 43 MB); the eager arm did the same for an
unfiltered `if`. A substr with a per-row start or length retained the
parent pool even when every result fitted inline and nothing pointed
into it.
`if` now builds a pool of exactly the chosen rows' bytes whenever the
sides come from different pools, a pooled scalar is involved, or the rows
keep a small share (under an eighth) of one shared pool; one shared pool
whose rows keep most of it is still pointed into as before, so a
derived-key expression over one column stays copy-free. A substr view
whose descriptors are all inline drops its pool reference; one that
points into the pool stays a view (the column is alive anyway, and a
sibling `if` over two such views can then pick either side without a
copy).
test/rfl/strop/str_view_pool_compact.rfl pins the retained bytes through
direct-bytes: an `if` under a where over two columns, over two substring
views, with a pooled scalar side, the eager `if` against the two-pool
sum, and an inline substr view outliving its table.
* test(string): cover the eager `if` over two STR pools
The case labelled eager in str_view_pool_compact went through the
selected path (a per-row condition column is always a row mask there),
so the eager arm's two-pool offset/copy had no regression coverage.
A scalar condition (a variable) is not a row mask: the selected path
declines and the eager fill runs over two different pools. New cases
take it at 500k rows (parallel fill) and 1000 rows (below the parallel
threshold, serial fill), for both sides, check every row after all the
sources are released, and bound the retained bytes to one side's pool.
On the pre-fix tree the retention checks fail. The old case is
relabelled as the per-row selected path it is.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Anton Kundenko <singaraiona@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(log): reject arguments to purge (#638)
* perf(vec): avoid rewriting copied text nulls during concat (#641)
Co-authored-by: singaraiona <3473381+singaraiona@users.noreply.github.com>
* fix(hnsw): reject invalid persisted graph ids (#642)
* fix(log): reject arguments to roll (#639)
* perf(expr): 32 registers per fused expression; nulls produced mid-program, narrow scratch (#636)
* perf(expr): 32 registers / 96 instructions per fused expression
A filter of four conditions, one of them a `within`, needs more than 16
expression registers; the compiler then bailed and the whole predicate
was evaluated column at a time over every row, without the chunk-zone
morsel decisions. A grouping under
`(and (== CounterID 62) (within EventDate [..]) (== IsRefresh 0) (== DontCountHits 0))`
on a 100M-row table took 71 ms warm / 540 ms cold, against 24 ms for
the same filter with three conditions.
The limits are now 32 registers and 96 instructions; the per-task
scratch is sized to the registers the expression uses instead of the
maximum. The same query: 24 ms warm, 130 ms cold, same answer.
Test: expr/wide_predicate (four and six conditions with arithmetic and
`within`, a grouping under the four-condition filter, a wide computed
column — each against the conditions applied one at a time; these
shapes bailed on the register limit before).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(expr): nulls produced mid-program survive; narrow scratch widened
Two wrong answers in the fused expression compiler, both reachable with
more registers (the wider programs that now compile are exactly the
shapes that chain a null-producing op into further arithmetic):
- A null produced by an instruction — x/0, overflow to Inf, sqrt(<0),
|INT64_MIN| — was only flagged when that instruction was the LAST one.
`(abs (/ h b))`, `(as 'I64 (/ h b))`, `(+ (div h b) 1)` and any wider
chain came back with the sentinel lanes but without HAS_NULLS, so sums
and averages were poisoned (0Nf) or, after a non-null-aware i64 kernel,
garbage. Registers now carry the origin of their nullability: a
column-derived flag keeps the conservative HAS_NULLS attr; an
op-generated one makes every downstream kernel null-aware and the
output is scanned for sentinels, whatever position the producer had.
A pure-finite result still leaves HAS_NULLS unset.
- An I32/I16 narrowing cast or a comparison result used as an operand
of an i64 op, a comparison or AND/OR was read as 8-byte lanes:
`(- a (as 'I32 -1))` gave [15 -4294967275 30], `(+ (> a 15) 1)`
gave [65793 3 3]. Such sources are widened to the lane type the
kernel reads (the widening casts map the narrow null sentinels); a
promotion that cannot be placed bails to the planner instead of
running the mismatched kernel.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Anton Kundenko <singaraiona@gmail.com>
* perf(select): filtered positional take stops early; sort by an unprojected column (#634)
* perf(select): filtered positional take stops once the answer is known
`(select {… from: T where: <pred> take: K})` with no ordering evaluated
the predicate over the whole table, gathered every passing row and cut
K afterwards; `take: -K` did the same for the last rows.
The fused path now answers it: the table is walked in 64k-row chunks
from the end the answer comes from, one pool task per chunk in that
order, each worker keeping at most |K| passing row ids. A worker that
holds |K| publishes the row beyond which nothing can be part of the
answer, and every later chunk returns at once. The lists are merged by
row id and only the |K| rows are gathered. `asc: key take: K` on a
column marked sorted (no nulls) takes the same path: its first K passing
rows are the K smallest, ties in table order.
Test: fused/fused_take_early_stop (first and last K over many chunks,
AND predicates, fewer matches than K, none, all columns, aliases, K
larger than one worker's share, the sorted-key form).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(select): sort by a source column the projection does not output
`(select {a: a from: T asc: c})` failed with `nyi` whenever the sort did
not go through the fused top-k: the sort runs over the projected table
and looked the key up there. A sort key that names a source column no
output carries is now projected too and dropped from the result after
the sort (an output with the same name is used as before).
Test: query/sort_key_not_projected (no filter, a filter the fused path
does not take, with and without take, two keys, computed outputs, a key
that is also an output).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(select): hidden sort key vs a renamed output; alter set clears sorted
- An output that renames the key column (`{b: a … asc: a}`) is a bare
scan the projection names by its source column; the key was added as
a hidden column under the same name and the drop took the output with
it. Such a scan now counts as the key being present, and hidden keys
are dropped by position (they are projected last).
- `alter … set` wrote into a vector marked sorted and kept the marker;
the ascending take on a sorted key trusts it. The marker is cleared
on write.
- Without a sort, the positional take path admits only a literal K, so
a `take:` expression is evaluated once, by the general path.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(select): fused take/top-k leave alias-chained projections to the planner
Projections resolve left to right: in `{b: a z: b …}` the output z reads
the alias b (column a), not the source column b. The fused positional
take and the fused top-k gather source columns by name, so z came back
as column b — `take: 2` gave [30 10] instead of [3 1], and the sorted
top-k had the same wrong answer before this branch. When an output
names an alias bound by an earlier output (other than itself), the
fused paths now decline and the planner resolves the projections.
Tests: fused/fused_take_early_stop (positive and negative take, sorted
take with and without a filter, an identity alias that stays fused).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(select): sort keys bind to projected columns by position, name source columns
The sort of a projection resolved its keys by NAME over the projected
table, whose columns are named by source column for scans and `_e<c>`
for expressions. Three wrong outcomes followed:
- a hidden key whose source column is called `_e0` collided with the
first computed output: `{z: (+ a 10) asc: _e0}` sorted by z;
- `{b: a asc: b}` treated the output alias b as the key, found no column
b in the projection (that scan is named a) and failed with `nyi`;
- hidden keys were capped at 16, so a 17-key sort failed with `nyi`.
A sort key now names a SOURCE column, like where: and by:. In the
planner it binds to the projected column that scans it — an output, or
the hidden key appended for it — and exec_sort reads a key that is one
of the SELECT's columns by position, so no name lookup can pick another
column. Output aliases are not consulted. The hidden-key capacity is
one slot per sort key name.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Anton Kundenko <singaraiona@gmail.com>
* perf(group): partitioned dense accumulate; derived-key grouping over the column's distinct values (#633)
* perf(group): partition the slots when the dense accumulate cannot afford one set per worker; pin the SYM vocabulary once per accumulate
The direct-array group path keeps one accumulator set per worker and bounds
that state by the row count. Past the bound it ran ranged tasks over as
many sets as the budget allowed — one set, i.e. a serial scan of every
row, whenever the slots times the per-group state exceeded the rows, which
a wide key with a nullable min and a strlen average over a large table
does at once (5M slots, three aggregates: budget 1 out of 24 workers).
Instead the slots are partitioned: one parallel pass computes each row's
group and buckets its row id by slot span per worker (no shared writes),
then one task per span drains every worker's bucket for it into the single
accumulator set. No two tasks touch a slot, nothing is merged, and the
state is one set however many workers run; the buckets cost four bytes a
row. Row-range tasks stay for the small-slot case and as the fallback
when the buckets cannot be allocated.
Inside the accumulate, strlen over a SYM column pinned the FILE-domain
vocabulary snapshot and ran ray_vec_is_null per row, and the lexical
min/max compare pinned it per compare. The snapshot is pinned once per
SYM aggregate column for the whole accumulate; null is the id-0 check.
test/rfl/group/dense_slot_split.rfl: 200k rows over 130k slots with a
nullable min/max/sum/avg/first/last, strlen over a SYM column with empty
cells, a two-key composite and a where: selection — each against the same
grouping forced onto the hash path by one far-away key.
* perf(group): decide a derived-key grouping over the column's distinct values when every aggregate reads that column
A grouping keyed by an expression over one SYM column C whose aggregates
are count(C), min(C), max(C), sum(strlen C) or avg(strlen C) is decided by
the distinct values of C: a group's count is the sum of its values' row
counts, its min the smallest of its values, its strlen sum the
length-weighted count. It used to be evaluated per row — a spread key over
every row and an accumulate of every row's length and lexical compare —
while the key expression itself already ran once per distinct value.
derived_key_vocab_aggs now groups the rows by C once (a column key and a
count, the engine's cheapest grouping, under the select's where:), runs the
key once per distinct value (derived_key_str_chunks) and rewrites the
aggregates over the per-value table: count → sum of counts (null rows
count, as before), avg(strlen) → weighted length sum over the non-null
count, min/max → over the values (null skipped). Sort and take clauses on
output aliases pass through; the outputs come back in the written order
under the computed key's usual name. Any other aggregate, a runtime-domain
column or a small table keeps the row path.
Interning the key strings cached dotted segments for every new value — a
host or URL interned every one of its dot-separated parts as a symbol too,
which was most of the cold-run cost. ray_sym_intern_batch_no_split interns
them as values; a value that later appears as an identifier gets its
segments on that intern. The per-value length pass runs on the workers.
An `if` whose two STR branches are cheap and total (scans, constants,
substrings, case/trim/length, add/sub/mul, comparisons, and/or/not) takes
the eager descriptor arm instead of compacting the rows for each branch.
test/rfl/group/derived_key_vocab_aggs.rfl: the five aggregates with and
without where:, against the row path (forced by one aggregate over
another column), the same table in memory, the null group's avg and min,
sort/take on an alias, and the key name yielding to an alias.
* fix(group): keep FIRST/LAST groupings past the slot budget on the serial row-order scan
Several FIRST/LAST aggregates share one first_row/last_row per slot (a
documented limitation of da_accum_row), so the answer for a null-mixed
pair depends on the order the rows arrive in; the slot-partitioned drain
walks the workers' buckets, not the rows, and made `last` differ by
worker count. Past the budget such a grouping stays on the single
accumulator in row order.
* fix(group): partitioned accumulate and vocabulary rewrite independent of core count
The slot-partitioned accumulate scattered rows by pool morsel into
per-worker buckets, so a slot received its rows in scheduler order and
float sums (and var/stddev) rounded differently from run to run. The
scatter now runs nw tasks over contiguous row ranges in order, bucket
row = task, and the drain walks them in order: each slot gets its rows
in table order, the serial scan's accumulation.
The vocabulary rewrite took its group order from the per-value count
grouping, whose order depends on the core count. The counts are now
put in domain-position order first (bitmap rank + scatter), so the
output order is the same at every core count.
Tests: dense_slot_split (1e16 + 1 - 1e16 + 1 per group sums to 1.0 in
table order), derived_key_vocab_aggs (unsorted output order pinned).
Both fail with their fix disabled.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* perf(csv): parallel splayed load, batched symfile interning, parallel hash index build (#632)
* perf(csv): scan, append and index the splayed load in parallel
Streaming a CSV into a splayed table ran its per-chunk row scan, the
per-column writes and the index build on the calling thread; only the
field parse used the pool.
- Row offsets for a chunk come from the parallel quote-parity scanner
over a byte window sized from the running average row length, widened
until it holds the chunk (falls back to the serial limited scan).
- The chunk's columns are appended by one task per column; SYM cells
go through a direct-mapped runtime-id to position cache in the writer,
so only a value's first occurrence probes the symfile domain.
- ray_splay_build_indexes builds one column per task.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* perf(csv): intern chunk dictionaries straight into the symfile domain
Every distinct string of a splayed load was interned twice, serially:
once into the runtime symbol table by the chunk materialisation, then
again into the table's symfile domain by the column writer, one cell at
a time under the global domain spinlock — the two together were more
than half of the load time, and the runtime table grew by the whole
vocabulary for nothing.
- csv_materialize_rows takes a target domain: SYM chunk columns are
built over the symfile domain, and the writer copies their positions.
- ray_sym_domain_intern_batch (domain.c): one batch per chunk over all
SYM columns. The existing vocabulary is probed in parallel against a
snapshot of the reverse index (replaced bucket tables are retired,
never freed), misses are deduplicated per hash partition, and the new
atoms are built in parallel into per-partition arena regions
(ray_arena_alloc_raw / ray_arena_str_at) before the count is
published. Index growth re-slots the old table instead of rehashing
every atom; the symfile flush packs records into a 1 MB buffer.
- SYM fields are hashed by the parse tasks; with a domain target each
SYM column's dedupe is split by hash partition across tasks.
- The window scan is preceded by a readahead hint and finished chunks
are dropped from the mapping, so a long file does not stay resident.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* perf(index): build hash indexes partition-parallel
ray_index_attach_hash grouped the rows in one serial open-addressing
pass; on a high-cardinality numeric column of a large table it was the
longest single stretch of a splayed load, and running the columns as
pool tasks left the build itself single-threaded.
Numeric and SYM keys above 64k rows now build in partition-parallel
passes: key words and hash-partition buckets over row ranges, a
per-partition dedupe that marks each group's first row in a bitmap,
then keys, counts, the row scatter and the key table (CAS on the empty
slot) per partition. A group's number is the rank of its first row
among the marked rows, so the layout is the serial walk's: groups in
first-occurrence order, rows ascending inside a group. Every array a
pass writes is faulted in beforehand from all workers in slices — a
fresh mapping faulted at random from every worker serialised on the
page-table locks. STR keys and small columns keep the serial walk.
ray_splay_build_indexes defers the columns that qualify for a hash
index out of the per-column tasks, builds them one after another with
the parallel path and writes all deferred columns together.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* chore(index,domain): cast the CAS targets in place
cppcheck cannot parse a declaration of `_Atomic(T)*` type (AST broken
on the `=`), which failed static analysis; the two batch inserts now
cast at the compare-exchange, the form the rest of domain.c uses.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(domain): retire the reverse index on external extend; batch intern tests
The batch intern probes a snapshot of a domain's reverse index outside
the lock, relying on replaced tables being retired rather than freed.
The external-growth path (a reader catching up with a symfile another
process appended to) still freed the table, so a probe running at that
moment could read freed memory. It now retires it like the rebuild
does.
A probe that meets an entry it cannot compare (no atom and no raw bytes,
possible only if publishing the raw snapshot failed) no longer counts
it as a miss: the whole batch then resolves under the lock, so no
string is appended twice.
The splayed writer's per-column dispatch goes through
ray_pool_par_dispatch_ok like the other sites. The batch-intern header
comment no longer claims appends in batch order: new strings get
positions grouped by hash partition.
Tests: domain/intern_batch (repeats across partitions, pre-interned
hits keep their positions, repeat batch appends nothing, "" is 0,
flush/reopen); csv_splayed_dedupe_overflow (140k distinct strings in
one SYM column overflow a partition's dictionary in debug builds and
take the row-by-row fallback); index/hash_large_parallel now creates
the pool so the parallel build is the one tested.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(csv): split every window as the whole file; symfile and index bytes independent of cores
A chunk's byte window chose the scanner's quote mode from its own bytes.
In a file with quotes elsewhere, a window without any took the
quote-free fast path, where a lone '\r' does not end a row: two rows
merged and a value was lost (the serial walk and .csv.read split them).
The window now scans in the file's quote mode.
The symfile no longer depends on the worker count. The chunk batch is
laid out the way the cell-by-cell writer meets the strings — columns in
order, each column's strings by first occurrence (partition runs merged
by the first row, recorded in the dictionary entry) — and the batch
intern appends new strings in batch order. The hash index key table is
filled in group order on the calling thread, as the serial build does,
so the persisted index is the same bytes on any core count.
A replaced reverse-index table is freed at once unless a batch intern is
probing a snapshot outside the lock (counted under the lock), so a
long-lived reader of a growing symfile keeps no extra tables.
Tests: splay/csv_symfile_order (positions follow the writer order over
three chunks, parallel batch and partitioned dedupe), and
splay/csv_quote_mode_per_file (quotes only in chunk 0, a lone '\r' in
chunk 2: same rows and values as .csv.read). Both fail with their fix
disabled.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Anton Kundenko <singaraiona@gmail.com>
* perf(select): fused top-k prunes chunks by zone extrema and visits the best chunks first (#643)
* perf(select): fused top-k skips chunks whose zone cannot beat the K-th key
The fused top-k evaluated the predicate over every row even when the
first sort key is a stored integer/temporal column whose chunk-zone
index shows most chunks cannot hold a top-K row (a table clustered by
another key, sorted within it, puts a time column's early values in a
few chunks).
Each worker whose heap is full publishes its worst first-key value; the
best of those bounds the final K-th key. A chunk whose zone minimum
(asc) is above the bound, or maximum (desc) below it, is skipped before
the predicate runs — its pages are never read. Chunks with nulls are
kept when nulls sort first.
On a 100M-row table ordering by a time column clustered under a
counter (52 of 1526 chunks can hold the top 10): cold 1205 ms -> 157 ms,
warm 30 ms -> 9 ms; with a second sort key 24 ms -> 5.5 ms.
Test: fused/topk_zone_prune (asc, desc, two keys, K near the match
count, all equal to an unindexed in-memory copy).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(select): publish the top-k pruning bound once per morsel
Publishing on every heap replacement made the bound a contended cache
line when rows arrive in the order being sought (a descending top-k over
a column stored ascending replaces the root on nearly every row, and
nothing is pruned there): 25-70% slower than without pruning. The bound
is now published once per morsel, on its own cache line. Those shapes
are back to the unpruned time; the prunable ones keep their gain.
Test: topk_zone_prune gains a nullable key with an all-null chunk,
asc/desc and two keys.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(select): fused top-k visits the most promising chunks first
With a chunk-zone index on the first sort key the K-th key only tightens
as fast as the physical order lets it: an `asc` over a descending column
scans everything before the bound becomes useful, and a key whose best
values sit in the middle prunes nothing until the workers get there.
Order the chunks by their zone extremum (smallest minimum for asc,
largest maximum for desc, null chunks first when nulls sort ahead) and
dispatch the workers over that virtual row space, mapping each slot back
to its chunk. The K-heaps then fill from the best chunks, the bound is
tight after the first few, and the rest are skipped whatever the layout.
Results are unchanged: the heap keeps the K best rows with a source-row
tie break, so the visiting order cannot alter them.
Local A/B on a 4M-row splayed table (release, 8 cores): asc over a
descending key K=100 43 -> 2.7 ms, K=8000 212 -> 104 ms, with a filter
22 -> 0.2-1.7 ms; already-favourable and unindexed shapes unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(select): `if` over two STR columns under a sparse filter stays on the selected path
Two cheap STR branches route `if` to the eager arm, which picks
descriptors for every row of the table in one pass. Under a `where:`
that keeps few rows that is a bad trade: the filtered result keeps the
whole intermediate — both parents' bytes — alive for its few rows
(`select {x: (if c S S2) where: (== st 1)}` over 500k rows held 20 MB
for 10k rows; the selected path holds nothing beyond them).
Take the eager arm only when the outer selection keeps more than a
quarter of the rows (the projection's pre-compaction threshold); a
sparse selection evaluates the branches over the selected rows only.
No filter and a filter keeping most rows are unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* perf(select): fused top-k re-tests the pruning bound before every morsel
The chunk test ran once at chunk entry, so a bound that tightened while
the chunk was being scanned (another worker's heap filled, or this one's)
only took effect from the next chunk. Re-test it before every morsel —
one relaxed load and a compare, no writes — and abandon the rest of the
chunk once it cannot beat the bound.
Local A/B on the 4M-row table (release, 8 cores): asc over a descending
key K=100 2.6-3.0 -> 0.4 ms, the same with a filter 0.8-1.4 -> 0.25 ms;
the other shapes unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(collection): hash int cells through f64 when a set meets a float probe (#646)
* fix(collection): hash int cells through f64 when a set meets a float probe
The row hashset behind find/except/union/sect/in hashed int cells with
ray_hash_i64 and float cells with ray_hash_f64. Numeric equality crosses
the two (atom_eq compares any two numerics through f64, and hs_eq_rows
already defers to it), but equal values landed in different buckets, so
the probe never reached the compare. A correct answer was a collision:
(find [1.0 2.0 3.0] [3 1]) [0Nl 0] -> [2 0]
(except [1 2 3] [1.0 3.0]) [1 2 3] -> [2]
(sect [1 2 3] [2.0 3.0]) [] -> [2 3]
(union [1 2] [2.0 3.0]) [1 2 2 3] -> [1 2 3]
(in [1 0Nl 3] [1.0 0Nf]) [false true false] -> [true true false]
The set gains a num_f64 mode in which int-family cells hash through f64.
The first probe of a new type goes to a cold helper; if it is the other
numeric class (or a list, whose numeric atoms already hash through f64),
the set rehashes once and keeps the mode. A cached last-probe type keeps
the per-probe cost to one compare, so same-type callers are unchanged.
hs_eq_rows compares mixed numeric typed cells as f64 directly instead of
boxing two atoms per probe. If the one-off rehash hits OOM, that probe
is answered by a full scan and the next probe retries.
1M x 500k, -c 8, release, 3 alternating runs:
except/sect i64-i64, f64-f64 within ~1% of dev
except i64-f64 153 ms -> 80 ms (and now correct)
sect i64-f64 155 ms -> 80 ms (and now correct)
find i64 hay / f64 needles 34.5 ms -> 0.1 ms (and now correct)
collection_branch_cov.rfl pinned the old miss -- its comment says the
probe "may miss" because the hashes differ -- and now expects 2. New
mixed_numeric.rfl covers find/except/union/sect/in in both directions,
narrow ints, fractional misses, both-sides-null in, and sizes that grow
the table before and after the switch.
Closes #645
* fix(collection): skip the no-op rehash of a float-built set; pin ints past 2^53
Review follow-ups on #646. Only int cells change hash in num_f64 mode, so
a float- or list-built set is already in the f64 layout: set the mode and
skip hashset_rehash unless the source is int-class (except with an i64
probe against a 500k f64 set, 79 -> 73 ms). mixed_numeric.rfl gains the
float-set/int-probe order and a note, with cases, that ints past 2^53
compare through f64 -- consistent with == -- while same-type ints sharing
a double image stay distinct.
* test(agg): pin the LLC in dense_cache_bound so the route is runner-independent (#647)
agg_contract/dense_cache_bound failed intermittently on the macOS debug CI
job (got strategy 2, expected 1) across unrelated PRs. The group route
bounds replicated task-local slabs by the probed last-level cache, and
switches to partition ownership when fewer than three slabs fit. A macOS
CI VM reports a few MB of L2, which fits fewer than three of the test's
2.5 MB slabs, so the route correctly chose partitioning while the test
asserted the task-local strategy.
ray_cache_llc_bytes gains a DEBUG-only pin, ray_cache_llc_set_for_test,
following ray_ipc_set_auto_journal_eval_for_test. The three platform
probes become static cache_llc_probe and one wrapper honours the pin. The
test pins 32 MB (the size the route assumes when none is reported) around
the routed query and restores it before any assert can return. It also
asserts the pinned bound actually bites (3..19 slabs).
Pinning 4 MB instead reproduces the CI failure exactly:
test_agg_contract.c:1759 got 2, expected 1. make test 3953/3953.
* fix(in): admit the typed kernel unless both sides carry a null (#644)
* fix(in): admit the typed kernel unless both sides carry a null
#607 opened the typed membership kernel to text columns but kept it behind
"both operands exactly null-free". That is stricter than the semantics
require: the kernel (null matches nothing) and the hashset fallback (null
equals null) disagree only on a null row against a null set element, which
needs a null on each side. One null row in the column therefore sent the
whole column to the per-row hashset probe.
351,393-row column, -c 8, one null row, null-free needles:
SYM (in col one) 4,820 us -> 60 us
I64 (in icol [0 1]) 5,730 us -> 80 us
SYM null-free (in col one) 145 us -> 55 us (set tested first: a
null-free set skips the
column's null scan)
It also fixes wrong answers: the hashset hashes i64 and f64 cells with
different functions, so mixed int/float membership with a null on one side
missed matches -- (in [1.0 0Nf 3.0] [1 3]) was [true false false]. Those
shapes now take the kernel's double promotion. The both-sides-null
mixed-type case still reaches the hashset and is still wrong; it is a
hashset bug, not a gate one, and is left for a separate fix.
A 23-case null edge matrix is byte-identical before and after apart from
that corrected row. The in.rfl comment claiming a needle-only null "must
still reject" is corrected, and one-sided nulls across every null-bearing
width are pinned.
* fix(in): temporal equals only its own type; hash large needle sets in the kernel
Review follow-ups on widening the `in` kernel gate.
Temporal types. The typed kernel classified DATE/TIME/TIMESTAMP as plain
int-family and compared raw payloads, so (in [2000.01.01 0Nd] [0]) was
[true false] while find, except and the hashset `in` -- all atom_eq, which
never matches a DATE to an int or to a TIMESTAMP -- said no. It also
compared a DATE's day count to a TIMESTAMP's nanoseconds:
(in [2000.01.02] [2000.01.01D00:00:00.000000001]) was [true]. A temporal
operand against a different type is now an empty probe, as SYM vs non-SYM
already was. Bare `in`, the fused where: and the hashset agree whichever
side carries a null.
Large needle sets. Past the SIMD small-set size the kernel scanned every
needle per row, O(rows x needles). The widened gate sent nullable int/float
columns there, and null-free ones already went: a 351k-row column against
100k needles took ~1.2 s. Sets larger than IN_SIMD_SET (8) now get an
open-addressing table over the live needles -- int64 values, or f64 bit
patterns with -0.0 folded to +0.0 and NaN never inserted -- built once per
call and shared by all workers. An allocation failure keeps the scan.
351,393 rows, -c 8, release, dev before #644 -> this commit:
I64 with a null x 3 needles 6,500 us -> 50 us
I64 with a null x 1k 10,600 us -> 1,400 us
I64 with a null x 100k 10,000 us -> 2,500 us
I64 null-free x 100k 1,155,000 us -> 1,500 us
F64 with a null x 100k 10,500 us -> 2,500 us
select where: in, null-free x 100k 1,155,500 us -> 2,000 us
in.rfl pins temporal vs int and distinct temporal types with nulls on
neither, either and both sides, the fused where:, and the 8/9-needle
boundary, -0.0, NaN, duplicates and narrow ints on the hash path.
make test 3952/3952; reverting exec.c fails in.rfl:169.
* fix(agg): integer mean divides the exact 128-bit sum; mapped index tail unmapped after drop (#648)
* fix(agg): integer mean divides the exact 128-bit sum
The mean of an integer column depended on where it was computed. The
group engines divided a WRAPPED int64 sum — a whole-table
`select {a: (avg id)}` over signed 64-bit ids answered -5.59e10 for a
true mean of 2.53e18, and any per-group total past 2^63 was garbage —
while the vector path accumulated in double and lost low bits past
2^53, so its last bits moved with the accumulation order. The
chunk-zone metadata could only answer a mean when every partial sum
stayed below 2^53.
Every integer mean now divides the exact 128-bit sum of its rows,
converted to double by one formula, so the vector reduction, the parted
column path, the columnar, hash-row, scatter and slice group engines,
the streaming engine, pivot and the metadata agree bit for bit:
- the chunk-zone aggregates keep the high word of each chunk's sum in a
third block (older two-block indexes still load; without high words
the metadata answers only when the extrema prove the total fits);
- the keyless reduction and the group accumulators carry a high word
beside the wrapped sum, allocated only when an integer avg over an
I64 / TIMESTAMP column (or a table of 2^31 rows or more) is present —
narrower inputs cannot leave int64 and keep their plain sum;
- the streaming engine keeps its 16-byte state: a signed running sum
plus the wrap count packed with the row count, exact while a group
has fewer than 2^32 rows (larger tables stay on the legacy engines);
narrow inputs use an exact int64 state;
- a bare I64 column sums in blocks whose mode (values below 2^53, below
2^53 in magnitude, or general) is escalated once.
`sum` keeps its int64 wraparound contract. The grouped combine over a
parted table (per-partition sums joined afterwards) is not changed here
and still divides the wrapped total.
Tests: group/avg_exact_i128 (exact means over values that overflow
int64 within every group, across the keyless, columnar, hash, scatter,
slice, streaming and pivot paths, nulls, narrow types, a splayed copy),
test_agg_engine avg_exact_i128_engines (v2 on and off),
store/splayed_zone_aggs (metadata means past 2^53 and past 2^63).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(store): a mapped column that drops its index still unmaps the index tail
A column loaded by mmap with an inline index is longer than its payload,
and ray_free sized the unmap from the attached index. Once the loaded
column itself dropped that index — an in-place edit of its only
reference detaches a mapped index without releasing it — or when
ray_index_inline_map had discarded a child that is still in the file,
the unmap covered the payload only and the index tail stayed mapped
for the life of the process.
A mapped column that carries an index now registers its region in the
file-map registry under its own address with the true mapped length,
the way a string column's pool already does, and ray_free consults the
registry for every mapped column before falling back to the size …
Problem
(select {from: T a: (avg id)})over signed 64-bit ids answered-5.59e10for a true mean of2.53e18, and any per-group total past 2^63 was garbage. The vector path accumulated in double, so its last bits past 2^53 moved with the accumulation order. The chunk-zone metadata could answer a mean only while every partial sum stayed below 2^53.ray_freesized the unmap from the attached index. Once the loaded column itself dropped that index (an in-place edit of its only reference), or whenray_index_inline_maphad discarded a child still present in the file, the index tail stayed mapped for the life of the process (follow-up to perf(select): whole-table aggregates from chunk-zone metadata; mapped index safety #635).Fix
sumkeeps its int64 wraparound contract. The grouped combine over a parted table is not changed here and still divides the wrapped total.ray_freeconsults the registry for every mapped column before falling back to the size it can derive.Tests:
group/avg_exact_i128(exact means over values that overflow int64 within every group, across the keyless, columnar, hash, scatter, slice, streaming and pivot paths, nulls, narrow types, a splayed copy),test_agg_engineavg_exact_i128_engines(v2 on and off),store/splayed_zone_aggs(metadata means past 2^53 and past 2^63),index/mapped_drop_unmaps_tail(the tail page is unmapped after the sole reference drops its index and is freed).