Repository navigation
Support N-D array columns when constructing a Dataset from Julia data #55
Copy link
Copy link
Closed
Description
Activity
Python
datasetsusage examplesReproducible snippets illustrating the conventions above (verified against the pinned
datasets4.x; outputs shown as comments).1. Inferred schema — a column of N-D arrays becomes nested
List, notArray2D:import numpy as np from datasets import Dataset ds = Dataset.from_dict({"x": [np.arange(6).reshape(2, 3), np.arange(6, 12).reshape(2, 3)]}) ds.features # {'x': List(List(Value('int64')))} ds[0]["x"] # default format → nested Python lists: [[0, 1, 2], [3, 4, 5]]
2.
numpyformat — per-rowndarray, and a range index stacks rows:dn = ds.with_format("numpy") dn[0]["x"].shape # (2, 3) → single row is a 2-D array dn[0:2]["x"].shape # (2, 2, 3) → batch stacked with the observation axis FIRST
3. Explicit fixed-shape schema —
Array2D/3D/4D/5D(shape is row-major):from datasets import Features, Array2D feats = Features({"x": Array2D(shape=(2, 3), dtype="int64")}) ds2 = Dataset.from_dict({"x": [np.arange(6).reshape(2, 3)]}, features=feats) ds2.features # {'x': Array2D(shape=(2, 3), dtype='int64')}
4. Ragged columns — fine per-row, but cannot stack (object array on a batch):
dr = Dataset.from_dict({"x": [[1, 2], [3, 4, 5]]}).with_format("numpy") dr[0]["x"] # array([1, 2]) → per-row OK dr[0:2]["x"].dtype # dtype('O') → object array, no numeric stack
5. Per-column formats — array column as
numpy, keep the rest as Python:dm = Dataset.from_dict({"x": [np.arange(4).reshape(2, 2)], "name": ["a"]}) dm.set_format("numpy", columns=["x"], output_all_columns=True) row = dm[0] type(row["x"]).__name__ # 'ndarray' → array column decoded as numpy row["name"] # 'a' → string column stays Python
How this maps onto the Julia wrapper
- Snippet 2 is the read mechanism we'd reuse: numpy format stacks (obs axis first),
thennumpy2jl(DLPack) reverses axes → a Julia(dims…, N)tensor with the obs axis
last (MLUtils convention). On the write sidejl2numpyapplies the same reversal, so
the two cancel andDataset(d)[i] == d["x"][i]. - Snippet 5 is what makes the julia format able to route only numeric array columns
through numpy while leaving strings/scalars on the python-default path (options R2/R3). - Snippets 1 vs 3 are the W2-vs-W3 write choice: infer
List(List)(no new exports) or
emit explicitArray2D…5D(needs exposing theArray*feature types). - Snippet 4 is why ragged numeric columns must fall back to vector-of-arrays on read.
- Snippet 2 is the read mechanism we'd reuse: numpy format stacks (obs axis first),
- added a commit that references this issue
on Jul 1, 2026 - added a commit that references this issue
on Jul 2, 2026
Metadata
Metadata
Assignees
Labels
No labels
Support N-D array columns when constructing a
Datasetfrom Julia dataPR #54 added
Dataset(::AbstractDict)/Dataset(::NamedTuple)but rejects array-valued(multi-dimensional) columns loudly rather than silently transposing them. This issue
scopes out how to lift that restriction, and pairs naturally with the batch-stacking gap
(item 7 in the internal review). Filing to collect design decisions before implementing.
Python /
datasetsconventionsFeature types for arrays:
Value(dtype)— scalarsList(feature)/Sequence(feature)— variable-length (or fixed-length) nested listsArray2D/Array3D/Array4D/Array5D(shape, dtype)— fixed-shape tensors;shapeisPython row-major, first axis may be
Nonefor a variable dimImage/Audio/Video— modality featuresVerified empirically against the pinned
datasets4.x:from_dictwith no explicitfeaturesinfersList(List(Value))for acolumn of 2-D arrays, not
Array2D. Rectangular data still works;Array2Donly addsshape/dtype enforcement.
jl2numpyreverses axes on write (
(2,3)→(3,2)),numpy2jl(DLPack) reverses them back; thepackage already guarantees
numpy2jl(jl2numpy(x)) == x.This is what our
"julia"format uses today, so an array column reads asVector{Vector{…}}, never a realMatrix.ndarrays, and a range indexstacks rows into an
(N, dims…)tensor.per-row.
set_format("numpy", columns=[…], output_all_columns=True)gives numpy for the array columns and leaves string/scalar columns as Python, so mixed
schemas are fine.
The core problem
Our current
"julia"format = python default format +py2jl, so array columns come backas nested lists (
Vector{Vector}), never a real N-D array — and a naivejl2numpywritepath does not line up with that nested read. Supporting N-D requires making write and
read symmetric.
The recipe that works (verified):
jl2numpy(axis-reversed numpy); features inferredList(List)or explicitArrayNDwith the reversed shape.numpy2jl. The two axisreversals cancel:
This one mechanism solves N-D tensor columns and batch-stacking together.
Design choices to settle
1. Read semantics under the
"julia"formatVector{Vector}); no true N-D.N-D + free stacking, but changes today's
Vector{Vector}batch semantics.with_format(ds, "julia"; stack=true)or a distinct format; defaultbehavior unchanged, incremental.
2. Write path (constructor)
jl2numpyeach element, inferList(List). Simplest; rectangular round-trips;ragged allowed but won't stack.
jl2numpy+ explicitArrayNDfeatures (reversed shape). Enforces shape/dtype,reliable stacking; needs exposing
Array2D…5D(ties into theFeatures/ClassLabelview work).
features=kwarg for user-supplied schema. Most flexible; larger surface.3. Input shapes to accept for an array column
Vector{Matrix}(element = one observation), and/or(dims…, N)array (last axis = observations) — mirrors exactly whatthe read path returns, giving clean symmetry
(
Dataset((; x=batch))[:]["x"] == batch).4. Ragged / non-numeric columns
Recommendation
Implement the symmetric write + read recipe (covers N-D construction and batch stacking
in one pass, sharing the axis-reversal logic), accept both input shapes, and gate the
stacking read behind an opt-in first (R3) so current
"julia"semantics don't change —promoting to default (R2) can follow once it's proven. Start with inferred
List(List)features (W2) and add explicit
ArrayND(W3) alongside the plannedFeaturesview.Context: follow-up to #54.