Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,10 @@ See [Getting started](docs/getting-started.md) and
Benchmark rounds preserve separate warmup and measurement artifacts.
- xDR retains buffers until GPU work completes, permits independent reads
during GDS waits, and exports benchmark results as JSON.
- xDR defaults to fused FITS pixel restoration and prefers native Gzip decoding, with
automatic selection of compatible decompression hardware. CLI flags and
Python options select separate restoration kernels, aligned raw DEFLATE,
or CUDA decompression explicitly; see [runtime choices](docs/components/xdr.md#runtime-choices).

See the [component guides](docs/README.md#workflow-guides),
[data contracts](docs/data-artifacts.md), and
Expand Down
3 changes: 3 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,9 @@ path and xDR supplies eligible GPU reads. Reader selection is separate from
the numerical backend. Read receipts record the requested and actual reader,
selected HDUs, and any fallback reason; see
[FITS reading in a workflow](docs/components/xdr.md#read-fits-images-in-a-workflow).
The [xDR runtime controls](docs/components/xdr.md#runtime-choices) select
postprocessing, Gzip decoding and decompression backends through the CLI,
Python APIs or input manifests.

## Installation profiles

Expand Down
3 changes: 2 additions & 1 deletion docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,8 @@ requirements.

- [Core](components/core.md): shared CLI and configuration behavior.
- [xDataReader](components/xdr.md): GPU FITS loading and the shared
Astropy/xDR reader interface; native GDS depends on the storage setup.
Astropy/xDR reader interface, with [runtime compression controls](components/xdr.md#runtime-choices)
for CLI, Python and manifests; native GDS depends on the storage setup.
- [xFit](components/xfit.md): batched nonlinear least-squares dipole fitting.
- [xPois](components/xpois.md): kernel fitting and image subtraction.
- [xScan](components/xscan.md): transient datasets, classification,
Expand Down
7 changes: 7 additions & 0 deletions docs/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,13 @@ Check command help for the applicable default and the
[FITS reader guide](components/xdr.md#read-fits-images-in-a-workflow) for
scaling, compression, and section-read rules.

Applicable FITS commands and `xdr benchmark-fits` also expose
`--xdr-postprocess`, `--xdr-gzip-decoder` and `--xdr-decompression-backend`.
Omitted flags preserve manifest choices; otherwise the xDR runtime defaults
are `auto`. These settings apply when xDR is the selected FITS reader. See
[runtime choices](components/xdr.md#runtime-choices) for values, Python
keywords and capability fallback.

## xDataReader: `cuphoton xdr`

`benchmark-fits` runs the GPU-native FITS loading benchmark for individual
Expand Down
190 changes: 160 additions & 30 deletions docs/components/xdr.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,42 +12,151 @@ Current scope:
- pipelined loading with `batch_to_device_stream`
- explicit `NotImplementedError` for unsupported compression formats


## Runtime choices

The default `postprocess="auto"` uses fused unshuffle, byte-order conversion,
and scatter for supported GZIP FITS tiles. `"fused"` selects that path explicitly;
`"separate"` runs those steps with individual kernels for comparison.
Both preserve FITS values, including integer masks and floating-point bits.
Both leave the decoded input buffer unchanged. The separate path allocates
another buffer the size of the decoded tiles and copies GZIP_1 input before
byte-order conversion. Its timings therefore include that copy and are not
an exact baseline for the former in-place GZIP_1 implementation.
Three independent controls select how xDR restores supported compressed FITS
pixels. Each defaults to `auto`. They preserve image values, mask bits and
floating-point bit patterns; choose explicit modes for comparisons or debugging.

| Direct Python keyword / `xdr_options` key | CLI flag | Values | Meaning of `auto` |
| --- | --- | --- | --- |
| `postprocess` | `--xdr-postprocess` | `auto`, `fused`, `separate` | Fuse unshuffle, byte-order conversion and scatter. |
| `gzip_decoder` | `--xdr-gzip-decoder` | `auto`, `gzip`, `deflate` | Use native Gzip when available; otherwise decode aligned raw DEFLATE. |
| `decompression_backend` | `--xdr-decompression-backend` | `auto`, `cuda` | Let nvCOMP choose a compatible decompression engine for native Gzip, with CUDA fallback. |

`postprocess="fused"` selects the same kernel as `auto`; `"separate"` runs
individual restoration kernels. `gzip_decoder="gzip"` requires native Gzip
support. `"deflate"` strips the Gzip framing and aligns payloads for raw
DEFLATE. `decompression_backend="cuda"` requires a native extension with
backend selection support and selects CUDA kernels explicitly. The native
raw DEFLATE decoder uses CUDA in either backend mode.

Both postprocessing modes leave the decoded input buffer unchanged. The
separate path allocates another buffer the size of the decoded tiles and
copies GZIP_1 input before byte-order conversion. Its timings therefore
include that copy and are not an exact baseline for the former in-place
GZIP_1 implementation.

### Direct Python APIs

Both batch APIs and `GpuCompImageReader.read` accept these named keywords:

```python
from cuphoton.xdr import batch_to_device

(images,) = batch_to_device(paths, postprocess="auto")
(images,) = batch_to_device(
["exposure-1.fits", "exposure-2.fits"],
hdu_indices=[1],
postprocess="auto",
gzip_decoder="auto",
decompression_backend="auto",
)
```

The same controls apply with `batch_to_device_stream` or `parallel=False`.
Python prefetching (`native_batcher=False`, or CLI `--native-batcher off`)
still uses the selected GPU decoder. It does not select CPU decompression.

### Workflow APIs and manifests

Workflow APIs accept the three keys in an `xdr_options` mapping. For example,
to require xDR and compare CUDA decompression with separate pixel restoration:

```python
from cuphoton.core.fits_io import read_fits_images

result = read_fits_images(
"exposure.fits", [1], reader="xdr", device=True,
xdr_options={
"postprocess": "separate",
"gzip_decoder": "gzip",
"decompression_backend": "cuda",
},
)
image = result.arrays[0]
```

Workflow FITS APIs accept `xdr_options={"postprocess": "separate"}` alongside
`fits_reader="xdr"` (or `reader="xdr"` for `read_fits_images`). FITS input
descriptors can persist the same `xdr_options` mapping. The corresponding
CLI flag is `--xdr-postprocess {auto,fused,separate}`, including on
`cuphoton xdr benchmark-fits`. Explicit CLI values override matching descriptor
fields; omitted flags preserve the descriptor's choices.

Gzip decoding defaults to `gzip_decoder="auto"`: use the native Gzip helper
when available, or aligned raw DEFLATE with an older extension or the Python
fallback. Select `gzip_decoder="gzip"` to require native Gzip support, or
`gzip_decoder="deflate"` to strip the wrapper and decode aligned raw payloads.
The same choice is available as `xdr_options={"gzip_decoder": "deflate"}` in
workflow APIs and descriptors, or `--xdr-gzip-decoder {auto,gzip,deflate}` on
the CLI. Both decoders preserve the FITS pixel values.

These controls apply when the existing reader policy selects xDR. Reader
selection and CPU fallback remain controlled by `--fits-reader`. Invalid options
fail during validation; decode, I/O, and CUDA failures propagate to the caller.
The mapping is also accepted by the FITS-consuming APIs in
[xFit](xfit.md), [xPois](xpois.md), [xRep](xrep.md), and [xScan](xscan.md).
Manifest-based workflows preserve their input choices through worker
serialization and execution. See the component guide for the mapping's
placement in its manifest or FITS descriptor.

An explicit API override or CLI flag replaces only the corresponding manifest
key. Omitted flags preserve manifest values; keys absent from both use `auto`.
Passing `--xdr-gzip-decoder auto` explicitly therefore replaces a manifest's
`gzip_decoder` value. Reader receipts include explicit `xdr_options`; they
record requested settings rather than which nvCOMP engine actually ran.

### Reader selection and fallback

`--fits-reader` (Python `reader` or `fits_reader`) selects Astropy versus xDR.
The three controls above apply after that selection and do not request a GPU
reader by themselves. See [reader policies below](#read-fits-images-in-a-workflow)
for automatic CPU fallback and FITS semantic restrictions.

Within xDR, native Gzip with `decompression_backend="auto"` may use CUDA when
a hardware engine or suitable buffers are unavailable. An older extension
without native Gzip support can use aligned raw DEFLATE. The low-level decoder's
Python fallback also runs on the GPU and retains nvCOMP's existing backend
defaults. Full batch FITS loading
still requires the native FITS planner; codec fallback does not remove that
requirement.

When xDR decodes compressed HDUs, explicit `gzip` or `cuda` requests fail clearly
if the installed extension cannot honor them. Rebuild an older source extension
using the [native build instructions](#native-extension-availability). The `auto`
choices select available capabilities before decoding. I/O or decoder exceptions
propagate; they do not trigger another reader, codec or postprocessing attempt.
The supported FITS compression formats remain `GZIP_1` and `GZIP_2`.

### Hardware eligibility and diagnostics

nvCOMP selects an engine for each native Gzip call. Its
[Decompression Engine FAQ](https://docs.nvidia.com/cuda/nvcomp/decompression_engine_faq.html)
lists B200, B300, GB200 and GB300 support. GB10 and RTX PRO 6000 Blackwell
use CUDA kernels; the Blackwell name alone does not imply engine support.
The compressed data, output and decoded-size buffers must all
use compatible allocations. B200 has a 4 MiB hardware chunk limit; the limit
on a device is available through `CU_DEVICE_ATTRIBUTE_MEM_DECOMPRESS_MAXIMUM_LENGTH`.
Use `decompression_backend="cuda"` when comparing execution paths or reading
tiles beyond the hardware limit.

Batch readers retain native pooled scratch buffers. The low-level non-pooled
path, including `GpuCompImageReader.read()` without an explicit stream, uses
`cudaMallocAsync` scratch, which is not hardware-decompression capable. That
path selects CUDA directly, including with `decompression_backend="auto"`,
to avoid a failed hardware attempt on each call. Custom CuPy allocators can
also affect eligibility in pooled calls.
An `auto` receipt and a compatible GPU therefore do not establish hardware use.
See [nvCOMP logging](../troubleshooting.md#confirm-the-decompression-engine)
to inspect the actual decoder calls.

The benchmark's `--output-json` report includes
`capabilities.hardware_decompression`: the current device's algorithm mask,
`supports_deflate`, and `max_chunk_bytes`. Unavailable queries leave those
values `null` and record an `error`. These device limits are collected after
the timed phases and do not identify the engine used by an individual call.

The native decoder expects valid compressed payloads and correct output sizes.
Header validation and a successful launch do not verify payload integrity;
nvCOMP's [C API](https://docs.nvidia.com/cuda/nvcomp/c_api.html)
does not guarantee safe decoding of corrupt Gzip or DEFLATE streams.

### Tile size and batch size

Gzip tiles provide independent work for the decoder. A batch with only a few
large tiles can leave much of the GPU idle, including on GPUs without a
hardware decompression engine. Increasing `decode_batch_files` can supply
more tiles per decode call when GPU memory allows it.

For newly written FITS files, start with the usual row-sized tiles or tiles
of a few tens of KiB, then measure the complete read with representative
data. In one 128 MiB workload, 8–64 KiB tiles gave similar read times;
multi-MiB tiles were much slower. That result is a starting point, not a
universal optimum. Tiles beyond the device's hardware chunk limit use CUDA
in automatic mode; changing the backend alone does not restore the missing
tile parallelism.

## Read FITS images in a workflow

Expand Down Expand Up @@ -221,6 +330,24 @@ cuphoton xdr benchmark-fits \
/path/to/file1.fits /path/to/file2.fits
```

All three [runtime controls](#runtime-choices) are available on this command.
For a comparison against separate restoration and the native CUDA DEFLATE path,
repeat the same workload with a distinct report:

```bash
cuphoton xdr benchmark-fits \
--hdu-indices 1,2,3 --output-json separate-deflate-cuda.json \
--xdr-postprocess separate --xdr-gzip-decoder deflate \
--xdr-decompression-backend cuda \
/path/to/file1.fits /path/to/file2.fits
```

Use the same files, HDUs, batching settings and storage mode for both runs.
These controls affect the `batch_to_device` and `batch_to_device_stream`
phases; planner and raw-read timings do not measure decompression. To isolate
one change, vary only that control. The Python `run_benchmark` API accepts
`postprocess`, `gzip_decoder` and `decompression_backend` directly.

Use `--dir` and `--max-files` to scan directories of FITS files. The benchmark
defaults to `--native-read-threads=4`; the loading APIs default to the available
CPU core count when `native_read_threads` is omitted.
Expand All @@ -243,8 +370,11 @@ The version 1 JSON report contains:
measured read used GDS, including when storage is mocked.
- `storage`: the effective `real`, `host`, or `device` mode, including an
ambient mock-storage context or environment setting.
- `options`: HDUs, iteration count, thread/queue settings, and native-batcher
selection. `native_batcher_enabled=null` records an invalid forced selection
- `options`: HDUs, iteration count, thread/queue settings, native-batcher
selection, and requested `postprocess`, `gzip_decoder` and
`decompression_backend`. An `auto` value does not identify which nvCOMP
engine ran or prove hardware decompression. `native_batcher_enabled=null`
records an invalid forced selection
with the reason in `native_batcher_error`.
- `workload`: ordered input files, counts, planned raw bytes, and decoded MiB.
Failed planning can leave these planned sizes at zero.
Expand Down
20 changes: 14 additions & 6 deletions docs/components/xfit.md
Original file line number Diff line number Diff line change
Expand Up @@ -182,12 +182,20 @@ Selecting xDR does not assert native GPUDirect Storage use. Run artifacts
record the manifest and source-file hashes, selected reader and any automatic
fallback. Existing NPZ loading is unaffected by this option.

`--xdr-postprocess auto|fused|separate` selects xDR postprocessing for FITS
reads. Python `load_xfit_dataset` accepts the equivalent
`xdr_options={"postprocess": "separate"}`. FITS manifests may specify
`xdr_options` at the top level and on individual image or `stamp_basis`
descriptors. Per-image choices override the manifest default; explicit
Python/CLI choices override only supplied keys and survive worker dispatch.
FITS commands also accept `--xdr-postprocess`, `--xdr-gzip-decoder`, and
`--xdr-decompression-backend`. These control xDR's pixel conversion, Gzip
decoder, and decompression backend when xDR is the selected reader. See
[xDR runtime choices](xdr.md#runtime-choices) for values, automatic defaults,
and capability fallback.

Python `load_xfit_dataset` accepts the corresponding `xdr_options` keys
`postprocess`, `gzip_decoder`, and `decompression_backend`. For example,
`xdr_options={"gzip_decoder": "deflate", "decompression_backend": "cuda"}`
selects raw Deflate decoding on CUDA. In a FITS manifest, `xdr_options` may
appear at the top level and on each entry in `images` or the `stamp_basis`
descriptor. Descriptor choices override top-level defaults; explicit Python
options or CLI flags override only their matching keys. Omitted flags preserve
the manifest's choices, including when dispatching work to MPI or Dragon.

Input archives contain candidate identifiers and exact image pixels. Fit
artifacts contain identifiers, hashes, parameters, uncertainties, covariance,
Expand Down
17 changes: 13 additions & 4 deletions docs/components/xpois.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,10 +48,19 @@ to require xDR. With `--backend cpu`, automatic reading uses Astropy.
The option also applies to batch, MPI and Dragon execution. NPY inputs keep
their existing loading path.

`--xdr-postprocess auto|fused|separate` selects xDR's postprocessing path
when xDR is the active reader. Python workflows and `BatchFitOptions` accept
`xdr_options={"postprocess": "separate"}`; batch workers retain these choices.
Omitting the option uses xDR's default without changing FITS reader selection.
FITS commands accept `--xdr-postprocess`, `--xdr-gzip-decoder`, and
`--xdr-decompression-backend` when xDR is the active reader. Python workflows
and `BatchFitOptions` accept the corresponding `xdr_options` keys
`postprocess`, `gzip_decoder`, and `decompression_backend`. For a CUDA-only
decoding comparison, pass `--xdr-decompression-backend cuda` or
`xdr_options={"decompression_backend": "cuda"}`. See
[xDR runtime choices](xdr.md#runtime-choices) for all values and automatic
capability selection.

Configure batch choices through the command flags or `BatchFitOptions`;
they apply to every pair and are retained by MPI and Dragon workers.
Omitted keys use xDR defaults. These controls leave the `fits_reader` policy
in charge of reader selection.

Image, variance and mask HDU selection stays the same. `summary.json` records
`fits_reader` and `fits_reads`, including the reader used and any fallback.
Expand Down
18 changes: 13 additions & 5 deletions docs/components/xrep.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,11 +32,19 @@ decoded device arrays directly. Output images retain the usual host-array and
FITS contracts. Summaries record the selected reader and any fallback under
`fits_reads`.

FITS-consuming commands also accept `--xdr-postprocess auto|fused|separate`
for xDR's pixel conversion path. Python FITS APIs and workflow helpers accept
`xdr_options={"postprocess": "separate"}` (or `auto`/`fused`) and carry that
choice through image, mask, and stack reads. Omitted options use xDR defaults;
CPU Astropy reads retain their existing behavior.
FITS-consuming commands, including benchmarks and backend comparisons,
accept `--xdr-postprocess`, `--xdr-gzip-decoder`, and
`--xdr-decompression-backend`. Python FITS APIs and workflow helpers accept
the corresponding `xdr_options` keys `postprocess`, `gzip_decoder`, and
`decompression_backend`. For example, use `--xdr-decompression-backend cuda`
or `xdr_options={"decompression_backend": "cuda"}` to select CUDA
decompression when the reader uses xDR. See
[xDR runtime choices](xdr.md#runtime-choices) for values and automatic
capability selection.

The same choices reach image, mask, and stack reads. Omitted keys use xDR
defaults; `fits_reader` continues to select the reader. Header-only WCS
inspection does not use these decompression controls.

WCS mapping and target-grid setup read headers and dimensions without
decompressing image pixels. Reading with xDR does not by itself imply native
Expand Down
Loading
Loading