A licence-aware registry and set of dataloaders for marine and coral-reef datasets, and the open image datasets built from it.
Marine computer vision draws on many datasets, with incompatible label schemes and a wide range of licence terms.
Some are public domain, some forbid commercial use, and some forbid derivative works, which rules out masks, crops and
augmentations. marine-data records the licence of every source before the source is used, so that "every label in
this training set was permitted for this purpose" can be checked by code instead of remembered.
The project has two parts:
- The registry and tooling (
src/marinedata,registry/): per-source licence tiers, access rules, label crosswalks, dataloaders, a dataset builder and a release pipeline. The registry lists 143 sources and defines 26 capabilities. - The published datasets on Hugging Face:
reefsupport/marine-data, 58,803 images with masks or bounding boxes, each row carrying its ownlicenceandattributionfields.
This is metadata, not legal advice. Tier assignments record Reef Support's reading of primary sources, cited per entry,
and are a starting point for your own review. See docs/LEGAL.md.
Release v1.0 (October 2026) is a set of four Parquet configs with images and masks embedded. Splits follow the upstream splits where a source defines them, and a deterministic 80/10/10 group hash otherwise.
| Config | Task | Images | Sources |
|---|---|---|---|
coral-masks |
Benthic semantic segmentation | 5,997 | Coralscapes, Reef Support benthic surveys, Seaview imagery with Reef Support masks |
scene-masks |
Underwater scene semantic segmentation | 1,598 | SUIM |
instance-masks |
Underwater instance segmentation | 25,296 | UIIS, UIIS10K, USIS10K |
fish-boxes |
Object detection (fish and fauna boxes) | 25,912 | UIIS, UIIS10K, USIS10K, Roboflow Aquarium Dataset |
Load a config with the datasets library, in one of three ways.
Quick look in a notebook: stream the split and read the first row.
from datasets import load_dataset
stream = load_dataset("reefsupport/marine-data", "coral-masks", split="train", streaming=True)
row = next(iter(stream))
print(row["licence"], row["attribution"])Scripts and training: download the config once, cached afterwards. coral-masks is 12.6 GB (all splits are fetched);
to try the pipeline first, use scene-masks (0.2 GB) or the one-source form below.
from datasets import load_dataset
ds = load_dataset("reefsupport/marine-data", "coral-masks", split="train")One source only: download just that source's shards (file names are <split>-<source>-<n>-of-<total>.parquet).
from datasets import load_dataset
ds = load_dataset(
"reefsupport/marine-data",
"coral-masks",
split="train",
data_files={"train": "data/coral-masks/train-reef-support-seaview-labels-*.parquet"},
)In a plain Python script, stopping a stream after a few rows can keep the process from exiting (an upstream issue in pyarrow's dataset scanner). Download the config or a single source instead.
Licences differ per source. Keep the attribution column when you share or publish results, and see
Licensing below. The dataset card documents fields, splits, annotation process and known limitations.
Requires Python 3.10 or newer and git. The package is not on PyPI; install the v1.0.0 release from GitHub:
pip install "marinedata[hf] @ git+https://github.com/reefsupport/marine-data@v1.0.0"To work on the code, clone the release and install it in editable mode:
git clone --branch v1.0.0 https://github.com/reefsupport/marine-data
cd marine-data
uv venv && uv pip install -e ".[dev]"Optional extras pull in heavier dependencies only when needed: pandas, torch,
tf, hf, croissant, geo, eval and quality.
Ask the registry which sources may be used for a purpose, and what the licence gate excluded:
import marinedata as md
# Which sources may train a closed-weights benthic segmenter?
result = md.find(task="benthic-segmentation", profile="ship-commercial")
print(result.summary()) # matched sources with tier and image count, then each exclusion
for source in result.sources:
print(source.id, source.licence.id)
# What did the gate exclude, and why?
for decision in result.excluded:
print(decision.reason)The same check from the command line:
marinedata check --profile ship-commercial # every source, admitted or refused, with the reason
marinedata list --profile research --task benthic-segmentation
marinedata show coralscapesThe masks on Hugging Face keep each source's native ids, and those ids differ between sources. marinedata.labels maps
them onto three fixed schemes, benthic-coarse, coral-binary and scene, with 255 as the ignore value. The example
needs the hf extra (uv pip install -e ".[hf]"), which adds datasets and Pillow:
from datasets import load_dataset
from marinedata import labels
stream = load_dataset("reefsupport/marine-data", "coral-masks", split="train", streaming=True)
row = labels.remap_row(next(iter(stream)), "benthic-coarse") # adds `label`, an HxW uint8 array
names = labels.class_names("benthic-coarse")
present = sorted({int(v) for v in row["label"].ravel()} - {labels.IGNORE_INDEX})
print(row["source"], [names[i] for i in present])In a script, download the config or one source as described under Published datasets and pass its
rows to remap_row the same way:
from datasets import load_dataset
from marinedata import labels
ds = load_dataset("reefsupport/marine-data", "scene-masks", split="train")
row = labels.remap_row(ds[0], "scene")from datasets import load_dataset
from marinedata import labels
ds = load_dataset(
"reefsupport/marine-data",
"coral-masks",
split="train",
data_files={"train": "data/coral-masks/train-reef-support-seaview-labels-*.parquet"},
)
row = labels.remap_row(ds[0], "benthic-coarse")The other configs use the scene scheme. scene-masks rows work with remap_row as above:
from datasets import load_dataset
from marinedata import labels
stream = load_dataset("reefsupport/marine-data", "scene-masks", split="train", streaming=True)
row = labels.remap_row(next(iter(stream)), "scene")
names = labels.class_names("scene")
present = sorted({int(v) for v in row["label"].ravel()} - {labels.IGNORE_INDEX})
print(row["source"], [names[i] for i in present])instance-masks rows carry a list of instances instead of one semantic mask, so map each instance label with
labels.map_label (it returns None for a label the scheme does not cover):
from datasets import load_dataset
from marinedata import labels
stream = load_dataset("reefsupport/marine-data", "instance-masks", split="train", streaming=True)
row = next(iter(stream))
ids = [labels.map_label(row["source"], inst["label_native"], "scene") for inst in row["instances"]]
print(row["source"], ids)fish-boxes rows have boxes and name their source in source_id:
from datasets import load_dataset
from marinedata import labels
stream = load_dataset("reefsupport/marine-data", "fish-boxes", split="train", streaming=True)
row = next(iter(stream))
ids = [labels.map_label(row["source_id"], box["label"], "scene") for box in row["boxes"]]
print(row["source_id"], ids)For a dataset built with marinedata.builder,
to_torch_dataset(dataset, decode_masks=True, scheme="benthic-coarse") remaps each decoded mask to the scheme ids.
Some sources annotate only a subset of the classes, and for them 255 means "not annotated" rather than background.
docs/LABELS.md lists the class ids per scheme, how every source maps, and the rules for training with
partial sources, dead and bleached coral, and the Coralscapes classes. For joint training, benthic-coarse and
coral-binary send dead coral and the Coralscapes unknown hard substrate class to 255 by default and map a few
Coralscapes labels the registry leaves open; exclude_conditions=() restores the registry behaviour, and the
label policy lists every difference.
Each source is described along three facets that can be queried together.
| Facet | Question it answers |
|---|---|
| Licence tier and flags | What may be done with the data? |
| Capability | What is the source useful for: benthic segmentation, fish detection, 3D, pretraining? |
| Coverage and domain shift | Will it work in the region where it is deployed? |
| Tier | Meaning |
|---|---|
T0_OWN |
Rights held outright |
T1_PERMISSIVE |
CC0, CC BY, Apache-2.0, MIT, US-Government public domain |
T2_COPYLEFT |
CC BY-SA, GPL: usable if the derivative is shared alike |
T3_NONCOMMERCIAL |
CC BY-NC and variants: research and internal use |
T4_TDM_ONLY |
No grant; lawful access under a statutory text-and-data-mining exception |
TX_PROHIBITED |
Provenance-defective or contract-blocked: never usable |
Tier alone is not enough, so licences also carry flags. no_derivatives marks terms that forbid masks, crops and
augmentations, which makes a source unusable for training and not merely non-commercial. contract_gated marks access
that required assent to terms, provenance_defective marks a grant the licensor appears not to hold, and
share_alike and attribution_required record the remaining obligations.
A profile declares what a build is for. The gate raises an error when a source is not admitted; it does not warn.
| Profile | Admits | Intended for |
|---|---|---|
ship-commercial |
T0, T1 | Closed weights in a paid product |
ship-open |
T0, T1, T2 | Weights released under a share-alike licence |
ship-noncommercial |
T0, T1, T2, T3 | Public non-commercial releases |
research |
T0, T1, T2, T3 | Papers and benchmarks that are not shipped |
pretrain-eu |
T0 to T4 | Self-supervised pretraining under the EU text-and-data-mining exception; requires a legal_opinion_ref |
Cross-region performance drops are a common deployment failure, and entries record them together with the classes a source lacks:
domain_shift:
trained_regions: [red-sea]
missing_classes: ["soft coral", "gorgonian / sea fan", "octocoral (any)"]Coralscapes, for example, has no soft-coral class. A model trained only on it has no valid label for octocorals, which make up a large share of annotations on Caribbean reefs.
Sources declare an on-disk layout instead of shipping bespoke code. Registered layouts include image-folder,
image-mask-pairs, coco-json, yolo-txt, csv-points, labelbox-ndjson, audio-clips and metadata-only.
Loaders never download data, and an empty result is always an error.
import marinedata as md
from marinedata.loaders import build_loader
registry = md.Registry.load()
loader = build_loader(registry.source("suim"), "/data/suim")
for sample in loader:
print(sample.licence_tier) # provenance travels with every sampleA declared layout is only a hypothesis until it is run against real data. marinedata fetch <source> --limit 100
downloads a bounded sample, and marinedata verify runs each declared layout against it and records the result.
Label schemes are mapped onto canonical axes (taxon, form, condition) by crosswalks, and each mapping records what
it loses as exact, coarsened or approximate. A source label with no canonical equivalent leaves the axis
unsupervised instead of being coerced into the nearest class. The registry holds 41 crosswalks.
import marinedata as md
registry = md.Registry.load()
dataset = md.DatasetBuilder(
registry,
profile="ship-commercial",
roots={"coralscapes": "/data/coralscapes"},
).build()
dataset.split(by="site") # group-wise split, the default
print(dataset.summary())
frame = dataset.to_pandas()
torch_ds = dataset.to_torch(split="train")- Splits are group-wise by default. Consecutive transect frames overlap heavily, so a random split places
near-duplicates on both sides and inflates metrics.
by="random"is available when requested explicitly. - Unsupervised axes encode as
-100, the defaultignore_indexof PyTorch'sCrossEntropyLoss. Datasets with different label sets can be combined without custom masking. The TensorFlow adapter emits an explicit mask instead. - Vocabularies are not mixed. A source without a crosswalk into the target schema raises an error unless
allow_unmapped=Trueis passed. class_weights(),class_counts()andsupervision_coverage()help with long-tailed class distributions.
Every build can emit a record of which sources contributed, under which tier and legal basis, what was excluded and why, and which obligations carry forward:
import marinedata as md
registry = md.Registry.load()
result = md.find(task="benthic-segmentation", profile="ship-commercial")
lineage = md.build_lineage(
list(result.sources),
registry.profile("ship-commercial"),
excluded=list(result.excluded),
)
with open("LINEAGE.json", "w") as handle:
handle.write(lineage.to_json())
print(lineage.attribution_text()) # for a model card or NOTICE filemarinedata --help lists every command. The main groups are:
| Purpose | Commands |
|---|---|
| Explore the registry | list, show, check, doctor, taxa, labels, taxonomy |
| Check sources against real data | fetch, verify, registry verify |
| Build datasets | release, splits, dedup, decon, export-wds, manifest, verify-release |
| Quality and evaluation | quality, privacy-scan, labelquality, eval, bench, captions |
| Provenance | lineage, mirror |
| Add or ingest sources | add, ingest, ingest-source, ingest-batch, metadata |
marinedata mirror <source> reports what may be copied into storage you control and what is refused, because copying
a source is a redistribution decision as well as a technical one.
Code. The code in this repository is released under the Apache License 2.0 (LICENSE).
Data. The code licence does not apply to data. Each source keeps its own licence, recorded in registry/ and in
the licence and attribution columns of every published row. The sources in the published release are:
| Source | What is included | Licence | Required attribution |
|---|---|---|---|
| Coralscapes | coral-masks: 2,075 images |
Apache-2.0 | Sauder et al. 2025; keep the licence text and note changes (Parquet re-encoding, class map) |
| Reef Support benthic surveys | coral-masks: 1,226 images |
CC BY 4.0 | Reef Support, https://reef.support |
| Seaview survey photographs with Reef Support masks | coral-masks: 2,696 images |
Images CC BY 3.0; masks CC BY 4.0 | Images: González-Rivero et al., University of Queensland, doi:10.14264/UQL.2019.930. Masks: Reef Support |
| SUIM | scene-masks: 1,598 images |
MIT | Islam et al. 2020; Copyright (c) 2020 Md Jahidul Islam |
| UIIS, UIIS10K, USIS10K | instance-masks: 25,296 images; part of fish-boxes |
Apache-2.0, as distributed by the authors | Lian et al. 2023; Li et al. 2025; Lian et al. 2024 |
| Roboflow Aquarium Dataset | fish-boxes: 637 images |
CC BY 4.0 | Roboflow, Aquarium Dataset, 2020 |
Images in the UIIS, UIIS10K and USIS10K configurations originate from datasets distributed by their authors under the Apache License 2.0. The original publications state that the images were gathered from public sources and earlier datasets, so copyright in individual images may rest with third parties. Rights holders can request removal by opening an issue or through the contact page.
Sources whose terms forbid redistribution are not part of the published datasets. The registry still records them, so that their status is visible and gated.
- Labels follow each upstream vocabulary. Native labels are kept and mapped to shared nodes only where a crosswalk exists.
- 35
coral-masksimages and 12instance-masksrows were dropped because no class map was available. - The mask configs have not been checked for near-duplicate images.
- Some frames may show divers or visitors.
- Registry coverage is uneven: some entries are verified from secondary sources and are blocked from shipping profiles until a primary source is checked.
Machine-readable metadata is in CITATION.cff. To cite the release and the software:
@misc{reefsupport2026marinedata,
title = {Reef Support Marine Data},
author = {{Reef Support}},
year = {2026},
version = {1.0},
howpublished = {\url{https://huggingface.co/datasets/reefsupport/marine-data}},
note = {v1.0 (October 2026). Licences are set per source; see the dataset card.}
}
@software{reefsupport2026marinedatacode,
title = {marine-data: a licence-aware registry and dataloaders for marine datasets},
author = {{Reef Support}},
year = {2026},
version = {1.0},
url = {https://github.com/reefsupport/marine-data},
license = {Apache-2.0}
}Please also cite the upstream sources you use.
@inproceedings{sauder2025coralscapesdatasetsemanticscene,
title = {The Coralscapes Dataset: Semantic Scene Understanding in Coral Reefs},
author = {Sauder, Jonathan and Domazetoski, Viktor and Banc-Prandi, Guilhem and Perna, Gabriela and Meibom, Anders and Tuia, Devis},
booktitle = {Proceedings of the International Conference on Computer Vision Joint Workshop on Marine Vision},
year = {2025}
}
@misc{gonzalezrivero2019seaview,
title = {Seaview Survey Photo-quadrat and Image Classification Dataset},
author = {Gonz{\'a}lez-Rivero, Manuel and others},
year = {2019},
publisher = {The University of Queensland},
doi = {10.14264/UQL.2019.930},
note = {CC BY 3.0}
}
@inproceedings{islam2020suim,
title = {Semantic Segmentation of Underwater Imagery: Dataset and Benchmark},
author = {Islam, Md Jahidul and Edge, Chelsey and Xiao, Yuyang and Luo, Peigen and Mehtaz, Muntaqim and Morse, Christopher and Enan, Sadman Sakib and Sattar, Junaed},
booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
year = {2020}
}
@inproceedings{lian2023watermask,
title = {WaterMask: Instance Segmentation for Underwater Imagery},
author = {Lian, Shijie and Li, Hua and Cong, Runmin and Li, Suqi and Zhang, Wei and Kwong, Sam},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
pages = {1305--1315},
year = {2023}
}
@article{li2025uwsam,
title = {Advancing Marine Research: UWSAM Framework and UIIS10K Dataset for Precise Underwater Instance Segmentation},
author = {Li, Hua and Lian, Shijie and others},
journal = {arXiv preprint arXiv:2505.15581},
year = {2025}
}
@inproceedings{lian2024usis10k,
title = {Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale Dataset},
author = {Lian, Shijie and Zhang, Ziyi and Li, Hua and Li, Wenjie and Yang, Laurence Tianruo and Kwong, Sam and Cong, Runmin},
booktitle = {Proceedings of the 41st International Conference on Machine Learning (ICML)},
series = {PMLR},
volume = {235},
pages = {29545--29559},
year = {2024}
}
@misc{roboflow2020aquarium,
title = {Aquarium Dataset},
author = {{Roboflow}},
year = {2020},
howpublished = {\url{https://public.roboflow.com/object-detection/aquarium}},
note = {CC BY 4.0}
}The most valuable contribution is a verified licence. Adding a source is a pull request that adds an entry to the registry, under two rules:
verified_bymust cite a primary source: a licence file, a dataset card, written permission, or a terms-of-use page with the date it was read. A paper's phrasing or a hosting platform's default is not enough.- Changes to a source's tier need a second reviewer.
See CONTRIBUTING.md for the workflow and CHANGELOG.md for release notes.
Questions, corrections to a licence entry and removal requests: open an issue or use the contact page. Built by Reef Support.