snaport downloads AWS EC2 EBS snapshots (standalone, or every volume of an AMI) using the official EBS Direct APIs — no volumes, no instances, no third-party snapshot tools — and stores them as sparse raw images: only allocated blocks consume disk space. Every block is checksum-verified against the value AWS returns, the whole image is hashed, and the result is compressed with zstd and restore-tested before the raw image is discarded.
Born Windows-first; since v0.2 Linux is an equally supported release target —
prebuilt binaries for windows-amd64, linux-amd64 and linux-arm64, and
it builds on macOS too. On Windows the images are NTFS sparse files; on
Linux they are plain sparse files you can loop-mount directly.
- No EBS volume or EC2 instance needed — the EBS Direct APIs
(
ListSnapshotBlocks/GetSnapshotBlock) read snapshot data directly. - Storage-efficient — a 500 GiB snapshot with 40 GiB of allocated blocks needs roughly 40 GiB during download and ~20-30 GiB after compression, never 500 GiB (see Disk usage below).
- Minimal IAM — two
ebs:*actions (+ec2:DescribeImagesfor AMIs,kms:Decryptfor encrypted snapshots). - Integrates into your own backup workflow — single static binary, JSON manifests, scriptable.
Download a prebuilt binary from the releases page — Windows amd64 zip, Linux amd64/arm64 tarballs, plus a SHA-256 checksums file.
Or build from source (requires Go 1.27+):
go build -o snaport ./cmd/snaport # Linux / macOS
go build -o snaport.exe ./cmd/snaport # Windows
or with a version stamp:
go build -ldflags "-X snaport/internal/cli.Version=0.2.0" -o snaport ./cmd/snaport
snaport download snap-0123456789abcdef0 # one snapshot -> ./snap-....img(.zst)
snaport download ami-0123456789abcdef0 -o D:\backups # all EBS volumes -> D:\backups\ami-...\*.img.zst
snaport download snap-0123456789abcdef0 --dry-run # plan only: sizes, blocks, cost, disk headroom
snaport verify D:\backups\ami-0123456789abcdef0 # re-verify a finished download
snaport selftest # offline end-to-end test, no AWS needed
Credentials and region resolve through the standard AWS chain (env vars,
~/.aws/credentials profiles, instance/container roles); override with
--profile / --region.
| Flag | Default | Meaning |
|---|---|---|
-o, --output |
./<id>.img or ./<ami-id>/ |
output file (snapshot) or directory (AMI) |
--concurrency |
32 | parallel GetSnapshotBlock requests (auto-reduced on throttling) |
--zstd-level |
6 | compression level 1-19 |
--keep-raw |
off | keep the sparse raw image next to the .zst after verified compression |
--no-compress |
off | keep only the raw sparse image |
--ntfs-compress |
off | also apply transparent NTFS compression to the raw image (Windows only) |
--paranoid-resume |
off | spot-check already-downloaded blocks against checksums when resuming |
--force |
off | discard manifest/journal/image and restart |
--dry-run |
off | list blocks; estimate sizes, AWS cost and free space; transfer nothing |
--root-only |
off | AMI mode: only the root volume |
--wait |
off | wait up to e.g. --wait 15m for pending snapshots to complete before downloading |
--base-manifest |
off | incremental sync: fetch only blocks changed since this completed base manifest (snapshot mode) |
--stream |
off | compress chunk-by-chunk during download into a seekable .img.szc; no raw image, no resume |
--state-dir |
alongside output | where resume journals live |
--egress-per-gb |
0.09 | USD/GB internet egress rate used in --dry-run cost estimates (0 = ignore egress, e.g. inside AWS) |
ListSnapshotBlockspages through the snapshot's allocated blocks (index + short-lived block token + token expiry time).- A worker pool calls
GetSnapshotBlockper block, verifying the base64 SHA-256 the API returns against the downloaded bytes before accepting them. - Each verified block is written at
blockIndex × blockSizein a sparse image file (marked viaFSCTL_SET_SPARSEon NTFS; plain sparse files on Linux/macOS). Unallocated regions and allocated-but-zero blocks are never written — they stay holes and read back as zeros, exactly matching EBS semantics. The file's logical length is set to exactlyvolumeGiB × 1 GiB. - A verification pass re-reads the image: every block's checksum is re-checked and the SHA-256 of the whole logical image is recorded in the manifest.
- The image is stream-compressed to
.img.zst; the archive is then restore-tested — decoded end-to-end and its decoded hash compared with the image hash. Only then is the raw image deleted (unless--keep-raw). - Artifacts per snapshot:
foo.img.zst+foo.manifest.json(geometry, per-block checksums, whole-image hash, compression metadata).
AMI mode resolves volumes via DescribeImages and produces one
image/manifest per EBS-backed device plus a top-level ami-manifest.json.
For a 500 GiB snapshot with 40 GiB of allocated blocks:
| Phase | On disk |
|---|---|
| Downloading | ~40 GiB (sparse raw image; holes are free) |
| Compress + verify peak | ~60-70 GiB (raw + .zst) |
| Final (default) | just the .zst (~20-30 GiB) |
--dry-run prints the allocated-block estimate, the projected AWS cost
and your free space before anything is transferred.
Downloaded images are raw disk images (GPT/MBR partitioned), so they need a little help to browse:
-
On Linux - no helper needed, loop-mount it read-only:
sudo losetup -fP --show snap-xxxx.img # prints the assigned /dev/loopN sudo mount -o ro /dev/loopNp1 /mnt # add ,norecovery for XFS imagesThe
norecoveryoption is for XFS images that were snapshotted while mounted. Detach withsudo umount /mnt && sudo losetup -D. -
Linux filesystems (ext4/XFS - typical EC2 root volumes), from Windows - use WSL2:
tools\mount-image.cmd C:\path\to\snap-xxxx.imgIt attaches a loop device, mounts partition 1 read-only (with an XFS
norecoveryfallback for volumes snapshotted while mounted) and prints the Explorer path, normally\\wsl.localhost\Ubuntu/mnt/snaport. The mount disappears when the WSL VM stops; re-run the script to restore it. Unmount withwsl -u root umount /mnt/snaport && wsl -u root losetup -D. ext4-only images can also be browsed directly in 7-Zip (Open archive), but 7-Zip does not understand XFS. -
NTFS/Windows volumes - mount with OSFMount or ImDisk to get a drive letter, or convert to VHD and attach via Disk Management.
For recurring backups of the same volume, --base-manifest downloads
only what changed. Give it the completed manifest of a previously
downloaded snapshot in the same volume lineage (snapshots of the
same EBS volume) and the same -o image path:
snaport download snap-B --base-manifest snap-A.manifest.json -o snap-A.img
ListChangedBlocksdiffs the snapshots server-side; changed blocks are fetched, unchanged blocks are reused from the base image, and blocks deallocated in the new snapshot are returned to sparse holes.- The base image is spot-checked (32 evenly-strided block checksums)
before anything builds on it; a mismatch aborts with guidance. If the
raw base was deleted after verified compression, it is rebuilt
(hash-checked) from the
.img.zstautomatically. - Interrupted deltas resume: the journal records the lineage base, and a partially-migrated image continues rather than restarting.
- Volume growth is supported (the image is extended); shrink is rejected, as EBS volumes cannot shrink.
- The final verification and (optional) compression stages are exactly
the same as a full download, so a delta result is bit-for-bit
identical to one - the test suite and
snaport selftestboth assert this.--dry-runprints the delta plan (blocks to fetch, blocks zeroed, reused count, cost).
--stream targets tight disks: blocks are compressed into the final
container as they arrive, so the raw image is never materialized and
peak disk usage is roughly the compressed output alone.
snaport download snap-... --stream # produces snap-....img.szc + manifest
snaport verify snap-....img.szc # self-contained verify (no manifest needed)
Trade-offs: an interrupted streamed run restarts from scratch (no journal) and incremental sync is not available in this mode - the default raw-image pipeline keeps both.
Container format (documented so the artifact is never a black box; all integers big-endian):
[4] magic "SZTC" [4] format version (u32)
... chunk records: [8] logical offset, [4] uncompressed length,
[4] compressed length, then one independent zstd frame
... trailer JSON: geometry, per-chunk index (offsets, sizes,
SHA-256 of each chunk's uncompressed data), whole-image SHA-256
[16] footer: trailer offset (u64), trailer length (u32),
CRC32 of the trailer bytes
Unallocated and all-zero blocks are simply not stored; readers
synthesize zeros, so the container is as storage-efficient as the
sparse raw image. Chunks compress independently, so any block can be
decoded by seeking to it alone - snaport verify decodes chunk-by-
chunk, and the format is directly usable for random access from Go via
img.NewSZCReader.
The EBS Direct APIs bill per request (about $0.003 per 1,000 requests each
for ListSnapshotBlocks and GetSnapshotBlock), and downloading to a
machine outside AWS adds internet data-transfer-out charges for the
allocated bytes. --dry-run prints an itemized estimate:
cost est: $4.08 (2 list + 81,920 get requests; egress 40.0 GB @ $0.09/GB = $3.60)
The estimate counts happy-path requests only — one GetSnapshotBlock per
allocated block plus the listing pages — so throttled retries and
token-refresh re-lists add marginally. Allocated bytes are an upper-bound
estimate (the listing exposes indexes, not lengths). Egress uses
--egress-per-gb (default $0.09/GB, roughly the us-east-1 internet-out
rate); set it to your region's rate, or 0 when running inside AWS or
through a VPC endpoint where egress is free. Verify current pricing at
https://aws.amazon.com/ebs/pricing/.
Interrupted runs (Ctrl-C, crash, reboot) resume safely:
- Progress is recorded in an append-only journal (
*.state.jsonl) next to the image: one line per completed block (ordinal, length, SHA-256). - The journal is only flushed after the image is fsynced, so anything the journal claims done is genuinely durable; anything not journaled is simply re-downloaded.
- Block tokens expire — resuming re-lists the snapshot for fresh tokens, and workers refresh tokens on demand mid-download.
- Snapshots are immutable, so an existing manifest is validated against
the fresh listing; mismatches (e.g. a different snapshot re-using the
same file name) refuse to proceed without
--force. --paranoid-resumeadditionally re-reads a sample of completed blocks and re-downloads any that no longer match their recorded checksum.
Re-running a completed download is a no-op (no re-transfer) and can be
followed by snaport verify at any time.
Minimal read-only policy for plain snapshots:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "SnaportReadBlocks",
"Effect": "Allow",
"Action": ["ebs:ListSnapshotBlocks", "ebs:GetSnapshotBlock", "ebs:ListChangedBlocks"],
"Resource": "arn:aws:ec2:*::snapshot/*"
}
]
}Add for AMI downloads:
{
"Sid": "SnaportResolveAmi",
"Effect": "Allow",
"Action": "ec2:DescribeImages",
"Resource": "*"
}Optional but recommended: ec2:DescribeSnapshots lets snaport detect
pending/erroring snapshots up front (without it, snaport proceeds and
surfaces the raw API error instead):
{
"Sid": "SnaportCheckState",
"Effect": "Allow",
"Action": "ec2:DescribeSnapshots",
"Resource": "*"
}Note that the EBS Direct read APIs only serve completed snapshots;
downloading one that is still pending fails fast with a clear message,
or waits with --wait 15m.
For encrypted snapshots the caller also needs kms:Decrypt on the
snapshot's KMS key (the EBS Direct APIs return decrypted block data):
{
"Sid": "SnaportKms",
"Effect": "Allow",
"Action": "kms:Decrypt",
"Resource": "arn:aws:kms:<region>:<account>:key/<key-id>"
}The EBS Direct APIs are subject to account quotas (throttling on
ListSnapshotBlocks / GetSnapshotBlock). snaport layers three defenses:
the AWS SDK's adaptive retryer, per-block exponential backoff with jitter,
and AIMD concurrency control — parallelism is halved while throttled and
ramps back up when calm. If you consistently throttle, lower
--concurrency.
- Per block: the SHA-256 returned by
GetSnapshotBlock(base64) must match the downloaded bytes; mismatches are refetched, then fail loudly. - Whole image: re-read after download — every block checksum plus a SHA-256 over the full logical image (holes as zeros) recorded in the manifest.
- Compressed artifact: decoded end-to-end after compression; the decoded stream's hash must equal the image hash (this is what gates deletion of the raw image).
snaport verifyre-runs raw-image or archive verification on demand.
- The default (non-stream) compression is a whole-image zstd stream
(universally decodable); use
--streamwhen random access matters. - The manifest records every block as JSON; multi-tens-of-millions of blocks (fully-allocated very large volumes) makes it large (~100 B/block).
- Windows: the console progress display works in Windows Terminal and conhost; redirected output gets periodic plain lines instead.
go test ./... # unit + integration tests (offline; no AWS account needed)
go vet ./...
snaport selftest # the same offline end-to-end gate as the test suite
internal/testutil provides an in-memory ebsx.Source fake that
exercises pagination, token expiry, throttling injection and crash/resume
paths against the real download pipeline.