s3gc finds and removes orphaned objects from ClickHouse S3 disks and other
S3-compatible storage. An object is a candidate only when it exists under the
configured bucket/prefix but is absent from ClickHouse
system.remote_data_paths for the configured disk.
- Collect object names, sizes, and timestamps into an auxiliary ClickHouse table.
- Anti-join that inventory with
system.remote_data_paths(or all replicas of a configured cluster). - Report candidates in dry-run mode, or delete them in batches and record confirmed deletion checkpoints in the auxiliary table.
The command-line script supports these actions directly. For Kubernetes, the repository supplies a one-shot Job runner that separates collection, review, and deletion.
Deleting an object is irreversible. Always run and review a dry-run before deletion, and scope the configured bucket and prefix as narrowly as possible.
- Use a unique collection-table prefix for each cleanup.
- For clustered ClickHouse, use the cluster name and expected replica count.
- A failed delete Job does not automatically retry. Successfully deleted batches remain checkpointed, so a replacement delete Job can resume safely.
- Never put credentials, customer manifests, or target-cluster details in Git.
- Python 3.11 for local development; the container image also uses Python 3.11.
- Network access to ClickHouse and the target S3-compatible endpoint.
- A ClickHouse user that can read
system.remote_data_pathsand manage the auxiliary table. - S3 permissions appropriate to the action: list for collection, plus delete for deletion.
Create a local environment and inspect the available options:
python3.11 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt -r requirements-dev.txt
.venv/bin/python s3gc.py --helpConfiguration can be supplied as command-line arguments or S3GC_*
environment variables. Set the ClickHouse connection, S3 endpoint/bucket/prefix,
region, disk name, and either static S3 keys or workload identity. Keep secrets
in your approved secret manager or environment, not in command history.
Run a dry-run first:
.venv/bin/python s3gc.py --verbose --dry-runFor a production or customer cleanup, use the Kubernetes procedure below rather than a one-line delete command.
The following is a non-secret target configuration. Replace every
<placeholder> value and do not commit this environment to Git:
export S3GC_CHHOST='<clickhouse-host>'
export S3GC_CHPORT=8123
export S3GC_CHUSER='<clickhouse-user>'
export S3GC_S3IP='s3.eu-central-1.amazonaws.com'
export S3GC_S3PORT=443
export S3GC_S3BUCKET='<bucket>'
export S3GC_S3PATH='<only-the-target-prefix>/'
export S3GC_S3REGION='eu-central-1'
export S3GC_S3SECURE_FLAG=true
export S3GC_S3DISKNAME=s3
export S3GC_CLUSTERNAME='<clickhouse-cluster>'
export S3GC_EXPECTED_REPLICAS=2
export S3GC_COLLECTTABLEPREFIX='s3gc_example_'
export S3GC_AGE=24
export S3GC_USEAGE=24Select one with S3GC_S3AUTH (or --s3auth):
| Mode | Credentials | Needs boto3 | Typical use |
|---|---|---|---|
static (default) |
S3GC_S3ACCESSKEY + S3GC_S3SECRETKEY, optionally S3GC_S3SESSIONTOKEN |
no | long-lived keys, or explicit temporary credentials |
aws |
boto3 credential chain, optionally S3GC_S3PROFILE |
yes | AWS SSO / named profiles on a workstation |
iam |
MinIO workload identity provider | no | EKS IRSA, EC2 instance profile, ECS task role |
S3GC_S3PROFILE implies aws. Contradictory combinations are rejected rather
than silently resolved.
Inject static credentials from a secret manager or interactive shell rather than saving them in a file:
export S3GC_CHPASS='<clickhouse-password>'
export S3GC_S3ACCESSKEY='<s3-access-key>'
export S3GC_S3SECRETKEY='<s3-secret-key>'Every S3GC_* boolean accepts true/false, yes/no, on/off, 1/0, or an
empty value for false. Unset also means false.
Authenticate with the AWS CLI first, then let s3gc resolve temporary
credentials through the boto3 chain:
aws sso login --profile my-sso-profile
export S3GC_S3AUTH=aws
export S3GC_S3PROFILE=my-sso-profile
export S3GC_S3IP=s3.amazonaws.com
export S3GC_S3PORT=443
export S3GC_S3REGION=us-east-1
export S3GC_S3SECURE_FLAG=true
.venv/bin/python ./s3gc.py --verbose --dry-runS3GC_S3ACCESSKEY and S3GC_S3SECRETKEY are unused in aws mode, and setting
them is an error rather than a silent override. The resolved credentials must
allow s3:ListBucket on the bucket for that prefix even for --dry-run —
collection lists objects. On failure s3gc prints the required permission and
the commands to verify it:
aws sts get-caller-identity --profile my-sso-profile
aws s3api list-objects-v2 --bucket <bucket> --prefix <prefix> --max-keys 1 --profile my-sso-profileexport S3GC_CHPASS='<clickhouse-password>'
export S3GC_S3AUTH=iamiam uses MinIO's AWS IAM credential provider and refreshes temporary
credentials from EKS IRSA/workload identity, an EC2 instance profile, or an ECS
task role. Prefer it over aws inside Kubernetes: it keeps boto3 out of the
request path and avoids credentials expiring during a long collect or delete.
It does not read AWS CLI profiles, aws sso login state, ~/.aws/config, or
AWS_PROFILE; use aws mode for that workstation workflow.
GCS has no batch DeleteObjects. s3gc detects a storage.googleapis.com
endpoint and falls back to per-object deletion automatically, warning that it is
slower; --use-remove-objects=false sets it explicitly. Note the disk name is
usually gcs, not s3, and GCS needs HMAC/interop keys:
export S3GC_S3ACCESSKEY='GOOG1...'
export S3GC_S3SECRETKEY='...'
export S3GC_S3IP=storage.googleapis.com
export S3GC_S3PORT=443
export S3GC_S3SECURE_FLAG=true
export S3GC_S3DISKNAME=gcs
.venv/bin/python ./s3gc.py --verbose --use-remove-objects=falseCollection makes an auxiliary table; the second command reads it and reports candidates without deleting objects:
.venv/bin/python s3gc.py --collectonly --keepdata
.venv/bin/python s3gc.py --usecollected --dry-runA crashed or interrupted --collectonly restarts its listing from the
beginning; there is no checkpoint. On a multi-million-object bucket that can
cost hours, and long runs are exactly where a rotating password or a dropped
connection tends to strike.
Shard the listing by prefix and re-run only the shards that failed. This is safe
to repeat: the auxiliary table is a ReplacingMergeTree keyed on objpath, so
re-listing a shard is idempotent.
# buckets laid out as <prefix>/<3-char hash>/<blob>
for shard in 0 1 2 3 4 5 6 7 8 9 a b c d e f g h i j k l m n o p q r s t u v w x y z; do
S3GC_S3PATH="<prefix>/${shard}" ./s3gc.py --collectonly --keepdata || \
echo "shard ${shard} FAILED — re-run just this one"
doneThe same variables can be passed as flags (for example,
--ch-host or --s3-bucket). Run .venv/bin/python s3gc.py --help for the
complete flag and environment-variable reference. Avoid direct deletion for
customer or production work; use the reviewed Kubernetes workflow instead.
Released images are public at ghcr.io/altinity/s3gc, so Kubernetes needs no
imagePullSecret. Always reference them by digest, never by tag — tags get
re-pushed and stop reproducing what you tested:
docker pull ghcr.io/altinity/s3gc@sha256:<digest>CI prints the exact IMAGE= line in its job summary; paste that into your
.env. render.py refuses anything not digest-pinned.
Build locally for a quick check:
docker build -f docker/Dockerfile -t s3gc:local .To publish by hand, both architectures are mandatory — ClickHouse node pools are frequently arm64, and an amd64-only image will not schedule there:
docker buildx build --platform linux/amd64,linux/arm64 \
-f docker/Dockerfile -t ghcr.io/altinity/s3gc:<tag> --push .
docker buildx imagetools inspect ghcr.io/altinity/s3gc:<tag> # expect amd64 AND arm64Buildx builder containers cache
/etc/resolv.confat creation time. A builder left running across a network or VPN change fails withlookup registry-1.docker.io: i/o timeoutwhile the host resolves fine. Recreate the builder, or create one with--driver-opt network=host.
The CI workflow runs tests on every pull request and publishes from pushes to
master and version tags, authenticating to GHCR with the automatic
GITHUB_TOKEN.
The Kubernetes runner lives in deploy/kubernetes/. It
uses a digest-pinned image and external Kubernetes Secrets; it does not create
or store credentials in the repository.
For customer and production work, follow:
collect → dry-run → approved delete → verify
The concise operator procedure, Secret requirements, and renderer configuration
are in deploy/kubernetes/README.md. A guarded
dev-automation phase is available only for non-production testing; it runs
collect, dry-run, and delete in one Job and still requires an explicit delete
confirmation.
Run all isolated unit tests:
.venv/bin/python -m pytest -vRun only the development-automation tests:
.venv/bin/python -m pytest -v -k dev_automationValidate that the example Kubernetes configuration renders without creating a cluster resource:
python3 deploy/kubernetes/render.py deploy/kubernetes/example.env > /tmp/s3gc-job.yaml
docker run --rm --entrypoint /kubeconform -v /tmp:/tmp:ro \
ghcr.io/yannh/kubeconform@sha256:85dbef6b4b312b99133decc9c6fc9495e9fc5f92293d4ff3b7e1b30f5611823c \
-strict -summary /tmp/s3gc-job.yamlThe unit suite does not contact ClickHouse, S3, or Kubernetes. The reserved
dev_cluster pytest marker is excluded from CI; any future tests using it must
be selected explicitly with .venv/bin/python -m pytest -m dev_cluster after
reviewing their fixture scope. The collect/dry-run/delete exercise is manual
because it can intentionally delete development objects.
s3gc.py— collection, anti-join, and deletion logic.docker/— Python 3.11 container image and Kubernetes entrypoint.deploy/kubernetes/— plain Job template, renderer, example configuration, and operator guide.tests/— pytest safety, renderer, and entrypoint tests.
See CHANGELOG.md for the full history, including why each
change was made and the evidence behind defects found in production use.
Planned: concurrency and asynchronous collection/deletion; a --collectafter
checkpoint so an interrupted collect can resume instead of re-listing.