Skip to content

feat: guard bootstrap against an existing App of Apps and announce the target context - #77

Merged
user-cube merged 2 commits into
mainfrom
feat/bootstrap-app-of-apps-safeguard
Aug 31, 2026
Merged

user-cube merged 2 commits into
mainfrom
feat/bootstrap-app-of-apps-safeguard

Conversation

@filipegalo

Copy link
Copy Markdown
Collaborator

Problem

bootstrap ran regardless of whether the cluster had already been bootstrapped, silently overwriting the app-of-apps root Application. It also never named the cluster it was about to modify unless --context was passed explicitly — so a wrong kubectl config use-context was enough to bootstrap the wrong cluster with no warning.

Changes

Two checks run after the Kubernetes client connects but before anything is written, so aborting leaves the cluster untouched.

Existing App of Apps safeguard — Client.GetAppOfApps reads the app-of-apps Application. When one exists, bootstrap reports what it found and stops:

  Existing App of Apps:
    Application:  app-of-apps (namespace argocd)
    Repository:   git@github.com:acme/repo.git
    Revision:     main
    Path:         apps
    Sync/Health:  Synced / Healthy

ERROR cluster already bootstrapped: App of Apps "app-of-apps" exists in namespace argocd on context microk8s
  hint: inspect the existing installation with: cluster-bootstrap-cli info <environment>
  tip: re-run with --force to overwrite the existing App of Apps

--force overwrites it instead, still reporting what was found. A missing ArgoCD CRD is treated as "absent" rather than an error, so first-time bootstraps are unaffected.

Target context countdown — the resolved context is announced with a 10 second grace period:

⚠  Bootstrap will modify the cluster on Kubernetes context: microk8s
    Starting in  7s... press Ctrl+C to abort

Skipped by --yes/-y, and automatically when stdout is not a terminal, so CI is not delayed. The context line still prints in that case.

k8s.ResolveContext resolves the context through client-go rather than shelling out to kubectl, so the Configuration stage and the JSON report now name the real target context even when --context is omitted (previously it printed nothing).

New flags

Flag Default Description
--force false Bootstrap even if the cluster already has an App of Apps, overwriting it
--yes, -y false Skip the countdown. Already skipped when stdout is not a terminal

Drive-by fixes the early abort exposed

  • The report rendered zero-valued resource entries, printing Application: (updated) on a run that aborted before touching anything — reading as if the App of Apps had been overwritten. Resource lines are now printed only for resources bootstrap actually reached.
  • Runtime errors dumped the full usage text and printed the error twice. SilenceUsage is now set inside runBootstrap (so flag-parse errors still show usage) and SilenceErrors on the root command, since Execute prints errors itself.

⚠️ Behaviour change for existing users

Re-running bootstrap on an already bootstrapped cluster now requires --force. The docs previously advertised bootstrap as "fully idempotent, safe to re-run" — the underlying operations remain idempotent, but the default run now stops first. Any CI that re-runs bootstrap needs --force --yes. Marked as a feat (minor) rather than a breaking change; say the word if you'd rather this land as a major.

Docs updated accordingly: new Safeguards section and flags in docs/cli/bootstrap.md, a troubleshooting entry for the new error, the README idempotency note, and the one place in subfolder-setup.md that told users to re-run bootstrap.

Testing

  • go build, go vet, gofmt clean; 288 tests pass
  • New unit tests: guard (absent / exists / --force / lookup error), countdown (interactive vs not), GetAppOfApps via a fake dynamic client, and ResolveContext (override, kubeconfig fallback, no current context)
  • Verified end-to-end against a live cluster that already had an app-of-apps: bootstrap aborted at the guard with the correct context name and no cluster mutation
  • Not exercised: the countdown against a live TTY-attached bootstrap, since reaching it requires --force on a real cluster (installs ArgoCD, overwrites the root Application). The logic is unit-tested and the TTY gate uses the same term.IsTerminal idiom as vault_token.go.
  • golangci-lint / gosec not installed locally — left to CI. The one integer conversion carries the same #nosec G115 annotation as the existing one in vault_token.go.

🤖 Generated with Claude Code

filipegalo and others added 2 commits August 31, 2026 16:54
…e target context

Bootstrap previously ran regardless of whether the cluster had already
been bootstrapped, silently overwriting the app-of-apps root Application.
It also never named the cluster it was about to modify unless --context
was passed explicitly.

Add two checks that run after the Kubernetes client connects but before
anything is written, so aborting leaves the cluster untouched:

- Client.GetAppOfApps reads the app-of-apps Application. When one exists,
  bootstrap reports its repository, revision, path and sync/health and
  stops. --force overwrites it instead, still reporting what was found.
- The resolved kubeconfig context is announced with a 10 second countdown
  so a run against the wrong cluster can be aborted with Ctrl+C. Skipped
  by --yes and automatically when stdout is not a terminal, so CI is not
  delayed.

k8s.ResolveContext resolves the context via client-go rather than
shelling out to kubectl, so the Configuration stage and the JSON report
now name the real target context even when --context is omitted.

Also fix two output problems the early abort exposed:

- The report rendered zero-valued resource entries, claiming the App of
  Apps was "updated" on a run that never reached it. Resource lines are
  now printed only for resources bootstrap actually touched.
- Runtime errors dumped the full usage text and printed the error twice.
  SilenceUsage is set inside runBootstrap, so flag parse errors still
  show usage, and SilenceErrors is set on the root command since Execute
  prints errors itself.

Note for existing users: re-running bootstrap on an already bootstrapped
cluster now requires --force. The underlying operations remain idempotent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
errcheck does not exclude fmt.Fprint* calls that write to a generic
io.Writer, only those targeting os.Stdout, os.Stderr or a bytes.Buffer.
The guard's output helpers take an io.Writer so they can be tested
against a buffer, so discard the return values explicitly.

golangci-lint's default max-same-issues of 3 meant CI reported only 3 of
the 9 Fprintf call sites; all of them are handled here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@user-cube
user-cube self-requested a review August 31, 2026 16:18
@user-cube
user-cube merged commit 28ea0be into main Aug 31, 2026
8 checks passed
@user-cube
user-cube deleted the feat/bootstrap-app-of-apps-safeguard branch August 31, 2026 16:18
devopsbuddyapp Bot pushed a commit that referenced this pull request Aug 31, 2026
# [1.10.0](v1.9.0...v1.10.0) (2026-08-31)

### Features

* guard bootstrap against an existing App of Apps and announce the target context ([#77](#77)) ([28ea0be](28ea0be))
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants