Skip to content

Develop -> main CAPI 1.14.2, K8s 1.36.4, CR 0.24.1 - #137

Merged
hrak merged 15 commits into
mainfrom
develop
Sep 22, 2026
Merged

hrak merged 15 commits into
mainfrom
develop

Conversation

@hrak

@hrak hrak commented Sep 22, 2026

Copy link
Copy Markdown
Member

Issue #, if available:

Description of changes:

Testing performed:

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

hrak and others added 15 commits September 15, 2026 11:33
Cluster API v1.14 raises the floor for controller-runtime (v0.24.x) and
the Kubernetes libraries (v1.36.x), so this bump pulls those in as well,
along with controller-tools v0.21.

The first part of the upstream code organization proposal landed in
v1.14 and accounts for most of the non-mechanical changes. The API types
now live in their own `sigs.k8s.io/cluster-api/api` Go module: the
import paths are unchanged, but the root and test/e2e modules need an
explicit require on it. The same reorganization moved the core provider
manifests from config/ to core/config/, so getFilePathToCAPICRDs in the
envtest helper had to follow, otherwise the controller integration
suite cannot install the CAPI CRDs.

Two APIs in use became deprecated. controller-runtime deprecates
pkg/scheme.Builder in favour of the apimachinery builder, so
groupversion_info.go switches to runtime.NewSchemeBuilder and the eight
types files append to objectTypes instead of calling
SchemeBuilder.Register. CAPI deprecates util/record; its three call
sites in pkg/cloud are dropped, since record.InitFromRecorder was never
called and they were writing to a FakeRecorder.

Unrelated to CAPI, k8s.io/apimachinery/pkg/util/dump became a deprecated
//go:fix inline shim in v0.36, which trips the govet inline analyzer.
pkg/scope/hash.go now imports k8s.io/utils/dump directly, which is what
the shim delegated to, so hashes are unchanged.

Also aligns the tool pins with CAPI v1.14.2 (envtest 1.37.0,
setup-envtest release-0.24, controller-gen v0.21.0, conversion-gen
v0.36.0), updates the e2e config to v1.14.2, and adds the missing 1.13
and 1.14 release series to the shared e2e metadata, without which
clusterctl cannot resolve the contract for the configured CAPI version.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Applies the patch bump across all three modules, so root, hack/tools and
test/e2e keep resolving the same Kubernetes library versions.

The bumped set is api, apiextensions-apiserver, apimachinery, apiserver,
client-go, cluster-bootstrap, component-base and streaming. k8s.io/streaming
is new to the group as of the CAPI 1.14.2 bump; it was split out of the core
libraries in Kubernetes 1.36 and arrives via k8s.io/apiserver. No broad
update was run, so no unrelated transitive drift is pulled in.

CAPI v1.14.2 itself requires v0.36.3, which module resolution simply takes
the higher of.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Update to CAPI 1.14.2, controller-runtime 0.24.1, controller-tools 0.21 and K8s 1.36
Bumps ginkgo to v2.32.2 and gomega to v1.43.0 in both the root module and
test/e2e. Keeping the two in step matters here because the Makefile derives
GINGKO_VER from the root go.mod, so a root-only bump builds the ginkgo
runner from one version while the e2e module compiles against another.

Also bumps two unrelated direct dependencies that only had patch or minor
updates: jellydator/ttlcache to v3.4.1 and golang.org/x/text to v0.42.0.
go mod tidy carries golang.org/x/mod, net, sync and tools along as
indirect requirements.

test/e2e required github.com/blang/semver v3.5.1+incompatible, a
pre-modules version, while already carrying blang/semver/v4 as an
indirect requirement. It is now on v4 alone. The only use is a single
semver.ParseTolerant call in common.go, whose signature is identical in
both versions and whose result is discarded, so the import path is the
whole change.

Deliberately not bumped:

  - sigs.k8s.io/controller-runtime v0.24.1 -> v0.25.1. Cluster API v1.14.2
    requires v0.24.1, and v0.25 pairs with CAPI v1.15. dependabot already
    ignores minor bumps here for that reason.
  - github.com/prometheus/client_golang v1.23.2 -> v1.24.1, which
    controller-runtime v0.24.1 pins at v1.23.2. dependabot groups
    prometheus/* with the dependencies upgraded manually alongside
    controller-runtime, so this rides along with the next bump of it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
chore(deps): Bump test libraries and consolidate on blang/semver/v4
github.com/pkg/errors was imported unaliased in all 29 files that use it,
so the bare identifier errors meant pkg/errors everywhere and the standard
library errors package was unreachable without a rename. The package a
reader sees as errors was not the one they would expect.

This adopts the convention upstream Cluster API uses. CAPI v1.14.2 imports
pkg/errors in 278 files and aliases it as pkgerrors at every one of those
sites, which frees the bare errors for the standard library. Doing the same
here keeps us aligned with upstream rather than diverging from it.

Migrating off pkg/errors entirely was considered and rejected. CAPI requires
it directly, so it stays in our module graph as an indirect dependency
whether or not we import it, and upstream shows no sign of moving: v1.14.2
still has 901 Wrapf, 415 Wrap, 451 Errorf and 317 New calls. A rewrite would
mean ~191 behaviour-bearing call sites, including a few where Wrap(nil)
returning nil is load-bearing, for no dependency reduction and a permanent
style split with upstream.

The rename itself is a pure identifier substitution: same package, same
functions, no behaviour change. Inverting it on the working tree reproduces
the previous revision byte for byte, and the per-function call counts are
unchanged at 54 New, 49 Wrap, 42 Errorf, 39 Wrapf, 5 Is, 1 Cause and 1 As.
The apierrors and kerrors aliases in the seven files that also import an
apimachinery errors package are untouched.

An importas entry pins the alias, so a future unaliased import fails lint
instead of quietly reintroducing the shadow. errorlint is enabled alongside
it and already passes with no findings; it guards the fmt.Errorf %w sites
that pkg/cloud/instance.go already has, and anything added later. Note it
only inspects fmt.Errorf, so it cannot see pkgerrors.Errorf or Wrapf.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… sentinel

With pkg/errors now aliased as pkgerrors, the bare errors identifier is free
for the standard library. These seven sites are the ones where pkg/errors
adds nothing, so they move over. This mirrors upstream Cluster API, which
uses stdlib errors.New, errors.Is and errors.As alongside pkgerrors.

pkg/errors' Is and As are thin forwarders to the standard library, so the
six call sites in cloudstackmachine_controller.go, pkg/cloud/instance.go,
pkg/cloud/client.go and pkg/metrics/metrics.go call exactly the same
function before and after.

cloud.ErrNotFound is returned bare and never wrapped, and every comparison
against it is by identity, so stdlib errors.New is equivalent. It also drops
a stack trace that nothing ever formatted.

pkg/metrics/metrics.go had errors.As as its only pkg/errors use, so it no
longer imports pkg/errors at all. The other three files keep it for Wrap,
Wrapf, Errorf and New.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CreateCloudStackClient validated the secret and, on failure, called
Fail with "Invalid secret: %+v, %s, %s, %s" against secret.Data, apiURL,
apiKey and secretKey. That wrote the entire secret map plus the CloudStack
api-key and secret-key in cleartext into the test output, which in CI means
the job log.

The failure path only needs to say which keys are missing, so it now names
them and prints nothing else. The secret's namespace and name are included
so the message still points at the right object.

This is also a better failure message: previously a missing verify-ssl or a
typo'd key gave you a wall of base64 to read, and now it names the offending
keys directly.

Swept the rest of the repo for other places credentials could reach logs or
test output and found none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
chore: Alias pkg/errors as pkgerrors, enable errorlint, and use stdlib errors where it suffices
fix(e2e): Stop printing CloudStack credentials on an invalid secret
ResolveZone appended an identical Wrapf of the GetZoneID error to retErr
both before and after the metrics call, so a failure to resolve a zone by
name reported the same "could not get Zone ID from <name>" message twice in
the returned multierror.

Keeps one append and puts the metrics call first, matching the equivalent
block in ResolveNetworkForZone directly below it.

No test asserted on the duplication, so nothing had pinned the behaviour.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
reconcileLBattachments' error was turned into a new error with
errors.Errorf("failed to reconcile LB attachment: %+v", err). Two problems.

Errorf does not wrap, so the original error was flattened into a string and
errors.Is and errors.As could not see through it. Any caller wanting to
match on a sentinel underneath, cloud.ErrNotFound for instance, silently
could not.

%+v on a pkg/errors value renders a full stack trace, so the returned
error's message carried that trace into the controller logs on every failed
load balancer detach.

errors.Wrap fixes both: it keeps the "failed to reconcile LB attachment:"
prefix, renders the cause without a stack trace, and implements Unwrap, so
Is and As now traverse the chain.

This was the only production call formatting an error into a new one. The
remaining %v cases are panics in envtest setup and one Info log, where
there is no error to propagate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ResolveNetworkForZone called the ACS error counter in the count != 1 branch
of the GetNetworkByID check, where err is nil by construction. The counter
ignores nil, so that call did nothing, while a genuine error from
GetNetworkByID returned early without ever being counted.

Moves the call into the err != nil branch. This matches both the
GetNetworkByName block at the top of the same function and the GetZoneByID
block in ResolveZone: count the ACS error, do not count a well-formed
response that simply did not match one network.

Found while fixing the duplicate append in ResolveZone; same copy-paste
origin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
fix: Correct error reporting in ResolveZone, ResolveNetworkForZone and the LB attachment path
clusterctl reads this file to map a provider version to a Cluster API
contract, so a version whose series is missing cannot be resolved and the
provider cannot be installed or upgraded at it. Tagging v0.15.0-develop-1
publishes a release in the 0.15 series, so the entry has to exist first.

v1beta1 is the correct contract: the CRD label key in
config/crd/kustomization.yaml is still cluster.x-k8s.io/v1beta1, and every
other series in this file declares the same. The move to the v1beta2
contract is separate and still in progress.

Note that the comment at the top says to update this file only when a new
major or minor version is released. Following that literally is what left
v0.14.0 shipping a metadata.yaml without its own series, which had to be
hotfixed in v0.14.1. The entry belongs in place before the first tag in a
series, not after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@hrak
hrak merged commit 0a79f0b into main Sep 22, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants