How Canonical Kubernetes CAPI Providers Handle In-Place Upgrades
Abstract
Upgrading the machines of a Kubernetes cluster provisioned by the Cluster API can be achieved by utilizing In-Place upgrades. In Canonical Kubernetes, the orchestrated in-place upgrades happen by applying custom annotations on objects and reconciling those objects to reach the desired state (i.e. an upgraded object/cluster). In this blog, we’re gonna take a deep look at the inner workings of this process.
Disclaimer: I make lots of mistakes on a daily basis. If you noticed one, please let me know. All credits regarding the in-place upgrades goes to my colleagues at Canonical (k8s crew). They did the heavy lifting, I only added some finishing touches. This blog post reflects my personal understanding and views. It is not an official publication by Canonical, nor does it represent the views of Canonical.
Introduction
Cluster API facilitates cluster provisioning by enabling us to manage a “workload” cluster through another cluster called the “management” (or “bootstrap”) cluster. The workload cluster, as its name suggests, is where our workload is going to be deployed, while the management cluster, helps us operate and administrate the workload cluster (e.g. scaling up/down, upgrading the Kubernetes version on the nodes, etc.).

Figure 1: Canonical Kubernetes Cluster API Overview
Just like a rolling upgrade, an in-place upgrade is yet another strategy to upgrade the machines of a workload cluster. Whether it’s a security update or a feature that we desperately need, upgrading the machines in the long run is (almost) inevitable. A rolling upgrade might turn out to be infeasible in certain situations. We might be constrained by the amount of resources that we have (i.e. can not spawn more machines) or new machines might need extensive manual configurations. Either way, in-place upgrades are the way to go.
Unfortunately in-place upgrades are not currently supported by the upstream (although there’s currently a proposal regarding that). In CAPI, machines are considered immutable. They can not be changed once they’re created. In order to change a machine, according to CAPI design decisions, we must replace it with a new one. Quoting from CAPI concepts:
From the perspective of Cluster API, all Machines are immutable: once they are created, they are never updated (except for labels, annotations and status), only deleted.
Canonical Kubernetes CAPI leverages the fact that “annotations” can be changed on an already created machine. As we will later see in details, we can perform an in-place upgrade by changing a machine’s annotations. Changing “labels” or “status” does not make much sense, as the former is mostly used as a grouping/organization mechanism and the latter is supplied and updated by Kubernetes.
That being said, the whole idea of doing an in-place upgrade might not be completely aligned with the upstream cluster API design decisions as we’re essentially changing a machine’s Kubernetes version, something that’s supposed to be achieved by doing a rolling upgrade (replacing old ones). Either way, having the ability to perform an in-place upgrade comes with many benefits, so we decided to design and implement a way to enable users perform these upgrades with minimum friction.
In-Place Upgrade Controllers
In Canonical Kubernetes CAPI, we have two main types of controllers that handle the process of performing in-place upgrades:
- Single Machine In-Place Upgrade Controller
- Orchestrated In-Place Upgrade Controller
The core component of performing an in-place upgrade is the “Single Machine Upgrader”. It watches for certain annotations on machines and reconciles them to make sure the upgrades happen as expected. As its name suggests, it reconciles each machine individually, considers the machines separately and does not assume any sort of relation between them.
The “Orchestrator” on the other hand, watches for certain annotations on machine owners, reconciles them and upgrades groups of owned machines. It’s responsible for making sure that all the machines owned by the reconciled object get upgraded successfully.
The main annotations that drive the upgrade process are as follows (for a complete and up-to-date list of these annotations and their values please refer to the official docs):
v1beta2.k8sd.io/in-place-upgrade-to: Instructs the controller to perform an upgrade with the specified option/method. This is the only annotation that we as the users need to put on the objects.v1beta2.k8sd.io/in-place-upgrade-status: As soon as the controller starts the upgrade process, the object will be marked with this annotation to indicate the status of the upgrade. It can either bein-progress,failedordone.v1beta2.k8sd.io/in-place-upgrade-release: When the upgrade is performed successfully, this annotation will indicate the current Kubernetes release/version installed on the machine.
Single Machine In-Place Upgrade Controller
The Machine objects can be marked with the upgrade-to annotation to trigger an in-place upgrade for that machine. As an example, let’s say we have a machine with 1.30/stable Kubernetes snap installed on it. To initiate an in-place upgrade for this machine, we can annotate it with v1beta2.k8sd.io/in-place-upgrade-to: channel=1.31/stable. The single machine upgrader (which is watching for changes on machines), notices this annotation and attempts to upgrade the Kubernetes version of that machine to the specified version.
Because Canonical Kubernetes is shipped in a snap package, performing an upgrade can be as easy as doing a snap refresh. Upgrade methods or options can be specified to upgrade to a snap channel, revision, or install a new snap from a file already placed on the machine to make up for air-gapped environments.
When the upgrade is finished successfully, we will notice (at least) the following annotations on the machine:
annotations:
v1beta2.k8sd.io/in-place-upgrade-release: "channel=1.31/stable"
v1beta2.k8sd.io/in-place-upgrade-status: "done"
If the upgrade fails, the controller will mark the machine with the following annotations and retry immediately:
annotations:
# the `upgrade-to` causes the retry to happen
v1beta2.k8sd.io/in-place-upgrade-to: "channel=1.31/stable"
v1beta2.k8sd.io/in-place-upgrade-status: "failed"
# orchestrator will notice this annotation and knows that the
# upgrade for this machine failed
v1beta2.k8sd.io/in-place-upgrade-last-failed-attempt-at: "Sat, 19 Oct 2024 20:30:00 +0400"
Disclaimer: Code snippets in this blog are simplified versions of the actual implementation. Make sure to check out Canonical Kubernetes Cluster API repository to see how things work exactly if I failed to describe them clearly. Also, the implementation might/will evolve and change over time so we’re just going to see how things are as of the time of writing this blog.
The controller is setup to watch for changes of individual machines:
func (r *SingleMachineUpgrader) SetupWithManager(mgr ctrl.Manager) {
ctrl.NewControllerManagedBy(mgr).For(&clusterv1.Machine{}).Build(r)
}
Let’s have a look at how the reconciliation loop is implemented:
func (r *SingleMachineUpgrader) Reconcile(ctx context.Context, req ctrl.Request) {
m := getMachine()
if m.Annotations["in-place-upgrade-change-id"] != "" {
// Starting a new upgrade or retrying a failed one
r.handleUpgradeRequest()
}
switch m.Annotations["in-place-upgrade-status"] {
case "in-progress":
r.handleUpgradeInProgress()
case "done":
r.handleUpgradeDone()
case "failed":
r.handleUpgradeFailed()
default:
r.markUpgradeFailed("invalid in-place upgrade status")
}
}
The change-id annotation is part of the snap refresh process (reference) and (as we will see in the future) is set on the machine when a snap refresh command is initiated by the controller. Let’s see the other bits and pieces that handle the upgrade process:
func (r *SingleMachineUpgrader) handleUpgradeRequest() {
m := getMachine()
delete(m.Annotations, "in-place-upgrade-status")
delete(m.Annotations, "in-place-upgrade-change-id")
delete(m.Annotations, "in-place-upgrade-release")
patchMachine(m)
changeID := refreshMachine(m, getUpgradeTo())
// We're setting the `change-id` annotation here
r.markUpgradeInProgress(changeID)
}
func (r *SingleMachineUpgrader) handleUpgradeInProgress() {
m := getMachine()
status := getRefreshStatusForMachine(m)
if !status.Completed {
return
}
switch status.Status {
case "Done":
r.markUpgradeDone()
case "Error":
r.markUpgradeFailed()
default:
r.markUpgradeFailed("invalid refresh status")
}
}
func (r *SingleMachineUpgrader) handleUpgradeDone() {
m := getMachine()
delete(m.Annotations, "in-place-upgrade-to")
delete(m.Annotations, "in-place-upgrade-change-id")
delete(m.Annotations, "in-place-upgrade-last-failed-attempt-at")
mAnnotations["in-place-upgrade-release"] = getUpgradeTo()
patchMachine(m)
}
func (r *SingleMachineUpgrader) handleUpgradeFailed() {
m := getMachine()
delete(m.Annotations, "in-place-upgrade-status")
delete(m.Annotations, "in-place-upgrade-change-id")
patchMachine(m)
}
By applying and removing annotations, the single machine upgrader determines the upgrade status of the machine it’s trying to reconcile and takes necessary actions to successfully complete an in-place upgrade. Let’s have a look at this diagram showing the flow of the in-place upgrade of a single machine:

Figure 2: Single Machine In-Place Upgrade Controller Flow
Machine Upgrade Process
Canonical Kubernetes has a daemon running in the background called the k8sd. It’s responsible for many things in the context of Canonical Kubernetes but one of them is to expose certain endpoints that can be used to interact with the cluster. The single machine upgrader calls the /snap/refresh endpoint on the machine that it’s trying to upgrade (the refreshMachine() function above). This endpoint call will trigger the “actual” upgrade process and in the meantime, the /snap/refresh-status will be called periodically by the single machine upgrader to see how things are going. It’s worth noting that ensuring secure communication between the single machine upgrader and the Canonical Kubernetes daemon (k8sd) is really important and something that is handled internally as well.

Figure 3: Single Machine In-Place Upgrade Controller Calling The k8sd Endpoints
Orchestrated In-Place Upgrade Controller
While the “Single Machine In-Place Upgrade Controller” is responsible for upgrading individual machines, if our workload cluster is made up of tens, hundreds, or thousands of nodes, annotating them one by one and keeping track of their statuses, failures, reasons and errors can become a daunting or even impossible task. “Orchestrated In-Place Upgrade Controller” to the rescue!
The main idea behind this class of controllers is to enable us upgrade multiple machines by only annotating a single object: the owner of those machines. In CAPI, we mostly have two main machine groups in our workload cluster: “control-plane-node machines” and “worker-node machines”.
Let’s say we want to upgrade all of our worker nodes. In this case, we only need to annotate the MachineDeployment object. The orchestrated upgrade controller watches for the upgrade-to annotation on the MachineDeployment and will trigger an in-place upgrade for all of its owned machines by annotating them one by one. The responsibility is then delegated to the single machine upgrader to make sure each individual machine is getting upgraded successfully. The orchestrator is only there to keep track of the upgrade status of these machines, reported by the single machine upgrader (by annotating the machine). If any of the machines fail to get upgraded, the orchestrator marks the owner (here, MachineDeployment) with the respective annotations (status: failed and last-failed-attempt-at), publishes helpful and informative events and trusts the single machine controller to do the retry and eventually (hopefully) succeed. When all the owned machines are upgraded as expected, the orchestrator considers the operation successful, marks the owner with release and status: done and steps down.
Let’s see how the orchestrator for MachineDeployment is implemented. We first define the object that we want to reconcile (MachineDeployment) and indicate that we also want to watch the machines that it owns:
func (r *Orchestrator) SetupWithManager(mgr ctrl.Manager) {
if err := ctrl.NewControllerManagedBy(mgr).
For(&clusterv1.MachineDeployment{}).
Owns(&clusterv1.Machine{}).
Complete(r)
}
Then, we define the steps we want to take in each reconciliation loop:
func (r *Orchestrator) Reconcile(ctx context.Context, req ctrl.Request) {
machineDeployment := getMachineDeployment()
ownedMachines := getOwnedMachines(machineDeployment)
var upgradedMachines int
for _, m := range ownedMachines {
if r.isMachineUpgraded(m) {
upgradedMachines++
continue
}
if r.isMachineUpgradeFailed(m) {
r.markUpgradeFailed()
return
}
if r.isMachineUpgrading(m) {
return
}
// Machine is not upgraded, mark it for upgrade
r.markMachineToUpgrade(m)
r.markUpgradeInProgress(m)
return
}
if upgradedMachines == len(ownedMachines) {
r.markUpgradeDone()
}
}
The single machine upgrader communicates with the orchestrator by setting annotations on the machines. Note that these two controllers will not and should not collide or get into a race condition. The orchestrator smoothly transfers the upgrade responsibility to the single machine upgrader and patiently observes. Let’s see how this communication and responsibility transfer happens:
func (r *Orchestrator) isMachineUpgraded(m *clusterv1.Machine) bool {
return m.Annotations["in-place-upgrade-release"] == getUpgradeTo()
}
func (r *Orchestrator) isMachineUpgrading(m *clusterv1.Machine) bool {
return m.Annotations["in-place-upgrade-status"] == "in-progress" ||
m.Annotations["in-place-upgrade-to"] != ""
}
func (r *Orchestrator) isMachineUpgradeFailed(m *clusterv1.Machine) bool {
return m.Annotations["in-place-upgrade-last-failed-attempt-at"] != ""
}
func (r *Orchestrator) markMachineToUpgrade(m *clusterv1.Machine) {
delete(m.Annotations, "in-place-upgrade-release")
delete(m.Annotations, "in-place-upgrade-status")
delete(m.Annotations, "in-place-upgrade-change-id")
delete(m.Annotations, "in-place-upgrade-last-change-attempt-at")
m.Annotations["in-place-upgrade-to"] = getUpgradeTo()
patchMachine()
publishEvent(
"Machine %q is upgrading to %q",
m.Name,
getUpgradeTo(),
)
}
To have a better understanding, let’s have a look the flow of the orchestrator:

Figure 4: Orchestrated In-Place Upgrade Controller Flow
Locking The Upgrade Process
There might be scenarios where we need to make sure that only a limited number of machines are going to get upgraded at the same time. An example might be upgrading control plane machines. If multiple nodes become unavailable due to getting upgraded, we might loose quorum, have a severe downtime, or end up in an undesirable state that is very costly to get out of.
Let’s say we only want to allow a single upgrade at any given point (no parallel upgrades). A lock or a semaphore can be implemented to ensure that the orchestrator is not going to trigger an upgrade for multiple machines, even if multiple instances of the orchestrator are reconciling the same object in parallel. Note that this implementation can be generalized to allow n upgrades at the same time, instead of only 1. Here is the interface that our lock would implement:
type UpgradeLock interface {
// IsLocked checks if the upgrade process is locked and (if locked) returns the machine that the process is locked for.
IsLocked() *clusterv1.Machine
// Lock tries to lock the upgrade process for the given machine.
Lock(m *clusterv1.Machine)
Unlock()
}
The idea behind this lock is to store the information around the machine that is currently being upgraded into an object (like a ConfigMap) and to make sure that the upgrade process is triggered for new machines only if that object (ConfigMap in our case) does not exist. Let’s see how the lock is implemented:
func (l *upgradeLock) IsLocked() *clusterv1.Machine {
cm, found := getConfigMap()
if !found {
return nil
}
return cm.getMachine()
}
func (l *upgradeLock) Unlock() {
deleteConfigMap()
}
func (l *upgradeLock) Lock(machine *clusterv1.Machine) {
cm := newConfigMap()
cm.setInformation(machine.Name, machine.Namespace)
create(cm)
}
In order to incorporate this lock into the upgrade process, we can change the reconciliation loop like below:
func (r *OrchestratedInPlaceUpgradeController) Reconcile(ctx context.Context, req ctrl.Request) {
upgradingMachine := r.lock.IsLocked()
// Upgrade is locked and a machine is already upgrading
if upgradingMachine != nil {
if inplace.IsUpgraded(upgradingMachine) {
r.lock.Unlock(ctx, scope.cluster)
return
}
if inplace.IsMachineUpgradeFailed(upgradingMachine) {
r.markUpgradeFailed(ctx, scope, upgradingMachine)
return
}
// machine is still upgrading
return
}
// Check if there are machines to upgrade
var upgradedMachines int
for _, m := range scope.ownedMachines {
if inplace.IsUpgraded(m) {
upgradedMachines++
continue
}
// Lock the process for the machine and start the upgrade
r.lock.Lock(m)
r.markMachineToUpgrade(m)
r.markUpgradeInProgress(m)
return
}
if upgradedMachines == len(scope.ownedMachines) {
r.markUpgradeDone(ctx, scope)
}
}
Conclusion
Provisioning a cluster can be challenging. Trying to upgrade the machines of that cluster without carefully engineered tools can get out of hand pretty quickly. In-Place Upgrade Controllers that come out of the box with Canonical Kubernetes CAPI providers can take a huge burden off our shoulders by taking care of the upgrade process in a responsive, self-healing and cloud-native way.
Hope you enjoyed this blog post. I’m Hue, a Software Engineer at Canonical, working in the Kubernetes team alongside amazing teammates. Feel free reach out to me on LinkedIn, Twitter (X) or any other platform mentioned in my Personal Website.
References
- Canonical Kubernetes Cluster API: https://github.com/canonical/cluster-api-k8s/
- Canonical Kubernetes: https://github.com/canonical/k8s-snap/
- Canonical Kubernetes Docs: https://documentation.ubuntu.com/canonical-kubernetes/latest/
- In-Place Upgrades in Cluster API: https://hackmd.io/@capi-in-place/rkWgYh74a