Kubernetes Deployment Strategies: Rolling Update, Recreate, Blue/Green and Canary

Key takeaways

A Deployment has only two built-in strategies, RollingUpdate and Recreate. Rolling updates are safe only when readiness probes and graceful shutdown are right; blue/green and canary are built on top with extra Deployments, Service selectors or weighted routing. Every example runs on a local minikube cluster.

Introduction

“Deployment strategy” sounds like a menu of options Kubernetes offers, but a Deployment has exactly two: RollingUpdate (the default) and Recreate. Blue/green and canary releases, which people often mean when they say “deployment strategy”, are patterns you assemble on top from more than one Deployment and some way of steering traffic.

What all of them have in common is that the strategy is only as safe as two things inside your Pods: a readiness probe that tells the truth about whether the Pod can serve, and a shutdown path that lets in-flight requests finish. Most of the rollout incidents I have seen were not caused by choosing the wrong strategy; they were a rolling update doing exactly what it was told with a readiness probe that said “ready” too early, or an app that died the instant it got SIGTERM.

This article sets up a small versioned service on a local minikube cluster, then walks through each strategy on it: what Kubernetes actually does, which fields control it, how to watch and undo it, and where it breaks.


Lab setup: a versioned service on minikube

To see a strategy work, you need two versions of something that tell you which one answered. A tiny Node.js server that returns its VERSION environment variable is enough.

server.mjs:

import http from 'http';
const port = Number(process.env.PORT || 3000);
const version = process.env.VERSION || 'dev';

const server = http.createServer((req, res) => {
  if (req.url === '/health') {
    res.writeHead(200, { 'Content-Type': 'application/json' });
    res.end(JSON.stringify({ status: 'ok', version }));
    return;
  }
  res.writeHead(200, { 'Content-Type': 'text/plain; charset=utf-8' });
  res.end(`hello from ${version}\n`);
});
server.listen(port, '0.0.0.0', () => console.log(`listening on ${port} (${version})`));

process.on('SIGTERM', () => {
  server.close(() => process.exit(0));        // stop accepting, finish in-flight requests
  setTimeout(() => process.exit(1), 10_000);  // safety net
});

Dockerfile:

FROM node:22-alpine
WORKDIR /app
COPY server.mjs .
ENV NODE_ENV=production
EXPOSE 3000
USER node
CMD ["node", "server.mjs"]

Start a cluster and build two tagged images where the minikube node can see them:

minikube start --driver=docker
kubectl config current-context        # should print minikube

docker build -t demo-api:1.0.0 .
docker build -t demo-api:2.0.0 .
minikube image load demo-api:1.0.0
minikube image load demo-api:2.0.0

The one local-cluster trap worth knowing before any rollout experiment: the node’s container runtime does not see your host’s Docker image cache, so an image you only built on the host ends in ImagePullBackOff. minikube image load copies it in (alternatively, eval $(minikube docker-env) builds straight into minikube’s daemon). Use fixed tags, not latest: with latest the default imagePullPolicy is Always, so the node tries Docker Hub and fails; with a fixed tag it is IfNotPresent and the loaded image is used. Reusing one tag for different builds has the opposite problem, the node keeps the old image, which makes rollout experiments lie to you.

The base Deployment and Service:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: demo-api
spec:
  replicas: 4
  selector:
    matchLabels:
      app: demo-api
  template:
    metadata:
      labels:
        app: demo-api
    spec:
      terminationGracePeriodSeconds: 30
      containers:
        - name: api
          image: demo-api:1.0.0
          imagePullPolicy: IfNotPresent
          env:
            - name: VERSION
              value: '1.0.0'
          ports:
            - containerPort: 3000
          readinessProbe:
            httpGet: { path: /health, port: 3000 }
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /health, port: 3000 }
            initialDelaySeconds: 10
            periodSeconds: 10
          lifecycle:
            preStop:
              exec:
                command: ["sleep", "5"]
---
apiVersion: v1
kind: Service
metadata:
  name: demo-api
spec:
  selector:
    app: demo-api
  ports:
    - port: 80
      targetPort: 3000
kubectl apply -f demo-api.yaml
kubectl rollout status deployment/demo-api

To watch which version answers during the experiments below, run a loop from inside the cluster. kubectl port-forward is not suitable here: it tunnels to one Pod and never spreads requests.

kubectl run probe --rm -it --image=busybox --restart=Never -- \
  sh -c 'while true; do wget -qO- http://demo-api/; sleep 0.5; done'

busybox wget opens a new connection per request, so each request is balanced independently. Remember that when reading canary numbers later: clients that keep connections alive do not get re-balanced per request.


RollingUpdate

A Deployment does not replace Pods itself. It owns ReplicaSets, one per revision of the Pod template. When you change the template (image, env, anything under spec.template), the Deployment creates a new ReplicaSet and then scales the new one up and the old one down in steps. The Service is not involved at all: it selects Pods by label, and a Pod starts receiving traffic only once its readiness probe passes.

Trigger a rollout to 2.0.0 and watch the two ReplicaSets trade replicas:

kubectl set image deployment/demo-api api=demo-api:2.0.0
kubectl set env deployment/demo-api VERSION=2.0.0
kubectl get rs -l app=demo-api -w

Running set image and set env as two commands creates two rollouts, the first of which runs the new image with the old VERSION label. In real use, change the manifest and kubectl apply it once; the commands are here only to keep the demo short.

maxSurge and maxUnavailable

spec:
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  • maxSurge: how many Pods may exist above replicas during the rollout. It costs extra capacity for a while.
  • maxUnavailable: how many Pods may be below replicas in the ready state. It costs serving capacity.

Both accept a number or a percentage and default to 25%; maxSurge rounds up and maxUnavailable rounds down. With 4 replicas, the default is surge 1, unavailable 1: Kubernetes may create one new Pod and take one old Pod out at the same time, so you briefly run on 3 ready Pods. With 2 replicas the defaults round to surge 1, unavailable 0, which is why small Deployments already behave cautiously.

maxSurge: 1, maxUnavailable: 0 is the setting I reach for on anything user-facing: an old Pod is removed only after a new one is ready, so capacity never drops, and a new version that never becomes ready stalls the rollout instead of taking the service down. The cost is speed; the rollout moves one Pod at a time. maxSurge: 100%, maxUnavailable: 0 is the fastest safe option if the cluster has room for a full second copy, and on a cluster with no spare room maxUnavailable: 1, maxSurge: 0 is the only setting that can make progress at all.

Two more fields shape the pace:

  • minReadySeconds (default 0): a new Pod must stay ready this long before it counts as available. It catches versions that pass the first probe and crash a few seconds later, which would otherwise be counted as successes while the rollout races on.
  • progressDeadlineSeconds (default 600): if the rollout makes no progress for this long, the Deployment’s Progressing condition becomes False with reason ProgressDeadlineExceeded, and kubectl rollout status exits non-zero. Kubernetes does not roll back automatically; this is a signal for CI or a human.

Watching, pausing and undoing

kubectl rollout status deployment/demo-api          # blocks until done or failed
kubectl rollout history deployment/demo-api         # revisions
kubectl rollout undo deployment/demo-api            # back to the previous revision
kubectl rollout undo deployment/demo-api --to-revision=3
kubectl rollout pause deployment/demo-api           # batch several changes...
kubectl rollout resume deployment/demo-api          # ...into one rollout

undo works by scaling an old ReplicaSet back up, so it only reaches as far back as revisionHistoryLimit (10 by default) keeps old ReplicaSets around. The CHANGE-CAUSE column in history is filled from the kubernetes.io/change-cause annotation; the old --record flag is deprecated, so set the annotation in CI if you want readable history. Also note that undo changes the live object: if your manifests live in Git and a GitOps tool syncs them, the next sync re-applies the bad version unless you revert the commit too.

The price: two versions at once

For the length of a rolling update, old and new Pods serve traffic side by side, and a client can hit v2 on one request and v1 on the next. That is harmless for a stateless text change and dangerous for a change in API shape or database schema. The standard answer is to make every release compatible with the one before it, the expand/contract pattern: first release code that works with both the old and new schema (expand), migrate data, then remove the old path in a later release (contract). If you cannot do that, you need Recreate or blue/green, not a faster rolling update.


Shutdown during a rollout

When a rollout removes an old Pod, two things start at roughly the same time and are not ordered: the Pod is removed from the Service’s endpoints (and each node’s kube-proxy rules are updated), and the kubelet runs the preStop hook and then sends SIGTERM to the container. If the app exits the moment it sees SIGTERM, requests that were routed to it just before the endpoint update propagated fail with connection errors. This shows up as a handful of 502s at every deploy, which is easy to dismiss as noise.

The lab manifest handles it in three layers:

  • preStop: sleep 5 delays SIGTERM for a few seconds, which is usually enough for endpoint removal to reach every node. (Recent Kubernetes versions also support a built-in sleep action in preStop; the exec form needs a sleep binary in the image, which Alpine has.)
  • The SIGTERM handler stops accepting new connections and lets in-flight requests finish. Node.js running as PID 1 ignores SIGTERM without a handler, so without it every Pod waits the full grace period and is then killed mid-request.
  • terminationGracePeriodSeconds: 30 is the hard deadline for the whole sequence, and it includes the preStop time. If your requests can run long, raise it; if preStop sleeps 25 seconds, the app gets only 5 to drain.

A readiness probe that returns 200 before the app has warmed its caches or opened its database pool is the mirror-image problem at startup: the new Pod receives traffic it cannot yet serve. Point readiness at something that reflects the ability to serve, but keep liveness trivial; a liveness probe that checks the database restarts every Pod at once during a database outage.


Recreate

spec:
  strategy:
    type: Recreate

Recreate scales the old ReplicaSet to zero, waits for its Pods to terminate, then creates the new ones. There is a gap with no Pods at all, so there is downtime, roughly the shutdown time of the old Pods plus the start and readiness time of the new ones.

It is the right choice when two versions must never run concurrently: a single-writer process that holds a lock or a file, an app that runs an incompatible schema migration at startup, or a workload attached to a ReadWriteOnce volume. That last case is a common surprise. With the default rolling update, the new Pod waits for a volume the old Pod still holds, the old Pod is not removed until the new one is ready, and the rollout sits stuck until the progress deadline. For a single-replica workload on a ReadWriteOnce volume, Recreate is what you want.


Blue/green

Blue/green runs the new version as a complete second Deployment, tests it while the old one serves all traffic, and then switches traffic in one step. In Kubernetes the switch is a Service selector.

# demo-api-blue: the current version
metadata:
  name: demo-api-blue
spec:
  replicas: 4
  selector:
    matchLabels: { app: demo-api, track: blue }
  template:
    metadata:
      labels: { app: demo-api, track: blue }
    # ... image demo-api:1.0.0
---
# demo-api-green: the new version, same shape, track: green, image demo-api:2.0.0

The public Service selects one track:

apiVersion: v1
kind: Service
metadata:
  name: demo-api
spec:
  selector: { app: demo-api, track: blue }
  ports:
    - port: 80
      targetPort: 3000

A second Service such as demo-api-preview selecting track: green lets you test the new version inside the cluster before anyone else sees it. When it looks good, switch:

kubectl patch service demo-api -p '{"spec":{"selector":{"app":"demo-api","track":"green"}}}'

Rollback is the same patch with blue, and it is instant because the blue Pods are still running. That is the main reason to pay for blue/green: rollback does not depend on old Pods starting up again.

The costs are real, though. You run two full copies during the release. The switch is not as atomic as it looks: new connections go to green once endpoints update, but existing keep-alive connections to blue Pods stay open until they close, so keep blue running for a while after the switch. And blue/green does not remove the database compatibility problem; if green migrates the schema in a way blue cannot read, “instant rollback” rolls back to a version that no longer works.


Canary

A canary sends a small share of real traffic to the new version, watches error rates and latency, and widens the share only if they stay healthy. The difference from a rolling update is control: the rollout pauses at a percentage you choose, for as long as you choose, and moves forward on evidence.

Replica-ratio canary with plain Services

The simplest version needs nothing beyond two Deployments that share the label the Service selects:

# demo-api-stable: 9 replicas, labels { app: demo-api, track: stable }, image 1.0.0
# demo-api-canary: 1 replica,  labels { app: demo-api, track: canary }, image 2.0.0
# Service selector: { app: demo-api }   # matches both

kube-proxy spreads connections across all ready endpoints, so about one in ten new connections reaches the canary. Run the probe loop from the setup section and you will see hello from 2.0.0 roughly one line in ten. Promoting is scaling canary up and stable down; aborting is scaling canary to zero.

The limits are why teams outgrow this quickly. The traffic share is tied to replica counts, so 1% needs 99 stable Pods. It is per connection, not per request, so a client with a long-lived keep-alive connection is either all-canary or not at all. And there is no way to route by header or cookie, for example to send only internal users to the canary.

Weighted routing

For percentages independent of replica counts, the routing layer has to do the split. With the Gateway API, an HTTPRoute can weight two backend Services directly:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: demo-api
spec:
  parentRefs:
    - name: my-gateway
  rules:
    - backendRefs:
        - name: demo-api-stable
          port: 80
          weight: 95
        - name: demo-api-canary
          port: 80
          weight: 5

This needs a Gateway API implementation installed in the cluster. Older setups often use the ingress-nginx controller’s canary annotations (nginx.ingress.kubernetes.io/canary: "true" with canary-weight or header-based rules) on a second Ingress. The Kubernetes project has announced the retirement of the community ingress-nginx controller, so for new clusters the Gateway API route is the safer thing to learn.

Automating the decision

Moving weights by hand and reading dashboards between steps works, but it is exactly the kind of repetitive judgment that gets skipped under pressure. Controllers such as Argo Rollouts and Flagger replace the Deployment (or wrap it) with a rollout object that shifts weight in steps and queries metrics such as error rate or latency from Prometheus at each step, aborting automatically when the analysis fails. The real work there is not the YAML but choosing metrics that actually move when the new version is broken; a canary that is judged on CPU usage will happily promote a version that returns 500s quickly.


Choosing a strategy

StrategyDowntimeExtra capacityRollback speedTwo versions live at onceFits
RollingUpdateNone if probes and shutdown are rightUp to maxSurgeMinutes (old Pods must start again)Yes, during the rolloutMost stateless services
RecreateYesNoneMinutes, with downtime againNoSingle writers, RWO volumes, incompatible migrations
Blue/greenNoneFull second copySeconds (selector switch)Only while drainingReleases that need a fast, clean rollback
CanaryNoneSmall to full copySeconds (weight to 0)Yes, deliberatelyHigh-traffic services where you want evidence before full exposure

In practice the order of adoption is usually the same: get RollingUpdate right first (readiness, graceful shutdown, maxUnavailable: 0), because blue/green and canary both depend on the same Pod behavior. Add blue/green when rollback time matters more than capacity, and canary when you have enough traffic and good enough metrics for a small percentage to tell you something.


Troubleshooting rollouts

SymptomWhat to check
rollout status hangs, then fails with ProgressDeadlineExceededNew Pods never become ready: kubectl describe pod events for probe failures, ImagePullBackOff or Pending
New Pods in ImagePullBackOff on minikubeImage not loaded into the node, or tag latest forcing a pull
A few 502s or connection resets at every deployMissing SIGTERM handling or preStop delay; app exits before endpoints update
Rollout stuck with one new Pod Pending or waiting for a volumeReadWriteOnce volume still attached to the old Pod; use Recreate
rollout undo worked, then the bad version came backGitOps re-synced the manifest; revert the commit as well
Canary gets more or less traffic than the replica ratioKeep-alive connections; per-connection balancing is not per-request
Service has no endpoints after a blue/green switchSelector typo; kubectl get endpoints demo-api must list Pod IPs

When a rollout misbehaves, kubectl describe deployment demo-api is the first stop: its Conditions and Events show which ReplicaSet is scaling, how many Pods are available, and whether the progress deadline has been hit. From there, kubectl describe pod on one of the new Pods usually says why it is not ready. For Pod-level failures in more depth (CrashLoopBackOff, OOMKilled, probe failures), see Debugging Kubernetes Pods.


Frequently Asked Questions (FAQ)

Q. Does kubectl apply with an unchanged image tag trigger a rollout?

A. No. A rollout starts only when something in spec.template changes. Rebuilding an image under the same tag changes nothing Kubernetes can see, which is one more reason to tag each build uniquely (a Git SHA in CI). kubectl rollout restart deployment/<name> forces a new rollout by adding a timestamp annotation to the template, which is useful to pick up changed ConfigMaps or Secrets that are read only at startup.