Learn / Kubernetes survival kit / Rollouts and rollbacks

Lesson 2 of 5 7 min

Rollouts and rollbacks

How a Deployment replaces Pods safely, readiness probes as the real safety net, and rolling back a bad release fast.

Rolling updates: gradual, not all-at-once

When you change a Deployment’s Pod spec (most commonly, a new container image tag), the default RollingUpdate strategy replaces old Pods with new ones incrementally, not by tearing everything down and starting over. Two settings control the pace:

spec:
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1          # how many EXTRA Pods above the desired replica count are allowed mid-rollout
      maxUnavailable: 0    # how many Pods are allowed to be unavailable mid-rollout

maxUnavailable: 0 with maxSurge: 1 is a common “zero downtime” pattern: Kubernetes creates one new Pod before removing an old one, so capacity never drops below the desired count, just briefly exceeds it while the swap happens.

The real safety net: readiness probes

A rolling update is only as safe as Kubernetes’ ability to know a new Pod is actually working. By default, a container is considered “ready” the moment its process starts - which is often well before your application has finished loading config, warming a cache, or connecting to a database. Without a readiness probe, the rollout happily routes live traffic to a Pod that’s still starting up, and requests fail during every deploy.

spec:
  containers:
    - name: hello
      image: myregistry/hello:1.5
      readinessProbe:
        httpGet:
          path: /healthz
          port: 8080
        initialDelaySeconds: 3
        periodSeconds: 5

With this in place, the Service (previous lesson) only routes traffic to Pods that are passing /healthz - and the rollout controller only considers a new Pod successfully rolled out once it’s ready, which is also what makes an automatic rollout failure detectable at all.

Watching and controlling a rollout

kubectl rollout status deployment/hello       # blocks until the rollout finishes or fails
kubectl rollout history deployment/hello      # lists revisions
kubectl rollout pause deployment/hello        # stop mid-rollout, e.g. to inspect a partial state
kubectl rollout resume deployment/hello       # continue a paused rollout

rollout status is the command to actually use in a deploy script or CI pipeline instead of polling kubectl get pods and eyeballing it - it exits non-zero if the rollout fails (e.g. new Pods never become ready), which is exactly the signal a pipeline needs to stop and alert.

Rolling back

Every update to a Deployment’s Pod template creates a new ReplicaSet revision; the old ReplicaSet is kept (scaled to zero) rather than deleted, which is what makes a fast rollback possible:

kubectl rollout undo deployment/hello                  # back to the previous revision
kubectl rollout undo deployment/hello --to-revision=3   # back to a specific revision

This performs the same rolling-update mechanism in reverse - gradually replacing current Pods with Pods matching the target revision’s spec, respecting the same maxSurge/maxUnavailable settings. It’s usually the fastest first response to “the new release is broken” - revert first with a single command, investigate the root cause after traffic is stable again, not before.

Key takeaways

  • A Deployment's default rolling update replaces old Pods with new ones gradually, controlled by maxSurge and maxUnavailable, instead of an all-at-once replacement.
  • A readiness probe is what makes a rollout actually safe - without one, Kubernetes considers a new Pod 'ready' the instant its container starts, even if the application inside hasn't finished booting.
  • kubectl rollout status watches a rollout to completion (or failure) instead of guessing from kubectl get pods; kubectl rollout undo reverts to the previous ReplicaSet immediately.
  • Every Deployment update is recorded as a new ReplicaSet revision, which is what a rollback actually reverts to - the old ReplicaSet's Pod spec, not a re-deploy of old code from scratch.

Quick check

3 questions - see how much stuck.

1. What controls how many Pods are replaced at once during a rolling update?
2. What does a readiness probe actually protect against?
3. What does 'kubectl rollout undo' actually do?