Learn / Kubernetes survival kit / Autoscaling, including KEDA

Lesson 5 of 5 7 min

Autoscaling, including KEDA

HPA scaling on CPU/memory, why that's not enough for many real workloads, and event-driven scaling with KEDA.

HPA: scale replicas on CPU/memory

The Horizontal Pod Autoscaler watches a metric (CPU and memory utilization by default) and adjusts a Deployment’s replica count to keep that metric near a target you set:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: hello-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: hello
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

This scales replicas up as average CPU utilization across Pods climbs past 70% of what they requested (not their limit), and back down as it drops - within the minReplicas/ maxReplicas bounds. Note the requirement this implies: resources.requests must be set on the Pod spec, because utilization is computed as a percentage of the requested amount - with no request defined, there’s nothing to compute the percentage against.

Where CPU/memory scaling falls short

Plenty of real workloads don’t correlate cleanly with CPU or memory usage. A worker consuming messages from a queue is the classic case: under a growing backlog, it might sit at modest, steady CPU usage the whole time - it’s I/O-bound waiting on the queue, not CPU-bound - while falling further and further behind. Plain HPA sees nothing alarming in CPU and never scales up, even as the actual, user-visible problem (backlog growing) gets worse.

KEDA: scale on the metric that actually reflects load

KEDA (Kubernetes Event-Driven Autoscaling) extends the autoscaling model with scalers for real event sources - queue depth (SQS, RabbitMQ, Azure Service Bus), Kafka consumer lag, Prometheus query results, and many others - so you can scale on the number that actually represents backlog, not a proxy metric that might not move with it.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: queue-worker-scaler
spec:
  scaleTargetRef:
    name: queue-worker
  minReplicaCount: 0
  maxReplicaCount: 20
  triggers:
    - type: aws-sqs-queue
      metadata:
        queueURL: https://sqs.ap-south-1.amazonaws.com/123456789012/my-queue
        queueLength: "5"          # target: ~5 messages per replica

Under the hood, KEDA works with HPA - it creates and manages an HPA object using its own custom metrics as the scaling signal, rather than replacing the HPA mechanism entirely.

Scale-to-zero: the capability plain HPA doesn’t have

minReplicaCount: 0 above is a real capability plain HPA doesn’t offer - HPA’s minReplicas can’t go below 1. KEDA can scale a Deployment all the way down to zero replicas when there’s genuinely no work (an empty queue) and back up the moment new events arrive, which is a meaningful cost saving for spiky or intermittent workloads that would otherwise sit idle at their minimum replica count around the clock.

Choosing between them

Use plain HPA for workloads where CPU/memory genuinely reflects load - a typical stateless web API serving synchronous requests is usually a good fit. Reach for KEDA when the real load signal is external to the Pod itself - queue depth, stream lag, a custom business metric - or when scale-to-zero is a real cost win for an intermittent workload.

Key takeaways

  • The Horizontal Pod Autoscaler (HPA) adjusts a Deployment's replica count based on observed metrics (CPU/memory by default) against a target you set - it changes replica count, not per-Pod resource limits.
  • HPA needs resources.requests set on the Pod spec to calculate utilization percentages against - without requests defined, CPU/memory-based autoscaling has nothing to compare current usage to.
  • Plain CPU/memory-based scaling is a poor fit for workloads whose load isn't CPU/memory-shaped - a queue consumer under a growing backlog might sit at low CPU while falling further behind.
  • KEDA extends autoscaling to scale on event-source metrics directly (queue depth, Kafka lag, a custom metric) - including scaling a Deployment down to zero when there's no work, which plain HPA cannot do.

Quick check

3 questions - see how much stuck.

1. What does the Horizontal Pod Autoscaler (HPA) actually change?
2. Why does HPA require resources.requests to be set on the Pod spec for CPU/memory-based scaling to work?
3. What can KEDA do that plain CPU/memory-based HPA cannot?