Learn / AWS Lambda for backend devs / Cold starts

Lesson 3 of 5 7 min

Cold starts

What actually happens during a cold start, what makes them worse, and the real levers to reduce their impact.

What a cold start actually is

When a request arrives and no existing execution environment can serve it (first invocation ever, a burst past current warm capacity, or a new deployment), AWS has to:

  1. Provision a new execution environment.
  2. Download your deployment package (or pull your container image).
  3. Start the language runtime.
  4. Run any code at your module’s top level - imports, client initialization, config loading - before your handler function is even called.

Only after all of that does your handler run for the first time on that environment. A warm invocation skips straight to step 5 (the handler call) because an already-initialized environment is reused - which is why warm invocations are typically dramatically faster than cold ones.

What makes it worse

  • Package/image size - more to download before startup.
  • Runtime choice - interpreted runtimes generally start faster than ones needing heavier VM/JIT warmup; the relative gap varies by runtime and changes over time as AWS improves startup performance, so treat this as a factor to measure for your specific stack rather than a fixed ranking.
  • Heavy module-level initialization - loading a large ML model, building a big in-memory structure, or opening several connections at import time all run on every cold start, before the handler even begins.
  • VPC attachment - functions inside a VPC used to add meaningful cold-start latency from network interface (ENI) setup; AWS has significantly reduced this over time with improvements to how ENIs are provisioned, but it’s still worth confirming current behavior for your runtime/region rather than assuming it’s identical to a non-VPC function.

Levers you actually control

  1. Trim the package (lesson 2) - fewer, lighter dependencies download and initialize faster.
  2. Move initialization out of the hot path where possible. Anything that must run once (a DB connection pool, a loaded config) belongs at module level so it’s reused across warm invocations - but anything not needed for every invocation shouldn’t run eagerly at import time either. Lazy-init what’s conditionally needed.
  3. Prefer a lighter runtime/architecture for latency-sensitive functions where your dependencies allow it, and measure - don’t assume - the actual difference for your specific package.
  4. Provisioned Concurrency for the functions where cold-start latency directly hits users (a synchronous, user-facing API call). It’s a standing cost (you pay for the reserved, pre-warmed capacity whether or not it’s invoked), so apply it to the specific functions that need it, not the whole application by default.

When to not bother

Cold starts matter most on the synchronous, user-facing path - an API Gateway-fronted function a person is waiting on. They matter far less for asynchronous or batch workloads (an S3-triggered processing function, a scheduled job) where an extra few hundred milliseconds of one invocation is invisible in the overall pipeline. Spend the optimization effort where a human is actually waiting.

Key takeaways

  • A cold start is the one-time cost of provisioning a fresh execution environment - downloading code, starting the runtime, and running your top-level (module-load-time) initialization - before your handler runs for the first time on that environment.
  • Package size, runtime choice, and how much work happens at module load time (not inside the handler) are the main levers you control; VPC-attached functions can add extra cold-start latency from ENI setup, though AWS has substantially improved this over the years.
  • Provisioned Concurrency keeps a set number of execution environments pre-initialized and warm, trading a standing cost for eliminated cold-start latency on that reserved capacity.
  • Not every function needs cold-start optimization - it matters most for latency-sensitive, user-facing paths (a synchronous API call), and matters far less for async/batch workloads where a few extra hundred milliseconds is noise.

Quick check

3 questions - see how much stuck.

1. What actually happens during a Lambda cold start?
2. Which of these is a lever a developer directly controls to reduce cold-start impact?
3. What does Provisioned Concurrency actually do?