Skip to content

Kubernetes probes and resource limits explained

By · Kubernetes & DevOps · 5 min read · Published

Two small sections of a Kubernetes Deployment cause a large share of production incidents: health probes and resource settings. Configured well, they let Kubernetes route traffic only to healthy pods, restart stuck ones and pack workloads efficiently onto nodes. Configured badly, they cause restart loops, dropped requests during deploys, throttled services and pods killed for using too much memory. This guide explains how each setting works and gives sensible starting values.

The three probes

Kubernetes can run three kinds of checks against each container. Each can use an HTTP GET, a TCP connection, a gRPC health check or a command.

Readiness probe: should this pod receive traffic?

When a readiness probe fails, Kubernetes removes the pod from the Service's endpoints, so no new requests are routed to it. The container is not restarted. Readiness covers temporary conditions: the app is still starting, warming a cache, or is overloaded and wants to shed traffic. Without a readiness probe, a pod receives traffic as soon as its container starts, which often means errors during every rollout.

Liveness probe: should this container be restarted?

When a liveness probe fails repeatedly, the kubelet kills and restarts the container. Use it only to detect states the app cannot recover from by itself, such as a deadlock or an event loop that stopped responding. A liveness probe is a blunt instrument, and misconfigured ones are dangerous.

Startup probe: has the app finished starting?

While a startup probe is defined and has not yet succeeded, liveness and readiness checks are disabled. It lets slow-starting applications, such as large Java services, take a few minutes to boot without loosening the liveness probe for the rest of their life.

Probe mistakes that cause outages

  • Checking dependencies in liveness. If the liveness endpoint checks the database and the database has a brief outage, every pod fails liveness and restarts at once. Restarting does not fix the database, and now the whole service is down and cold. Liveness should check only the process itself.
  • Timeouts that are too tight. The default timeoutSeconds is 1. Under load, a healthy app may take longer to answer, fail liveness, and get restarted, which increases load on the remaining pods and spreads the failure.
  • No startup allowance. A liveness probe that begins before the app is ready kills it during startup, producing a CrashLoopBackOff that looks like an application bug.
  • Same endpoint for everything. Readiness and liveness answer different questions; they usually deserve different endpoints, such as /ready and /live.

Reasonable probe settings

startupProbe:
  httpGet: { path: /live, port: 8080 }
  periodSeconds: 5
  failureThreshold: 30        # up to 150 s to start
readinessProbe:
  httpGet: { path: /ready, port: 8080 }
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 3
livenessProbe:
  httpGet: { path: /live, port: 8080 }
  periodSeconds: 10
  timeoutSeconds: 3
  failureThreshold: 3

The readiness endpoint may check critical dependencies, so a pod that cannot reach its database stops receiving traffic. The liveness endpoint should return 200 as long as the process can serve a trivial request.

Requests and limits

Each container can declare resource requests and limits for CPU and memory:

resources:
  requests:
    cpu: 250m
    memory: 256Mi
  limits:
    memory: 512Mi
  • Requests are what the scheduler reserves. A pod is placed only on a node with enough unreserved capacity to cover its requests. Requests also determine the container's share of CPU when the node is busy.
  • Limits are the maximum the container may use.

CPU is measured in cores or millicores: 250m is a quarter of a core. Memory uses binary units: 256Mi is 256 mebibytes.

CPU and memory behave differently

CPU is compressible

If a container tries to use more CPU than its limit, it is throttled: the kernel pauses it until the next scheduling period. Nothing crashes, but latency spikes. Throttling can occur even when average usage looks low, because limits are enforced in 100 ms windows and a burst of work can exhaust the quota early in the window. For latency-sensitive services, many teams set CPU requests but no CPU limit, letting containers use idle CPU on the node.

Memory is not compressible

If a container exceeds its memory limit, the kernel kills it, and the pod shows OOMKilled as the reason for the last termination. If the node itself runs out of memory, the kubelet evicts pods, starting with those using the most memory relative to their requests. Always set a memory limit, and set the memory request close to it so the scheduler does not overcommit the node.

Quality of Service classes

Kubernetes assigns each pod a QoS class from its settings, which decides eviction order under memory pressure:

ClassWhenEviction priority
GuaranteedEvery container has requests equal to limits for both CPU and memoryLast
BurstableAt least one request or limit is set, but not GuaranteedMiddle
BestEffortNo requests or limits at allFirst

Pods with no resource settings are the first to go when a node is under pressure, and they also make cluster capacity impossible to plan.

Choosing values

  1. Start with an estimate: run the service under realistic load and observe peak memory and typical CPU.
  2. Set the memory request to typical usage and the memory limit to peak usage plus headroom of 25 to 50 percent.
  3. Set the CPU request to typical usage. Add a CPU limit only if you need to protect neighbours from a noisy workload.
  4. Tell the runtime about the limit. JVMs since Java 10 respect container memory limits, but heap settings like -XX:MaxRAMPercentage still need tuning; Node.js may need --max-old-space-size.
  5. Watch metrics after deploying and adjust. A Vertical Pod Autoscaler in recommendation mode can suggest values.

Putting it together

A production-ready Deployment has a readiness probe on every container, a simple liveness probe, a startup probe for slow starters, a memory limit, and CPU and memory requests based on measurements. The Kubernetes YAML Generator produces a Deployment and Service with these fields already wired up, and you can validate any manifest you edit by hand with the YAML Formatter.

More Kubernetes & DevOps guides