not accepting clients
← back to blog

graceful shutdown in kubernetes

rolling updates drop connections when kubelet sends sigterm before the endpoint is removed from routing. the fix spans the pod spec and the application.

a deployment rolls out. replicas are healthy, capacity is more than sufficient, the new pods pass their readiness probes. and yet the error rate shows a spike of connection resets for the fifteen seconds the rollout takes, and adding replicas does not change it.

the failures come from ordering. when a pod is deleted, kubernetes does two things at the same time: it tells kubelet to terminate the container and it tells every endpoint consumer to stop routing to that pod. the two are independent and eventually consistent. a request accepted in the gap between them arrives at a process that is already shutting down.

the short version

pod deletion is not a sequence. the api server marks the pod terminating, and from that moment two paths run in parallel:

                   ┌─ kubelet ─────────▶ preStop hook ─▶ SIGTERM ─▶ grace period ─▶ SIGKILL
pod marked deleted ┤
                   └─ endpointslice updated ─▶ kube-proxy / ingress / service mesh
                                               reload their routing tables

the second path involves multiple hops - the endpointslice controller writes, every kube-proxy or ingress controller watches and reloads - and it is not fast. the first path can reach the application in milliseconds. the application therefore stops accepting connections before the thing sending it connections has learned to stop.

closing that gap takes two things: delaying the shutdown until routing has caught up, and then shutting down in the right order once it has.

the propagation delay is real

the endpointslice update itself is quick. what takes time is every consumer acting on it. kube-proxy in iptables mode rewrites rules on each node. an ingress controller reloads or reconfigures its upstreams. a cloud load balancer pointed at node ports has its own deregistration delay, often measured in tens of seconds. an external client with a connection pool holds an established tcp connection that no routing table change affects at all.

none of these are under the pod’s control, and none of them report back. the only mechanism a pod has is to stay alive and keep serving during the window.

that is what a preStop sleep buys:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: payments
spec:
  template:
    spec:
      terminationGracePeriodSeconds: 60
      containers:
        - name: payments
          image: ghcr.io/kontrolplane/payments:1.4.2
          lifecycle:
            preStop:
              sleep:
                seconds: 10

sleep as a native lifecycle action landed in kubernetes 1.29 and is stable in 1.32. before that the same thing was written as exec: command: ["sleep", "10"], which requires a sleep binary in the image - a real problem for distroless and scratch images, and the reason so many services shipped a shell they did not otherwise need.

during those ten seconds the container has not been signalled yet. it serves traffic normally while the routing tables converge. SIGTERM is sent only after the hook returns.

the grace period is a budget, not a delay

terminationGracePeriodSeconds is the total wall clock allowed between the delete and SIGKILL, and the preStop hook spends from that same budget. a 10 second hook against the default 30 second grace period leaves the process 20 seconds to drain, not 30. if the hook alone outlasts the grace period, kubelet sends SIGKILL while the hook is still running and the application never receives SIGTERM at all.

the budget needs to cover the longest legitimate request the service handles, plus the hook. a service whose p99 is 2 seconds is fine on 30. a service that streams responses for two minutes needs a grace period longer than two minutes, or its longest requests are guaranteed to be killed mid-flight on every rollout.

handling the signal

once SIGTERM arrives, the application has to stop accepting new connections while finishing the ones it already has. in go the http server does exactly this with Shutdown:

package main

import (
	"context"
	"errors"
	"log/slog"
	"net/http"
	"os/signal"
	"syscall"
	"time"
)

func main() {
	ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM, syscall.SIGINT)
	defer stop()

	server := &http.Server{
		Addr:    ":8080",
		Handler: routes(),
	}

	go func() {
		if err := server.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
			slog.Error("listen failed", "error", err)
		}
	}()

	<-ctx.Done()
	slog.Info("sigterm received, draining")

	drainCtx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
	defer cancel()

	if err := server.Shutdown(drainCtx); err != nil {
		slog.Error("drain timed out, closing forcefully", "error", err)
		server.Close()
	}

	slog.Info("shutdown complete")
}

Shutdown closes the listeners immediately, so no new connection is accepted, then waits for in-flight handlers to return before releasing. the 20 second drain timeout must sit inside the grace period minus the prestop hook - here 60 minus 10 leaves 50, comfortably.

the ordering after the http server drains matters too. connections to databases, message brokers, and caches are closed after Shutdown returns, never before, because in-flight handlers still need them.

pid 1 and signal delivery

kubelet sends SIGTERM to pid 1 in the container. if pid 1 is a shell, the signal usually goes nowhere useful.

# wrong - the shell is pid 1 and does not forward signals to the child
CMD /app/payments --config /etc/payments.yaml
# better - the binary is pid 1 and receives sigterm directly
CMD ["/app/payments", "--config", "/etc/payments.yaml"]

the shell form runs /bin/sh -c "...", and sh does not forward signals to its children unless it has exec’d them. the application never sees SIGTERM, sits idle until the grace period expires, and is then SIGKILLed. a rollout that looks like it takes exactly terminationGracePeriodSeconds per pod is nearly always this.

entrypoint scripts have the same problem, solved by handing the process over rather than spawning it:

#!/bin/sh
set -e
/app/migrate --wait
exec /app/payments --config /etc/payments.yaml

readiness during shutdown

a common suggestion is to fail the readiness probe on SIGTERM so the endpoint is removed. by the time SIGTERM arrives the endpoint removal is already in flight, and the probe interval - default 10 seconds - is slower than the deletion path anyway. it adds nothing that the preStop sleep does not already cover.

readiness matters at the other end of the lifecycle. minReadySeconds holds a new pod as unavailable for a fixed period after it first reports ready, which prevents a rollout from advancing on pods that are ready but not yet warm:

spec:
  minReadySeconds: 15
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0
      maxSurge: 1

maxUnavailable: 0 is the setting that makes a rollout additive: a new pod is created and becomes ready before an old one is removed. with the default of 25% a rollout starts by deleting a quarter of the fleet.

non-http workloads

a queue consumer has no listener to close and no in-flight http request to finish. draining means stopping the fetch loop and settling what is already in hand:

func (c *Consumer) Run(ctx context.Context) error {
	for {
		select {
		case <-ctx.Done():
			return c.drain()
		default:
		}

		msgs, err := c.fetch(ctx, 10)
		if err != nil {
			return err
		}

		for _, msg := range msgs {
			if err := c.handle(msg); err != nil {
				msg.Nak()
				continue
			}
			msg.Ack()
		}
	}
}

on shutdown, unacknowledged messages should be negatively acknowledged explicitly or left unacked. a message that is neither acked nor nacked is redelivered after its ack wait expires, which is correct but slow. an explicit nak returns it to the broker immediately.

what to watch out for

the grace period includes the prestop hook. the two are not additive. a 30 second hook under a 30 second grace period means the process is killed the instant the hook returns, without ever handling SIGTERM.

shell-form CMD swallows the signal. any container that consistently takes the full grace period to terminate should be checked for a shell at pid 1 before anything else is tuned.

preStop does not run on node failure. the hook is a kubelet action. if the node is gone, nothing runs, and the pod is force-deleted after the node controller’s eviction timeout. graceful shutdown handles rollouts and evictions, not hardware.

long-lived connections ignore endpoint changes. websockets, grpc streams, and http/2 connections stay pinned to the pod they were established against. the server has to close them itself - GOAWAY for grpc, a close frame for websockets - and the client has to reconnect.

a pod disruption budget does not slow a rollout. pdbs constrain voluntary evictions such as node drains, not the deployment controller. a rollout that respects availability is configured through maxUnavailable on the deployment.

sidecars can outlive or predecease the app. a proxy sidecar that exits on SIGTERM before the main container finishes draining takes the network path with it. native sidecars - init containers with restartPolicy: Always, stable since kubernetes 1.29 - are terminated after the regular containers, which is the ordering that makes draining work.

references

[1] kubernetes documentation. “pod lifecycle: termination of pods.”
kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle

[2] kubernetes documentation. “container lifecycle hooks.”
kubernetes.io/docs/concepts/containers/container-lifecycle-hooks

[3] kubernetes documentation. “sidecar containers.”
kubernetes.io/docs/concepts/workloads/pods/sidecar-containers

[4] kubernetes documentation. “deployments: rolling update.”
kubernetes.io/docs/concepts/workloads/controllers/deployment

[5] go documentation. “net/http: server.shutdown.”
pkg.go.dev/net/http#Server.Shutdown

[6] kubernetes documentation. “endpointslices.”
kubernetes.io/docs/concepts/services-networking/endpoint-slices

# ask the author

a question
about this
post?

direct line