Writing a production-ready Dockerfile: multi-stage builds, layer caching and non-root users
A Dockerfile that works takes five minutes. A Dockerfile that is ready for production takes a few more decisions. It should rebuild in seconds when only code changes, produce a small image, run without root, shut down cleanly when the orchestrator asks, and contain nothing an attacker could use. This guide takes a typical first attempt for a Node.js service and fixes it one problem at a time. The techniques carry over directly to Python, Java and Go.
The starting point
FROM node:22
WORKDIR /app
COPY . .
RUN npm install
RUN npm run build
EXPOSE 3000
CMD npm start
It works on a laptop, and it has eight problems. The base image is over a gigabyte and full of compilers. Any source change invalidates the dependency install. COPY . . may copy .env, .git and the host's node_modules. Dev dependencies and source ship to production. npm install may resolve different versions than the lockfile. The process runs as root. The shell-form CMD swallows shutdown signals. And nothing tells Docker whether the app is healthy.
1. Keep the build context clean
Everything in the build directory is sent to the builder and can end up in a layer. A .dockerignore is the cheapest improvement you can make:
node_modules
.git
.env*
dist
coverage
*.log
Dockerfile
.dockerignore
Besides speed, this protects you from the worst Docker mistake: a secret copied into a layer stays in the image even if a later RUN rm deletes it. Anyone who pulls the image can recover it with docker history or by unpacking the layers.
2. Order layers for caching
Each instruction produces a cached layer, and a change to one layer rebuilds every layer after it. Copy the dependency manifest first, install, then copy the source:
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
Now a code change reuses the installed dependencies and the build takes seconds instead of minutes. npm ci installs exactly what the lockfile says and fails if the lockfile and package.json disagree. That is what you want in a build. The same order applies to requirements.txt, go.mod/go.sum and pom.xml.
With BuildKit, which is the default builder in current Docker, you can also keep the package manager's download cache between builds, so even a lockfile change does not re-download everything:
RUN --mount=type=cache,target=/root/.npm npm ci
The cache lives on the build machine, not in the image, so it costs nothing in image size. On CI runners that start fresh each time, use your CI's cache export instead (--cache-to/--cache-from with buildx).
3. Split build and runtime with a multi-stage build
# syntax=docker/dockerfile:1
FROM node:22-slim AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN --mount=type=cache,target=/root/.npm npm ci
COPY . .
RUN npm run build && npm prune --omit=dev
FROM node:22-slim
ENV NODE_ENV=production
WORKDIR /app
COPY --from=build --chown=node:node /app/node_modules ./node_modules
COPY --from=build --chown=node:node /app/dist ./dist
COPY --chown=node:node package.json ./
# the image's "node" user, by UID
USER 1000
EXPOSE 3000
HEALTHCHECK --interval=30s --timeout=5s --start-period=15s --retries=3 \
CMD ["node", "-e", "fetch('http://localhost:3000/health').then(r => process.exit(r.ok ? 0 : 1), () => process.exit(1))"]
CMD ["node", "dist/server.js"]
The first stage has the full toolchain, dev dependencies and source. The final stage receives only production node_modules and the compiled output. Nothing else from the build stage exists in the image you ship. For compiled languages the effect is even bigger. A Go binary built with CGO_ENABLED=0 can run from scratch or a distroless image of a few megabytes. A Java service needs only a JRE image, not a JDK plus Maven. Moving from the starting point to a slim multi-stage build typically takes a Node.js image from over 1 GB to a couple of hundred megabytes; measure yours with docker images.
4. Pick the base image deliberately
- slim (Debian with fewer packages) is a safe default with glibc, so native modules behave as on most servers.
- alpine is smaller, but uses musl libc. Native modules need musl builds, DNS resolution behaves differently, and some workloads run slower. Fine for many apps; test before switching.
- distroless contains the runtime and nothing else: no shell, no package manager. An attacker who gets code execution has very little to work with. The trade-off is that you cannot
docker execa shell to debug, and health checks must be written in the app's own language. - scratch is empty, for static binaries. Remember to copy CA certificates if the binary makes TLS calls.
Pin versions (node:22.11-slim) for reproducible builds. For critical services, pin by digest (node:22-slim@sha256:…) and let Renovate or Dependabot open pull requests when the base image updates, so you get security patches deliberately rather than by accident.
5. Do not run as root
Containers run as root unless told otherwise. Isolation limits the damage, but root inside a container plus a vulnerability in your app or the runtime is a much worse day than the same bug running as an unprivileged user. Official Node images ship a node user with UID 1000, used above. Elsewhere, create one:
RUN useradd --system --uid 10001 --no-create-home app
USER 10001
Write USER with a numeric UID. Kubernetes' runAsNonRoot: true check cannot resolve a user name, and refuses to start the container with CreateContainerConfigError. Use COPY --chown so the user can read its files, keep anything it must write to a specific directory or a mounted volume, and listen on a port above 1024.
6. Keep secrets out of layers
Runtime secrets belong in environment variables or mounted files supplied when the container starts, never in ENV or ARG instructions. Build-args are recorded in the image history. For secrets needed during the build, such as a private registry token, use a BuildKit secret mount, which is never written to a layer:
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ci
# docker build --secret id=npmrc,src=$HOME/.npmrc .
7. Shut down properly
When Docker stops a container, or Kubernetes replaces a pod, the main process gets SIGTERM, and after a grace period (10 seconds in Docker, 30 in Kubernetes) a SIGKILL. Two things commonly go wrong.
First, shell form. CMD npm start runs /bin/sh -c "npm start", and the signal goes to the shell, which does not pass it on. Use exec form, CMD ["node", "dist/server.js"], so your process is PID 1 and receives the signal itself. Skip npm start in production, since npm adds another process layer.
Second, the app must actually handle the signal. A Node.js process running as PID 1 does not get the default "exit on SIGTERM" behaviour, so without a handler it ignores the signal until the SIGKILL arrives. Handle it by finishing in-flight requests:
process.on('SIGTERM', () => {
server.close(() => process.exit(0)); // stop accepting, finish in-flight requests
setTimeout(() => process.exit(1), 10_000).unref(); // hard stop if something hangs
});
In a test, a request that was 200 ms into a 1.5-second response when the signal arrived still completed, and the process exited with code 0 about 1.5 seconds later. If you cannot change the app, run it under a minimal init with docker run --init or tini as the entrypoint, which forwards signals and reaps zombie processes.
8. Health checks
The HEALTHCHECK above uses Node's built-in fetch because slim and distroless images have no curl. Alpine images include BusyBox wget, which is why the Dockerfile Generator uses wget -qO- for its Alpine-based output. Docker and Compose use the result. Kubernetes ignores it and uses its own probes, explained in the probes guide. The same /health endpoint serves both.
9. Verify in CI
docker build --check .runs Docker's built-in Dockerfile checks without building, flagging things like shell-formCMDand secrets inENV.- Scan the image with Trivy, Grype or Docker Scout, and fail the build on fixable critical vulnerabilities.
- Rebuild on a schedule even when code has not changed, so base image patches reach production.
- Run the container in CI as it will run in production: read-only root filesystem, non-root, and a stop to confirm it exits cleanly.
The Dockerfile Generator produces multi-stage, non-root Dockerfiles with a matching .dockerignore for Node.js, Python, Java, Go and static sites. Use it as a starting point and apply the rest of this list. If you deploy to Kubernetes, continue with Deployment, Service and Ingress: a minimal working example.