Docker Security in Production: Non-Root Images, Secrets, Capabilities and the Mistakes That Expose Hosts

Key takeaways

A container is a process with namespaces and cgroups around it, not a VM, so container security is mostly about not handing that process more than it needs. This guide goes layer by layer — image, build, secrets, runtime, network, Kubernetes — and spends most of its time on the mistakes that actually expose hosts and data.

The Container Security Threat Model

A container is not a virtual machine. It is an ordinary Linux process that the kernel places in its own namespaces (so it sees its own filesystem, process tree and network) and cgroups (so its resources are limited). Every container on a host shares that host’s kernel. That one fact explains most of the guidance below: the goal is to give the process as little as possible, so that if the application is compromised, the attacker has very little to work with and no easy path to the kernel or the host.

The realistic threats, in the order they tend to cause incidents:

  1. Exposed services — a database or admin port published to the internet by accident.
  2. Leaked secrets — credentials baked into image layers or sitting in environment variables that show up in logs and docker inspect.
  3. Vulnerable dependencies — outdated OS packages and libraries with known CVEs.
  4. Over-privileged containers — root, --privileged, extra capabilities, or the Docker socket mounted inside.
  5. Supply chain — base images, packages or CI actions that change underneath you.

Container escapes via kernel exploits get the headlines, but the incidents I have seen in practice were almost always the first two: something reachable that should not have been, or a credential somewhere it should not have been.


Dockerfile Security

Run as a non-root user

# Runs as root by default
FROM node:22-alpine
WORKDIR /app
COPY . .
RUN npm install
CMD ["node", "server.js"]
# Non-root, with application files owned by root (read-only to the app)
FROM node:22-alpine
WORKDIR /app

COPY package*.json ./
RUN npm ci --omit=dev

COPY . .

# The official Node images already include a 'node' user (UID 1000).
# Use the numeric UID so Kubernetes can verify it is non-root.
USER 1000
CMD ["node", "server.js"]

Two choices here differ from many tutorials.

The files stay owned by root. A common pattern is COPY --chown=appuser:appgroup . ., which makes the application user the owner of its own code. That means an attacker who gets code execution as that user can modify the application files, for example to plant a backdoor that survives in a long-running container. If the app only needs to read its code, leave the files root-owned and give the app user write access only to the specific directories it writes to.

USER 1000, not USER node. Kubernetes’ runAsNonRoot: true check needs a numeric UID. With a user name, the kubelet refuses to start the pod with container has runAsNonRoot and image has non-numeric user (node), cannot verify user is non-root. I have seen this take down a deploy that passed every local test, because plain Docker does not care.

npm ci --omit=dev replaces the deprecated --only=production, which newer npm versions warn about.

Minimal base images

FROM node:22                                   # Debian-based, full toolchain, largest attack surface
FROM node:22-slim                              # Debian with most extras removed
FROM node:22-alpine                            # musl-based, small
FROM gcr.io/distroless/nodejs22-debian12       # no shell, no package manager

Smaller images have fewer packages, which means fewer CVEs to track and fewer tools for an attacker to use. The trade-offs are real, though:

  • Alpine uses musl instead of glibc. Most Node, Python and Go code does not notice, but native modules and prebuilt binaries sometimes fail or behave differently (DNS resolution and some locale handling are the usual surprises). If you hit that, -slim images are a reasonable middle ground.
  • Distroless has no shell. That removes sh, curl and a package manager from an attacker’s toolkit, but it does not prevent code execution — a remote code execution bug in your Node app still runs attacker-controlled JavaScript. It also makes docker exec -it ... sh impossible when debugging; distroless publishes :debug variants with a busybox shell for that.

Multi-stage builds

# Build stage: dev dependencies and build tools live here only
FROM node:22-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build && npm prune --omit=dev

# Runtime stage: only what is needed to run
FROM node:22-alpine
WORKDIR /app
COPY --from=builder /app/dist ./dist
COPY --from=builder /app/node_modules ./node_modules
COPY --from=builder /app/package.json ./
USER 1000
EXPOSE 3000
CMD ["node", "dist/server.js"]

The npm prune --omit=dev step matters. Copying node_modules straight from a builder that ran a full npm ci ships every dev dependency — test frameworks, bundlers, linters — to production, which defeats much of the point. More on layer layout in the multi-stage build guide.

Pinning base images

FROM node:latest                                   # changes without notice
FROM node:22.11-alpine3.20                         # pinned version, still re-tagged for patches
FROM node:22.11-alpine3.20@sha256:<digest>         # exactly one image, forever

Pinning to a digest makes builds reproducible, but it also freezes security patches: that image will never receive an OS package update. Pinning only works if something updates the pin — Dependabot or Renovate can open pull requests when a new digest is published. A pinned digest that nobody updates is how images end up running two-year-old OpenSSL.


Image Scanning

# macOS
brew install trivy

# Scan an image
trivy image --severity HIGH,CRITICAL myapp:latest

# Fail CI on critical findings that have a fix available
trivy image --exit-code 1 --severity CRITICAL --ignore-unfixed myapp:latest

# Dockerfile / IaC misconfigurations
trivy config ./Dockerfile

# Filesystem: dependencies and misconfigurations in a repo
trivy fs --scanners vuln,misconfig .

--scanners replaced the older --security-checks flag, which newer Trivy versions reject.

--ignore-unfixed is what makes scanning usable as a CI gate. A base image routinely has CVEs with no available fix; failing the build on those only teaches the team to disable the scanner. Gate on what you can actually fix, and review the rest on a schedule.

In CI

# .github/workflows/security.yml
- name: Scan image
  uses: aquasecurity/trivy-action@<full-commit-sha>   # pin actions to a commit, not @master
  with:
    image-ref: 'myapp:${{ github.sha }}'
    exit-code: '1'
    severity: 'CRITICAL,HIGH'
    ignore-unfixed: true

Referencing a third-party action by @master or even a version tag means running whatever that ref points to at build time, with access to your repository secrets. In March 2025 the popular tj-actions/changed-files action was compromised and its tags were repointed to code that dumped CI secrets into build logs; workflows pinned to a commit SHA were not affected. Security scanning steps are not exempt from this.


Secrets Management

What leaks, and where

# Visible to anyone who can pull the image: docker history / docker inspect
ENV API_KEY=sk-production-key-123
ARG NPM_TOKEN
RUN echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > .npmrc && npm ci && rm .npmrc

# Still in the image: the file lives on in the COPY layer even after rm
COPY .env .
RUN rm .env

The ARG example is the one that catches careful people. It looks safe because the .npmrc is removed in the same RUN, but build arguments used in a RUN are recorded in the image history, so docker history --no-trunc shows the token. Each Dockerfile instruction creates a layer, and deleting a file in a later layer does not remove it from earlier ones.

Build secrets with BuildKit

# syntax=docker/dockerfile:1
FROM node:22-alpine
WORKDIR /app
COPY package*.json ./
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ci --omit=dev
docker build --secret id=npmrc,src=$HOME/.npmrc .

The secret is mounted only for that one RUN step and is never written into a layer or the build cache.

Runtime secrets

Environment variables are the most common way to pass secrets at runtime, and they are fine for many setups, but know where they leak: docker inspect shows them to anyone with Docker access, they are inherited by every child process, and they tend to end up in crash reports and debug logs that dump the environment. File-based secrets are safer, because the app reads them explicitly and they do not travel with the process.

# compose.yaml
services:
  app:
    image: myapp
    secrets:
      - db_password
    environment:
      DB_PASSWORD_FILE: /run/secrets/db_password   # app reads the file, not the value

secrets:
  db_password:
    file: ./secrets/db_password.txt

With plain Docker Compose, file: secrets are bind-mounted files — not encrypted, just kept out of the environment and out of the image. Encrypted, distributed secrets need Swarm (external: true), Kubernetes Secrets (with encryption at rest enabled), or an external manager such as Vault or AWS Secrets Manager, which also give you rotation and an audit trail.


Runtime Security

Never mount the Docker socket (unless you mean it)

# Equivalent to giving the container root on the host
volumes:
  - /var/run/docker.sock:/var/run/docker.sock

Anything that can talk to the Docker socket can start a new container with --privileged and the host’s root filesystem mounted, which is root on the host in one command. It shows up often in CI runners, monitoring agents and “container management” UIs. If a tool genuinely needs it, treat that container as part of your trusted host, or use a socket proxy that allows only the specific API calls it needs. For the same reason, adding a user to the docker group is effectively giving them root.

Likewise, --privileged disables most isolation at once: all capabilities, access to host devices, no seccomp or AppArmor profile. It is almost never the right fix for a permission error.

Drop capabilities

services:
  app:
    image: myapp
    cap_drop:
      - ALL
    security_opt:
      - no-new-privileges:true

Docker already drops the most dangerous capabilities (SYS_ADMIN, SYS_PTRACE, NET_ADMIN, SYS_MODULE) by default. What it leaves in the default set still includes CHOWN, DAC_OVERRIDE, SETUID, SETGID and NET_RAW (which allows crafting raw packets, useful for spoofing on the container network). A typical web application needs none of them, so cap_drop: ALL usually costs nothing. If something breaks, add back the single capability it needs rather than removing the line.

You also rarely need NET_BIND_SERVICE anymore: since Docker 20.10, containers get net.ipv4.ip_unprivileged_port_start=0 by default, so a non-root process can bind port 80 inside the container. Simpler still, listen on 8080 and map it.

no-new-privileges stops a process from gaining privileges through setuid binaries (sudo, su, ping on some systems), which closes a common escalation path from a non-root user.

Read-only root filesystem

services:
  app:
    image: myapp
    read_only: true
    tmpfs:
      - /tmp
    volumes:
      - app-data:/data     # the one place the app persists data

A read-only root filesystem means an attacker cannot drop tools, modify binaries, or persist changes inside the container. The practical work is finding every path your app writes to. The first run usually fails with EROFS: read-only file system on some cache or temp directory you did not know about — nginx, for example, writes to /var/cache/nginx and /var/run. Run the container once with read_only: true, fix each write location with a tmpfs or volume, and you have an accurate map of what the app writes.

Resource limits

services:
  app:
    image: myapp
    deploy:
      resources:
        limits:
          cpus: '1.0'
          memory: 512M
    pids_limit: 200        # caps the number of processes: stops fork bombs

Without a memory limit, one leaking container can push the whole host into the OOM killer, which may kill a different, healthy process. For fork bombs, use pids_limit; the nproc ulimit sometimes recommended for this is counted per UID across the whole host, so containers running as the same UID share it and it does not isolate them.


Network Security

Published ports bypass ufw

This is the most common way I have seen a database exposed to the internet on a server that “had a firewall”:

ufw default deny incoming
ufw allow 22
docker run -d -p 5432:5432 postgres    # reachable from the internet anyway

Docker implements -p by inserting its own iptables rules (in the DOCKER and DOCKER-USER chains) that are evaluated before the rules ufw manages. The host looks locked down in ufw status, and the port is open. Automated scanners find exposed Postgres, Redis and Elasticsearch instances within hours.

The fixes, from simplest:

services:
  db:
    image: postgres:17
    # no 'ports:' at all — other services reach it over the Compose network as db:5432

  admin:
    image: adminer
    ports:
      - "127.0.0.1:8080:8080"    # reachable only from the host, e.g. via SSH tunnel

If you need filtering on published ports, rules go in the DOCKER-USER chain, which Docker leaves for administrators and evaluates first.

Isolate services with networks

services:
  nginx:
    image: nginx
    ports: ["80:80", "443:443"]
    networks: [frontend]
  app:
    image: myapp
    networks: [frontend, backend]
  db:
    image: postgres:17
    networks: [backend]

networks:
  frontend:
  backend:
    internal: true     # no route to the outside world

Only nginx is published. The app can reach both nginx and the database, and the database cannot reach, or be reached from, the internet. internal: true also blocks outbound traffic, which limits what a compromised database container can exfiltrate — but it means the database cannot reach external services either, so check that nothing on that network needs to call out.


Kubernetes Security Context

apiVersion: apps/v1
kind: Deployment
metadata:
  name: app
spec:
  selector:
    matchLabels: { app: app }
  template:
    metadata:
      labels: { app: app }
    spec:
      automountServiceAccountToken: false   # unless the app calls the Kubernetes API
      securityContext:
        runAsNonRoot: true
        runAsUser: 1000
        runAsGroup: 1000
        fsGroup: 1000
        seccompProfile:
          type: RuntimeDefault
      containers:
        - name: app
          image: myapp:1.0.0
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities:
              drop: ["ALL"]
          resources:
            requests: { cpu: "100m", memory: "128Mi" }
            limits: { memory: "512Mi" }
          volumeMounts:
            - name: tmp
              mountPath: /tmp
      volumes:
        - name: tmp
          emptyDir: {}

This matches the Docker settings above, plus two Kubernetes-specific points:

  • automountServiceAccountToken: false. By default every pod gets a token for its service account mounted at /var/run/secrets/kubernetes.io/serviceaccount. If the app does not call the Kubernetes API, that token is only useful to an attacker.
  • seccompProfile: RuntimeDefault. Plain Docker applies its default seccomp profile automatically; Kubernetes historically ran pods unconfined unless you asked. Setting it explicitly blocks a few dozen rarely used syscalls that show up in kernel exploits.

Rather than reviewing every manifest by hand, enable Pod Security Admission with the restricted level on application namespaces (pod-security.kubernetes.io/enforce: restricted label). Pods that request root, privilege escalation or extra capabilities are then rejected at admission time.


.dockerignore

.git
node_modules
.env
.env.*
*.pem
secrets/
coverage
*.log
.github

.dockerignore controls what is sent as the build context. Without it, COPY . . sends everything, including .git (full history, which may contain secrets committed and later removed) and local .env files. It is also a performance fix: a large node_modules or .git makes every build upload hundreds of megabytes to the builder.


Layer by layer: what to do and the mistake to look for

LayerDoCommon mistake
ImageNumeric non-root USER, root-owned app files, slim or distroless baseUSER appuser rejected by runAsNonRoot; app user owns its own code
BuildMulti-stage, prune dev deps, --mount=type=secretTokens in ARG/ENV, visible in docker history
DependenciesTrivy in CI with --ignore-unfixed, automated digest updatesPinned digests nobody updates; actions pinned to @master
SecretsFile-based secrets or a secrets managerSecrets in env vars dumped by crash reports
Runtimecap_drop: ALL, no-new-privileges, read-only root, memory and pids limitsMounting docker.sock; --privileged to fix a permission error
NetworkPublish only the proxy; bind admin ports to 127.0.0.1Trusting ufw to block -p ports
KubernetessecurityContext, no automounted token, Pod Security restrictedDefault service account token in every pod

None of these requires changing application code, and most are a line or two of configuration. They are also far easier to add at the start of a project than to retrofit onto containers that already depend on running as root or writing all over their filesystem.