Docker & Kubernetes Beginners Guide

Key takeaways

Docker packages your app and all its dependencies into a portable container. Kubernetes orchestrates those containers across many machines. Together they are the backbone of modern cloud deployment.

Why Containers?

The classic developer problem:

"It works on my machine!"
→ Push to staging → crashes
→ Different OS, different Python version, missing library

Docker solves this by packaging your app and all its dependencies into a container — a portable, isolated unit that runs identically everywhere.

Traditional:
App → OS dependency hell → "works on my machine"

With Docker:
App + Dependencies → Image → Container (same everywhere)

What Containers Actually Change

Containers have become the default packaging format for server software, and managed Kubernetes is offered by every major cloud. What they change day to day:

For developers, Docker means:

  • A new developer runs docker compose up instead of following a page of install instructions
  • Deploy the same image from laptop → staging → production
  • Run the full stack (database, cache, queue) locally without installing each one

For operations, Kubernetes means:

  • Rolling updates that stop when new Pods fail their health checks, and one-command rollbacks
  • Scaling by changing a replica count (or letting an autoscaler do it), provided the app is stateless
  • Self-healing — crashed containers restart and Pods on a failed node are rescheduled

“Runs identically everywhere” has limits worth knowing from the start. A container shares the host’s kernel, so an image built for linux/amd64 will not run natively on an ARM machine (Apple Silicon Macs and AWS Graviton hosts) without emulation or a multi-architecture build — the error is exec format error. And the image fixes your dependencies, not your configuration: environment variables, mounted files, and network access still differ between laptop and production, which is where most remaining “works on my machine” bugs live.


Docker Concepts

Image     → Blueprint (read-only, like a class)
Container → Running instance of an image (like an object)
Registry  → Storage for images (Docker Hub, ECR, GCR)
Dockerfile → Instructions to build an image

VM vs Container

Virtual Machine:          Docker Container:
┌─────────────────┐       ┌─────────────────┐
│   App A │ App B │       │   App A │ App B │
├─────────┼───────┤       ├─────────┼───────┤
│  OS A   │ OS B  │       │   Docker Engine  │ ← shared kernel
├─────────────────┤       ├─────────────────┤
│   Hypervisor    │       │   Host OS        │
├─────────────────┤       ├─────────────────┤
│  Physical Server│       │  Physical Server │
└─────────────────┘       └─────────────────┘
GBs, minutes to start     MBs, milliseconds to start

The diagram simplifies one thing: on macOS and Windows there is no Linux kernel to share, so Docker Desktop runs a small Linux VM and your containers run inside it. That is why file-system-heavy work in bind-mounted folders (for example npm install into a mounted node_modules) can be much slower on a Mac than on Linux. The isolation is also weaker than a VM’s: containers are separated by kernel namespaces and cgroups, so a kernel vulnerability affects every container on the host. For running untrusted code, that difference matters; for packaging your own services, it rarely does.


Installation

Docker Desktop (macOS / Windows)

Download from docker.com/get-started. Includes Docker Engine, Docker CLI, and Docker Compose.

Linux

# Ubuntu / Debian
curl -fsSL https://get.docker.com | sh
sudo usermod -aG docker $USER
# Log out and back in

Adding yourself to the docker group avoids typing sudo for every command, and until you log out and back in you will keep seeing permission denied while trying to connect to the Docker daemon socket. Be aware of what it grants: anyone who can talk to the Docker daemon can start a container that mounts / from the host, so membership in the docker group is effectively root access. On shared machines, rootless Docker or Podman avoids that.

Verify

docker --version      # Docker version 26.x.x
docker compose version # Docker Compose version v2.x.x

Docker Basics

Your First Container

# Pull and run Nginx
docker run -d -p 8080:80 --name my-nginx nginx
# → open http://localhost:8080

# List running containers
docker ps

# View logs
docker logs my-nginx

# Stop and remove
docker stop my-nginx
docker rm my-nginx

-p 8080:80 reads as host:container: port 8080 on your machine forwards to port 80 inside the container, where nginx listens. Getting the order backwards is the most common first mistake, and the symptom is simply “connection refused” on the port you expected. -d runs the container in the background, and --name gives it a fixed name so later commands don’t need the random ID; running the same command twice fails with Conflict. The container name "/my-nginx" is already in use, because stopped containers still exist until removed (docker run --rm removes them automatically on exit).

Essential Commands

# Images
docker pull nginx:alpine          # Download image
docker images                     # List local images
docker rmi nginx:alpine           # Remove image

# Containers
docker run -d -p 3000:3000 myapp  # Run detached
docker run -it ubuntu bash        # Interactive terminal
docker ps                         # List running containers
docker ps -a                      # List all (including stopped)
docker stop <id>                  # Stop gracefully
docker rm <id>                    # Remove container
docker rm -f <id>                 # Force remove running container

# Debugging
docker logs <id>                  # View logs
docker logs -f <id>               # Follow logs (live)
docker exec -it <id> bash         # Shell into running container
docker inspect <id>               # Full container details

docker exec -it <id> bash fails on many small images, including alpine and distroless variants, with exec: "bash": executable file not found in $PATH — use sh instead, or nothing at all for distroless images (debug them with docker debug or an ephemeral container). docker stop sends SIGTERM and, if the process has not exited after 10 seconds, SIGKILL; if your containers always take exactly ten seconds to stop, the app is not handling SIGTERM, usually because it runs under a shell that doesn’t forward signals (see the note on CMD below).


Writing a Dockerfile

Python / FastAPI Example

# Dockerfile
FROM python:3.12-slim

# Set working directory
WORKDIR /app

# Install dependencies first (layer cache optimization)
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# Copy application code
COPY . .

# Expose port
EXPOSE 8000

# Run the application
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]

The order of instructions is the whole trick. Docker caches each layer and reuses it until something that layer depends on changes. Copying only requirements.txt first means the slow pip install layer is rebuilt only when dependencies change; your code changes invalidate just the final COPY . .. Put COPY . . first and every one-line code edit reinstalls all dependencies.

Two details in the last line are easy to get wrong. --host 0.0.0.0 is required: uvicorn’s default 127.0.0.1 listens only on the container’s own loopback interface, so the port mapping reaches nothing and curl reports Empty reply from server or connection reset by peer while the logs look perfectly healthy. And CMD is written in exec form (a JSON array). The shell form, CMD uvicorn main:app ..., runs the app as a child of /bin/sh, which becomes PID 1 and does not forward SIGTERM — so docker stop and Kubernetes rollouts wait for the kill timeout instead of shutting down gracefully. EXPOSE is documentation only; it does not publish anything by itself.

# Build the image
docker build -t my-fastapi-app .

# Run it
docker run -d -p 8000:8000 --name api my-fastapi-app

# Test
curl http://localhost:8000/

Node.js Example

FROM node:20-alpine

WORKDIR /app

# Install dependencies
COPY package*.json ./
RUN npm ci --omit=dev

# Copy source
COPY . .

EXPOSE 3000

CMD ["node", "server.js"]

Multi-stage Build (Smaller Images)

# Stage 1: Build
FROM node:20-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build
RUN npm prune --omit=dev   # drop devDependencies before copying node_modules

# Stage 2: Production (no dev dependencies, no source)
FROM node:20-alpine AS production
WORKDIR /app
COPY --from=builder /app/dist ./dist
COPY --from=builder /app/node_modules ./node_modules
EXPOSE 3000
CMD ["node", "dist/server.js"]

The production image only contains the compiled output — much smaller than including all source files.

The first stage has everything needed to build: TypeScript, bundlers, test tools. The second stage starts from a fresh base image and copies in only what is needed to run. The npm prune --omit=dev line matters: without it, the node_modules copied from the builder still contains every devDependency, and the “smaller image” is mostly an illusion. (In the single-stage Node example, --omit=dev replaces the older --only=production flag, which npm now warns about.) Alpine-based images are small because they use musl instead of glibc; most pure-JavaScript apps don’t care, but packages with native binaries sometimes fail with errors such as Error relocating ... symbol not found, in which case a -slim Debian-based image is the pragmatic choice. Finally, both Node images run as root by default; adding USER node before CMD is a cheap hardening step.


.dockerignore

node_modules/
.git/
.env
*.log
dist/
__pycache__/
.pytest_cache/

Exclude these to keep your build context small and fast.

The build context is everything in the directory you pass to docker build, and it is sent to the builder before the first instruction runs. Without a .dockerignore, COPY . . also copies your local node_modules (built for your OS, possibly overwriting the Linux ones installed in the image), the entire .git history, and — worst — .env files with real credentials, which then live in an image layer anyone with pull access can extract. Deleting the file in a later layer does not remove it from the earlier one.


Docker Compose — Multi-Container Apps

Docker Compose runs multiple containers together as a single service.

Web App + Database + Cache

# docker-compose.yml (the old top-level "version" key is obsolete in Compose v2)
services:
  api:
    build: .
    ports:
      - "8000:8000"
    environment:
      - DATABASE_URL=postgresql://user:pass@db:5432/mydb
      - REDIS_URL=redis://redis:6379
    depends_on:
      db:
        condition: service_healthy
      redis:
        condition: service_started
    restart: unless-stopped

  db:
    image: postgres:16-alpine
    environment:
      POSTGRES_USER: user
      POSTGRES_PASSWORD: pass
      POSTGRES_DB: mydb
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U user -d mydb"]
      interval: 10s
      timeout: 5s
      retries: 5

  redis:
    image: redis:7-alpine
    volumes:
      - redis_data:/data

volumes:
  postgres_data:
  redis_data:
# Start all services
docker compose up -d

# View logs for all services
docker compose logs -f

# View logs for one service
docker compose logs -f api

# Stop all services
docker compose down

# Stop and remove volumes (⚠️ deletes data)
docker compose down -v

# Rebuild after code changes
docker compose up -d --build

Compose creates a private network for the project and registers each service under its name, which is why DATABASE_URL points at host db rather than localhost. Inside the api container, localhost is the container itself; connecting to localhost:5432 there fails with ECONNREFUSED even though the database is running. The depends_on with condition: service_healthy waits for the Postgres healthcheck to pass, not just for the container to start — without it, the API often starts first and crashes on its initial connection. Even so, the app should retry database connections, because Compose only orders startup and does nothing if the database restarts later.

Named volumes (postgres_data) persist data across docker compose down and up; down -v deletes them. The credentials in this file are fine for local development, but for anything shared, move them into an .env file (excluded from Git) or Docker secrets.


Kubernetes Basics

Kubernetes (K8s) orchestrates containers across a cluster of machines.

You describe desired state → K8s makes it happen and keeps it that way

"I want 3 replicas of my API, always"
→ K8s runs 3 pods
→ If one crashes → K8s starts a new one automatically
→ If traffic spikes → K8s scales up (with an autoscaler configured)

This “desired state” model is the key idea. You never tell Kubernetes “start a container”; you submit an object describing what should exist, and controllers continuously compare that description with reality and act on the difference. That is why deleting a Pod managed by a Deployment just produces a new Pod a moment later, and why kubectl apply of the same file twice is harmless. Scaling on traffic is not automatic out of the box, though: it requires a HorizontalPodAutoscaler, a metrics source, and resource requests on the containers.

Core Objects

ObjectWhat it does
PodSmallest deployable unit — one or more containers
DeploymentManages Pods — rolling updates, scaling, self-healing
ServiceStable network endpoint to reach Pods (load balancing)
IngressRoutes external HTTP traffic to Services
ConfigMapStore non-secret config (env vars, config files)
SecretStore sensitive data (passwords, API keys)

Local Setup (for learning)

# Install kubectl
brew install kubectl  # macOS
# Windows: winget install Kubernetes.kubectl

# minikube (local single-node cluster)
brew install minikube
minikube start

Pod & Deployment

Pod (basic unit)

# pod.yaml
apiVersion: v1
kind: Pod
metadata:
  name: my-api-pod
  labels:
    app: my-api
spec:
  containers:
    - name: api
      image: my-fastapi-app:latest
      ports:
        - containerPort: 8000
      env:
        - name: DATABASE_URL
          valueFrom:
            secretKeyRef:
              name: db-secret
              key: DATABASE_URL

Pods are ephemeral — don’t use them directly. Use Deployments.

A bare Pod is not recreated if its node fails or it is deleted, and it cannot be updated in place except for a few fields. The secretKeyRef must name a key that actually exists in the Secret (here DATABASE_URL, matching the Secret defined later); a mismatch doesn’t fail at apply time but leaves the Pod in CreateContainerConfigError with an event like couldn't find key url in Secret default/db-secret. kubectl describe pod shows those events and is the first command to run when a Pod will not start.

# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-api
spec:
  replicas: 3                          # Run 3 copies
  selector:
    matchLabels:
      app: my-api
  template:
    metadata:
      labels:
        app: my-api
    spec:
      containers:
        - name: api
          image: myregistry/my-api:1.0.0
          ports:
            - containerPort: 8000
          resources:
            requests:
              memory: "128Mi"
              cpu: "250m"
            limits:
              memory: "256Mi"
              cpu: "500m"
          readinessProbe:
            httpGet:
              path: /health
              port: 8000
            initialDelaySeconds: 5
            periodSeconds: 10
kubectl apply -f deployment.yaml

# Check status
kubectl get deployments
kubectl get pods
kubectl describe deployment my-api

# Scale up/down
kubectl scale deployment my-api --replicas=5

# Rolling update (zero downtime)
kubectl set image deployment/my-api api=myregistry/my-api:1.1.0

# Rollback
kubectl rollout undo deployment/my-api

The selector.matchLabels must match the labels in template.metadata.labels; that is how the Deployment knows which Pods it owns, and a mismatch is rejected with selector does not match template labels. Resources deserve more attention than beginners usually give them. requests are what the scheduler reserves when placing the Pod on a node; limits are enforced at runtime, and the two behave differently. Exceeding the CPU limit only throttles the container, but exceeding the memory limit kills it — kubectl get pods shows OOMKilled and a growing restart count, and a service that is slow to start can end up in CrashLoopBackOff. Set the memory limit from observed usage plus headroom, not from a guess.

The readiness probe is what makes the “zero downtime” comment true. During a rolling update, Kubernetes sends traffic to a new Pod only after its readiness probe succeeds, and removes old Pods only as new ones become ready. Without a probe, a Pod counts as ready as soon as the process starts, so requests reach an app that is still loading and fail. The probe needs a real /health endpoint in your app. A liveness probe is a different thing — it restarts the container when it fails — and making it too aggressive (short timeouts, checking the database) is a classic way to turn a slow dependency into a restart storm.


Service & Ingress

Service (internal load balancer)

# service.yaml
apiVersion: v1
kind: Service
metadata:
  name: my-api-service
spec:
  selector:
    app: my-api          # Routes to Pods with this label
  ports:
    - port: 80
      targetPort: 8000
  type: ClusterIP        # Internal only (default)
  # type: LoadBalancer   # External IP (cloud providers)
  # type: NodePort       # Port on each node

A Service gives a stable virtual IP and DNS name (my-api-service within the namespace, my-api-service.default.svc.cluster.local in full) in front of a changing set of Pods selected by label. port is what clients connect to; targetPort is the container port. If requests to the Service hang or are refused, check kubectl get endpoints my-api-service: an empty list means the selector matches no ready Pods — usually a label typo or failing readiness probes.

Ingress (external HTTP routing)

# ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: my-ingress
  annotations:
    nginx.ingress.kubernetes.io/rewrite-target: /
spec:
  rules:
    - host: api.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: my-api-service
                port:
                  number: 80

An Ingress object does nothing by itself: it is configuration for an ingress controller (ingress-nginx, Traefik, a cloud load balancer controller) that you must install separately — on minikube, minikube addons enable ingress. Clusters with more than one controller also need spec.ingressClassName to say which one should handle it. The rewrite-target: / annotation is specific to ingress-nginx and is harmless with a single / path, but copied into a config with path: /api, it rewrites every request to /, a frequent source of “every route returns the home page”. Note too that the Ingress API is feature-frozen; the newer Gateway API is where new routing features are being added.


Secrets & ConfigMaps

# secret.yaml ("data" values must be base64 encoded; "stringData" takes plain text)
apiVersion: v1
kind: Secret
metadata:
  name: db-secret
type: Opaque
stringData:                        # kubectl auto-encodes
  DATABASE_URL: "postgresql://user:pass@db:5432/mydb"
  API_KEY: "sk-..."

---
# configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: app-config
data:
  LOG_LEVEL: "info"
  MAX_CONNECTIONS: "100"
kubectl apply -f secret.yaml
kubectl apply -f configmap.yaml

# Use in a Deployment
# env:
#   - name: DATABASE_URL
#     valueFrom:
#       secretKeyRef:
#         name: db-secret
#         key: DATABASE_URL

Base64 is an encoding, not encryption: anyone who can read the Secret object can decode it with base64 -d. Kubernetes Secrets are protected by access control (RBAC) and, if the cluster is configured for it, encryption at rest in etcd — so restrict who can get secrets, and never commit a manifest like this with real values to Git. Teams usually generate Secrets from an external store (a cloud secret manager via External Secrets Operator) or commit only encrypted versions (Sealed Secrets, SOPS). Also note that Pods read environment variables only at startup; changing a Secret or ConfigMap does not update env vars in running Pods until they restart (kubectl rollout restart deployment/my-api).


Essential kubectl Commands

# Context (cluster) management
kubectl config get-contexts
kubectl config use-context my-cluster

# Resources
kubectl get pods
kubectl get pods -n my-namespace       # Specific namespace
kubectl get all                        # All resource types
kubectl describe pod <name>            # Detailed info + events

# Debugging
kubectl logs <pod-name>
kubectl logs -f <pod-name>             # Follow
kubectl exec -it <pod-name> -- bash    # Shell into pod

# Apply / delete
kubectl apply -f manifest.yaml
kubectl delete -f manifest.yaml
kubectl delete pod <name> --force

# Port forwarding (local testing)
kubectl port-forward pod/<name> 8080:8000
kubectl port-forward service/<name> 8080:80

When something is broken, the order that finds the cause fastest is almost always: kubectl get pods (what state?), kubectl describe pod <name> (the Events section at the bottom explains scheduling failures, image pull errors, probe failures, and OOM kills), then kubectl logs <name> — and kubectl logs <name> --previous for a container that already crashed and restarted, since plain logs shows the new, often empty, instance. Be careful with kubectl delete pod --force: it removes the object from the API immediately without waiting for the node to confirm the container stopped, which for stateful workloads can leave two copies running briefly. For a deeper walkthrough of Pod states, see Kubernetes Pod troubleshooting.


When to Use What

ScenarioTool
Local developmentDocker + docker compose
Single-server deploymentDocker + docker compose
Multi-server, auto-scalingKubernetes
Managed cloud (AWS/GCP/Azure)EKS / GKE / AKS
Simple side projectDocker alone

Do you need Kubernetes yet?

  1. Start with Docker — containerize your app, run it locally with Docker Compose
  2. Graduate to Kubernetes when you need horizontal scaling, rolling updates, or self-healing across multiple nodes
  3. Use managed Kubernetes (GKE, EKS, AKS) in production — avoid managing the control plane yourself

The step people most often skip is the middle one: deciding whether they need Kubernetes at all. A single VM running Docker Compose, or a managed container platform (Cloud Run, App Runner, Fly.io, Azure Container Apps), handles a surprising amount of traffic with far less to operate. Kubernetes pays for itself when you have many services, need a common deployment model across teams, or need its scheduling and self-healing across many nodes — and its cost is real: networking, ingress, certificates, secrets, monitoring, and upgrades all become your problem, even on a managed service.


Frequently Asked Questions (FAQ)

Q. I built my image locally, but the Pod is stuck in ErrImagePull or ImagePullBackOff. Why?

A. Cluster nodes pull images from a registry; they cannot see an image that exists only in your local Docker daemon. With a :latest tag, as in the Pod example, the default imagePullPolicy is Always, so Kubernetes tries to download the image even if a copy happens to be present. Push a versioned tag to a registry the cluster can reach, as the Deployment example does with myregistry/my-api:1.0.0, or on a local cluster load the image directly with kind load docker-image or minikube image load.