Docker & Kubernetes Beginners Guide
Key takeaways
Docker packages your app and all its dependencies into a portable container. Kubernetes orchestrates those containers across many machines. Together they are the backbone of modern cloud deployment.
Why Containers?
The classic developer problem:
"It works on my machine!"
→ Push to staging → crashes
→ Different OS, different Python version, missing library
Docker solves this by packaging your app and all its dependencies into a container — a portable, isolated unit that runs identically everywhere.
Traditional:
App → OS dependency hell → "works on my machine"
With Docker:
App + Dependencies → Image → Container (same everywhere)
What Containers Actually Change
Containers have become the default packaging format for server software, and managed Kubernetes is offered by every major cloud. What they change day to day:
For developers, Docker means:
- A new developer runs
docker compose upinstead of following a page of install instructions - Deploy the same image from laptop → staging → production
- Run the full stack (database, cache, queue) locally without installing each one
For operations, Kubernetes means:
- Rolling updates that stop when new Pods fail their health checks, and one-command rollbacks
- Scaling by changing a replica count (or letting an autoscaler do it), provided the app is stateless
- Self-healing — crashed containers restart and Pods on a failed node are rescheduled
“Runs identically everywhere” has limits worth knowing from the start. A container shares the host’s kernel, so an image built for linux/amd64 will not run natively on an ARM machine (Apple Silicon Macs and AWS Graviton hosts) without emulation or a multi-architecture build — the error is exec format error. And the image fixes your dependencies, not your configuration: environment variables, mounted files, and network access still differ between laptop and production, which is where most remaining “works on my machine” bugs live.
Docker Concepts
Image → Blueprint (read-only, like a class)
Container → Running instance of an image (like an object)
Registry → Storage for images (Docker Hub, ECR, GCR)
Dockerfile → Instructions to build an image
VM vs Container
Virtual Machine: Docker Container:
┌─────────────────┐ ┌─────────────────┐
│ App A │ App B │ │ App A │ App B │
├─────────┼───────┤ ├─────────┼───────┤
│ OS A │ OS B │ │ Docker Engine │ ← shared kernel
├─────────────────┤ ├─────────────────┤
│ Hypervisor │ │ Host OS │
├─────────────────┤ ├─────────────────┤
│ Physical Server│ │ Physical Server │
└─────────────────┘ └─────────────────┘
GBs, minutes to start MBs, milliseconds to start
The diagram simplifies one thing: on macOS and Windows there is no Linux kernel to share, so Docker Desktop runs a small Linux VM and your containers run inside it. That is why file-system-heavy work in bind-mounted folders (for example npm install into a mounted node_modules) can be much slower on a Mac than on Linux. The isolation is also weaker than a VM’s: containers are separated by kernel namespaces and cgroups, so a kernel vulnerability affects every container on the host. For running untrusted code, that difference matters; for packaging your own services, it rarely does.
Installation
Docker Desktop (macOS / Windows)
Download from docker.com/get-started. Includes Docker Engine, Docker CLI, and Docker Compose.
Linux
# Ubuntu / Debian
curl -fsSL https://get.docker.com | sh
sudo usermod -aG docker $USER
# Log out and back in
Adding yourself to the docker group avoids typing sudo for every command, and until you log out and back in you will keep seeing permission denied while trying to connect to the Docker daemon socket. Be aware of what it grants: anyone who can talk to the Docker daemon can start a container that mounts / from the host, so membership in the docker group is effectively root access. On shared machines, rootless Docker or Podman avoids that.
Verify
docker --version # Docker version 26.x.x
docker compose version # Docker Compose version v2.x.x
Docker Basics
Your First Container
# Pull and run Nginx
docker run -d -p 8080:80 --name my-nginx nginx
# → open http://localhost:8080
# List running containers
docker ps
# View logs
docker logs my-nginx
# Stop and remove
docker stop my-nginx
docker rm my-nginx
-p 8080:80 reads as host:container: port 8080 on your machine forwards to port 80 inside the container, where nginx listens. Getting the order backwards is the most common first mistake, and the symptom is simply “connection refused” on the port you expected. -d runs the container in the background, and --name gives it a fixed name so later commands don’t need the random ID; running the same command twice fails with Conflict. The container name "/my-nginx" is already in use, because stopped containers still exist until removed (docker run --rm removes them automatically on exit).
Essential Commands
# Images
docker pull nginx:alpine # Download image
docker images # List local images
docker rmi nginx:alpine # Remove image
# Containers
docker run -d -p 3000:3000 myapp # Run detached
docker run -it ubuntu bash # Interactive terminal
docker ps # List running containers
docker ps -a # List all (including stopped)
docker stop <id> # Stop gracefully
docker rm <id> # Remove container
docker rm -f <id> # Force remove running container
# Debugging
docker logs <id> # View logs
docker logs -f <id> # Follow logs (live)
docker exec -it <id> bash # Shell into running container
docker inspect <id> # Full container details
docker exec -it <id> bash fails on many small images, including alpine and distroless variants, with exec: "bash": executable file not found in $PATH — use sh instead, or nothing at all for distroless images (debug them with docker debug or an ephemeral container). docker stop sends SIGTERM and, if the process has not exited after 10 seconds, SIGKILL; if your containers always take exactly ten seconds to stop, the app is not handling SIGTERM, usually because it runs under a shell that doesn’t forward signals (see the note on CMD below).
Writing a Dockerfile
Python / FastAPI Example
# Dockerfile
FROM python:3.12-slim
# Set working directory
WORKDIR /app
# Install dependencies first (layer cache optimization)
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Copy application code
COPY . .
# Expose port
EXPOSE 8000
# Run the application
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
The order of instructions is the whole trick. Docker caches each layer and reuses it until something that layer depends on changes. Copying only requirements.txt first means the slow pip install layer is rebuilt only when dependencies change; your code changes invalidate just the final COPY . .. Put COPY . . first and every one-line code edit reinstalls all dependencies.
Two details in the last line are easy to get wrong. --host 0.0.0.0 is required: uvicorn’s default 127.0.0.1 listens only on the container’s own loopback interface, so the port mapping reaches nothing and curl reports Empty reply from server or connection reset by peer while the logs look perfectly healthy. And CMD is written in exec form (a JSON array). The shell form, CMD uvicorn main:app ..., runs the app as a child of /bin/sh, which becomes PID 1 and does not forward SIGTERM — so docker stop and Kubernetes rollouts wait for the kill timeout instead of shutting down gracefully. EXPOSE is documentation only; it does not publish anything by itself.
# Build the image
docker build -t my-fastapi-app .
# Run it
docker run -d -p 8000:8000 --name api my-fastapi-app
# Test
curl http://localhost:8000/
Node.js Example
FROM node:20-alpine
WORKDIR /app
# Install dependencies
COPY package*.json ./
RUN npm ci --omit=dev
# Copy source
COPY . .
EXPOSE 3000
CMD ["node", "server.js"]
Multi-stage Build (Smaller Images)
# Stage 1: Build
FROM node:20-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build
RUN npm prune --omit=dev # drop devDependencies before copying node_modules
# Stage 2: Production (no dev dependencies, no source)
FROM node:20-alpine AS production
WORKDIR /app
COPY --from=builder /app/dist ./dist
COPY --from=builder /app/node_modules ./node_modules
EXPOSE 3000
CMD ["node", "dist/server.js"]
The production image only contains the compiled output — much smaller than including all source files.
The first stage has everything needed to build: TypeScript, bundlers, test tools. The second stage starts from a fresh base image and copies in only what is needed to run. The npm prune --omit=dev line matters: without it, the node_modules copied from the builder still contains every devDependency, and the “smaller image” is mostly an illusion. (In the single-stage Node example, --omit=dev replaces the older --only=production flag, which npm now warns about.) Alpine-based images are small because they use musl instead of glibc; most pure-JavaScript apps don’t care, but packages with native binaries sometimes fail with errors such as Error relocating ... symbol not found, in which case a -slim Debian-based image is the pragmatic choice. Finally, both Node images run as root by default; adding USER node before CMD is a cheap hardening step.
.dockerignore
node_modules/
.git/
.env
*.log
dist/
__pycache__/
.pytest_cache/
Exclude these to keep your build context small and fast.
The build context is everything in the directory you pass to docker build, and it is sent to the builder before the first instruction runs. Without a .dockerignore, COPY . . also copies your local node_modules (built for your OS, possibly overwriting the Linux ones installed in the image), the entire .git history, and — worst — .env files with real credentials, which then live in an image layer anyone with pull access can extract. Deleting the file in a later layer does not remove it from the earlier one.
Docker Compose — Multi-Container Apps
Docker Compose runs multiple containers together as a single service.
Web App + Database + Cache
# docker-compose.yml (the old top-level "version" key is obsolete in Compose v2)
services:
api:
build: .
ports:
- "8000:8000"
environment:
- DATABASE_URL=postgresql://user:pass@db:5432/mydb
- REDIS_URL=redis://redis:6379
depends_on:
db:
condition: service_healthy
redis:
condition: service_started
restart: unless-stopped
db:
image: postgres:16-alpine
environment:
POSTGRES_USER: user
POSTGRES_PASSWORD: pass
POSTGRES_DB: mydb
volumes:
- postgres_data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U user -d mydb"]
interval: 10s
timeout: 5s
retries: 5
redis:
image: redis:7-alpine
volumes:
- redis_data:/data
volumes:
postgres_data:
redis_data:
# Start all services
docker compose up -d
# View logs for all services
docker compose logs -f
# View logs for one service
docker compose logs -f api
# Stop all services
docker compose down
# Stop and remove volumes (⚠️ deletes data)
docker compose down -v
# Rebuild after code changes
docker compose up -d --build
Compose creates a private network for the project and registers each service under its name, which is why DATABASE_URL points at host db rather than localhost. Inside the api container, localhost is the container itself; connecting to localhost:5432 there fails with ECONNREFUSED even though the database is running. The depends_on with condition: service_healthy waits for the Postgres healthcheck to pass, not just for the container to start — without it, the API often starts first and crashes on its initial connection. Even so, the app should retry database connections, because Compose only orders startup and does nothing if the database restarts later.
Named volumes (postgres_data) persist data across docker compose down and up; down -v deletes them. The credentials in this file are fine for local development, but for anything shared, move them into an .env file (excluded from Git) or Docker secrets.
Kubernetes Basics
Kubernetes (K8s) orchestrates containers across a cluster of machines.
You describe desired state → K8s makes it happen and keeps it that way
"I want 3 replicas of my API, always"
→ K8s runs 3 pods
→ If one crashes → K8s starts a new one automatically
→ If traffic spikes → K8s scales up (with an autoscaler configured)
This “desired state” model is the key idea. You never tell Kubernetes “start a container”; you submit an object describing what should exist, and controllers continuously compare that description with reality and act on the difference. That is why deleting a Pod managed by a Deployment just produces a new Pod a moment later, and why kubectl apply of the same file twice is harmless. Scaling on traffic is not automatic out of the box, though: it requires a HorizontalPodAutoscaler, a metrics source, and resource requests on the containers.
Core Objects
| Object | What it does |
|---|---|
| Pod | Smallest deployable unit — one or more containers |
| Deployment | Manages Pods — rolling updates, scaling, self-healing |
| Service | Stable network endpoint to reach Pods (load balancing) |
| Ingress | Routes external HTTP traffic to Services |
| ConfigMap | Store non-secret config (env vars, config files) |
| Secret | Store sensitive data (passwords, API keys) |
Local Setup (for learning)
# Install kubectl
brew install kubectl # macOS
# Windows: winget install Kubernetes.kubectl
# minikube (local single-node cluster)
brew install minikube
minikube start
Pod & Deployment
Pod (basic unit)
# pod.yaml
apiVersion: v1
kind: Pod
metadata:
name: my-api-pod
labels:
app: my-api
spec:
containers:
- name: api
image: my-fastapi-app:latest
ports:
- containerPort: 8000
env:
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: db-secret
key: DATABASE_URL
Pods are ephemeral — don’t use them directly. Use Deployments.
A bare Pod is not recreated if its node fails or it is deleted, and it cannot be updated in place except for a few fields. The secretKeyRef must name a key that actually exists in the Secret (here DATABASE_URL, matching the Secret defined later); a mismatch doesn’t fail at apply time but leaves the Pod in CreateContainerConfigError with an event like couldn't find key url in Secret default/db-secret. kubectl describe pod shows those events and is the first command to run when a Pod will not start.
Deployment (recommended)
# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-api
spec:
replicas: 3 # Run 3 copies
selector:
matchLabels:
app: my-api
template:
metadata:
labels:
app: my-api
spec:
containers:
- name: api
image: myregistry/my-api:1.0.0
ports:
- containerPort: 8000
resources:
requests:
memory: "128Mi"
cpu: "250m"
limits:
memory: "256Mi"
cpu: "500m"
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 5
periodSeconds: 10
kubectl apply -f deployment.yaml
# Check status
kubectl get deployments
kubectl get pods
kubectl describe deployment my-api
# Scale up/down
kubectl scale deployment my-api --replicas=5
# Rolling update (zero downtime)
kubectl set image deployment/my-api api=myregistry/my-api:1.1.0
# Rollback
kubectl rollout undo deployment/my-api
The selector.matchLabels must match the labels in template.metadata.labels; that is how the Deployment knows which Pods it owns, and a mismatch is rejected with selector does not match template labels. Resources deserve more attention than beginners usually give them. requests are what the scheduler reserves when placing the Pod on a node; limits are enforced at runtime, and the two behave differently. Exceeding the CPU limit only throttles the container, but exceeding the memory limit kills it — kubectl get pods shows OOMKilled and a growing restart count, and a service that is slow to start can end up in CrashLoopBackOff. Set the memory limit from observed usage plus headroom, not from a guess.
The readiness probe is what makes the “zero downtime” comment true. During a rolling update, Kubernetes sends traffic to a new Pod only after its readiness probe succeeds, and removes old Pods only as new ones become ready. Without a probe, a Pod counts as ready as soon as the process starts, so requests reach an app that is still loading and fail. The probe needs a real /health endpoint in your app. A liveness probe is a different thing — it restarts the container when it fails — and making it too aggressive (short timeouts, checking the database) is a classic way to turn a slow dependency into a restart storm.
Service & Ingress
Service (internal load balancer)
# service.yaml
apiVersion: v1
kind: Service
metadata:
name: my-api-service
spec:
selector:
app: my-api # Routes to Pods with this label
ports:
- port: 80
targetPort: 8000
type: ClusterIP # Internal only (default)
# type: LoadBalancer # External IP (cloud providers)
# type: NodePort # Port on each node
A Service gives a stable virtual IP and DNS name (my-api-service within the namespace, my-api-service.default.svc.cluster.local in full) in front of a changing set of Pods selected by label. port is what clients connect to; targetPort is the container port. If requests to the Service hang or are refused, check kubectl get endpoints my-api-service: an empty list means the selector matches no ready Pods — usually a label typo or failing readiness probes.
Ingress (external HTTP routing)
# ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: my-ingress
annotations:
nginx.ingress.kubernetes.io/rewrite-target: /
spec:
rules:
- host: api.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: my-api-service
port:
number: 80
An Ingress object does nothing by itself: it is configuration for an ingress controller (ingress-nginx, Traefik, a cloud load balancer controller) that you must install separately — on minikube, minikube addons enable ingress. Clusters with more than one controller also need spec.ingressClassName to say which one should handle it. The rewrite-target: / annotation is specific to ingress-nginx and is harmless with a single / path, but copied into a config with path: /api, it rewrites every request to /, a frequent source of “every route returns the home page”. Note too that the Ingress API is feature-frozen; the newer Gateway API is where new routing features are being added.
Secrets & ConfigMaps
# secret.yaml ("data" values must be base64 encoded; "stringData" takes plain text)
apiVersion: v1
kind: Secret
metadata:
name: db-secret
type: Opaque
stringData: # kubectl auto-encodes
DATABASE_URL: "postgresql://user:pass@db:5432/mydb"
API_KEY: "sk-..."
---
# configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: app-config
data:
LOG_LEVEL: "info"
MAX_CONNECTIONS: "100"
kubectl apply -f secret.yaml
kubectl apply -f configmap.yaml
# Use in a Deployment
# env:
# - name: DATABASE_URL
# valueFrom:
# secretKeyRef:
# name: db-secret
# key: DATABASE_URL
Base64 is an encoding, not encryption: anyone who can read the Secret object can decode it with base64 -d. Kubernetes Secrets are protected by access control (RBAC) and, if the cluster is configured for it, encryption at rest in etcd — so restrict who can get secrets, and never commit a manifest like this with real values to Git. Teams usually generate Secrets from an external store (a cloud secret manager via External Secrets Operator) or commit only encrypted versions (Sealed Secrets, SOPS). Also note that Pods read environment variables only at startup; changing a Secret or ConfigMap does not update env vars in running Pods until they restart (kubectl rollout restart deployment/my-api).
Essential kubectl Commands
# Context (cluster) management
kubectl config get-contexts
kubectl config use-context my-cluster
# Resources
kubectl get pods
kubectl get pods -n my-namespace # Specific namespace
kubectl get all # All resource types
kubectl describe pod <name> # Detailed info + events
# Debugging
kubectl logs <pod-name>
kubectl logs -f <pod-name> # Follow
kubectl exec -it <pod-name> -- bash # Shell into pod
# Apply / delete
kubectl apply -f manifest.yaml
kubectl delete -f manifest.yaml
kubectl delete pod <name> --force
# Port forwarding (local testing)
kubectl port-forward pod/<name> 8080:8000
kubectl port-forward service/<name> 8080:80
When something is broken, the order that finds the cause fastest is almost always: kubectl get pods (what state?), kubectl describe pod <name> (the Events section at the bottom explains scheduling failures, image pull errors, probe failures, and OOM kills), then kubectl logs <name> — and kubectl logs <name> --previous for a container that already crashed and restarted, since plain logs shows the new, often empty, instance. Be careful with kubectl delete pod --force: it removes the object from the API immediately without waiting for the node to confirm the container stopped, which for stateful workloads can leave two copies running briefly. For a deeper walkthrough of Pod states, see Kubernetes Pod troubleshooting.
When to Use What
| Scenario | Tool |
|---|---|
| Local development | Docker + docker compose |
| Single-server deployment | Docker + docker compose |
| Multi-server, auto-scaling | Kubernetes |
| Managed cloud (AWS/GCP/Azure) | EKS / GKE / AKS |
| Simple side project | Docker alone |
Do you need Kubernetes yet?
- Start with Docker — containerize your app, run it locally with Docker Compose
- Graduate to Kubernetes when you need horizontal scaling, rolling updates, or self-healing across multiple nodes
- Use managed Kubernetes (GKE, EKS, AKS) in production — avoid managing the control plane yourself
The step people most often skip is the middle one: deciding whether they need Kubernetes at all. A single VM running Docker Compose, or a managed container platform (Cloud Run, App Runner, Fly.io, Azure Container Apps), handles a surprising amount of traffic with far less to operate. Kubernetes pays for itself when you have many services, need a common deployment model across teams, or need its scheduling and self-healing across many nodes — and its cost is real: networking, ingress, certificates, secrets, monitoring, and upgrades all become your problem, even on a managed service.
Related Articles
- Docker Compose Guide
- Docker Multi-Stage Build Optimization
- Kubernetes Pod Troubleshooting
- FastAPI in Production
- GitHub Actions CI/CD Guide
Frequently Asked Questions (FAQ)
Q. I built my image locally, but the Pod is stuck in ErrImagePull or ImagePullBackOff. Why?
A. Cluster nodes pull images from a registry; they cannot see an image that exists only in your local Docker daemon. With a :latest tag, as in the Pod example, the default imagePullPolicy is Always, so Kubernetes tries to download the image even if a copy happens to be present. Push a versioned tag to a registry the cluster can reach, as the Deployment example does with myregistry/my-api:1.0.0, or on a local cluster load the image directly with kind load docker-image or minikube image load.