Kubernetes in Practice: Pods, Deployments, Services, Ingress, HPA and kubectl Troubleshooting

Key takeaways

Kubernetes orchestrates containers at scale — handling deployment, scaling, networking, and self-healing. This guide covers the core objects you'll work with every day, from Pods to Ingress, plus Helm and real troubleshooting commands.

How Kubernetes Works

Kubernetes (K8s) is a container orchestration platform. You describe your desired state in YAML manifests; Kubernetes continuously reconciles actual state with desired state.

You:        kubectl apply -f deployment.yaml  (desired state)
Kubernetes: "I see 0 pods, you want 3. Creating 3 pods."
            → Node failure: "Pod died. Creating replacement."
            → CPU spike: "HPA says scale to 5. Creating 2 more pods."

Architecture:

Control Plane                    Worker Nodes
─────────────────────────        ──────────────────────────
API Server  ←── kubectl          Node 1: Pod A, Pod B
etcd (state store)               Node 2: Pod C, Pod D
Scheduler                        Node 3: Pod E
Controller Manager

The single most important mental model for everything that follows: Kubernetes is a reconciliation loop, not an imperative task runner. You never tell it “start 3 pods now” the way you’d run a script — you declare “I want 3 replicas of this pod running” in a manifest, and a controller continuously compares that desired state against the cluster’s actual state, taking corrective action (creating, deleting, or replacing pods) whenever they diverge, forever, until the manifest changes or is deleted. This is why a node dying doesn’t require anyone to notice and manually restart anything — the reconciliation loop notices the discrepancy (desired: 3, actual: 2) and corrects it on its own, and it’s also why understanding “what’s the desired state right now” (via kubectl get/describe) is the starting point for debugging almost anything that goes wrong on a cluster.


Core Objects

ObjectPurpose
PodSmallest deployable unit — one or more containers
DeploymentManages Pod replicas, rolling updates, rollbacks
ServiceStable network endpoint for a set of Pods
ConfigMapNon-sensitive configuration data
SecretSensitive data (passwords, tokens, TLS certs)
IngressHTTP/HTTPS routing from outside the cluster
PersistentVolumePersistent storage that outlives Pods
HPAHorizontal Pod Autoscaler — scale based on CPU/memory

These objects layer on top of each other deliberately, and understanding the layering is what makes the rest of this guide click: a Pod is the actual running unit, but you almost never create one directly (covered below) — a Deployment wraps Pods to manage their lifecycle, a Service gives that ever-changing set of Pods a stable network address (since Pods themselves get replaced and get new IPs constantly), and Ingress sits in front of Services to route external traffic in. ConfigMaps/Secrets inject configuration into Pods without baking it into the container image, and PersistentVolumes give Pods durable storage that survives a Pod being replaced — each object exists to solve a specific problem the layer below it can’t.


kubectl — Essential Commands

# Context and cluster
kubectl config get-contexts                   # List clusters
kubectl config use-context my-cluster         # Switch cluster
kubectl cluster-info

# Resource inspection
kubectl get pods                              # List pods (default namespace)
kubectl get pods -n kube-system               # Specific namespace
kubectl get pods -A                           # All namespaces
kubectl get pods -o wide                      # Show node, IP
kubectl get all                               # All resources in namespace

kubectl describe pod my-pod                   # Detailed pod info + events
kubectl describe deployment my-app

# Logs
kubectl logs my-pod                           # Pod logs
kubectl logs my-pod -c my-container           # Specific container
kubectl logs my-pod --previous                # Previous container (after crash)
kubectl logs -f my-pod                        # Follow logs
kubectl logs -l app=my-app --all-containers   # All pods matching label

# Exec into a pod
kubectl exec -it my-pod -- /bin/bash
kubectl exec -it my-pod -c my-container -- sh

# Apply and delete
kubectl apply -f deployment.yaml              # Create/update
kubectl apply -f ./k8s/                       # Apply all files in directory
kubectl delete -f deployment.yaml
kubectl delete pod my-pod
kubectl delete pod my-pod --force --grace-period=0  # Force delete (stuck pods)

# Port forwarding (for testing)
kubectl port-forward pod/my-pod 8080:8080
kubectl port-forward svc/my-service 8080:80

# Copy files
kubectl cp my-pod:/app/logs/app.log ./app.log

A few of these are worth understanding rather than just memorizing. kubectl describe (versus get) is where the actual debugging information lives — it includes the Events section at the bottom, a chronological log of everything the scheduler and kubelet did (and failed to do) with that resource, which is almost always the first place to look when something isn’t working as expected, covered more in the Troubleshooting section below. --previous on logs is easy to forget but essential after a crash: once a container restarts, kubectl logs without --previous shows the new container’s logs (likely empty or just starting up), not the crashed one’s — the actual error that caused the crash is only visible in the previous container’s log.


Namespaces

Namespaces partition cluster resources — use them to separate environments or teams.

kubectl create namespace staging
kubectl get namespaces

# Set default namespace (saves typing -n flag)
kubectl config set-context --current --namespace=production

Namespaces are a logical partition within a single cluster, not physical isolation — resources in different namespaces still share the same underlying nodes and compute capacity unless you additionally set up resource quotas per namespace, and by default, Pods in one namespace can reach Services in another (via the full service.namespace.svc.cluster.local DNS name) unless network policies restrict it. This matters when reasoning about “is staging actually isolated from production” — namespace separation alone answers “can I list/manage these resources separately,” not “is this environment network-isolated” or “can this environment consume all of production’s resources under load.”


Pods

A Pod wraps one or more containers that share network and storage.

# pod.yaml
apiVersion: v1
kind: Pod
metadata:
  name: my-app
  labels:
    app: my-app
    version: v1
spec:
  containers:
    - name: app
      image: my-app:1.0.0
      ports:
        - containerPort: 8080
      env:
        - name: PORT
          value: "8080"
      resources:
        requests:
          memory: "64Mi"
          cpu: "100m"         # 100 millicores = 0.1 CPU
        limits:
          memory: "128Mi"
          cpu: "500m"
      readinessProbe:
        httpGet:
          path: /health
          port: 8080
        initialDelaySeconds: 5
        periodSeconds: 10
      livenessProbe:
        httpGet:
          path: /health
          port: 8080
        initialDelaySeconds: 15
        periodSeconds: 20

Pods are ephemeral — don’t run Pods directly. Use Deployments.

The two probes here answer genuinely different questions, and confusing them is a common source of production incidents: readinessProbe asks “is this Pod ready to receive traffic right now” — failing it removes the Pod from a Service’s routing without killing it, which is exactly the mechanism that keeps traffic away from a Pod that’s still warming up (loading a large cache, establishing a database connection) during the initialDelaySeconds window. livenessProbe asks “is this Pod alive at all” — failing it kills and restarts the container, which is meant for detecting a genuinely hung/deadlocked process, not slow startup. Setting livenessProbe too aggressively (a short initialDelaySeconds on an app with a slow startup) is a classic self-inflicted outage: Kubernetes kills and restarts a Pod that was simply still starting up, which can trigger a crash-restart loop that never lets the app finish initializing. requests and limits serve different purposes too — requests is what the scheduler uses to decide which node has room for this Pod (a scheduling-time guarantee), while limits is a hard runtime ceiling enforced by the kernel/cgroups; exceeding a memory limit gets a Pod OOMKilled, covered in the Troubleshooting section.


Deployments

Deployments manage Pod replicas with rolling updates and rollback support.

# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
  namespace: production
spec:
  replicas: 3
  selector:
    matchLabels:
      app: my-app
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1           # Max extra pods during update
      maxUnavailable: 0     # Never go below 3 pods during update
  template:
    metadata:
      labels:
        app: my-app
    spec:
      containers:
        - name: app
          image: my-app:1.2.0
          ports:
            - containerPort: 8080
          env:
            - name: DATABASE_URL
              valueFrom:
                secretKeyRef:
                  name: db-secret
                  key: url
            - name: APP_ENV
              valueFrom:
                configMapKeyRef:
                  name: app-config
                  key: environment
          resources:
            requests:
              memory: "128Mi"
              cpu: "100m"
            limits:
              memory: "256Mi"
              cpu: "1000m"
          readinessProbe:
            httpGet:
              path: /health
              port: 8080
            initialDelaySeconds: 10
            periodSeconds: 5

maxSurge/maxUnavailable together control exactly how a rolling update plays out, and setting maxUnavailable: 0 here is a deliberate production-safety choice: it guarantees the Deployment never drops below the full 3 replicas during an update (it can temporarily run more than 3, per maxSurge: 1, but never fewer), trading a slightly slower rollout for zero capacity loss during the deploy. The readinessProbe from the earlier Pod example is what makes this actually safe rather than just configured correctly — Kubernetes only considers a new Pod “up” for the purposes of maxUnavailable accounting once it passes its readiness check, so a rolling update won’t route traffic to a new version that’s still starting up, even if the container process itself is already running.

# Deployment operations
kubectl rollout status deployment/my-app     # Watch rollout progress
kubectl rollout history deployment/my-app    # View revision history
kubectl rollout undo deployment/my-app       # Rollback to previous version
kubectl rollout undo deployment/my-app --to-revision=2  # Rollback to specific revision

# Scale
kubectl scale deployment my-app --replicas=5

# Update image
kubectl set image deployment/my-app app=my-app:1.3.0

rollout undo works because every Deployment change creates a new revision in its rollout history, not just an in-place mutation — this is what makes rollback genuinely fast and reliable rather than “redeploy the old config and hope”: Kubernetes already has the previous ReplicaSet’s exact Pod template retained and just needs to scale it back up while scaling the current one down, following the same rolling-update mechanics as a forward deploy. This is also why kubectl rollout history is worth checking before an undo on a system with frequent deploys — without a specific --to-revision, undo reverts to the immediately previous revision, which may not be the last known-good one if multiple bad deploys happened in a row.


Services

Services provide a stable network endpoint for a set of Pods (selected by labels).

# service.yaml
apiVersion: v1
kind: Service
metadata:
  name: my-app
spec:
  selector:
    app: my-app         # Matches Pods with this label
  ports:
    - port: 80          # Service port
      targetPort: 8080  # Pod port
  type: ClusterIP       # Internal only (default)

A Service’s whole reason for existing is solving a problem the Deployment layer alone can’t: individual Pods are constantly being created and destroyed (rolling updates, crash restarts, scaling), each getting a brand-new internal IP address every time — a Service provides one stable name and IP that never changes, continuously updated (via the selector label match) to route to whichever Pods currently exist and are ready. This is why Services select Pods by label, not by directly referencing specific Pod names or a Deployment — labels are the loose-coupling mechanism that lets Services, Deployments, and NetworkPolicies all independently target “Pods matching app: my-app” without hardcoded references to each other.

Service types:

# ClusterIP — internal access only (default)
type: ClusterIP

# NodePort — exposed on every node's IP at a static port (30000-32767)
type: NodePort
ports:
  - port: 80
    targetPort: 8080
    nodePort: 30080     # Access: any-node-ip:30080

# LoadBalancer — creates cloud load balancer (AWS ELB, GCP LB)
type: LoadBalancer

These three exist on a spectrum of exposure, and picking the wrong one is a common early-Kubernetes mistake: ClusterIP is correct for anything only other Pods should reach (a backend API, a database) and should be the default choice unless you specifically need external access. NodePort opens a port directly on every node’s IP — functional for quick testing, but rarely the right production choice since it ties external access to node IPs directly and needs a separate load balancer or DNS strategy in front of it anyway. LoadBalancer is the standard production entry point on any major cloud provider, but it provisions a real, billed cloud load balancer resource per Service — a cluster running LoadBalancer Services for many separate apps can accumulate real infrastructure cost, which is one of the practical reasons Ingress (routing many hostnames/paths through one shared entry point) is usually preferred over one LoadBalancer Service per app.


ConfigMaps and Secrets

# configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: app-config
data:
  environment: "production"
  log_level: "info"
  max_connections: "100"
# secret.yaml — values are base64 encoded
apiVersion: v1
kind: Secret
metadata:
  name: db-secret
type: Opaque
data:
  url: cG9zdGdyZXNxbDovL3VzZXI6cGFzc3dAZGI6NTQzMi9teWRi  # base64
  password: c3VwZXJzZWNyZXQ=

The security-relevant detail worth being explicit about: base64 in a Secret is encoding, not encryption — it’s trivially reversible (echo c3VwZXJzZWNyZXQ= | base64 -d), and its only real purpose is letting arbitrary binary data live safely inside a YAML text field. Anyone with get/describe access to Secrets in a namespace can read the plaintext values instantly; a Kubernetes Secret by itself is a mechanism for distributing sensitive config to Pods via RBAC-controlled access, not a substitute for genuine encryption at rest, which needs to be configured separately (etcd encryption, or an external secrets manager like Vault/AWS Secrets Manager integrated via a controller). This is also why committing raw Secret YAML files with real base64 values to source control is exactly as dangerous as committing a plaintext .env file — the encoding provides no actual protection.

# Create secrets from literal values (easier than base64)
kubectl create secret generic db-secret \
  --from-literal=url=postgresql://user:pass@db:5432/mydb \
  --from-literal=password=supersecret

# Create from file
kubectl create secret generic tls-secret \
  --from-file=tls.crt=./cert.pem \
  --from-file=tls.key=./key.pem

Use secrets as environment variables or volume mounts:

containers:
  - name: app
    # As environment variables
    env:
      - name: DB_URL
        valueFrom:
          secretKeyRef:
            name: db-secret
            key: url

    # As a volume (file)
    volumeMounts:
      - name: secrets-vol
        mountPath: /etc/secrets
        readOnly: true

volumes:
  - name: secrets-vol
    secret:
      secretName: db-secret

The environment-variable and volume-mount approaches to consuming a Secret have a real, practical difference: environment variables are visible via kubectl describe pod (to anyone with pod-read access) and to anything that can inspect the process’s environment inside the container, and they’re set once at container start — a Secret rotated afterward won’t update a running container’s environment without a restart. Mounting a Secret as a volume avoids exposing it via describe, and Kubernetes automatically updates the mounted file’s contents when the underlying Secret changes (typically within about a minute), which application code can pick up by watching the file for changes — a meaningfully better fit for credentials that need to rotate without a full Pod restart.


Ingress — HTTP Routing

Ingress routes external HTTP/HTTPS traffic to Services based on host and path rules.

# Install Nginx Ingress Controller (common choice)
kubectl apply -f https://raw.githubusercontent.com/kubernetes/ingress-nginx/main/deploy/static/provider/cloud/deploy.yaml
# ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: my-app-ingress
  annotations:
    nginx.ingress.kubernetes.io/rewrite-target: /
    cert-manager.io/cluster-issuer: "letsencrypt-prod"  # Auto SSL
spec:
  ingressClassName: nginx
  tls:
    - hosts:
        - myapp.com
      secretName: myapp-tls   # cert-manager creates this
  rules:
    - host: myapp.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: frontend
                port:
                  number: 80
          - path: /api
            pathType: Prefix
            backend:
              service:
                name: api-service
                port:
                  number: 80

Ingress is what lets a single, shared entry point (one LoadBalancer, one public IP) route to many different backend Services based on hostname and URL path — without it, exposing multiple HTTP services externally would mean one LoadBalancer Service each, with its own cloud load balancer cost, as noted earlier. The ingress-nginx controller installed above isn’t built into Kubernetes itself; an Ingress resource (the YAML manifest) is just a routing specification, and it needs an actual Ingress controller running in the cluster to read those rules and configure a real reverse proxy accordingly — this is a genuinely common source of confusion for newcomers, since applying an Ingress manifest with no controller installed silently does nothing. cert-manager.io/cluster-issuer is a similarly separate piece: it’s an annotation cert-manager (a separate controller, also not built in) watches for, automatically provisioning and renewing a Let’s Encrypt TLS certificate for the hosts listed and storing it in the secretName referenced under tls — genuinely free, automated HTTPS without manually managing certificates.


Resource Limits and HPA

Set resource requests (scheduling guarantee) and limits (hard cap), then let HPA scale automatically:

# hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: my-app-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-app
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70   # Scale when avg CPU > 70%
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 80
kubectl get hpa                          # View current HPA status
kubectl describe hpa my-app-hpa          # Detailed metrics and events

HPA reads CPU/memory utilization percentage relative to each Pod’s requests value, not an absolute number — this is precisely why setting accurate requests (from the Pod/Deployment examples earlier) isn’t just a scheduling formality, it directly determines when autoscaling kicks in: a Pod with an artificially low CPU request will hit “70% utilization” and trigger scaling much sooner than one with a request that accurately reflects real usage, even under identical actual load. minReplicas: 2 (rather than 1) as a floor matters independent of HPA’s scaling logic — it guarantees baseline availability even at the lowest traffic, since a single-replica Deployment has no redundancy if that one Pod crashes or its node fails.


PersistentVolumes for Stateful Apps

# pvc.yaml — request storage
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: postgres-pvc
spec:
  accessModes:
    - ReadWriteOnce      # Single node read-write
  storageClassName: gp2  # AWS EBS SSD (cloud-specific)
  resources:
    requests:
      storage: 20Gi

ReadWriteOnce here is a real constraint worth understanding before it causes a confusing scheduling failure: it means the underlying volume can only be mounted read-write by Pods on a single node at a time — which is fine for a database with one replica, but means a naive attempt to scale that same PVC’s consumer to multiple replicas across different nodes will leave some Pods stuck unable to mount the volume. ReadWriteMany exists for genuinely shared-across-nodes storage, but it requires a storage backend that supports it (like NFS or a cloud-provider equivalent) — not every storageClassName does, which is why checking what access modes your storage class actually supports matters before designing around it.

# Use PVC in a StatefulSet
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: postgres
spec:
  replicas: 1
  selector:
    matchLabels:
      app: postgres
  template:
    metadata:
      labels:
        app: postgres
    spec:
      containers:
        - name: postgres
          image: postgres:16
          env:
            - name: POSTGRES_PASSWORD
              valueFrom:
                secretKeyRef:
                  name: postgres-secret
                  key: password
          volumeMounts:
            - name: postgres-data
              mountPath: /var/lib/postgresql/data
  volumeClaimTemplates:
    - metadata:
        name: postgres-data
      spec:
        accessModes: ["ReadWriteOnce"]
        storageClassName: gp2
        resources:
          requests:
            storage: 20Gi

volumeClaimTemplates is the mechanism that makes StatefulSets genuinely different from a Deployment for storage purposes: rather than one shared PVC (as in the plain pvc.yaml example above), each replica of a StatefulSet gets its own, automatically-created PVC from this template, and — crucially — that identity is stable across restarts. A StatefulSet Pod named postgres-0 always gets rescheduled back onto its own specific PVC, not a fresh one, which is exactly the durable, stable-identity guarantee a database replica needs and a Deployment (where Pods are explicitly interchangeable and replaceable) doesn’t provide. This is the concrete answer to the FAQ’s “Deployment vs StatefulSet” question above — it’s not a stylistic choice, StatefulSet’s storage semantics are a hard requirement for genuinely stateful workloads.


Helm — Kubernetes Package Manager

Helm manages complex Kubernetes applications with templating and versioning.

# Install Helm
brew install helm

# Add a chart repository
helm repo add bitnami https://charts.bitnami.com/bitnami
helm repo update

# Install a chart
helm install my-postgres bitnami/postgresql \
  --set auth.postgresPassword=mypassword \
  --set primary.persistence.size=20Gi

# Install with custom values file
helm install my-app ./my-chart -f values-prod.yaml

# Upgrade
helm upgrade my-app ./my-chart -f values-prod.yaml

# Rollback
helm rollback my-app 1    # Rollback to revision 1

# List releases
helm list

# Uninstall
helm uninstall my-postgres

Every helm command here maps onto a concept plain kubectl apply doesn’t have: a release — a named, versioned instance of a chart’s manifests applied to the cluster, which is what helm upgrade/helm rollback operate on. This is genuinely more than syntactic sugar over YAML templating: helm rollback my-app 1 reverts every resource the chart manages back to exactly how it looked at that release revision in one atomic operation, where doing the equivalent with plain kubectl apply -f would mean manually tracking and reapplying old YAML for every affected resource yourself.

Create a Helm Chart

helm create my-app    # Scaffold chart structure
my-app/
├── Chart.yaml            # Chart metadata
├── values.yaml           # Default values
└── templates/
    ├── deployment.yaml   # Template using {{ .Values.* }}
    ├── service.yaml
    └── ingress.yaml
# values.yaml
replicaCount: 3
image:
  repository: my-app
  tag: "1.0.0"
resources:
  requests:
    cpu: 100m
    memory: 128Mi
  limits:
    cpu: 500m
    memory: 256Mi
ingress:
  enabled: true
  host: myapp.com

values.yaml is what lets the exact same chart deploy differently across environments — a values-staging.yaml with replicaCount: 1 and a values-prod.yaml with replicaCount: 5 can share one chart’s templates entirely, with only the per-environment values file changing (helm install my-app ./my-chart -f values-prod.yaml, shown earlier). This is the concrete answer to “why Helm over plain YAML” from the FAQ above: without templating, managing the same app config differences across dev/staging/prod typically means either duplicated near-identical YAML files per environment, or a separate templating tool bolted on independently.

# templates/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: {{ .Release.Name }}-app
spec:
  replicas: {{ .Values.replicaCount }}
  template:
    spec:
      containers:
        - name: app
          image: {{ .Values.image.repository }}:{{ .Values.image.tag }}
          resources:
            {{- toYaml .Values.resources | nindent 12 }}

Helm’s templates are Go’s text/template syntax, which is why {{ .Release.Name }} and {{ .Values.replicaCount }} look different from anything in plain Kubernetes YAML — .Release is metadata Helm itself injects about the current release (its name, revision, whether it’s an install or upgrade), while .Values pulls from the merged result of values.yaml and whatever -f/--set overrides were passed at install time. nindent 12 deserves a specific callout since it trips up newcomers writing their first chart: because Helm templates are pure text substitution before the result is parsed as YAML, indentation has to be produced manually by the template itself — nindent 12 inserts a newline and then indents every line of the resources block by 12 spaces so it lines up correctly as valid YAML once substituted in; getting the indent count wrong produces YAML that fails to parse, not a helpful template error.


Troubleshooting

# Pod won't start — check events
kubectl describe pod <pod-name>
# Look for: "Back-off restarting failed container", image pull errors, OOMKilled

# Check pod logs (including previous container after crash)
kubectl logs <pod-name> --previous

# Pod stuck in Pending — check node resources
kubectl describe pod <pod-name>  # Look for "Insufficient cpu/memory"
kubectl describe nodes           # Check node capacity and allocations

# Check all events in namespace (sorted by time)
kubectl get events --sort-by='.metadata.creationTimestamp'

# Container OOMKilled — increase memory limit
kubectl describe pod <pod-name>
# Look for: "OOMKilled", "Last State: Terminated: Reason: OOMKilled"

# Service not reachable — verify selector matches pod labels
kubectl get svc my-service -o yaml    # Check selector
kubectl get pods --show-labels        # Check pod labels
kubectl describe endpoints my-service # See which pods are registered

# Debug network with a temporary pod
kubectl run debug --image=curlimages/curl -it --rm -- sh
# Then: curl http://my-service/health

These commands map onto the two most common failure categories, and it’s worth knowing which one you’re looking at before reaching for the wrong fix. A Pod stuck Pending means the scheduler can’t place it at all — usually insufficient node resources for the requested requests values, and the fix is either scaling the cluster or lowering the requests, not restarting anything (there’s nothing running yet to restart). A Pod in CrashLoopBackOff means it did get scheduled and started, but the container keeps exiting — --previous logs are the way to see why, and the fix is in the application or its config, not the cluster. The debug-Pod trick (kubectl run debug --image=curlimages/curl -it --rm) is worth remembering specifically for “Service not reachable” issues: it lets you test connectivity from inside the cluster’s network, which is the only way to distinguish “the Service/Pod networking is broken” from “the app itself is returning an error” — testing from your own machine outside the cluster conflates both possibilities into one symptom.


Production Checklist

# deployment.yaml production settings
spec:
  replicas: 3                      # Always > 1 for availability
  strategy:
    rollingUpdate:
      maxUnavailable: 0            # Never take all pods down at once
  template:
    spec:
      containers:
        - resources:
            requests:              # Set requests — required for scheduling
              cpu: 100m
              memory: 128Mi
            limits:                # Set limits — prevent noisy neighbors
              cpu: 1000m
              memory: 512Mi
          readinessProbe:          # Required — prevents sending traffic to unready pods
            httpGet:
              path: /health
              port: 8080
          livenessProbe:           # Required — restarts crashed containers
            httpGet:
              path: /health
              port: 8080
      affinity:
        podAntiAffinity:           # Spread pods across nodes
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              podAffinityTerm:
                labelSelector:
                  matchLabels:
                    app: my-app
                topologyKey: kubernetes.io/hostname

podAntiAffinity is the setting most teams discover they needed only after an incident: without it, nothing stops Kubernetes from scheduling all 3 replicas of a Deployment onto the same physical node, which quietly defeats the whole point of running multiple replicas — a single node failure would take down every replica simultaneously, exactly the scenario replicas: 3 was meant to protect against. preferredDuringSchedulingIgnoredDuringExecution (a soft preference the scheduler tries to honor but won’t refuse to schedule over, versus the stricter required variant) is the right default for most cases — it spreads Pods across nodes when possible without risking Pods staying unschedulable entirely on a small cluster that genuinely doesn’t have enough distinct nodes to satisfy a hard requirement.