Deploying a Python Web App: Gunicorn, Docker, Heroku and EC2 with Nginx

Key takeaways

The development server that runs with flask run or manage.py runserver is not meant for production. A production Python deployment runs the app under a WSGI server such as Gunicorn, reads config from the environment, and sits behind a reverse proxy. This post walks through that setup on Docker, Heroku and a plain EC2 server, and the errors you are likely to hit.

Introduction: what changes between development and production

On your laptop, flask run or python manage.py runserver starts a server that reloads on every change and shows full tracebacks in the browser. That is exactly what you do not want in production. Both frameworks’ documentation says the development server is not meant for production use: it is not designed for concurrent load or hardened against hostile clients, and debug mode can expose code and configuration to anyone who triggers an error.

A production deployment of a Python web app almost always has the same three layers, whatever the hosting platform:

  1. A WSGI server (Gunicorn in this post; uWSGI and Waitress are alternatives) that runs several copies of your app and handles the HTTP protocol with the outside world.
  2. Configuration from the environment: secret keys, database URLs and the debug flag come from environment variables, never from committed code.
  3. A reverse proxy or platform router in front (Nginx, Caddy, a cloud load balancer, or Heroku’s router) that terminates TLS, serves static files, and buffers slow clients.

The sections below build those layers first, then show how Docker, Heroku and EC2 each provide them. If you use an async framework such as FastAPI, the structure is the same, with an ASGI server (Uvicorn, or Gunicorn with Uvicorn workers) instead of a WSGI one.


Preparing the app

Pin dependencies

python -m pip freeze > requirements.txt
Flask==3.1.0
gunicorn==23.0.0
python-dotenv==1.0.1

Pinning exact versions makes the server install what you tested. Run pip freeze inside the project’s virtual environment, not your global Python, or you ship every package you ever installed. Watch for platform-specific packages (for example Windows-only ones) that fail to install on a Linux server. Larger projects often keep direct dependencies in pyproject.toml or requirements.in and generate the pinned file with a tool such as pip-tools or uv, which separates “what I need” from “exact versions”.

Read configuration from the environment

# app.py
import os
from dotenv import load_dotenv
from flask import Flask

load_dotenv()  # reads .env if present; existing environment variables take precedence

app = Flask(__name__)
app.config["SECRET_KEY"] = os.environ["SECRET_KEY"]          # fail fast if missing
app.config["DEBUG"] = os.getenv("DEBUG", "false").lower() == "true"
# .env (local development only; add it to .gitignore)
SECRET_KEY=dev-only-not-secret
DATABASE_URL=postgresql://user:pass@localhost/db
DEBUG=true

Using os.environ["SECRET_KEY"] instead of os.getenv is deliberate: if the variable is missing, the app crashes at startup with a KeyError you see immediately, rather than running with SECRET_KEY = None and failing in confusing ways later (Flask sessions, for example, refuse to work without a key). Parse booleans explicitly. os.getenv("DEBUG") returns the string "False", which is truthy, so if os.getenv("DEBUG"): turns debug mode on in production.

In production, the platform supplies the variables (Heroku config vars, systemd EnvironmentFile, Docker --env-file or orchestrator secrets); the .env file is a development convenience.


Gunicorn

pip install gunicorn
gunicorn app:app                              # module "app", variable "app"
gunicorn -w 4 -b 127.0.0.1:8000 app:app       # 4 workers, local interface only
gunicorn myproject.wsgi                       # Django: uses myproject/wsgi.py's "application"

The argument is module:variable. With the Flask app factory pattern, you can call the factory directly: gunicorn "app:create_app()". If the name is wrong, Gunicorn exits with an error such as Failed to find attribute 'app' in 'app'., which usually means the variable has a different name or the module is being run from the wrong working directory.

Configuration file

# gunicorn.conf.py (picked up automatically from the working directory)
import multiprocessing

bind = "127.0.0.1:8000"
workers = multiprocessing.cpu_count() * 2 + 1
worker_class = "sync"
timeout = 30
graceful_timeout = 30
accesslog = "-"   # log to stdout
errorlog = "-"

How it works. Gunicorn starts a master process, which forks the configured number of workers. Each worker imports your app and handles one request at a time (for the default sync worker class). The master monitors the workers: if one stops responding for longer than timeout seconds, the master kills it and starts a replacement, logging [CRITICAL] WORKER TIMEOUT (pid:1234).

Choosing workers. With sync workers, the number of workers is the number of requests you can serve simultaneously. The (2 × cores) + 1 suggestion from the Gunicorn docs assumes requests spend part of their time waiting on I/O. The real limit is often memory: each worker holds a full copy of the app, so a 300 MB Django process with nine workers needs close to 3 GB. If requests mostly wait on external APIs, threaded workers (--threads 4) or an async worker class serve more concurrent requests with less memory.

Timeouts. A WORKER TIMEOUT almost never means the timeout is too small. It means one request blocked a worker: a slow query without an index, an HTTP call to a third party with no timeout, or a report generated inline. Raising timeout to 300 hides the problem while every worker slowly ends up stuck on the same slow endpoint, and the whole site stops responding. Set timeouts on outbound calls and move long jobs to a task queue.

I have seen a deployment go down because of a single missing timeout on a requests.get() call to a partner API. When that API hung, each request to the affected endpoint tied up one sync worker until Gunicorn killed it, and with four workers it took only four concurrent users to make the site return 502 to everyone. Nothing in the app’s code was “wrong” from a correctness point of view, which is why this failure is easy to miss in review: always pass timeout= to requests, since its default is to wait forever.


Docker

Dockerfile

FROM python:3.12-slim

ENV PYTHONDONTWRITEBYTECODE=1 \
    PYTHONUNBUFFERED=1

WORKDIR /app

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY . .

RUN useradd --create-home appuser
USER appuser

EXPOSE 8000
CMD ["gunicorn", "-w", "4", "-b", "0.0.0.0:8000", "app:app"]

Copying requirements.txt and installing it before copying the rest of the code means Docker reuses the cached dependency layer when only your code changes, so rebuilds take seconds instead of minutes. PYTHONUNBUFFERED=1 makes print and logging output appear in docker logs immediately instead of sitting in a buffer. Inside a container Gunicorn must bind to 0.0.0.0; binding to 127.0.0.1 makes it reachable only from inside the container itself, and curl localhost:8000 from the host fails with “connection reset” or “empty reply”. Add a .dockerignore with .env, .git, venv/ and __pycache__/, so secrets and local environments never end up in the image.

docker-compose.yml

services:
  web:
    build: .
    ports:
      - "8000:8000"
    env_file: .env.production
    depends_on:
      db:
        condition: service_healthy

  db:
    image: postgres:16
    environment:
      POSTGRES_USER: user
      POSTGRES_PASSWORD: pass
      POSTGRES_DB: mydb
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U user -d mydb"]
      interval: 5s
      retries: 10

volumes:
  postgres_data:
docker compose up -d

The top-level version: key found in older examples is obsolete; current Docker Compose ignores it and prints a warning. A plain depends_on: [db] only waits for the database container to start, not for PostgreSQL to accept connections, so the web container often crashes on its first boot with connection refused. The healthcheck plus condition: service_healthy fixes that ordering. The named volume keeps the data when containers are recreated; docker compose down -v deletes it.


Heroku

# Procfile
web: gunicorn app:app
# .python-version
3.12
heroku login
heroku create myapp
heroku config:set SECRET_KEY="$(python -c 'import secrets; print(secrets.token_hex(32))')"
git push heroku main
heroku logs --tail

Heroku detects a Python app from requirements.txt (or other supported package manager files), installs dependencies, and runs the web process from the Procfile. You do not pass -b: Heroku sets a PORT environment variable and Gunicorn binds to 0.0.0.0:$PORT automatically when it is set. The Python version is selected with a .python-version file in current versions of the buildpack; older guides use runtime.txt, which has been deprecated.

Things that behave differently from a server: the file system is ephemeral, so uploaded files disappear on every restart or deploy (store them in object storage such as S3); dynos restart at least daily; and Heroku’s router cuts off requests that take longer than 30 seconds with an H12 Request timeout error, regardless of your Gunicorn timeout. Heroku no longer has a free tier, so check current pricing before choosing it for a hobby project. Run migrations as part of the release, for example with a release: python manage.py migrate line in the Procfile.


AWS EC2 with Nginx and systemd

Server setup

ssh -i key.pem ubuntu@<server-ip>
sudo apt update
sudo apt install -y python3-venv nginx

sudo mkdir -p /var/www/myapp && sudo chown ubuntu: /var/www/myapp
git clone https://github.com/user/myapp.git /var/www/myapp
cd /var/www/myapp
python3 -m venv venv
venv/bin/pip install -r requirements.txt

Recent Ubuntu and Debian releases mark the system Python as externally managed, so sudo pip install fails with error: externally-managed-environment. That is intentional: installing into a virtual environment keeps the app’s packages separate from the ones the OS depends on.

systemd service

# /etc/systemd/system/myapp.service
[Unit]
Description=myapp Gunicorn
After=network.target

[Service]
User=www-data
Group=www-data
WorkingDirectory=/var/www/myapp
EnvironmentFile=/etc/myapp.env
ExecStart=/var/www/myapp/venv/bin/gunicorn -w 4 -b 127.0.0.1:8000 app:app
Restart=always

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now myapp
sudo systemctl status myapp
journalctl -u myapp -f

Running Gunicorn from an SSH session with gunicorn ... & works until you log out or the server reboots. systemd starts it at boot, restarts it if it crashes, and collects its output in the journal. Calling the virtualenv’s gunicorn by absolute path makes activation unnecessary. /etc/myapp.env holds KEY=value lines; make it readable only by root (chmod 600), since systemd reads it before dropping privileges. www-data must be able to read the app directory.

Nginx

# /etc/nginx/sites-available/myapp
server {
    listen 80;
    server_name example.com;

    location /static/ {
        alias /var/www/myapp/static/;
    }

    location / {
        proxy_pass http://127.0.0.1:8000;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}
sudo ln -s /etc/nginx/sites-available/myapp /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
sudo apt install -y certbot python3-certbot-nginx && sudo certbot --nginx -d example.com

Gunicorn binds to 127.0.0.1 so only Nginx can reach it; the security group then only needs ports 80 and 443 open. Nginx serves static files directly from disk, which is far cheaper than passing them through Python, and buffers slow clients so a user on a bad connection does not hold a Gunicorn worker for the whole upload.

The forwarded headers tell the app the original scheme and client address. The app must be told to trust them: in Flask, wrap the app with werkzeug.middleware.proxy_fix.ProxyFix; in Django, set SECURE_PROXY_SSL_HEADER = ("HTTP_X_FORWARDED_PROTO", "https"). Without that, the app thinks every request is plain HTTP, and HTTPS redirects loop forever or generated URLs use http://.

When Nginx returns 502 Bad Gateway

A 502 means Nginx could not get a response from Gunicorn. The Nginx error log (/var/log/nginx/error.log) says why:

  • connect() failed (111: Connection refused) while connecting to upstream: Gunicorn is not running or listens on a different address. Check systemctl status myapp and the journal; a crash at import time (missing environment variable, missing package) is the usual cause.
  • upstream prematurely closed connection: a worker died mid-request, often because of a WORKER TIMEOUT or the kernel’s out-of-memory killer (look for Out of memory: Killed process in dmesg).
  • upstream timed out: the request took longer than Nginx’s proxy_read_timeout (60 seconds by default).

Before you go live

For Django, run python manage.py check --deploy, which flags insecure settings. Then make sure of the following:

  • DEBUG is off, and Django’s ALLOWED_HOSTS lists your domain (with DEBUG = False and an empty list, every request returns 400 Bad Request).
  • Static files are collected (python manage.py collectstatic) and served by Nginx, a CDN or WhiteNoise.
  • Migrations run as part of each deploy, before the new code serves traffic.
  • Logs go to stdout or the journal and are shipped somewhere you can search.
  • A health-check endpoint exists for your load balancer or platform.
  • Backups of the database exist and a restore has actually been tested.

The mistake I would single out is DEBUG = True left on a public server. It feels harmless until the first error, when Django or Flask renders a page listing settings, local variables and file paths to whoever triggered it, and Werkzeug’s interactive debugger goes further by letting a visitor run Python code on the server if its PIN protection is bypassed. Parse the flag strictly from the environment and default it to off, so a missing variable means production-safe behaviour.


Options at a glance

ApproachYou manageGood fit
PaaS (Heroku, Render, Fly.io, …)App code and configSmall teams, fast iteration
VM + systemd + Nginx (EC2, Lightsail, any VPS)OS, TLS, processes, monitoringPredictable costs, full control
Containers on an orchestrator (ECS, Kubernetes)Images, manifests, clusterMany services, larger teams

Further reading: Gunicorn settings, the Django deployment checklist, and Flask’s deployment docs.

Next in the series: Pandas for data analysis.