Running Node.js in Production with PM2: Cluster Mode, Ecosystem Files, Logs and Zero-Downtime Reloads

Key takeaways

PM2 is a production process manager for Node.js. It provides clustering, monitoring, logging, zero-downtime reloads, and automatic restarts for reliable Node.js deployments.

Introduction

PM2 is a production-grade process manager for Node.js applications. It keeps your app running 24/7, handles crashes, enables clustering, and provides monitoring and logging.

Node.js itself does none of this. When you run node app.js directly, the moment that process throws an uncaught exception, runs out of memory, or the machine reboots, your application is simply gone — nothing brings it back. Node’s single-threaded event loop also means a plain node app.js invocation uses exactly one CPU core, no matter how many cores the host actually has. PM2 exists to close both gaps: it wraps your application in a supervisor process that restarts it on failure, and it uses Node’s built-in cluster module under the hood to fork multiple worker processes that share one listening port, so a multi-core machine actually gets used.

It’s worth being precise about what PM2 is not. It is not a container runtime, not a load balancer across machines, and not a deployment pipeline. It solves one layer of the problem — keeping a Node.js process (or a fleet of them on one host) alive and observable — and composes with the rest of your infrastructure (Docker, Nginx, Kubernetes, CI/CD) rather than replacing it.

Why PM2?

Without PM2:

node app.js
# App crashes → stays down
# Server restarts → app doesn't start
# 1 process → can't use all CPU cores

With PM2:

pm2 start app.js
# App crashes → auto-restart
# Server restarts → app auto-starts
# Cluster mode → uses all CPU cores
# Built-in monitoring and logs

The difference above is not cosmetic — it’s the difference between an app that survives an unhandled promise rejection at 3 a.m. and one that takes your service down until someone notices and SSHes in. PM2’s restart supervisor watches the process’s exit code and signal; a crash (non-zero exit, or death by SIGSEGV/SIGABRT) triggers an automatic restart with a short backoff, while a clean pm2 stop leaves the process down on purpose. That distinction — “died unexpectedly” vs. “was told to stop” — is exactly what a naive while true; do node app.js; done shell loop gets wrong, since it restarts indiscriminately and offers no visibility into why the process died.

Installation

# Global install (recommended)
npm install -g pm2

# Or with yarn
yarn global add pm2

# Verify
pm2 --version

A global install is the standard approach because PM2 is meant to manage processes for the whole machine, not just one project — you’ll typically run several unrelated apps under the same PM2 daemon. If the global install fails with an EACCES permission error, resist the urge to sudo npm install -g; that creates root-owned files inside your user-level npm prefix and causes confusing permission errors later. The cleaner fix is to point npm’s global prefix at a directory your user owns (or use a Node version manager like nvm, which sidesteps the problem entirely by installing Node and global packages under your home directory). Also note that PM2 runs as a background daemon (pm2 itself is a client that talks to a PM2 God daemon process) — the first command you run spins that daemon up, and it persists independently of your current shell session, which is exactly the property that lets pm2 start app.js keep the app running after you close your terminal.

Basic Usage

# Start app
pm2 start app.js

# Start with name
pm2 start app.js --name my-app

# Start with watch (auto-reload on file change)
pm2 start app.js --watch

# Start with environment
pm2 start app.js --env production

# List all processes
pm2 list

# Stop app
pm2 stop my-app

# Restart app
pm2 restart my-app

# Delete app from PM2
pm2 delete my-app

# Stop all
pm2 stop all

# Restart all
pm2 restart all

# Delete all
pm2 delete all

The --name flag matters more than it looks like it should. Without it, PM2 names the process after the entry script’s filename, so two different projects that both happen to have an app.js will collide in pm2 list and you’ll restart the wrong one. Get in the habit of naming every process explicitly, even for throwaway scripts.

stop and delete are easy to conflate but behave differently: pm2 stop halts the process but keeps its entry in PM2’s process table (visible in pm2 list with status stopped), so pm2 restart my-app still works afterward. pm2 delete removes the entry entirely — PM2 forgets the process existed, and you’d need to pm2 start it again with the original script path and flags. This distinction bites people most often when they run pm2 delete all intending to just stop everything for a deploy; if they haven’t also run pm2 save beforehand (covered in section 7), a subsequent reboot won’t bring anything back because there’s nothing left in the saved process list to resurrect.

The --watch flag is convenient in development (PM2 restarts the process whenever a file changes) but is a footgun in production: it restarts on any file write under the watched directory, including log files your own app writes, temp files, or .git metadata changes during a deploy — which can put a process into a restart loop at exactly the moment you’re trying to ship a change. Reserve --watch for local development, and rely on pm2 reload/pm2 restart triggered explicitly from your deploy script in production.

Cluster Mode

# Start in cluster mode (uses all CPU cores)
pm2 start app.js -i max

# Start with specific number of instances
pm2 start app.js -i 4

# Scale up/down
pm2 scale my-app 8   # Scale to 8 instances
pm2 scale my-app +2  # Add 2 more instances
pm2 scale my-app -2  # Remove 2 instances

Cluster mode is PM2’s signature feature, and it’s worth understanding what actually happens under the hood rather than treating -i max as magic. PM2 forks N worker processes using Node’s built-in cluster module, and a lightweight round-robin balancer inside PM2 distributes incoming connections across them. Each worker is a fully independent Node.js process with its own event loop, its own heap, and — critically — its own memory. That last point is the one people trip over most: cluster mode does not give you shared in-memory state across instances. If your app keeps session data, rate-limit counters, or an in-memory cache in a plain JavaScript object, each worker has its own separate copy, and a user’s second request can land on a different worker that has no idea who they are. The fix is to move any state that must be shared across requests into an external store — Redis is the usual choice — rather than relying on process memory.

The same limitation applies to WebSocket connections and anything relying on “sticky sessions.” A client’s long-lived socket is pinned to whichever worker accepted it; if you need session affinity at the load-balancer level (for example, fronting PM2 with Nginx), you need to configure sticky sessions there, because PM2’s round-robin balancer alone won’t guarantee repeat requests land on the same worker.

-i max uses every core the host reports via os.cpus().length. On a container with a CPU limit set lower than the host’s actual core count (common on shared Kubernetes nodes or constrained Docker containers), max can badly over-provision workers relative to the CPU quota actually available, leading to excessive context-switching rather than real parallelism — this is one of the more common reasons a “clustered” app performs worse than a single instance in a constrained environment. Pinning an explicit instance count (-i 2, tied to your container’s actual CPU request) is usually the safer choice in orchestrated environments.

flowchart LR
    C[Incoming request] --> LB["PM2 cluster balancer<br/>(round-robin)"]
    LB --> W1["Worker 1<br/>separate heap"]
    LB --> W2["Worker 2<br/>separate heap"]
    LB --> W3["Worker 3<br/>separate heap"]
    W1 --> R["(Redis / shared store)"]
    W2 --> R
    W3 --> R

Example: Cluster-Ready Express App

// app.js
const express = require('express');
const app = express();

app.get('/', (req, res) => {
  res.send(`Hello from process ${process.pid}`);
});

const PORT = process.env.PORT || 3000;
app.listen(PORT, () => {
  console.log(`Server running on port ${PORT}, PID: ${process.pid}`);
});
# Start 4 instances
pm2 start app.js -i 4

# Each request handled by different process
curl http://localhost:3000  # Process 1234
curl http://localhost:3000  # Process 1235
curl http://localhost:3000  # Process 1236
curl http://localhost:3000  # Process 1237

Notice this example needs no special code to become cluster-aware — it’s a plain Express app that calls app.listen() normally. That’s the appeal of PM2’s cluster mode over hand-rolling Node’s cluster API yourself: Node’s raw cluster module requires you to branch your own code on cluster.isPrimary/cluster.isWorker and manually fork workers, while PM2 handles all of that transparently outside your application code. The trade-off is that your app must be genuinely stateless per request (or externalize its state, as noted above) for the illusion of “one server” to hold up across four separate processes.

Ecosystem File

# Generate ecosystem file
pm2 ecosystem
// ecosystem.config.js
module.exports = {
  apps: [
    {
      name: 'api',
      script: './app.js',
      instances: 'max',
      exec_mode: 'cluster',
      env: {
        NODE_ENV: 'development',
        PORT: 3000,
      },
      env_production: {
        NODE_ENV: 'production',
        PORT: 8080,
      },
      error_file: './logs/err.log',
      out_file: './logs/out.log',
      log_date_format: 'YYYY-MM-DD HH:mm:ss Z',
      max_memory_restart: '500M',
      watch: false,
      ignore_watch: ['node_modules', 'logs'],
    },
    {
      name: 'worker',
      script: './worker.js',
      instances: 2,
      exec_mode: 'cluster',
      cron_restart: '0 0 * * *', // Restart at midnight
    },
  ],
};
# Start from ecosystem file
pm2 start ecosystem.config.js

# Start with specific environment
pm2 start ecosystem.config.js --env production

# Restart all apps in ecosystem
pm2 restart ecosystem.config.js

# Stop all apps in ecosystem
pm2 stop ecosystem.config.js

An ecosystem file exists to solve the same problem infrastructure-as-code solves elsewhere: a long pm2 start app.js -i 4 --max-memory-restart 500M --name api --env production command line is not reviewable, not diffable, and easy to get subtly wrong when someone retypes it from memory during an incident. Check the ecosystem file into version control alongside the application code, and every deploy — by any team member, from any machine — starts the app with identical settings.

There is one gotcha in the config above that catches nearly everyone the first time: env_production does not merge with env. When you run pm2 start ecosystem.config.js --env production, PM2 uses env_production as the complete environment for that process — any key defined only in env (say, a LOG_LEVEL you forgot to duplicate) is simply absent, not inherited. If your production environment block is missing a variable your code depends on, you’ll find out at runtime, not at deploy time, because Node doesn’t validate environment variables up front. The safe habit is to either duplicate every key across every env_* block, or better, load shared, non-environment-specific config from a .env file (with dotenv) and reserve the env_* blocks in the ecosystem file strictly for values that actually change between environments (NODE_ENV, PORT, connection strings).

Monitoring

# Real-time monitoring
pm2 monit

# Show detailed info
pm2 show my-app

# Show logs
pm2 logs

# Show logs for specific app
pm2 logs my-app

# Show only errors
pm2 logs --err

# Flush all logs
pm2 flush

# Dashboard (web UI)
pm2 plus

pm2 monit is a terminal dashboard that shows live CPU and memory usage per instance — genuinely useful for spotting a runaway worker during a load test, but it only reflects what’s happening on the single machine you’re SSHed into, and it stops recording the moment you close the terminal, so it’s not a substitute for real monitoring history. pm2 show <app> is the command to reach for when a process keeps restarting: it prints the restart count, uptime, exec mode, and the exact script path and arguments PM2 launched it with, which is usually enough to spot a misconfigured ecosystem.config.js entry without digging through logs first.

Logging

# View logs
pm2 logs

# View last 100 lines
pm2 logs --lines 100

# Real-time logs
pm2 logs --raw

# JSON logs
pm2 logs --json

# Log file locations
~/.pm2/logs/app-out.log    # stdout
~/.pm2/logs/app-error.log  # stderr

By default, PM2 writes every worker’s stdout/stderr into the same log file per app, which means logs from cluster instance 0 and instance 3 are interleaved line by line with no separator. That’s fine for console.log calls that are naturally single-line, but if your app logs multi-line stack traces or pretty-printed JSON, interleaving from concurrent workers can shuffle lines from different requests together and make a stack trace unreadable. pm2 logs --json switches to structured, single-line-per-entry output specifically to avoid this — it’s the format you want if logs are shipped to something like Loki, CloudWatch, or an ELK stack, since structured logs parse reliably even when multiple workers write concurrently, whereas free-text multi-line logs generally don’t.

A separate, easy-to-miss operational risk: PM2 does not rotate or cap these log files on its own. A busy production app can produce gigabytes of ~/.pm2/logs/*.log within weeks, and disk-full is a real failure mode that takes down every process on the host, not just the noisy one — this is exactly what section “Log Rotation” below addresses, and it’s worth setting up before you need it rather than after a disk-full incident.

Custom Log Files

// ecosystem.config.js
module.exports = {
  apps: [{
    name: 'api',
    script: './app.js',
    error_file: './logs/api-error.log',
    out_file: './logs/api-out.log',
    merge_logs: true,
    log_date_format: 'YYYY-MM-DD HH:mm:ss',
  }],
};

merge_logs: true is what keeps multiple cluster instances writing to the same out_file/error_file pair instead of PM2’s default behavior of appending an instance index to the filename (api-out-0.log, api-out-1.log, and so on). Merging is usually what you want operationally — one file to tail per app — but remember it reintroduces the interleaving concern from above, so pair it with structured (JSON) log lines if you can.

Log Rotation

# Install PM2 log rotate
pm2 install pm2-logrotate

# Configure rotation
pm2 set pm2-logrotate:max_size 10M        # Max size per file
pm2 set pm2-logrotate:retain 7            # Keep 7 files
pm2 set pm2-logrotate:compress true       # Compress old logs
pm2 set pm2-logrotate:rotateInterval '0 0 * * *'  # Rotate daily

pm2-logrotate is itself a PM2 module — it runs as another managed process under the same daemon, watching log file sizes and the cron schedule you configure, and it’s the standard answer to the “logs fill the disk” problem above. max_size caps each individual file before rotation kicks in regardless of the schedule, retain bounds how many rotated files accumulate (older ones are deleted), and compress gzips rotated files to shrink their footprint further. Treat this as a required step for any PM2 deployment that will run for more than a few days, not an optional extra — it’s cheap to set up once and expensive to discover missing during an incident.

Auto-Restart on Server Reboot

# Save current PM2 process list
pm2 save

# Generate startup script
pm2 startup

# Follow instructions to run the generated command
# Example: sudo env PATH=$PATH:/usr/bin pm2 startup systemd -u user --hp /home/user

# Now PM2 will auto-start on server reboot

# To remove startup script
pm2 unstartup

pm2 startup generates and registers an OS-level service definition (systemd on most modern Linux distributions, launchd on macOS) that starts the PM2 daemon on boot; pm2 save is what tells that daemon which processes to bring back — it serializes the current process list (name, script, args, working directory, cluster settings) to ~/.pm2/dump.pm2. These two steps are independent and both required: pm2 startup alone gets you an empty PM2 daemon after a reboot with nothing running, and pm2 save without pm2 startup means the saved list is never consulted because nothing launches PM2 in the first place. A very common operational mistake is adding a new app with pm2 start and forgetting to re-run pm2 save — the app runs fine today, but disappears silently after the next server reboot because it was never captured in the dump file.

One more gotcha worth flagging if you manage Node versions with nvm: the startup script pm2 startup generates hardcodes the current Node binary’s absolute path at generation time. If you later switch Node versions with nvm use (or nvm picks a different default), the systemd unit can still point at the old, now-stale path, and a reboot can fail to start PM2 at all with a confusing “command not found” in the systemd journal. Re-run pm2 unstartup followed by pm2 startup after any Node version change on a machine that relies on boot-time auto-start.

Zero-Downtime Reload

# Reload app (zero downtime, cluster mode only)
pm2 reload my-app

# Graceful reload (waits for existing requests)
pm2 reload my-app --wait-ready

# Restart (downtime)
pm2 restart my-app

This is the section most worth reading carefully, because reload and restart sound interchangeable and are not. pm2 restart kills the process (or all instances, for a clustered app) and starts fresh replacements — simple, but every in-flight request to that instance is dropped, and for a single-instance (fork mode) app there’s a real gap with zero listeners on the port. pm2 reload only produces zero downtime in cluster mode: PM2 brings up one new worker, waits for it to be listening, then tears down one old worker, and repeats this one-at-a-time rolling replacement across all instances — at every point in the sequence, at least one worker is still accepting connections. In fork mode (a single instance, no exec_mode: 'cluster'), there is no second worker to hold traffic during the swap, so pm2 reload on a fork-mode app silently degrades to the same behavior as pm2 restart — this is a frequent source of confusion when someone assumes reload always means zero downtime and it quietly didn’t.

--wait-ready makes the rolling handoff explicit rather than time-based: PM2 will not consider a new worker “up” and proceed to kill the next old worker until the new process calls process.send('ready') itself (see the graceful shutdown example below for the matching half of this handshake). Without --wait-ready, PM2 uses a fixed delay heuristic instead, which can prematurely kill an old worker before the new one has finished expensive startup work like warming a cache or establishing a database connection pool — a real gap in availability that looks like a zero-downtime deploy on paper but isn’t one in practice under load.

sequenceDiagram
    participant D as Deploy script
    participant P as PM2 daemon
    participant W1 as Worker (old)
    participant W2 as Worker (new)
    D->>P: pm2 reload app --wait-ready
    P->>W2: spawn replacement
    W2-->>P: process.send('ready')
    P->>W1: SIGINT (drain)
    W1-->>W1: finish in-flight requests
    W1-->>P: exit(0)
    P->>P: repeat for next worker

Graceful Shutdown

// app.js
const express = require('express');
const app = express();

const server = app.listen(3000);

// Listen for shutdown signal
process.on('SIGINT', () => {
  console.log('Received SIGINT, graceful shutdown...');
  
  server.close(() => {
    console.log('Closed all connections');
    process.exit(0);
  });
  
  // Force shutdown after 30s
  setTimeout(() => {
    console.error('Forced shutdown');
    process.exit(1);
  }, 30000);
});

This handler is not optional boilerplate — it’s the piece of application code that makes pm2 reload actually zero-downtime rather than merely appearing to be. Without it, a SIGINT (which is what PM2 sends when retiring an old worker during a reload) kills the process immediately, terminating whatever request happened to be in-flight on that connection mid-response — the client sees a reset connection. server.close() tells Node’s HTTP server to stop accepting new connections but keep serving requests already in progress, and only exit once those finish naturally, which is exactly the “drain” behavior a load balancer expects during a rolling deploy. The 30-second setTimeout is a safety valve: if a request hangs indefinitely (a slow downstream call, a stuck database query), the process still exits eventually rather than blocking a reload forever — tune that duration to be comfortably longer than your slowest legitimate request, not shorter.

Environment Variables

# Set environment variables
pm2 start app.js --env production

# Or via ecosystem file
pm2 start ecosystem.config.js --env production
// ecosystem.config.js
module.exports = {
  apps: [{
    name: 'api',
    script: './app.js',
    env: {
      NODE_ENV: 'development',
      PORT: 3000,
      DB_HOST: 'localhost',
    },
    env_production: {
      NODE_ENV: 'production',
      PORT: 8080,
      DB_HOST: 'prod-db.example.com',
    },
    env_staging: {
      NODE_ENV: 'staging',
      PORT: 8080,
      DB_HOST: 'staging-db.example.com',
    },
  }],
};

Named env_* blocks like env_staging are selected with a matching --env staging flag on the CLI — the suffix after env_ is exactly the string you pass. This pattern is convenient for values that genuinely differ per environment (hostnames, ports, feature flags), but it’s a poor fit for secrets. ecosystem.config.js is ordinary JavaScript that typically lives in version control, so a database password or API key written directly into an env_production block ends up in your git history — readable by anyone with repository access, and effectively permanent even if you later delete the line, since it remains in earlier commits. Prefer loading secrets at runtime from a .env file excluded via .gitignore (read with dotenv), from your platform’s secret manager, or from CI/CD pipeline variables injected at deploy time, and keep the ecosystem file itself limited to non-sensitive configuration.

TypeScript Support

# Install ts-node
npm install -D ts-node typescript

# Start TypeScript app
pm2 start app.ts --interpreter ts-node

# Or compile first (recommended for production)
npm run build
pm2 start dist/app.js
// ecosystem.config.js for TypeScript
module.exports = {
  apps: [{
    name: 'api',
    script: './src/app.ts',
    interpreter: 'ts-node',
    watch: ['src'],
    ignore_watch: ['node_modules', 'dist'],
    env: {
      TS_NODE_PROJECT: './tsconfig.json',
    },
  }],
};

Both approaches run the same application; the difference is when TypeScript gets compiled. Running through ts-node type-checks and transpiles on every process start — convenient in development where you’re restarting constantly and want changes reflected instantly, but that transpilation cost is paid again on every single PM2 restart in production, including automatic crash-restarts, which adds real latency to your restart-to-serving-traffic window exactly when you can least afford it (a process crash-looping under ts-node is slower to recover than one running precompiled JavaScript). The “compile first” path — npm run build then pm2 start dist/app.js — pays the compilation cost once at build/deploy time, and every restart afterward launches plain, already-checked JavaScript with no additional overhead. Reserve --interpreter ts-node for local development and treat a build step as a required part of any production deploy pipeline.

Memory Management

// ecosystem.config.js
module.exports = {
  apps: [{
    name: 'api',
    script: './app.js',
    max_memory_restart: '500M', // Restart if memory exceeds 500MB
    max_restarts: 10,           // Max restarts within restart_delay
    min_uptime: '10s',          // Min uptime before considered stable
    restart_delay: 4000,        // Delay between restarts (ms)
  }],
};

It’s important to be clear-eyed about what max_memory_restart actually does: it is a circuit breaker, not a fix. It restarts a process once its resident memory crosses the threshold, which reclaims the leaked memory and keeps the host from running out of RAM — but it does nothing to address whatever is leaking in the application code. If you find yourself relying on max_memory_restart to keep a process alive, that’s a signal to profile the app (heap snapshots via --inspect, or a tool like clinic.js) rather than a signal that the configuration is doing its job. It’s a legitimate safety net for genuinely bounded, expected memory growth (caches with eviction, for instance) but a band-aid for an actual leak.

min_uptime and max_restarts exist together to prevent the opposite failure mode: a crash loop. If a process keeps dying within min_uptime of starting (the default treats anything under a few seconds as “didn’t really start”), PM2 counts that toward max_restarts, and once that limit is hit within the window, PM2 stops trying and marks the app errored rather than restarting forever. Without this guard, a genuinely broken deploy (bad config, missing environment variable, syntax error slipping past a lint step) would otherwise consume CPU in an infinite fork-and-crash loop, which is its own kind of outage. restart_delay adds a fixed pause between restart attempts, which is what actually slows the loop down — without it, PM2 restarts as fast as the process can die, and a broken app can spin through max_restarts attempts in well under a second.

Monitoring with PM2 Plus

# Link to PM2 Plus (free tier available)
pm2 plus

# Or with specific bucket
pm2 link <secret_key> <public_key>

# Disconnect
pm2 unlink

PM2 Plus provides:

  • Real-time dashboard
  • Error tracking
  • Transaction tracing
  • Custom metrics
  • Email/Slack alerts

PM2 Plus is a separate, hosted SaaS product from the same team behind open-source PM2 (now marketed as PM2.io) — pm2 plus or pm2 link opts a machine into sending its process metrics to that external service over the network, which is worth being deliberate about rather than running reflexively on a production host, especially in environments with data-residency or vendor-approval constraints. The open-source CLI (pm2 monit, pm2 show, pm2 logs) covers point-in-time debugging on a single box without any external dependency; PM2 Plus adds what the CLI structurally can’t — historical trend graphs, alerting when nobody’s watching a terminal, and a unified view across a fleet of machines — at the cost of a third party holding that telemetry. If sending metrics off-host isn’t acceptable for your environment, the free CLI tools plus a self-hosted stack (Prometheus scraping via a PM2 metrics exporter, visualized in Grafana) covers most of the same ground without leaving your infrastructure.

Docker Integration

# Dockerfile
FROM node:18-alpine

WORKDIR /app

COPY package*.json ./
RUN npm ci --only=production

COPY . .

RUN npm install -g pm2

EXPOSE 3000

CMD ["pm2-runtime", "ecosystem.config.js"]
# docker-compose.yml
version: '3.8'
services:
  api:
    build: .
    ports:
      - "3000:3000"
    environment:
      - NODE_ENV=production
    restart: unless-stopped

Running PM2 inside a container is a genuinely debated choice, and it’s worth understanding the trade-off rather than cargo-culting it into every Dockerfile. If your orchestrator already provides process supervision and horizontal scaling — Kubernetes restarting a crashed pod, or Docker Compose’s restart: unless-stopped bringing a container back — then PM2’s restart supervisor is partially redundant with what the platform already does, and running it adds another moving part and a bit of memory overhead per container. Where PM2 still earns its place in a container is cluster mode within a single container: if that container is granted more than one CPU, a single node app.js process still only uses one core, and PM2’s -i max lets one container use all the CPU it was allocated instead of requiring you to run one container per core and load-balance across them at the orchestrator level.

The one detail in the Dockerfile above that’s easy to get wrong is pm2-runtime versus plain pm2 start as the container’s CMD. pm2 start is designed for interactive/daemon use — it forks the PM2 daemon into the background and the launching command exits immediately, which in a container means the process Docker was watching as PID 1 just exited, so the container stops right after starting (or, if you background it wrong, the container’s PID 1 becomes something that doesn’t correctly forward Unix signals to your app). pm2-runtime is the variant built specifically for containers: it runs the PM2 process manager in the foreground as PID 1, correctly forwards SIGTERM/SIGINT from docker stop down to your application (which is what lets the graceful shutdown handler from section 8 actually run before the container is killed), and exits with your app’s exit code so orchestrator-level restart policies see accurate failure signals. Using pm2 start as a container CMD is a common enough mistake that it’s worth calling out explicitly — always use pm2-runtime for containerized deployments.

Keep singleton workers in fork mode

module.exports = {
  apps: [{
    name: 'queue-worker',
    script: './worker.js',
    instances: 1,          // Single instance
    exec_mode: 'fork',     // Fork mode
  }],
};

Fork mode isn’t just “cluster mode with one instance” — it’s the right default for anything that shouldn’t be duplicated. A queue consumer, a scheduled job runner, or anything else where running N copies in parallel would mean N workers all racing to pull the same job off a queue (or worse, all running the same cron job N times) belongs in fork mode with instances: 1, precisely because cluster mode’s whole purpose is duplicating identical stateless workers behind a shared port — a purpose that actively works against a singleton worker’s correctness.

Common Commands

# Status
pm2 status          # List all processes
pm2 list            # Same as status
pm2 show <app>      # Detailed info

# Logs
pm2 logs            # All logs
pm2 logs <app>      # App logs
pm2 logs --lines 200  # Last 200 lines
pm2 flush           # Clear logs

# Control
pm2 start <app>
pm2 stop <app>
pm2 restart <app>
pm2 reload <app>    # Zero-downtime (cluster only)
pm2 delete <app>

# Monitoring
pm2 monit           # Real-time monitor
pm2 plus            # Web dashboard

# Maintenance
pm2 save            # Save process list
pm2 startup         # Auto-start on boot
pm2 resurrect       # Restore saved processes
pm2 update          # Update PM2 in-memory
pm2 reset <app>     # Reset restart counter

Keep this section as a quick-reference cheat sheet rather than a set of commands to memorize in order — in day-to-day operation you’ll mostly reach for pm2 logs <app> and pm2 show <app> to diagnose a problem, then one of restart/reload to fix it, with save as the habit to remember any time the running process list changes.

Real-World Example

// ecosystem.config.js
module.exports = {
  apps: [
    // API server (cluster mode)
    {
      name: 'api',
      script: './dist/api/server.js',
      instances: 'max',
      exec_mode: 'cluster',
      env_production: {
        NODE_ENV: 'production',
        PORT: 8080,
        DB_HOST: process.env.DB_HOST,
        REDIS_URL: process.env.REDIS_URL,
      },
      error_file: './logs/api-error.log',
      out_file: './logs/api-out.log',
      max_memory_restart: '1G',
      min_uptime: '10s',
      max_restarts: 10,
    },
    
    // Background worker (fork mode)
    {
      name: 'worker',
      script: './dist/worker/index.js',
      instances: 2,
      exec_mode: 'fork',
      env_production: {
        NODE_ENV: 'production',
        REDIS_URL: process.env.REDIS_URL,
      },
      error_file: './logs/worker-error.log',
      out_file: './logs/worker-out.log',
      max_memory_restart: '500M',
      cron_restart: '0 3 * * *', // Restart at 3 AM daily
    },
    
    // Cron job
    {
      name: 'cleanup',
      script: './dist/jobs/cleanup.js',
      instances: 1,
      cron_restart: '0 0 * * *', // Run at midnight
      autorestart: false,
    },
  ],
};
# Deploy
npm run build
pm2 start ecosystem.config.js --env production
pm2 save
pm2 startup

# Update
git pull
npm run build
pm2 reload ecosystem.config.js --env production

This config is a reasonable template for a typical Node.js service and it’s worth walking through why each process is configured the way it is, since the three entries deliberately don’t share settings. The api process is in cluster mode with instances: 'max' because it’s a stateless HTTP server — exactly the shape cluster mode is built for — and its environment pulls DB_HOST/REDIS_URL from process.env rather than hardcoding them, so the same ecosystem file works unchanged across machines with different connection strings, as long as those variables are set in the deploy environment (see the secrets discussion in section 9). The worker process explicitly uses exec_mode: 'fork' with instances: 2 — two independent workers rather than a duplicated cluster pool, appropriate if the workload is, say, two separate queue consumers rather than N identical replicas of the same singleton. Note also worker’s cron_restart: '0 3 * * *': a periodic restart isn’t there to recover from a crash, it’s a deliberate, scheduled recycle — useful for clearing accumulated memory fragmentation or rotating long-lived connections during a low-traffic window, on a service where you’ve decided that’s cheaper than chasing down the root cause.

The cleanup entry is the one most likely to confuse someone reading this config for the first time: autorestart: false combined with cron_restart means this process is expected to run to completion and exit normally every time, and PM2 should not treat that exit as a crash requiring a restart — it should simply wait for the next scheduled cron_restart trigger. Leaving autorestart at its default (true) on a cron-style job that’s designed to finish and exit would make PM2 immediately relaunch it in a tight loop right after each successful, expected exit — a textbook case of a restart policy fighting the job’s own intended lifecycle rather than protecting it.



Frequently Asked Questions (FAQ)

Q. Should I still use PM2 if my app already runs inside Kubernetes?

A. Usually not for restart supervision — Kubernetes already restarts a crashed pod based on its own liveness/readiness probes, and running PM2’s restart logic on top of that is redundant. Where PM2 can still help inside a pod is cluster mode: if a pod is allocated more than one CPU, a bare node process still only uses one core, so pm2-runtime start app.js -i max inside the container can use the full CPU allocation without needing to run multiple pods per node just to spread load across cores. If you do run PM2 in a pod, always launch it with pm2-runtime, not pm2 start, so it stays PID 1 and forwards signals correctly.

Q. My pm2 reload still causes dropped connections — why?

A. The two most common causes are running in fork mode (where reload silently degrades to a hard restart, since there’s no second instance to hold traffic — see section 8) and missing a graceful shutdown handler in the application itself. Even in cluster mode, if your app doesn’t listen for SIGINT/SIGTERM and call server.close() before exiting, PM2 has no way to know an in-flight request should be allowed to finish — the process just dies mid-response. Pair exec_mode: 'cluster' with the graceful shutdown pattern from section 8, and consider --wait-ready if your app has slow startup work that needs to finish before it should receive traffic.

Q. Is it safe to put database passwords directly in ecosystem.config.js?

A. No — treat the ecosystem file as code, not as a secrets store. It’s ordinary JavaScript that’s typically committed to version control, so anything written into an env or env_production block ends up in git history, readable by anyone with repo access, indefinitely (deleting the line later doesn’t remove it from earlier commits). Load secrets at runtime instead — a .env file excluded via .gitignore, your platform’s secret manager, or CI/CD-injected variables — and keep the ecosystem file limited to configuration that isn’t sensitive.