Deploying Node.js on EC2: Security Groups, IAM Roles, IMDSv2, systemd and Nginx

Key takeaways

A single EC2 instance is still a reasonable home for a Node.js API, but most of the work is not installing Node. It is network rules, credentials, instance metadata, disk type, CPU credits and a process supervisor. This guide walks through each decision and the errors you hit when one is wrong.

This guide covers a Node.js API on a single EC2 instance. Installing Node takes five minutes. The rest of the work is a handful of decisions (network rules, credentials, instance metadata, disk, CPU credits, the process supervisor) that decide whether the box stays up, stays cheap and stays out of an incident report. I go through them in the order you meet them, with the errors you see when one is wrong.

One instance is a real architecture for internal tools, early-stage products and APIs with steady, modest traffic. It has no redundancy, though. If the instance or its Availability Zone has a bad day, you are down. Once you need that redundancy, the same pieces go behind an Application Load Balancer and an Auto Scaling group, or into containers on ECS.


Network rules: security groups first, NACLs rarely

Two different firewalls sit in front of the instance, and they behave differently:

Security groupNetwork ACL
Attached toNetwork interface (instance)Subnet
StateStateful: return traffic is allowed automaticallyStateless: return traffic needs its own rule
RulesAllow onlyAllow and deny, evaluated in rule-number order
Typical usePer-application access controlCoarse subnet-wide blocks

Do almost all of your filtering in security groups. The default NACL allows everything, and that is fine for most setups. People who add custom NACLs usually add an inbound allow for 443 and forget that NACLs are stateless. The response then goes back to the client’s ephemeral port (1024–65535), finds no outbound rule, and the connection hangs. From the outside it looks exactly like a broken security group.

A web instance needs this security group:

  • Inbound 443 and 80 from 0.0.0.0/0 (and ::/0 if you serve IPv6). Port 80 only redirects to HTTPS and answers the Let’s Encrypt HTTP-01 challenge.
  • No inbound 3000. Node listens on localhost, and Nginx is the only thing talking to it.
  • No inbound 22 if you use Session Manager (next section). Otherwise, only your own /32.

Security groups can also reference other security groups. For example, “allow 5432 from the security group of the web instances” on the database. Use that instead of IP ranges when both ends are in the same VPC. The rule then keeps working when instance IPs change.


Getting a shell: Session Manager instead of an open port 22

Scanners hit a port 22 open to the internet within minutes of launch. Key-only authentication stops them from logging in, but it still leaves a key file to manage, a port to watch and no central log of who connected.

AWS Systems Manager Session Manager gives you a shell with no inbound rule at all. The SSM agent on the instance makes an outbound HTTPS connection to the SSM service, and you authenticate with IAM:

  1. Use an AMI that ships the SSM agent (Amazon Linux 2023 and the official Ubuntu AMIs both include it).
  2. Attach an instance role that has the managed policy AmazonSSMManagedInstanceCore.
  3. Make sure the instance can reach the SSM endpoints on 443, through a public IP, a NAT gateway, or VPC interface endpoints for ssm, ssmmessages and ec2messages in a private subnet.
# Requires the Session Manager plugin for the AWS CLI on your machine
aws ssm start-session --target i-0123456789abcdef0

# Port-forward the app's port to your laptop for debugging, still with no inbound rule
aws ssm start-session --target i-0123456789abcdef0 \
  --document-name AWS-StartPortForwardingSession \
  --parameters '{"portNumber":["3000"],"localPortNumber":["3000"]}'

If the instance does not appear in Fleet Manager, it is almost always one of those three: no role, the wrong policy, or no outbound path to the endpoints.

When you do use SSH: “Permission denied (publickey)”

This error means the TCP connection worked and authentication failed. (A timeout instead points to the security group, route table or NACL.) Common causes, in order of frequency:

  • Wrong username. ec2-user on Amazon Linux, ubuntu on Ubuntu, admin on Debian. SSH tries your key for a user that has no such key and reports it as a key failure.
  • Wrong key. The instance only trusts the key pair it was launched with (or keys you added later to ~/.ssh/authorized_keys). ssh -v shows which identities were offered.
  • Key file permissions. WARNING: UNPROTECTED PRIVATE KEY FILE! followed by bad permissions means the client refused to use the key at all. Run chmod 400 key.pem.
  • Broken authorized_keys on the server, often after someone edited it or changed ownership of the home directory. EC2 Instance Connect or Session Manager is the way back in.

Credentials: an instance role, never access keys on the box

A common first deployment runs aws configure on the instance, or puts AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY into .env. Those keys are long-lived. They end up in backups, AMIs, shell history, and sometimes a git commit. If one leaks, it works from anywhere until someone revokes it.

Attach an IAM role to the instance instead (through an instance profile). The AWS SDK for JavaScript v3 finds it automatically through the default credential provider chain. It fetches temporary credentials from the instance metadata service and refreshes them before they expire. The application code needs no credential configuration:

import { S3Client, GetObjectCommand } from '@aws-sdk/client-s3';

// No keys here: on EC2 the default provider chain uses the instance role
const s3 = new S3Client({ region: 'ap-northeast-2' });

Scope the role’s policy to what the app actually does, such as specific bucket ARNs or specific parameter paths. Do not attach AdministratorAccess to get past a permissions error.

Application secrets (database password, JWT signing key) belong in SSM Parameter Store (SecureString) or Secrets Manager. The same role reads them:

aws ssm get-parameter --name /my-app/prod/database-url \
  --with-decryption --query Parameter.Value --output text

A plain .env file with chmod 600 is acceptable for a small setup, as long as it is created on the box and never committed. The parameter store adds rotation, auditing, and one place to change a value without logging in.


IMDSv2: require tokens, and know about the hop limit

The instance metadata service at 169.254.169.254 hands out the role credentials above. With IMDSv1, any process that can make an HTTP GET from the instance can read them. That includes an SSRF bug in your app that fetches user-supplied URLs. IMDSv2 requires a session token obtained with a PUT first, and that blocks most SSRF-style attacks:

TOKEN=$(curl -sX PUT "http://169.254.169.254/latest/api/token" \
  -H "X-aws-ec2-metadata-token-ttl-seconds: 21600")
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
  http://169.254.169.254/latest/meta-data/instance-id

Many newer AMIs, Amazon Linux 2023 among them, already launch with tokens required. Check and enforce it on existing instances:

aws ec2 modify-instance-metadata-options \
  --instance-id i-0123456789abcdef0 \
  --http-tokens required --http-endpoint enabled

Once tokens are required, a request without one returns 401 Unauthorized. Current AWS SDKs speak IMDSv2 already. Very old SDK versions and hand-written curl scripts are what break.

The hop limit trips people up. The token response has an IP TTL (HttpPutResponseHopLimit), 1 by default in many configurations. A process in a Docker container on a bridge network is one hop further away, so its PUT succeeds but the response never arrives, and the SDK times out looking for credentials. Set the limit to 2 on instances that run containers:

aws ec2 modify-instance-metadata-options \
  --instance-id i-0123456789abcdef0 --http-put-response-hop-limit 2

I have watched a team enforce IMDSv2 across an account as a security task on a Friday, and on Monday their containerized workers were all failing with credential errors from the SDK. The processes running directly on the host were fine. It took a while to connect the two, because nothing in the error mentions metadata or hop limits. If you run containers on EC2, check that setting the same day you enforce tokens.


Instance type, disk and IP address: the cost traps

Burstable T instances (t3, t3a, t4g) are cheap because they assume you idle most of the time. Each size has a baseline CPU percentage. Below it you earn CPU credits, above it you spend them. What happens when credits run out depends on the credit mode:

  • Standard: the instance is throttled to the baseline. An API that was fast yesterday suddenly has high latency under the same load, and top inside the box looks confusing, because the hypervisor is doing the throttling.
  • Unlimited (the default for T3/T3a/T4g): no throttling, but sustained use above baseline is billed as surplus credits. At that point you might as well use a non-burstable type.

Watch CPUCreditBalance and CPUSurplusCreditBalance in CloudWatch. A balance that trends to zero every afternoon means the workload has outgrown the T family. For CPU-heavy Node work (image processing, heavy JSON, SSR), an m7g or c7g class instance is often more predictable. Graviton (arm64) types are fine for Node.js as long as your native dependencies publish arm64 builds.

EBS volumes: choose gp3, not gp2. gp2 performance scales with size (3 IOPS per GiB, with burst credits for small volumes), so a small gp2 root volume can run out of I/O burst during a big npm ci or log flood. gp3 gives a baseline of 3,000 IOPS and 125 MiB/s at any size, you can provision more independently of capacity, and AWS prices it lower per GiB than gp2. Existing volumes convert in place without detaching:

aws ec2 modify-volume --volume-id vol-0123456789abcdef0 --volume-type gp3

Public IPv4 is billed. Since February 2024, AWS charges for every public IPv4 address, both auto-assigned ones and Elastic IPs, whether attached or not ($0.005 per hour per address at the time of the announcement). An Elastic IP you allocated for a test instance and forgot keeps billing. An Elastic IP is still useful on a single instance so the address survives stop/start (auto-assigned public IPs change), but release it when the instance goes away. Behind a load balancer, the instances usually need no public IP at all.


Installing Node.js and the app

On Ubuntu 24.04, install a current LTS line from NodeSource. Node 20 reached end of life in April 2026, so use 22 or 24:

sudo apt-get update && sudo apt-get upgrade -y
curl -fsSL https://deb.nodesource.com/setup_24.x | sudo -E bash -
sudo apt-get install -y nodejs git
node --version

Run the app as its own unprivileged user, not ubuntu and never root:

sudo useradd --system --create-home --home-dir /srv/my-app --shell /usr/sbin/nologin myapp
sudo -u myapp git clone https://github.com/yourorg/my-app.git /srv/my-app/current
cd /srv/my-app/current
sudo -u myapp npm ci --omit=dev     # --only=production is the deprecated spelling

If the app needs a build step (TypeScript, bundling), either build in CI and ship the artifact, or install dev dependencies, build, then prune. Building on a t3.micro with 1 GiB of RAM is a common way to get the OOM killer in the middle of a deploy.

Make Node listen on 127.0.0.1, not 0.0.0.0. Then even a security group mistake does not expose the unproxied port:

app.listen(process.env.PORT ?? 3000, '127.0.0.1');

Keeping it running: systemd or PM2

Something has to start the app at boot and restart it when it crashes. Pick one supervisor.

systemd (already on the box)

# /etc/systemd/system/my-app.service
[Unit]
Description=my-app Node.js API
After=network-online.target
Wants=network-online.target

[Service]
User=myapp
WorkingDirectory=/srv/my-app/current
EnvironmentFile=/etc/my-app/env
ExecStart=/usr/bin/node server.js
Restart=on-failure
RestartSec=2
# Basic sandboxing
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/srv/my-app/current/tmp
PrivateTmp=true

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now my-app
journalctl -u my-app -f          # logs, with rotation handled by journald

ProtectSystem=strict makes the filesystem read-only for the service except the paths you list. If the app writes uploads or a cache somewhere, add that path, or you will see EROFS: read-only file system in the logs. The weak spot is deploys: systemctl restart stops the old process before the new one listens, so you get a second or two of 502s from Nginx.

PM2

PM2 runs several Node workers in cluster mode and can pm2 reload them one at a time, so a deploy causes no errors. The trade-off is another global npm package to keep updated, and its own log files to rotate (the pm2-logrotate module). Use pm2 startup so that systemd starts PM2 at boot, and pm2 save after changing the process list. Otherwise the app is gone after the next reboot. The PM2 production guide covers ecosystem files and reload behaviour in detail.

On a 2-vCPU instance, cluster mode gives you two workers, and on a burstable instance both of them draw from the same CPU credit balance. Cluster mode is mostly about deploys without errors and surviving a single crashing worker. Do not expect it to double throughput on small instances.


Nginx in front, and HTTPS

Nginx terminates TLS, serves static files, buffers slow clients and proxies the rest to Node on localhost:

# /etc/nginx/conf.d/websocket-map.conf
map $http_upgrade $connection_upgrade {
    default upgrade;
    ''      close;
}
# /etc/nginx/sites-available/my-app
server {
    listen 80;
    server_name api.example.com;

    location /static/ {
        alias /srv/my-app/current/public/;
        expires 7d;
    }

    location / {
        proxy_pass http://127.0.0.1:3000;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection $connection_upgrade;
        proxy_set_header Host $host;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

The map block avoids a common copy-paste mistake. Hard-coding Connection 'upgrade' sends an upgrade header on every request, not only WebSocket handshakes. Use $connection_upgrade so that normal requests get proper connection handling.

Only set a long cache lifetime such as expires 1y with immutable on files whose names contain a content hash. On /static/logo.png it means browsers keep the old logo for a year.

sudo ln -s /etc/nginx/sites-available/my-app /etc/nginx/sites-enabled/
sudo nginx -t && sudo systemctl reload nginx

sudo apt-get install -y certbot python3-certbot-nginx
sudo certbot --nginx -d api.example.com
sudo certbot renew --dry-run

Behind Nginx, Express sees every request coming from 127.0.0.1 unless you tell it to trust the proxy. Set app.set('trust proxy', 'loopback') so that req.ip, req.protocol and rate limiters keyed by IP use the forwarded values. Trust only the loopback proxy, not everything, or clients can spoof X-Forwarded-For. The Node.js + Nginx reverse proxy guide goes deeper on timeouts and buffering.

If you later put an Application Load Balancer in front, you can terminate TLS there with a free ACM certificate. Nginx then becomes optional.


Deploying updates without surprises

The minimum safe deploy on a single box:

cd /srv/my-app/current
sudo -u myapp git fetch --all && sudo -u myapp git checkout <release-tag>
sudo -u myapp npm ci --omit=dev
sudo systemctl restart my-app        # or: pm2 reload my-app
curl -fsS http://127.0.0.1:3000/health || echo "health check failed"

Deploy a tag or commit, not git pull on whatever main is, so you can roll back by checking out the previous tag. Better still, deploy into release directories (/srv/my-app/releases/<sha>) and switch a current symlink. Rollback is then a symlink swap and a restart, with no reinstall.

Also take an EBS snapshot, or better, keep the instance reproducible from a launch template plus user data. The most common single-instance disaster I know of is not a hack. It is a box that was configured by hand over two years, and nobody can rebuild it when it stops booting after a kernel update. Test the rebuild path once while the old instance still works.


When a single EC2 instance is the wrong choice

  • You need no downtime during instance failures or deploys: use an ALB plus an Auto Scaling group across at least two AZs. Everything above still applies, baked into a launch template.
  • Traffic is spiky or near zero most of the day: Lambda behind API Gateway or a Lambda function URL costs nothing when idle. The trade-offs are cold starts, a 15-minute execution limit and no long-lived connections. See the Lambda guide.
  • You want someone else to patch the OS: Elastic Beanstalk, App Runner or ECS on Fargate take over instance management, at the cost of less control and their own configuration quirks.

The EC2 settings behind most incidents

  • Filter with security groups. NACLs are stateless and need ephemeral-port rules for return traffic.
  • Use Session Manager instead of an open port 22. If you use SSH, “Permission denied (publickey)” is usually the wrong username or key, and a timeout is a network problem.
  • Give the instance an IAM role. Access keys do not belong on the box. Keep secrets in Parameter Store or Secrets Manager.
  • Require IMDSv2, and set the hop limit to 2 where containers need credentials.
  • Watch T-instance CPU credits, use gp3 volumes, and release unused Elastic IPs, because public IPv4 is billed.
  • Run Node as an unprivileged user on 127.0.0.1 under one supervisor (systemd or PM2), with Nginx and HTTPS in front and trust proxy configured narrowly.