First Steps on AWS: Launching EC2, Storing Files in S3 and Connecting RDS
Key takeaways
A beginner's path through the three AWS services most apps start with: an EC2 server you can SSH into, a private S3 bucket for files, and an RDS PostgreSQL database, plus the security group, IAM and billing details that cause most first-week problems.
What this post covers
Most applications on AWS start with the same three pieces: a server to run code (EC2), a place to keep files (S3) and a database (RDS). This post walks through setting each one up from the console, connecting them from a Node.js app, and the details that cause most first-week problems: security groups, IAM credentials, public access and surprise bills.
It deliberately stays small. Load balancers, auto scaling, containers and infrastructure as code come later, once the basic pieces make sense. Console layouts change often, so treat menu paths as approximate and use the names of the settings to find them.
Regions, Availability Zones and what you pay for
A Region is a geographic area such as us-east-1 (N. Virginia) or ap-northeast-2 (Seoul). Resources live in one Region, and most services do not share anything across Regions unless you configure it. Each Region contains several Availability Zones (AZs): separate data center groups with independent power and networking, close enough for low-latency links. Spreading instances or a database standby across AZs is how you survive one of them failing.
Pick the Region close to your users and keep everything for one project in it. A classic beginner mistake is creating an instance in one Region, switching the console Region selector by accident, and concluding that the instance has disappeared. The selector at the top right is the first thing to check whenever something seems missing.
AWS bills most resources by the hour or second while they exist, not while they are “used”. A stopped EC2 instance does not bill for compute, but its EBS disk, snapshots and any Elastic IP still cost money. An RDS instance bills while it is running whether or not anything connects. That billing model is why cleanup and budget alerts matter more on AWS than on a flat-rate VPS.
Free Tier and budgets
AWS changed its Free Tier for accounts created from mid-2025 onward to a credit-based model with a time-limited free plan, while older accounts kept the previous 12-month allowances. The exact terms, eligible instance types and limits change, so read the current Free Tier page for your account rather than relying on numbers from a tutorial.
Whatever the plan, set up an AWS Budget before you create anything: Billing and Cost Management → Budgets → create a monthly cost budget with an email alert at a low threshold. Budgets evaluate a few times a day, so they warn you about a forgotten resource, not about the first cent.
Account and IAM setup
The email address you sign up with becomes the root user, which can do anything, including closing the account. Enable MFA on it immediately and then stop using it for daily work.
For people, the current recommendation is IAM Identity Center (formerly AWS SSO): create a user there, give it a permission set, and sign in through the access portal. The AWS CLI v2 supports this directly:
aws configure sso
# follow the prompts: start URL, region, account, permission set, profile name
aws s3 ls --profile my-dev
This gives the CLI short-lived credentials that refresh through the browser. The older approach, an IAM user with an access key pasted into aws configure, still works, but that key is valid until someone deletes it. Access keys committed to Git repositories are a well-known way AWS accounts get abused, typically to run expensive compute. If you must create one, never put it in code, and never create one for the root user.
For code running on AWS, do not use keys at all: EC2 instances get an instance profile (an IAM role attached to the instance), Lambda functions get an execution role, and CI systems like GitHub Actions can assume a role through OIDC (shown later in this post). The SDKs find these credentials automatically.
Install AWS CLI v2 from the official installer. On Ubuntu, sudo apt install awscli may give you the older v1 CLI, which lacks aws configure sso and aws logs tail; check with aws --version.
EC2: a virtual server
An EC2 instance is a virtual machine: you choose an operating system image (AMI), a size (instance type), a disk (EBS volume) and a firewall (security group).
Launching an instance
- EC2 console → Launch instance.
- Name:
my-web-server. - AMI: Ubuntu Server LTS (the latest LTS listed) or Amazon Linux.
- Instance type: a small burstable type such as
t3.micro; check which types your account’s free plan covers. - Key pair: create one and download the
.pemfile. AWS does not keep a copy of the private key, so if you lose it you cannot recover it. - Network settings: allow SSH from My IP, and HTTP/HTTPS from anywhere if it will serve a website.
- Storage: the default root volume size is fine to start.
Connecting with SSH
chmod 400 my-key.pem
ssh -i my-key.pem ubuntu@<public-ip>
The username depends on the AMI: ubuntu for Ubuntu, ec2-user for Amazon Linux. Using the wrong one produces Permission denied (publickey), which looks like a key problem but is not.
OpenSSH refuses keys that other users can read:
WARNING: UNPROTECTED PRIVATE KEY FILE!
Permissions 0644 for 'my-key.pem' are too open.
chmod 400 fixes this on Linux and macOS. On Windows, chmod does nothing useful; remove inherited permissions instead, for example icacls my-key.pem /inheritance:r /grant:r "%USERNAME%:R" in cmd, or through the file’s Security properties.
If SSH simply hangs until it times out, the problem is network-level, almost always the security group. “My IP” is your address at the moment you created the rule; home and mobile connections change addresses, and the rule silently stops matching. EC2 Instance Connect and Systems Manager Session Manager are alternatives that avoid exposing port 22 at all.
Two related surprises: the auto-assigned public IP changes whenever you stop and start the instance, and AWS charges an hourly fee for every public IPv4 address. If the instance needs a stable address, allocate an Elastic IP; if it is a throwaway test server, terminate it when you are done rather than just stopping it.
Running an app
Once you are in, running a Node.js app looks like any other Linux server: install a current Node.js LTS, clone your code, run it under a process manager such as PM2 or systemd, and put Nginx in front as a reverse proxy.
sudo apt update && sudo apt install -y nginx
# install Node.js LTS via NodeSource or nvm, then:
git clone https://github.com/your-org/app.git && cd app
npm ci && npm run build
sudo npm install -g pm2
pm2 start npm --name my-app -- start
pm2 startup # prints a sudo command: copy and run it
pm2 save
pm2 startup does not configure boot startup by itself; it prints a sudo env PATH=... pm2 startup systemd ... command that you must run. Skipping that step is why apps often “vanish” after the first reboot. A full walkthrough of this part, including systemd, IAM roles and IMDSv2, is in Deploying Node.js on EC2.
S3: object storage
S3 stores objects (a file plus metadata) under keys in a bucket. It is not a filesystem: there are no real directories (images/cat.jpg is just a key containing a slash), no appends and no partial overwrites. You replace whole objects. In return you get storage that grows without provisioning and is designed for very high durability.
Creating a private bucket
- S3 console → Create bucket.
- Name: globally unique across all AWS accounts, lowercase, e.g.
myapp-uploads-<something-unique>. - Region: the same as your EC2 instance, to avoid cross-Region transfer charges and latency.
- Block Public Access: leave all four settings on.
Keeping Block Public Access on is the most important decision here. Public buckets are how private data ends up exposed, and you rarely need one: browsers can get files through CloudFront or presigned URLs instead.
CLI basics
aws s3 cp image.jpg s3://myapp-uploads/images/
aws s3 sync ./dist s3://myapp-site/ --delete
aws s3 ls s3://myapp-uploads/images/
aws s3 cp s3://myapp-uploads/images/image.jpg ./
sync --delete removes objects in the bucket that are not in your local folder. That is what you want for deploying a static site and exactly what you do not want if you point it at the wrong bucket or the wrong local directory. Add --dryrun the first time.
From Node.js (AWS SDK v3)
npm install @aws-sdk/client-s3 @aws-sdk/s3-request-presigner
import { S3Client, PutObjectCommand, GetObjectCommand } from '@aws-sdk/client-s3';
import { getSignedUrl } from '@aws-sdk/s3-request-presigner';
import { createReadStream, createWriteStream } from 'node:fs';
import { stat } from 'node:fs/promises';
import { pipeline } from 'node:stream/promises';
const s3 = new S3Client({ region: 'ap-northeast-2' });
const Bucket = 'myapp-uploads';
export async function uploadFile(filePath, key, contentType) {
const { size } = await stat(filePath);
await s3.send(new PutObjectCommand({
Bucket,
Key: key,
Body: createReadStream(filePath),
ContentLength: size,
ContentType: contentType,
}));
}
export async function downloadFile(key, outputPath) {
const res = await s3.send(new GetObjectCommand({ Bucket, Key: key }));
await pipeline(res.Body, createWriteStream(outputPath));
}
// A link the browser can use for 5 minutes, without making anything public
export function downloadUrl(key) {
return getSignedUrl(s3, new GetObjectCommand({ Bucket, Key: key }), { expiresIn: 300 });
}
Notice there are no credentials in this code. On EC2 with an instance profile, or locally after aws configure sso, the SDK finds credentials by itself. If it cannot, you get CredentialsProviderError: Could not load credentials from any providers.
Two details in the download function matter. res.Body is a stream, and using pipeline waits until the file is completely written and propagates errors; a bare stream.pipe() returns immediately, so a caller that reads the file right afterward can see a partial file. And set ContentType on upload: S3 does not sniff file types, so an image uploaded without it is served as application/octet-stream and browsers download it instead of displaying it.
When S3 returns AccessDenied, it is often not about the object at all. S3 returns AccessDenied rather than NoSuchKey for missing objects when the caller lacks s3:ListBucket permission, so as not to reveal which keys exist. Check the key spelling and the IAM policy together.
Serving files to browsers: CloudFront
The old way to host a static site on S3 was to enable “static website hosting” and attach a public-read bucket policy. That endpoint only speaks HTTP and requires the bucket to be public. The current approach is a CloudFront distribution with the bucket as origin and Origin Access Control (OAC): CloudFront gets permission to read the bucket, the bucket stays private, and you get HTTPS, a custom domain via an ACM certificate (which must be issued in us-east-1 for CloudFront), and edge caching. Older tutorials use Origin Access Identity (OAI), which AWS now treats as legacy.
After each deployment, cached files at the edge are still the old ones until they expire. Either invalidate (aws cloudfront create-invalidation --distribution-id E123 --paths "/*") or, better, use content-hashed filenames for assets and invalidate only index.html.
RDS: a managed relational database
RDS runs PostgreSQL, MySQL, MariaDB and other engines for you. AWS handles installation, minor version patching, automated backups and, if you enable it, failover to a standby in another AZ. You still own schema design, query performance, connection management and choosing when to do major version upgrades.
Creating a PostgreSQL instance
- RDS console → Create database → Standard create.
- Engine: PostgreSQL, a current major version.
- Templates: Free tier or Dev/Test for learning.
- Master username: the default
postgresis fine; set a strong password or let RDS manage it in Secrets Manager. - Instance class: a small burstable class such as
db.t3.microordb.t4g.micro. - Connectivity: the same VPC as your EC2 instance, Public access: No, and a new security group.
With public access off, the database has no internet-reachable address, and only resources in the VPC can connect. That is what you want for anything that holds real data. For local development, connect through the EC2 instance (an SSH tunnel) rather than exposing the database.
Security groups: allow the app, not the world
The RDS security group needs an inbound rule for port 5432 whose source is the EC2 instance’s security group, not an IP range. Referencing a security group means “any instance in that group”, so the rule keeps working when instances are replaced or their private IPs change.
When the connection times out rather than failing with an authentication error, it is almost always the security group, the database being in a different VPC, or public access being off while you connect from your laptop. An authentication error, on the other hand, proves the network path works.
Connecting from Node.js
npm install pg
import pg from 'pg';
import { readFileSync } from 'node:fs';
const pool = new pg.Pool({
host: process.env.DB_HOST, // my-database.xxxx.ap-northeast-2.rds.amazonaws.com
port: 5432,
user: process.env.DB_USER,
password: process.env.DB_PASSWORD,
database: 'postgres',
max: 10,
ssl: {
ca: readFileSync('./global-bundle.pem', 'utf8'), // RDS CA bundle
},
});
const { rows } = await pool.query('SELECT now()');
console.log(rows[0]);
Many tutorials use ssl: { rejectUnauthorized: false }. It makes the connection encrypted but skips checking that the server is really your database, which defeats half the point of TLS. AWS publishes a certificate bundle for RDS (global-bundle.pem in the RDS documentation); download it and pass it as ca. Recent RDS PostgreSQL versions enforce SSL by default (rds.force_ssl), so turning SSL off entirely is not an option either.
Keep credentials in environment variables or Secrets Manager, not in the source. And size the pool with the database in mind: each connection is a server process in PostgreSQL, and small instance classes have a low max_connections. Several app servers each opening a large pool can exhaust it, which shows up as sorry, too many clients already.
Backups and restores
RDS takes automated daily snapshots and keeps transaction logs, so you can restore to any point within the retention period (configurable, commonly 7 days). Manual snapshots stay until you delete them. Either way, a restore creates a new instance with a new endpoint; it does not rewind the existing one. Plan for updating the app’s DB_HOST, and practice a restore once before you need it.
You can stop an RDS instance to save money in development, but AWS automatically starts it again after seven days so it does not miss maintenance. People who stop a dev database and forget about it are regularly surprised by the bill a week later. Deleting it (with a final snapshot) is the reliable way to stop paying.
The first time a beginner project connects EC2 to RDS, I expect three things to go wrong, and in roughly this order: the RDS security group allows the wrong source, SSL verification gets disabled “temporarily” to make an error go away, and the dev database keeps running after the project is over. All three are cheap to prevent at setup time and annoying to discover later.
Deploying from CI without access keys
GitHub Actions can assume an IAM role through OpenID Connect, so no long-lived AWS key is stored in GitHub. You create an OIDC identity provider for token.actions.githubusercontent.com in IAM, a role whose trust policy allows your repository, and then:
name: Deploy site
on:
push:
branches: [main]
permissions:
id-token: write # needed to request the OIDC token
contents: read
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci && npm run build
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::123456789012:role/github-deploy
aws-region: ap-northeast-2
- run: aws s3 sync ./dist s3://myapp-site --delete
- run: aws cloudfront create-invalidation --distribution-id E123456 --paths "/index.html"
Scope the role’s trust policy to the specific repository and branch (the sub claim), and its permissions to the one bucket and distribution. Forgetting id-token: write produces an error that the action could not fetch an OIDC token.
Least privilege in practice
IAM policies are easiest to understand as “which actions on which resources”:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::myapp-uploads/*"
},
{
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::myapp-uploads"
}
]
}
Note the two different resource ARNs. Object actions apply to bucket/*; ListBucket applies to the bucket itself. Putting ListBucket on bucket/* is a very common reason aws s3 ls fails with AccessDenied while uploads work.
For security groups, the rule of thumb is: HTTP and HTTPS may be open to 0.0.0.0/0 on a web server, SSH should be limited to your address or replaced by Session Manager, and database ports should only ever allow the application’s security group.
Cleaning up and keeping costs visible
- Terminate test EC2 instances instead of stopping them, and check that their EBS volumes were deleted too (the root volume is by default; extra volumes may not be).
- Release Elastic IPs you no longer use.
- Delete RDS instances you are done with, keeping a final snapshot if needed.
- Add S3 lifecycle rules for data that should expire or move to cheaper storage:
{
"Rules": [
{
"ID": "expire-logs",
"Status": "Enabled",
"Filter": { "Prefix": "logs/" },
"Expiration": { "Days": 30 }
},
{
"ID": "archive-old-data",
"Status": "Enabled",
"Filter": { "Prefix": "archive/" },
"Transitions": [{ "Days": 90, "StorageClass": "GLACIER" }]
}
]
}
aws s3api put-bucket-lifecycle-configuration \
--bucket myapp-uploads --lifecycle-configuration file://lifecycle.json
Use the Cost Explorer grouped by service after the first few days. Unexpected line items such as NAT Gateway, public IPv4 or data transfer usually point to a resource you did not realize you created. The Region selector matters here too: Cost Explorer shows all Regions, but the console for each service shows only the one you are looking at.
When I look at a surprising first AWS bill, I do not start with the EC2 instance someone meant to run. I start with what was created alongside it by a wizard or a tutorial step and then forgotten: a NAT Gateway, an unattached Elastic IP, a database left running or an old snapshot. That is also why a budget alert and Cost Explorer grouped by service are the first things I set up on a new account.
Frequently Asked Questions
Q. Is the AWS Free Tier really free?
A. Usage within the limits of your account’s free plan is not charged, but the limits are specific (instance types, hours, storage) and some things that tutorials create, such as NAT Gateways and public IPv4 addresses, are billed regardless. A budget alert is the safety net.
Q. Should I use EC2 or Lambda for my backend?
A. EC2 fits long-running processes, WebSocket servers and anything that needs a persistent local disk or background workers. Lambda fits request/response APIs and event handlers with spiky or low traffic, as long as each invocation finishes within 15 minutes. Serverless with AWS Lambda covers the trade-offs in detail.
Q. How do I share S3 files with browsers without a public bucket?
A. Serve the bucket through CloudFront with Origin Access Control, or generate presigned URLs from your backend for time-limited downloads and uploads. Both keep Block Public Access on.
Q. Can I connect to a private RDS instance from my laptop?
A. Not directly. Use an SSH tunnel through an EC2 instance in the same VPC (ssh -L 5432:<rds-endpoint>:5432 ubuntu@<ec2-ip>), or Session Manager port forwarding, and point your client at localhost:5432.