Terraform on AWS in Practice: Resources, Variables, Remote State, Modules and Workspaces

Key takeaways

Terraform turns cloud infrastructure into version-controlled code. This guide covers everything from first resource to production-grade setup: variables, state management, modules, workspaces, and CI/CD integration.

Why Terraform?

Without Infrastructure as Code, cloud setups are:

  • Manual (click in the console → can’t reproduce exactly)
  • Undocumented (no record of what was created or why)
  • Error-prone (different settings in dev vs prod)
  • Slow to recreate (hours of clicking vs minutes of terraform apply)

Terraform solves all of this by describing your infrastructure in code:

Manual:            Click → Create → Forget → Can't reproduce

With Terraform:    Write HCL → terraform plan → terraform apply → Git commit
                   Reproducible, documented, version-controlled, team-reviewable

I’ve inherited infrastructure that existed only as tribal knowledge and console clicks, and the actual cost isn’t abstract — it’s the day someone needs to recreate a staging environment that matches production and discovers nobody actually knows every setting that was manually configured over the past two years, some of it by people who’ve since left the team. Terraform’s real value isn’t “automation” in the abstract sense, it’s that the .tf files themselves are the documentation, kept honest by the fact that terraform plan will show a drift the moment reality and the code disagree — a wiki page describing infrastructure can silently go stale forever, but code that’s actually re-applied regularly can’t drift unnoticed the same way.


Installation

# macOS
brew install terraform

# Linux
sudo apt update && sudo apt install terraform

# Windows
choco install terraform

# Verify
terraform version
# Terraform v1.7.x

Core Concepts

Configuration (.tf files)
  ↓ terraform init      Downloads providers
  ↓ terraform plan      Shows what will change
  ↓ terraform apply     Creates/modifies resources
  ↓ terraform destroy   Destroys all resources

State file (.tfstate)
  Tracks what Terraform actually created
  Stores resource IDs, attributes, dependencies
  Must be stored remotely (S3) for teams
TermWhat it is
ProviderPlugin for a cloud platform (AWS, GCP, Azure)
ResourceA cloud resource to create (EC2, S3, VPC)
Data sourceRead existing resources without managing them
VariableInput parameter for reusable configs
OutputValue exposed after apply (IP address, ARN)
ModuleReusable group of resources
StateRecord of what Terraform has created

The state file is the piece most newcomers underestimate, and it’s worth understanding precisely what problem it solves: Terraform doesn’t re-scan your cloud account on every run to figure out what exists — it trusts the state file as the source of truth for what it believes it created, and diffs your .tf configuration against that recorded belief, not against live reality. This is exactly why manually deleting or modifying a Terraform-managed resource through the AWS console is dangerous — Terraform’s state still says that resource exists in a certain configuration, and the next plan/apply either tries to “fix” a resource that’s already gone (erroring or recreating it unexpectedly) or, worse, silently overwrites a manual change nobody told Terraform about. State being authoritative-but-not-live is the single concept that explains most of the confusing Terraform behavior newcomers hit.


Your First Terraform Config

# main.tf

terraform {
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.0"
    }
  }
}

provider "aws" {
  region = "us-east-1"
}

resource "aws_s3_bucket" "my_bucket" {
  bucket = "my-terraform-bucket-unique-12345"

  tags = {
    Name        = "My Terraform Bucket"
    Environment = "dev"
    ManagedBy   = "Terraform"
  }
}
# Initialize — downloads the AWS provider
terraform init

# Preview what will be created
terraform plan

# Create the resources
terraform apply          # Prompts for confirmation
terraform apply -auto-approve  # Skip confirmation (use in CI only)

# Check current state
terraform show

# Destroy everything (⚠️ irreversible)
terraform destroy

The bucket name here (my-terraform-bucket-unique-12345) isn’t a placeholder to swap out casually — S3 bucket names are globally unique across every AWS account on the planet, not just within your own account, which is a genuinely unusual constraint compared to most other AWS resources and a common first-run failure (BucketAlreadyExists) for anyone copying an example literally. terraform apply -auto-approve skipping the confirmation prompt is explicitly flagged “CI only” for good reason — the interactive prompt showing the plan one more time before applying is a real, human safety check against a plan that looks different than expected, and automating past it entirely is a deliberate tradeoff only worth making once a pipeline already has other safeguards (a required plan review, a separate approval gate) in place.


Variables

Define Variables

# variables.tf

variable "region" {
  description = "AWS region to deploy into"
  type        = string
  default     = "us-east-1"
}

variable "environment" {
  description = "Deployment environment (dev, staging, prod)"
  type        = string

  validation {
    condition     = contains(["dev", "staging", "prod"], var.environment)
    error_message = "Must be dev, staging, or prod."
  }
}

variable "instance_type" {
  description = "EC2 instance type"
  type        = string
  default     = "t3.micro"
}

variable "common_tags" {
  description = "Tags to apply to all resources"
  type        = map(string)
  default = {
    ManagedBy = "Terraform"
    Project   = "MyApp"
  }
}

The validation block on environment is worth taking seriously as a real safeguard, not boilerplate — without it, a typo like environment = "produciton" doesn’t fail loudly at plan time, it just silently creates a new, unexpected environment’s worth of resources tagged with a value nothing else recognizes, which is a genuinely confusing thing to discover after the fact. Constraining input variables this way, everywhere a typo could otherwise produce a hard-to-notice wrong configuration, is cheap insurance against exactly the kind of mistake that’s invisible until someone’s debugging why staging traffic is hitting the wrong resources.

Use Variables

# main.tf

provider "aws" {
  region = var.region
}

resource "aws_instance" "web" {
  ami           = data.aws_ami.ubuntu.id
  instance_type = var.instance_type

  tags = merge(var.common_tags, {
    Name        = "web-server"
    Environment = var.environment
  })
}

merge() here is doing more than syntactic convenience — it’s what lets every resource share a common tag baseline (common_tags) while still layering on resource-specific tags, without every single resource block having to repeat the shared tags by hand and risk them drifting out of sync across dozens of resources. This matters in practice more than it sounds: consistent tagging is frequently what makes cost-allocation reports and “who owns this resource” audits actually usable months later, and it’s a lot easier to get consistent tagging right by centralizing it once in common_tags than to enforce it by convention across a growing set of resource blocks.

Set Variable Values

# CLI
terraform apply -var="environment=prod" -var="instance_type=t3.medium"
# terraform.tfvars (auto-loaded)
environment   = "prod"
instance_type = "t3.medium"
region        = "us-east-1"

common_tags = {
  ManagedBy = "Terraform"
  Project   = "MyApp"
  Team      = "Platform"
}

Terraform reads variable values from several sources with a defined precedence, and knowing that order matters the first time a value doesn’t take effect the way you’d expect: CLI -var flags override .tfvars files, which override environment variables (TF_VAR_x), which override the default in the variable’s own declaration. This is exactly why the “Best Practices” section later in this guide warns against committing .tfvars with secrets — a value quietly sitting in a committed terraform.tfvars file will silently apply for every team member and every CI run unless something with higher precedence explicitly overrides it, which is easy to forget is happening once a project has several config sources layered together.


Outputs

# outputs.tf

output "web_server_public_ip" {
  description = "Public IP of the web server"
  value       = aws_instance.web.public_ip
}

output "s3_bucket_arn" {
  description = "ARN of the S3 bucket"
  value       = aws_s3_bucket.my_bucket.arn
}

output "rds_endpoint" {
  description = "RDS connection endpoint"
  value       = aws_db_instance.main.endpoint
  sensitive   = true   # Hides value in terminal output
}
terraform output                        # All outputs
terraform output web_server_public_ip   # Specific output
terraform output -json                  # JSON format (for scripts)

sensitive = true on the RDS endpoint output is worth understanding precisely: it hides the value from terraform apply’s terminal output and from terraform output unless explicitly requested, but — and this catches people out — the actual value is still stored in plaintext in the state file itself. This is exactly why “never commit the state file” (covered in the Remote State section next, and in the FAQ above) isn’t a redundant warning once you’ve marked things sensitive — the sensitive flag protects against accidental exposure in logs and terminal scrollback, not against the state file itself being read by anyone with access to it, which is what remote state’s own access controls (S3 bucket policies, encryption) actually have to handle.


Real-World: VPC + EC2 + Security Group

# vpc.tf

resource "aws_vpc" "main" {
  cidr_block           = "10.0.0.0/16"
  enable_dns_hostnames = true
  enable_dns_support   = true

  tags = { Name = "${var.environment}-vpc" }
}

resource "aws_subnet" "public" {
  count = 2

  vpc_id            = aws_vpc.main.id
  cidr_block        = "10.0.${count.index + 1}.0/24"
  availability_zone = data.aws_availability_zones.available.names[count.index]
  map_public_ip_on_launch = true

  tags = { Name = "${var.environment}-public-${count.index + 1}" }
}

resource "aws_internet_gateway" "main" {
  vpc_id = aws_vpc.main.id
  tags   = { Name = "${var.environment}-igw" }
}

resource "aws_route_table" "public" {
  vpc_id = aws_vpc.main.id

  route {
    cidr_block = "0.0.0.0/0"
    gateway_id = aws_internet_gateway.main.id
  }
}

resource "aws_route_table_association" "public" {
  count          = 2
  subnet_id      = aws_subnet.public[count.index].id
  route_table_id = aws_route_table.public.id
}

# ec2.tf

resource "aws_security_group" "web" {
  name        = "${var.environment}-web-sg"
  description = "Web server security group"
  vpc_id      = aws_vpc.main.id

  ingress {
    from_port   = 80
    to_port     = 80
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }

  ingress {
    from_port   = 443
    to_port     = 443
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }

  ingress {
    from_port   = 22
    to_port     = 22
    protocol    = "tcp"
    cidr_blocks = ["YOUR_IP/32"]   # Restrict SSH to your IP
    # This one line is the difference between a reasonably secured server
    # and one that shows up in automated internet-wide scanning within
    # minutes of going live — 0.0.0.0/0 on port 22 (SSH open to the entire
    # internet) is one of the most common real-world misconfigurations
    # that leads to a compromised EC2 instance, since it's trivial for
    # bots to scan for open port 22 and attempt credential/key brute-force
    # against anything they find. The /32 suffix matters too: it's a CIDR
    # notation for "exactly this one IP address," not a range — dropping
    # it (or using a wider range like YOUR_IP/24) reopens exactly the
    # exposure this line exists to close.
  }

  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }
}

resource "aws_instance" "web" {
  ami                    = data.aws_ami.ubuntu.id
  instance_type          = var.instance_type
  subnet_id              = aws_subnet.public[0].id
  vpc_security_group_ids = [aws_security_group.web.id]
  key_name               = aws_key_pair.deployer.key_name

  user_data = <<-EOF
    #!/bin/bash
    apt update -y
    apt install -y nginx
    systemctl start nginx
    systemctl enable nginx
    echo "Hello from Terraform — ${var.environment}" > /var/www/html/index.html
  EOF

  tags = { Name = "${var.environment}-web-server" }
}

user_data running as a shell script on first boot is a real, one-time-only mechanism worth understanding the limits of — it executes once when the instance is first created, not on every terraform apply, so changing the script and reapplying does not re-run it against an already-running instance (Terraform will typically want to replace the instance entirely to apply a user_data change, since it can’t retroactively re-execute a boot script). This heredoc-based inline provisioning is genuinely fine for a quick demo or a minimal setup, but it doesn’t scale well to anything more complex than installing a package or two — a real production setup typically reaches for a pre-baked AMI (via Packer) or a proper configuration management tool layered on top, specifically because debugging a multi-step user_data script that failed partway through, with no easy way to re-run just the failed part, is a genuinely painful troubleshooting experience.


Remote State (Required for Teams)

Never use local state for team projects. Use S3 + DynamoDB for atomic, encrypted remote state:

# First, create the state bucket manually (one-time setup)
aws s3 mb s3://my-terraform-state-bucket --region us-east-1
aws s3api put-bucket-versioning \
  --bucket my-terraform-state-bucket \
  --versioning-configuration Status=Enabled

# Create DynamoDB table for state locking
aws dynamodb create-table \
  --table-name terraform-state-lock \
  --attribute-definitions AttributeName=LockID,AttributeType=S \
  --key-schema AttributeName=LockID,KeyType=HASH \
  --billing-mode PAY_PER_REQUEST

I’ve seen a local-state Terraform setup cause a genuinely painful incident firsthand: two people on a small team both had valid AWS credentials and both ran terraform apply within a few minutes of each other, each working from their own local .tfstate file that had already drifted from the other’s — the result was Terraform confidently “fixing” resources that weren’t actually broken, based on each person’s stale, disconnected view of what should exist. Remote state with locking (the DynamoDB table here) solves both halves of that problem at once: the state itself lives in one shared, versioned location instead of N divergent local copies, and the lock means a second apply started while one is already running waits or fails cleanly instead of racing against it — this genuinely isn’t optional infrastructure-as-code hygiene for a team, it’s the difference between Terraform being a coordination tool and Terraform actively working against itself.

# backend.tf

terraform {
  backend "s3" {
    bucket         = "my-terraform-state-bucket"
    key            = "prod/terraform.tfstate"   # Path within bucket
    region         = "us-east-1"
    dynamodb_table = "terraform-state-lock"     # Prevents concurrent applies
    encrypt        = true
  }
}
# State management commands
terraform state list                           # List all resources in state
terraform state show aws_instance.web          # Show resource details
terraform state mv aws_instance.old aws_instance.new  # Rename resource
terraform state rm aws_instance.unwanted       # Remove from state (keeps real resource)
terraform import aws_instance.web i-1234567890  # Import existing resource

terraform state mv/rm are worth understanding as state-only operations that never touch the real cloud resource — rm specifically removes Terraform’s tracking of a resource while leaving the actual EC2 instance, S3 bucket, or whatever it is completely untouched and still running, which is exactly the tool for “we’re migrating this resource to be managed by a different Terraform config” or “we want to stop managing this resource without destroying it.” state mv is the fix for a genuinely common refactor headache: renaming a resource in your .tf code (or moving it into a module) without state mv makes Terraform think the old resource was deleted and a brand-new one needs creating — for something like a database, that’s a plan showing “destroy and recreate,” which is a very different, much scarier outcome than what you actually intended (just a rename). terraform import is the reverse direction — bringing a resource that already exists in AWS (created manually, or by some other tool) under Terraform’s management without recreating it, which is the standard on-ramp for adopting Terraform on infrastructure that predates it.


Modules — Reusable Infrastructure

Define a Module

# modules/vpc/variables.tf
variable "environment" { type = string }
variable "cidr_block" { type = string }
variable "public_subnet_cidrs" { type = list(string) }

# modules/vpc/main.tf
resource "aws_vpc" "main" {
  cidr_block           = var.cidr_block
  enable_dns_hostnames = true
  tags = { Name = "${var.environment}-vpc" }
}

# modules/vpc/outputs.tf
output "vpc_id" { value = aws_vpc.main.id }
output "public_subnet_ids" { value = aws_subnet.public[*].id }

Use the Module

# main.tf

module "vpc" {
  source = "./modules/vpc"

  environment         = "production"
  cidr_block          = "10.0.0.0/16"
  public_subnet_cidrs = ["10.0.1.0/24", "10.0.2.0/24"]
}

resource "aws_instance" "web" {
  subnet_id = module.vpc.public_subnet_ids[0]  # Use module output
}

A module is really just a parameterized, reusable chunk of the exact same HCL covered throughout this guide — its own variables, its own resources, its own outputs — which is the whole appeal: the VPC setup from the “Real-World” section earlier, written once as a module with environment/cidr_block/public_subnet_cidrs as inputs, can be instantiated for dev, staging, and prod without copy-pasting the same 40 lines of VPC/subnet/route-table configuration three times and risking them drifting apart from each other as each environment gets tweaked independently over time. The module’s own outputs.tf is what makes composition possible — module.vpc.public_subnet_ids only exists because the module explicitly exposed it, the same way any other Terraform output works, just scoped to that module’s boundary instead of the root configuration’s.

Terraform Registry Modules

# Use a community module from registry.terraform.io
module "vpc" {
  source  = "terraform-aws-modules/vpc/aws"
  version = "~> 5.0"

  name = "production-vpc"
  cidr = "10.0.0.0/16"

  azs             = ["us-east-1a", "us-east-1b"]
  public_subnets  = ["10.0.1.0/24", "10.0.2.0/24"]
  private_subnets = ["10.0.11.0/24", "10.0.12.0/24"]

  enable_nat_gateway = true
}

The community-registry modules deserve one caution worth stating plainly: source = "terraform-aws-modules/vpc/aws" pulls in code you didn’t write, maintained by someone else, and while terraform-aws-modules specifically is a well-established, widely-used, actively-maintained organization, that’s not universally true of every registry module — pin the version explicitly (as shown), review what a module actually creates before pointing it at a real AWS account (terraform plan will show you, but it’s worth reading rather than trusting blindly), and treat an unfamiliar module with the same scrutiny you’d apply to any third-party dependency with the ability to provision or destroy real infrastructure.


Workspaces — Multiple Environments

# Create workspaces
terraform workspace new dev
terraform workspace new staging
terraform workspace new prod

# Switch workspaces
terraform workspace select prod

# Show current workspace
terraform workspace show    # prod

# List workspaces
terraform workspace list
# * prod
#   dev
#   staging

Use workspace in configuration:

locals {
  env = terraform.workspace

  instance_config = {
    dev     = { type = "t3.micro",  count = 1 }
    staging = { type = "t3.small",  count = 2 }
    prod    = { type = "t3.medium", count = 3 }
  }
}

resource "aws_instance" "web" {
  count         = local.instance_config[local.env].count
  instance_type = local.instance_config[local.env].type
}

Workspaces are worth understanding as a narrower tool than they initially sound like — each workspace gets its own separate state file under the same configuration, but the .tf code itself is entirely shared across all of them, which is exactly right for environments that are structurally identical and only differ in a handful of values (the instance_config lookup above is the idiomatic pattern for that). The FAQ above already flags the real limitation: the moment dev and prod need genuinely different resources — prod has a read replica and a CDN in front of it that dev doesn’t, say — cramming that difference into one shared codebase via more and more conditional logic gets unwieldy fast, and separate directories per environment (each with its own, potentially-diverging .tf files) become the more maintainable choice, which is exactly why the FAQ notes most teams end up preferring that approach once environments genuinely diverge in more than just sizing.


CI/CD Integration (GitHub Actions)

# .github/workflows/terraform.yml
name: Terraform

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

env:
  TF_VAR_environment: ${{ github.ref == 'refs/heads/main' && 'prod' || 'dev' }}

jobs:
  terraform:
    runs-on: ubuntu-latest
    permissions:
      id-token: write    # For OIDC authentication
      contents: read
      pull-requests: write

    steps:
      - uses: actions/checkout@v4

      - uses: hashicorp/setup-terraform@v3
        with:
          terraform_version: 1.7.x

      - name: Configure AWS Credentials (OIDC — no long-lived keys)
        uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::123456789:role/GitHubActionsRole
          aws-region: us-east-1
          # OIDC federation here means GitHub Actions authenticates to AWS
          # by proving its identity via a short-lived, cryptographically
          # signed token exchanged for temporary credentials — no AWS
          # access key/secret pair is ever generated, stored as a GitHub
          # secret, or capable of being leaked from a compromised
          # workflow log. This is a meaningfully better security posture
          # than the older pattern of long-lived IAM user credentials
          # stored as repo secrets, which, once leaked (a common class of
          # supply-chain incident), remain valid until someone notices and
          # manually revokes them — OIDC-issued credentials simply expire.

      - name: Terraform Init
        run: terraform init

      - name: Terraform Plan
        id: plan
        run: terraform plan -no-color -out=tfplan
        continue-on-error: true

      - name: Post Plan as PR Comment
        uses: actions/github-script@v7
        if: github.event_name == 'pull_request'
        with:
          script: |
            const output = `#### Terraform Plan 📖\`${{ steps.plan.outcome }}\`
            <details><summary>Show Plan</summary>

            \`\`\`
            ${{ steps.plan.outputs.stdout }}
            \`\`\`

            </details>`;
            github.rest.issues.createComment({
              issue_number: context.issue.number,
              owner: context.repo.owner,
              repo: context.repo.repo,
              body: output
            })

      - name: Terraform Apply (main branch only)
        if: github.ref == 'refs/heads/main' && github.event_name == 'push'
        run: terraform apply tfplan

The PR-comment step is worth treating as a genuine review gate, not a nice-to-have — the whole point of posting the plan output where reviewers actually look (the PR itself, not a CI log someone has to go dig up separately) is that infrastructure changes are frequently harder to review by reading a diff of .tf files alone than by reading what Terraform actually says it’s about to do, especially a plan showing an unexpected “destroy and recreate” on something that should have been a simple in-place update. terraform apply tfplan at the end applies the saved plan file from the earlier step in the same run, so nothing is recomputed between plan and apply, and if the state changed in between, Terraform refuses the stale plan instead of applying it. Note what it does not guarantee: the push to main is a new workflow run with a new plan, not the plan that was posted on the pull request, so if the infrastructure or another merge changed things in the meantime, what gets applied can differ from what was reviewed. Tools such as Atlantis or HCP Terraform close that gap by applying the reviewed plan before merge. Also, continue-on-error: true on the plan step keeps the job going so the comment can be posted, but it also means a failed plan does not fail the check; add a later step that fails when steps.plan.outcome == 'failure'.


Refactoring, locking and secrets: details that bite later

Prefer moved and import blocks over one-off state commands. Since Terraform 1.1, a moved block records a rename in the code itself:

moved {
  from = aws_instance.old
  to   = aws_instance.new
}

Unlike terraform state mv, which one person runs by hand against one state file, a moved block goes through review with the rename and is applied by every workspace and every teammate’s next plan, which shows the move instead of a destroy-and-create. import blocks (Terraform 1.5+) do the same for adoption: the import appears in plan before anything is written to state, and terraform plan -generate-config-out=generated.tf can draft the matching resource block.

Locking on S3 no longer needs DynamoDB. Terraform 1.10 added use_lockfile = true to the S3 backend, which stores the lock as an object next to the state file, and 1.11 deprecated the dynamodb_table argument. The DynamoDB setup shown earlier still works, but on a current Terraform version new setups can skip the table.

sensitive = true hides values from output, not from state. Marking a variable or output sensitive keeps it out of plan output and logs, but the value is still written in plain text to the state file. That is why access to the state bucket matters as much as access to the secrets themselves: enable encryption, restrict the bucket policy, and prefer having resources read secrets from AWS Secrets Manager or SSM at runtime over passing them through Terraform variables.

Commit .terraform.lock.hcl. terraform init records the exact provider versions and checksums it selected in this file. Committing it means CI and every teammate use the same provider build; without it, a version constraint like ~> 5.0 lets a new minor release arrive silently on the next fresh init.