Serverless with AWS Lambda: API Gateway, DynamoDB, S3 Events and Cold Starts

Key takeaways

How Lambda actually runs your code, then an API Gateway + DynamoDB API, an S3-triggered image resizer that does not trigger itself, EventBridge schedules, cold start trade-offs and presigned uploads, with the errors each step tends to produce.

What this post covers

AWS Lambda runs a function in response to an event (an HTTP request, a file landing in S3, a message on a queue, a schedule) and bills for the requests and the compute time used. You never provision or patch a server. This post starts with the execution model, because most Lambda surprises come from not knowing it, then builds a small HTTP API on API Gateway and DynamoDB, an S3-triggered image resizer, a scheduled job and a presigned upload endpoint.

The deployment examples use the Serverless Framework’s serverless.yml because it is compact to read. AWS SAM, the CDK and Terraform express the same resources.

When Lambda is a good fit (and when it is not)

Lambda works well for request/response APIs with uneven traffic, event handlers (S3, SQS, DynamoDB streams, EventBridge), scheduled jobs and glue between AWS services. You pay nothing for idle time, and scaling from one request to hundreds of concurrent ones requires no capacity planning.

It fits less well when work runs longer than the 15-minute maximum, when you need long-lived connections such as WebSocket servers you manage yourself, or when traffic is steady and high enough that an always-on container or instance is cheaper. Lambda’s per-request pricing is excellent at low and bursty volume, and the comparison changes as utilization rises, so estimate both with the AWS pricing calculator for your real traffic rather than assuming serverless is always cheaper.


The execution model

When an event arrives and no idle instance of your function is available, Lambda creates a new execution environment: it downloads your code, starts the runtime, and runs everything at the top level of your module (the init phase). That is a cold start. It then calls your handler. After the handler returns, the environment is frozen and kept around for a while; the next event may reuse it, skipping init entirely (a warm start).

Each environment handles one request at a time. If ten requests arrive at once, Lambda runs ten environments in parallel. This has several practical consequences:

  • Put expensive setup at module scope. SDK clients, database pools and parsed config created outside the handler are reused on warm starts.
  • Do not rely on module state across requests. A global counter or cache works within one environment and is invisible to the others, and environments disappear without notice.
  • /tmp persists between warm invocations in the same environment. That makes it a cache, and a way to fill up disk if you write files and never delete them.
  • Concurrency is a shared limit. Every function in an account and Region draws from the same concurrency quota. One runaway function can throttle all the others, which is what reserved concurrency guards against.

A minimal handler:

// index.mjs
export const handler = async (event, context) => {
  console.log('Request ID:', context.awsRequestId);
  return {
    statusCode: 200,
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ message: 'Hello from Lambda!' }),
  };
};

The Python equivalent is a function lambda_handler(event, context) in lambda_function.py. Anything logged with console.log or print goes to CloudWatch Logs under /aws/lambda/<function-name>; aws logs tail /aws/lambda/my-fn --follow with CLI v2 is the fastest way to watch it.

The default timeout for a function created in the console is 3 seconds. That is too short for many real handlers that call a database or another API, and the resulting error (Task timed out after 3.00 seconds) is often mistaken for a network problem.


An HTTP API with API Gateway and DynamoDB

# serverless.yml
service: users-api

provider:
  name: aws
  runtime: nodejs20.x
  region: us-east-1
  environment:
    TABLE_NAME: ${self:service}-users
  iam:
    role:
      statements:
        - Effect: Allow
          Action: [dynamodb:GetItem, dynamodb:PutItem, dynamodb:Query, dynamodb:Scan]
          Resource:
            Fn::GetAtt: [UsersTable, Arn]

functions:
  getUsers:
    handler: handlers/users.getUsers
    events:
      - http: { path: users, method: get, cors: true }
  getUser:
    handler: handlers/users.getUser
    events:
      - http: { path: 'users/{id}', method: get, cors: true }
  createUser:
    handler: handlers/users.createUser
    events:
      - http: { path: users, method: post, cors: true }

resources:
  Resources:
    UsersTable:
      Type: AWS::DynamoDB::Table
      Properties:
        TableName: ${self:service}-users
        BillingMode: PAY_PER_REQUEST
        AttributeDefinitions:
          - { AttributeName: id, AttributeType: S }
        KeySchema:
          - { AttributeName: id, KeyType: HASH }

The iam block matters. Each function runs with an execution role, and without the DynamoDB permissions every call fails with AccessDeniedException: User: arn:aws:sts::...:assumed-role/... is not authorized to perform: dynamodb:PutItem. Scoping the permission to this table’s ARN keeps a bug in one service from touching another service’s data.

// handlers/users.mjs
import { randomUUID } from 'node:crypto';
import { DynamoDBClient } from '@aws-sdk/client-dynamodb';
import { DynamoDBDocumentClient, ScanCommand, GetCommand, PutCommand } from '@aws-sdk/lib-dynamodb';

// Created once per execution environment, reused on warm starts
const docClient = DynamoDBDocumentClient.from(new DynamoDBClient({}));
const TableName = process.env.TABLE_NAME;

const json = (statusCode, body) => ({
  statusCode,
  headers: {
    'Content-Type': 'application/json',
    'Access-Control-Allow-Origin': '*',
  },
  body: JSON.stringify(body),
});

export const getUsers = async () => {
  const result = await docClient.send(new ScanCommand({ TableName, Limit: 50 }));
  return json(200, { items: result.Items, next: result.LastEvaluatedKey ?? null });
};

export const getUser = async (event) => {
  const { id } = event.pathParameters;
  const result = await docClient.send(new GetCommand({ TableName, Key: { id } }));
  return result.Item ? json(200, result.Item) : json(404, { error: 'User not found' });
};

export const createUser = async (event) => {
  let body;
  try {
    body = JSON.parse(event.body ?? '');
  } catch {
    return json(400, { error: 'Body must be JSON' });
  }
  if (typeof body.name !== 'string' || typeof body.email !== 'string') {
    return json(400, { error: 'name and email are required' });
  }

  const item = {
    name: body.name,
    email: body.email,
    id: randomUUID(),
    createdAt: new Date().toISOString(),
  };
  await docClient.send(new PutCommand({
    TableName,
    Item: item,
    ConditionExpression: 'attribute_not_exists(id)',
  }));
  return json(201, item);
};

A few choices here fix bugs that the naive version has:

  • Every response goes through one helper with the CORS header. cors: true in serverless.yml handles the preflight OPTIONS request, but with the Lambda proxy integration, the actual response headers come from your function. If only the 200 path sets Access-Control-Allow-Origin, the browser reports a CORS error for every 404 or 400, hiding the real status.
  • Parsing is guarded. An unhandled exception in a proxy-integrated function turns into a 502 with {"message": "Internal server error"} from API Gateway. That message tells the client nothing; the real stack trace is only in CloudWatch.
  • The client cannot choose the ID. Spreading the request body into the item after setting id would let a caller overwrite any existing record by sending its ID. Picking fields explicitly and adding attribute_not_exists(id) prevents that.
  • randomUUID() instead of Date.now(). Two requests in the same millisecond on different environments would get the same timestamp ID.

Scan is used for the list endpoint only because this is a demo. A scan reads the entire table (billed by data read), returns at most 1 MB per call, and signals more data through LastEvaluatedKey. Code that ignores LastEvaluatedKey works in development and silently returns partial results once the table grows. Real access patterns should be served by Query on a key or a secondary index.


S3 events without an infinite loop

functions:
  processImage:
    handler: handlers/images.process
    memorySize: 1024
    timeout: 30
    events:
      - s3:
          bucket: my-images-bucket
          event: s3:ObjectCreated:*
          rules:
            - prefix: uploads/
            - suffix: .jpg
// handlers/images.mjs
import { S3Client, GetObjectCommand, PutObjectCommand } from '@aws-sdk/client-s3';
import sharp from 'sharp';

const s3 = new S3Client({});

export const process = async (event) => {
  for (const record of event.Records) {
    const bucket = record.s3.bucket.name;
    // Keys in S3 events are URL-encoded, with spaces as '+'
    const key = decodeURIComponent(record.s3.object.key.replace(/\+/g, ' '));

    const { Body } = await s3.send(new GetObjectCommand({ Bucket: bucket, Key: key }));
    const input = await Body.transformToByteArray();

    const resized = await sharp(input)
      .resize(800, 600, { fit: 'inside' })
      .jpeg({ quality: 80 })
      .toBuffer();

    await s3.send(new PutObjectCommand({
      Bucket: bucket,
      Key: key.replace(/^uploads\//, 'thumbnails/'),
      Body: resized,
      ContentType: 'image/jpeg',
    }));
  }
};

The prefix filter is the important line. Without it, the function is triggered by every .jpg in the bucket, including the resized file it just wrote, which triggers it again, and so on. This recursive loop keeps running and billing until you notice. AWS has added recursive loop detection for some event sources, but the reliable fix is structural: trigger on uploads/ and write to thumbnails/, or write to a different bucket.

The other two fixes are about data, not configuration. The object key in an S3 event is URL-encoded, so my photo.jpg arrives as my+photo.jpg; passing it straight to GetObject fails with NoSuchKey only for files with spaces or special characters, which makes the bug look random. And Body is a stream, not a buffer; sharp expects a buffer or a file path, so the code reads it fully with transformToByteArray() first.

sharp contains native binaries. If you run npm install on macOS or Windows and upload the result, the function fails at import with an error that sharp could not load the linux-x64 (or linux-arm64) runtime. Install for the Lambda platform, for example npm install --os=linux --cpu=x64 sharp, or build inside a Linux container, and match the CPU architecture you configured for the function.

S3 invokes Lambda asynchronously. If the handler throws, Lambda retries the event (twice by default) and then drops it unless you configure a failure destination or dead-letter queue. Because of retries, the handler must be safe to run more than once for the same object, which this one is: it overwrites the same thumbnail key.


Scheduled jobs with EventBridge

functions:
  dailyReport:
    handler: handlers/reports.daily
    events:
      - schedule: cron(0 9 * * ? *)

EventBridge cron expressions have six fields (minutes, hours, day of month, month, day of week, year), and one of day-of-month or day-of-week must be ?. The standard five-field Unix syntax is rejected. More importantly, schedule rules run in UTC: cron(0 9 * * ? *) is 09:00 UTC, which is 18:00 in Seoul and 04:00 or 05:00 in New York depending on daylight saving. EventBridge Scheduler, the newer service, accepts a time zone per schedule if you need local times.

Scheduled invocations are also asynchronous and can, rarely, be delivered more than once. A report job that sends emails should record that it ran for a given date and skip duplicates.


Cold starts: what actually helps

Cold start time is mostly the init phase: loading the runtime, loading your code and its dependencies, and running module-level setup. The levers, roughly in order of how often they matter:

  1. Smaller packages and fewer imports. Bundling with esbuild and importing only the SDK clients you use (@aws-sdk/client-s3, not a full SDK) reduces what Node has to load and parse. The Node.js 18+ runtimes include AWS SDK v3, so you can mark @aws-sdk/* as external to shrink the bundle, at the cost of running whichever SDK version the runtime ships. The old advice to exclude aws-sdk refers to SDK v2, which those runtimes do not include.
  2. Less work at init. Fetching secrets, warming caches or opening many connections at module load all happen during the cold start. Do what the first request needs and defer the rest.
  3. More memory. Lambda allocates CPU in proportion to memory. Raising memory from the 128 MB default often shortens both init and handler time enough that the cost per request barely changes. Measure with a few settings; AWS Lambda Power Tuning is an open-source tool that automates this.
  4. Provisioned concurrency. Keeps a number of environments initialized at all times. It removes cold starts for that many concurrent requests, but you pay for those environments whether or not they are used, and requests beyond that number still cold start.

Runtime choice also affects init time. Interpreted runtimes such as Node.js and Python start quickly with small packages; JVM-based functions typically have heavier initialization, which is what features like Lambda SnapStart for Java target. Rather than trusting generic numbers, look at the Init Duration field that Lambda prints in the REPORT log line of each cold invocation; that is your actual cold start cost.

In my experience reviewing slow Lambda APIs, the cold start is often blamed for latency that is really caused by something in the handler, most commonly a new database connection per request or a secret fetched from Secrets Manager on every invocation. Checking Init Duration against Duration in the REPORT lines settles which one it is before anyone buys provisioned concurrency.

Databases and concurrency

Because each environment handles one request, 200 concurrent requests means up to 200 environments, and each one may open its own database connection. PostgreSQL and MySQL on a small RDS instance cannot accept that many, and you start seeing too many connections / sorry, too many clients already. Keep the per-environment pool at one or two connections and put RDS Proxy between Lambda and the database to multiplex them. Setting reserved concurrency on the function also caps how many environments can exist at once.


Presigned uploads: keep large files out of Lambda

API Gateway limits request payloads (10 MB for REST APIs), and streaming a large upload through a function wastes compute time you pay for. The standard pattern is for the function to hand out a presigned URL and let the browser upload directly to S3:

// handlers/upload.mjs
import { randomUUID } from 'node:crypto';
import { S3Client, PutObjectCommand } from '@aws-sdk/client-s3';
import { getSignedUrl } from '@aws-sdk/s3-request-presigner';

const s3 = new S3Client({});
const ALLOWED = new Set(['image/jpeg', 'image/png', 'image/webp']);

export const getUploadUrl = async (event) => {
  const { contentType } = JSON.parse(event.body ?? '{}');
  if (!ALLOWED.has(contentType)) {
    return { statusCode: 400, body: JSON.stringify({ error: 'Unsupported type' }) };
  }

  const key = `uploads/${randomUUID()}`;
  const uploadUrl = await getSignedUrl(
    s3,
    new PutObjectCommand({ Bucket: process.env.UPLOAD_BUCKET, Key: key, ContentType: contentType }),
    { expiresIn: 300 },
  );
  return { statusCode: 200, body: JSON.stringify({ uploadUrl, key }) };
};

The key is generated on the server, so clients cannot overwrite each other’s files or write outside uploads/. The content type is part of the signature, so the browser must send exactly the same Content-Type header with its PUT; a mismatch fails with SignatureDoesNotMatch. The bucket also needs a CORS configuration allowing PUT from your site’s origin, otherwise the browser blocks the upload even though the URL is valid. A short expiry (five minutes here) limits how long a leaked URL is useful.

The upload lands in uploads/, which is exactly the prefix the image resizer above listens on, so the two pieces compose into a complete upload-and-thumbnail pipeline.


Frequently Asked Questions

Q. Lambda or EC2: which should I use?

A. Lambda suits event-driven work and APIs with uneven traffic, and removes server management. EC2 or containers suit long-running processes, persistent connections and steady high load where paying for always-on capacity is cheaper. Many systems use both.

Q. Is there an execution time limit?

A. Yes, 15 minutes per invocation. For longer jobs, split the work into steps orchestrated by Step Functions, or run it on ECS/Fargate or AWS Batch.

Q. Why does my S3-triggered Lambda keep invoking itself?

A. It writes output to the same bucket and prefix that triggers it, so every write produces a new event. Restrict the trigger with a prefix such as uploads/ and write to a different prefix or bucket.

Q. Why do I get Internal server error from API Gateway?

A. With the proxy integration, an unhandled exception or a response that is not in the expected { statusCode, headers, body } shape (for example a non-string body) becomes a 502. Check the function’s CloudWatch logs for the actual error.