Background Jobs and Workflows with Inngest: Steps, Fan-Out, Retries, Concurrency and Cron
Key takeaways
Inngest runs reliable background jobs and multi-step workflows without managing queues or workers. Functions are plain TypeScript, retries are automatic, and the local dev server shows every execution. This guide covers everything from simple jobs to complex AI pipelines.
What is Inngest?
Most backend teams eventually hit the same wall: some piece of work — sending an email, charging a card, generating a report — takes too long to run inside an HTTP request, but it still needs to happen reliably. The traditional answer is a message queue (Redis, RabbitMQ, SQS) plus a worker process that polls it. That works, but it means you now own a second piece of infrastructure: something has to keep the worker alive, restart it when it crashes, scale it under load, and expose metrics so you know when it falls behind. On a serverless platform like Vercel or Cloudflare Pages, “keep a worker alive” isn’t even an option — there’s no long-running process to poll anything.
Inngest sidesteps this by inverting the relationship. Instead of your app pushing work onto a queue that a worker pulls from, your app sends an event, and Inngest’s own infrastructure calls your function back over HTTP when it’s ready to run it — with retries, backoff, and state tracking handled entirely on Inngest’s side. The queue still exists, but it lives in Inngest’s cloud, not in your infrastructure.
Traditional: Your app → Push to Redis queue → Worker polls queue → Run job
Inngest: Your app → Send event → Inngest calls your function → Done
The part that’s easy to underestimate the first time you read the docs is step. A step isn’t just “a chunk of code that gets retried” — it’s a checkpoint. After a step completes successfully, Inngest persists its return value. If a later step throws and the function is retried, Inngest replays the function from the top, but for every step it already has a recorded result for, it returns that cached value immediately instead of re-executing the code. This is what the Inngest docs call durable execution, and it’s the single most important mental model to internalize before you write your first function — it changes how you think about idempotency, side effects, and what belongs inside a step.run() versus outside one.
In addition to retries and durable state, Inngest handles:
- Retries with exponential backoff
- Concurrency limits
- Rate limiting
- Scheduling (cron)
- Fan-out (parallel steps, and multiple functions listening to one event)
- Step state (pause/resume between steps, including multi-day sleeps)
flowchart LR
A["Your app"] -->|"1. inngest.send(event)"| B["Inngest Cloud"]
B -->|"2. HTTP POST"| C["/api/inngest endpoint"]
C -->|"3. runs step.run()"| D["Your function code"]
D -->|"4. step result"| B
B -->|"5. next step call<br/>(if more steps remain)"| C
Notice that steps 2 through 5 can be separate HTTP requests. That’s the key architectural difference from a traditional worker: your function isn’t one continuously running process — it’s a series of short-lived invocations, each of which Inngest resumes exactly where the last one left off, using the persisted step results as its memory.
Installation
npm install inngest
# Start local dev server (separate terminal)
npx inngest-cli@latest dev
The inngest package is the SDK you import into your app code — it has no runtime dependency on anything else. The inngest-cli dev server, on the other hand, is a small local service that stands in for Inngest’s cloud during development: it receives events from your app, calls your function endpoint the same way production Inngest would, and gives you a dashboard to inspect every run. Running both side by side lets you develop step functions with full visibility before you ever touch a production API key.
Basic Setup
Next.js App Router
// app/api/inngest/route.ts
import { serve } from 'inngest/next'
import { inngest } from '@/lib/inngest'
import { sendWelcomeEmail, processPayment } from '@/lib/functions'
export const { GET, POST, PUT } = serve({
client: inngest,
functions: [sendWelcomeEmail, processPayment],
})
// lib/inngest.ts
import { Inngest } from 'inngest'
export const inngest = new Inngest({
id: 'my-app',
// eventKey and signingKey come from environment variables in production
})
This single route handler is the entire “server” Inngest needs from you — there’s no separate worker process to deploy. serve() does three jobs at once: it exposes a GET endpoint Inngest uses to introspect which functions exist (this is how the dashboard knows your function list without you manually registering anything), a PUT endpoint used once at deploy time to sync your function definitions, and a POST endpoint Inngest calls, one invocation per step, to actually execute your code.
The signingKey deserves more attention than the snippet gives it. Because this route is a public HTTP endpoint, anyone who finds the URL could, in principle, POST a fake payload and trigger your function early or with forged data. Inngest signs every request it sends with an HMAC derived from your signing key, and the SDK verifies that signature before running anything. If you forget to set INNGEST_SIGNING_KEY in production, the endpoint still works during local development (the dev server skips verification), which is exactly the kind of thing that passes code review and then quietly fails — or worse, quietly accepts unsigned requests — once deployed. Always confirm the signing key is present in your production environment variables before shipping a function that does anything sensitive, like processPayment above.
Express / Node.js
import { serve } from 'inngest/express'
import express from 'express'
import { inngest } from './inngest'
import { functions } from './functions'
const app = express()
app.use('/api/inngest', serve({ client: inngest, functions }))
app.listen(3000)
The Express adapter is functionally identical to the Next.js one — same three HTTP verbs, same signature verification — it just mounts on an Express router instead of a file-based route. This matters if your app already runs as a long-lived Node process (say, on Railway or a bare VM): you don’t need Inngest’s HTTP-call model to get durable execution and retries. You still benefit from the checkpointed step state even though your server could, in theory, run the whole function in memory — the difference is that a crash mid-function no longer means starting over from step one.
Simple Background Functions
import { inngest } from './inngest'
// Define an event-triggered function
export const sendWelcomeEmail = inngest.createFunction(
{ id: 'send-welcome-email' }, // unique function ID
{ event: 'user/signed-up' }, // trigger event
async ({ event, step }) => {
const { userId, email, name } = event.data
// Each step.run() is independently retried
await step.run('send-email', async () => {
await emailService.send({
to: email,
subject: `Welcome, ${name}!`,
template: 'welcome',
data: { name },
})
})
await step.run('update-crm', async () => {
await crm.users.update(userId, { emailSent: true })
})
return { success: true, userId }
}
)
Look closely at why this function is split into two step.run() calls instead of one block of code. If emailService.send() succeeds but crm.users.update() throws (say, the CRM API times out), Inngest retries the whole function, but because send-email already has a recorded result, it is not re-executed — only update-crm runs again. Without steps, a naive retry would re-send the welcome email every time the CRM call failed, which is exactly the kind of bug that looks fine in testing and then annoys real users in production. As a rule of thumb: anything with an external side effect (sending mail, charging money, writing to a third-party API) belongs inside its own step.run(), both so it gets independently retried and so it doesn’t get needlessly repeated when something else in the function fails.
The corollary is that the code you put inside step.run() must itself be safe to run more than once for the same logical event — Inngest guarantees the step’s result is cached after a success, but if the step itself throws partway through a side effect (e.g., the email provider accepts the request but the response times out on your end), a retry really will re-run it. Idempotency keys on the calls you make to external services (most email and payment providers support one) are the standard defense here, not something Inngest can do for you automatically.
Send an event from your app
// In your signup handler
await inngest.send({
name: 'user/signed-up',
data: {
userId: user.id,
email: user.email,
name: user.name,
},
})
// Send multiple events at once
await inngest.send([
{ name: 'user/signed-up', data: { userId: '1', email: '[email protected]', name: 'Alice' } },
{ name: 'analytics/track', data: { event: 'signup', userId: '1' } },
])
inngest.send() returns as soon as the event has been accepted by Inngest — it does not wait for any function to run, which is what makes this safe to call from inside a request handler without adding latency to the response your user sees. This also means events are fire-and-forget from your app’s perspective: if you need to know the outcome of the background work (did the welcome email actually go out?), you either query Inngest’s API for the run status, or have the function itself emit a follow-up event or write a status field the frontend can poll. Treat event names (user/signed-up) as a public contract, the same way you’d treat a REST API shape — multiple functions can subscribe to the same event, so renaming it later means updating every subscriber, not just the sender.
Step Functions — Multi-Step Workflows
export const onboardNewUser = inngest.createFunction(
{ id: 'onboard-new-user' },
{ event: 'user/signed-up' },
async ({ event, step }) => {
const { userId, email, plan } = event.data
// Step 1: Create workspace
const workspace = await step.run('create-workspace', async () => {
return await db.workspaces.create({ ownerId: userId, plan })
})
// Step 2: Send welcome email (runs even if step 3 fails)
await step.run('send-welcome-email', async () => {
await sendEmail(email, 'welcome', { workspaceId: workspace.id })
})
// Step 3: Wait 1 day, then send onboarding tips
await step.sleep('wait-for-onboarding', '1d')
// Checks if user completed setup before sending tips
const userActivity = await step.run('check-activity', async () => {
return await db.users.getActivity(userId)
})
if (!userActivity.completedSetup) {
await step.run('send-tips-email', async () => {
await sendEmail(email, 'onboarding-tips')
})
}
// Step 4: After 7 days, check if user upgraded
await step.sleep('wait-for-conversion', '7d')
const user = await step.run('check-plan', async () => {
return await db.users.findById(userId)
})
if (user.plan === 'free') {
await step.run('send-upgrade-nudge', async () => {
await sendEmail(email, 'upgrade-offer')
})
}
return { onboarded: true }
}
)
This is the example that makes the “no worker infrastructure” claim concrete, because step.sleep('wait-for-onboarding', '1d') blocks the logical function for a full day without holding open a connection, a container, or a billed invocation for those 24 hours. Under the hood, when execution hits step.sleep, the current invocation simply ends, and Inngest schedules the next HTTP call to your endpoint for one day later. From your infrastructure’s point of view, nothing is running during that gap — no server, no cron job you wrote, no idle Lambda burning execution time. This is the pattern that makes multi-day drip campaigns, trial-expiration reminders, and abandoned-cart nudges genuinely cheap to build on a serverless platform, where a setTimeout spanning a week is simply not possible.
A subtlety worth flagging: because each step boundary can be a fresh invocation on a fresh instance, regular JavaScript variables that aren’t step results don’t survive unless they were derived entirely from data already returned by a step.run(). workspace in the snippet above works because it’s the return value of step.run('create-workspace', ...), which Inngest persists and replays. If you tried to read some in-memory cache or a module-level variable set earlier in the function, it would be gone by the time wait-for-onboarding resumes a day later on a different instance. This is the most common class of bug developers hit when they’re new to Inngest: code that “works” in a quick local test (because the dev server often reuses the same process) and then behaves unpredictably once step gaps span real time in production.
Parallel Steps with fan-out
export const generateReport = inngest.createFunction(
{ id: 'generate-monthly-report' },
{ event: 'reports/requested' },
async ({ event, step }) => {
const { reportId, sections } = event.data
// Run all sections in parallel
const results = await step.run('fetch-all-data', async () => {
return await Promise.all([
fetchSalesData(reportId),
fetchUserMetrics(reportId),
fetchRevenueData(reportId),
fetchChurnData(reportId),
])
})
const [sales, users, revenue, churn] = results
// Generate PDF after all data is ready
const pdfUrl = await step.run('generate-pdf', async () => {
return await pdfService.generate({ sales, users, revenue, churn })
})
await step.run('notify-requester', async () => {
await sendEmail(event.data.requestedBy, 'report-ready', { pdfUrl })
})
return { pdfUrl }
}
)
There are two different kinds of parallelism at play in this guide, and it’s worth being precise about which one this example uses. Here, the four fetch*Data calls run concurrently inside a single step.run() via ordinary Promise.all — from Inngest’s perspective this is still one step with one recorded result. That’s the right choice when the parallel work is fast and tightly coupled (you need all four results before you can generate the PDF anyway), because it avoids the overhead of four separate step round-trips. The trade-off is that if any one of the four promises rejects, the entire step fails and all four are retried together, even the ones that already succeeded — Promise.all doesn’t give you partial credit. If the individual fetches were slow, expensive, or independently likely to fail, you’d instead give each one its own step.run() call (Inngest can execute independent steps concurrently across separate invocations) so a failure in fetchChurnData doesn’t force a redundant retry of fetchSalesData.
Fan-out — trigger multiple functions from one event
// One event → multiple functions run in parallel
export const processOrder = inngest.createFunction(
{ id: 'process-order' },
{ event: 'order/created' },
async ({ event }) => { /* fulfill order */ }
)
export const sendOrderConfirmation = inngest.createFunction(
{ id: 'send-order-confirmation' },
{ event: 'order/created' }, // same event!
async ({ event }) => { /* send email */ }
)
export const updateInventory = inngest.createFunction(
{ id: 'update-inventory' },
{ event: 'order/created' }, // same event!
async ({ event }) => { /* update stock */ }
)
// All three run in parallel when order/created is sent
This is the second, structurally different kind of parallelism: three entirely separate functions subscribed to the same event name. This is the pattern that replaces what would otherwise be a pub/sub topic with multiple consumers — think of order/created as a domain event that different parts of your system care about for different reasons. The advantage over cramming all three concerns into one function is failure isolation: if updateInventory throws because the warehouse API is down, Inngest retries only that function. processOrder and sendOrderConfirmation already succeeded and are unaffected, so the customer still gets their confirmation email on time even while inventory sync is being retried in the background. The cost is that you lose the implicit ordering a single function would give you — if fulfillment genuinely must happen before the confirmation email is sent, that dependency needs to be expressed explicitly (for example, sendOrderConfirmation waiting on a step, or listening for a second event that processOrder emits once fulfillment finishes) rather than relying on execution order between unrelated functions, which Inngest does not guarantee.
Retries and Error Handling
export const processPayment = inngest.createFunction(
{
id: 'process-payment',
retries: 5, // default: 3, max: 20
},
{ event: 'payment/initiated' },
async ({ event, step, attempt }) => {
console.log(`Attempt ${attempt + 1}`) // 0-indexed
const result = await step.run('charge-card', async () => {
const charge = await stripe.charges.create({
amount: event.data.amount,
currency: 'usd',
source: event.data.token,
})
if (charge.status !== 'succeeded') {
throw new Error(`Payment failed: ${charge.failure_message}`)
// Throwing causes automatic retry with exponential backoff
}
return charge
})
// Non-retriable errors — don't retry
// throw new NonRetriableError('Card permanently declined')
return { chargeId: result.id }
}
)
Payment processing is the example the docs reach for here because it forces a question every retry system has to answer: not every failure deserves a retry. A network blip talking to Stripe is transient — retrying with backoff is correct. A card that’s permanently declined for insufficient funds is not transient — retrying five more times just delays telling the user the truth, and in the worst case (if the charge call isn’t idempotent on Stripe’s side) risks a duplicate charge. That’s what the commented-out NonRetriableError is for: throwing it tells Inngest “stop, don’t retry this, mark the run as permanently failed” instead of walking through the exponential backoff schedule. In real payment code you’d branch on Stripe’s specific decline codes — insufficient_funds and card_declined are terminal, processing_error or a timeout is worth retrying — and throw the right error type for each.
Stripe’s charges.create call in this snippet is itself only safe to retry because Stripe supports idempotency keys; without one, a network timeout after the charge actually succeeded on Stripe’s end would cause a naive retry to charge the customer twice. This is the same idempotency concern from the earlier section on simple functions, just with higher stakes — anywhere money moves, the step’s retry safety depends on the third-party API’s own idempotency guarantees, not on anything Inngest adds automatically. The attempt value (0-indexed, as the comment notes) is useful for logging and for deliberately changing behavior on later attempts — for example, falling back to a secondary payment provider only once the primary has failed a couple of times.
Concurrency and Rate Limiting
export const processImport = inngest.createFunction(
{
id: 'process-csv-import',
concurrency: {
limit: 3, // max 3 concurrent executions globally
// Or per-user concurrency:
key: 'event.data.userId', // 1 import per user at a time
},
rateLimit: {
limit: 10, // max 10 events
period: '1m', // per minute
key: 'event.data.userId',
},
},
{ event: 'import/started' },
async ({ event, step }) => {
// Process CSV rows
const rows = await step.run('parse-csv', async () => {
return parseCsv(event.data.fileUrl)
})
for (const chunk of chunkArray(rows, 100)) {
await step.run(`process-chunk-${chunk[0].id}`, async () => {
await db.records.createMany({ data: chunk })
})
}
}
)
concurrency and rateLimit solve two different problems that are easy to conflate. Concurrency caps how many instances of a function can be running at the same time — useful when the bottleneck is a resource downstream, like a database connection pool that falls over if forty CSV imports hit it simultaneously. Rate limiting caps how many events are accepted in a time window, which protects against a different failure mode: a buggy client (or a malicious one) firing the same event hundreds of times in a burst. The key field is what makes both of these genuinely useful in a multi-tenant app — a global limit: 3 would mean one very large customer’s imports could starve every other customer’s imports of the three available slots, while key: 'event.data.userId' scopes the limit per user, so one customer’s heavy CSV upload can’t degrade the experience for everyone else. This “noisy neighbor” isolation is the main reason to reach for a per-key limit instead of a flat global one whenever the function handles work on behalf of multiple tenants.
The chunking loop at the bottom (chunkArray(rows, 100)) is also worth calling out: each 100-row chunk gets its own step.run() with a unique step ID (process-chunk-${chunk[0].id}). If row 350 fails to insert, only that chunk’s step is retried — the 300 rows that already succeeded are not re-inserted, which both avoids duplicate rows and avoids re-paying the cost of the chunks that already went through. This pattern (one step per unit of batched work) is generally preferable to a single step.run() wrapping the entire loop whenever the input size is unbounded, since a single giant step means a failure on the very last row forces a full retry of everything before it.
Scheduled Functions (Cron)
// Run every day at 9 AM UTC
export const dailyDigest = inngest.createFunction(
{ id: 'daily-digest-email' },
{ cron: '0 9 * * *' },
async ({ step }) => {
const activeUsers = await step.run('get-active-users', async () => {
return await db.users.findActive()
})
// Send digest to each user (fan-out)
for (const user of activeUsers) {
await step.run(`send-digest-${user.id}`, async () => {
const digest = await buildDigest(user.id)
await sendEmail(user.email, 'daily-digest', digest)
})
}
return { sentTo: activeUsers.length }
}
)
// Other cron examples
// '*/5 * * * *' → every 5 minutes
// '0 0 * * 1' → every Monday at midnight
// '0 0 1 * *' → first day of every month
Cron functions on Inngest are triggered centrally by Inngest’s own scheduler rather than by anything running in your deployment, which sidesteps the classic serverless cron problem of a platform “waking up” your function on a schedule only while it happens to be deployed and healthy. Two things are worth double-checking whenever you write one of these: the schedule is evaluated in UTC unless you configure a timezone, so '0 9 * * *' is 9 AM UTC, not 9 AM in your users’ local time — for a “daily digest” this is usually fine, but for anything time-sensitive to a specific region, confirm the timezone explicitly rather than assuming. Second, exactly like the earlier chunked-import example, the per-user digest loop wraps each user’s email in its own step.run() with a unique ID (send-digest-${user.id}). If the function crashes partway through a run with ten thousand active users, retrying resumes at whichever user it left off on rather than re-sending digests (and burning email-provider quota) to everyone who already received theirs.
It’s also worth knowing what happens if a cron trigger fires while an earlier run of the same function is still executing — long-running fan-out over a large user base can genuinely take longer than the interval between triggers. Inngest doesn’t automatically serialize these for you; if overlapping runs would be a problem (e.g., two runs both trying to increment a “digest sent” counter), add a concurrency limit on the function the same way you would for an event-triggered one.
AI Workflow (LLM Pipeline)
export const generateBlogPost = inngest.createFunction(
{
id: 'generate-blog-post',
timeouts: { finish: '10m' }, // LLM calls can be slow
},
{ event: 'content/requested' },
async ({ event, step }) => {
const { topic, userId } = event.data
// Step 1: Generate outline
const outline = await step.run('generate-outline', async () => {
const response = await openai.chat.completions.create({
model: 'gpt-4',
messages: [
{ role: 'user', content: `Create an outline for a blog post about: ${topic}` }
],
})
return response.choices[0].message.content
})
// Step 2: Generate each section in parallel
const sections = outline.split('\n').filter(Boolean)
const content = await step.run('generate-content', async () => {
return await Promise.all(
sections.map(section =>
openai.chat.completions.create({
model: 'gpt-4',
messages: [{ role: 'user', content: `Write the "${section}" section...` }],
}).then(r => r.choices[0].message.content)
)
)
})
// Step 3: Save to database
const post = await step.run('save-post', async () => {
return await db.posts.create({
userId,
topic,
content: content.join('\n\n'),
status: 'draft',
})
})
// Step 4: Notify user
await step.run('notify-user', async () => {
await sendEmail(userId, 'post-ready', { postId: post.id })
})
return { postId: post.id }
}
)
This example is where the durable-execution model earns its keep the most, because LLM API calls fail in exactly the ways that make naive retries expensive: they’re slow (multi-second round trips per call), they cost real money per token whether or not the response is used, and providers occasionally return 500s or time out under load even when nothing is wrong with the request. Because generate-outline is its own step, a failure in the content generation step never causes the outline to be regenerated — you don’t pay for a second outline API call, and you don’t risk getting a materially different outline on retry that no longer matches the sections that were already planned around it. Splitting a single logical “write a blog post” operation into four checkpointed steps turns an expensive, flaky, four-stage pipeline into one where each stage’s cost is paid exactly once on success.
The timeouts: { finish: '10m' } configuration matters more here than it would for a typical CRUD function, because LLM calls are one of the few operations whose latency is genuinely unpredictable — a gpt-4 completion can take anywhere from one second to well over a minute depending on load and output length, and the section-generation step is making several such calls concurrently via Promise.all. Without an explicit timeout tuned for this reality, you’re stuck with whatever the platform default is, which is usually calibrated for fast API handlers, not multi-call LLM pipelines. It’s also worth noting the same caveat from the fan-out section applies to Promise.all here: if one section’s completion call fails, the whole generate-content step fails and retries regenerate every section, not just the failed one. For a pipeline with many sections or an expensive model, giving each section its own step.run() trades a bit of code verbosity for materially lower retry cost.
Local Development
# Start your app
npm run dev # runs on port 3000
# Start Inngest dev server (separate terminal)
npx inngest-cli@latest dev
# Open Inngest dashboard
open http://localhost:8288
# The dashboard shows:
# - All events received
# - Function executions
# - Each step's input/output
# - Retry history
# - Timeline view
The dev server dashboard is genuinely worth building the habit of keeping open while you write functions, because it’s the fastest way to catch the two most common early mistakes: a function that never registers (usually a missing entry in the functions array passed to serve(), or the route handler not matching the URL the dev server expects), and a step function whose steps silently re-run more than expected because a value that should have been wrapped in step.run() was computed outside of it. The timeline view makes both of these visible immediately — you’ll see a function missing from the list entirely in the first case, or a step re-executing with the same “attempt” markers in the second, long before either bug would be obvious from application logs alone.
One gotcha specific to local development: the dev server bypasses signature verification, which is convenient for quick iteration but means a function that only works because the signing key check is disabled will pass every local test and then fail — or worse, silently accept unsigned requests if you also forgot to set the key in production — the moment it’s deployed. Treat a clean local run as necessary but not sufficient; verify the signing key is configured before considering a function production-ready.
Frequently Asked Questions (FAQ)
Q. When would I use Inngest instead of just queuing a job with BullMQ or SQS?
A. Reach for Inngest when the deployment target has no place to run a persistent worker — Vercel, Cloudflare Pages, or any platform where functions are short-lived by design — or when the workflow itself needs to pause for hours or days between steps, which is awkward to express with a traditional queue-plus-worker setup. If you already run a long-lived Node process and just need a simple FIFO job queue with no multi-day pauses, BullMQ against a Redis you already operate can be simpler to reason about, since there’s no third-party HTTP round trip in the execution path. The trade-off is infrastructure ownership either way: Inngest removes the worker but adds a dependency on an external service calling your endpoint; BullMQ keeps everything in your own infrastructure but means you’re the one who scales, monitors, and restarts the worker.
Related Articles
- Next.js App Router: SSR vs SSG vs ISR
- Next.js 15 Internals: App Router vs Pages Router, RSC, Rendering Decisions and the Four Caches
- Redis Internals and Usage: The Event Loop, Encodings, RDB vs AOF, Replication and Cluster