Background Jobs with Bull and Redis: Retries, Priorities, Delays and Cron Jobs

Key takeaways

Bull is a Redis-based queue for Node.js that handles job processing, retries, priorities, and delayed jobs. It's battle-tested and used by thousands of production apps.

Introduction

Bull is a Redis-based queue for Node.js that handles distributed job processing. It’s perfect for tasks that are too slow or unreliable to run in API request handlers.

The core problem Bull solves is temporal decoupling. An HTTP request handler and the work it triggers do not have to live in the same execution window. When a user signs up, the response to “your account was created” does not need to wait for a welcome email to actually leave your SMTP relay, a thumbnail to be resized, or a webhook to fire on a third-party system that might be slow or down. If you run that work inline, the request thread is at the mercy of every downstream dependency’s latency and failure modes. A queue absorbs that risk: the API call becomes “durably record that this work needs to happen,” which is a fast, local, almost-never-fails operation, and the actual work happens asynchronously, with retries, on infrastructure you can scale independently of your web tier.

This matters more than it looks once you think about failure. Without a queue, a downstream outage becomes an upstream outage — if your email provider degrades, your signup endpoint degrades with it. With a queue, the job simply waits in Redis until a worker can process it successfully. The user-facing latency budget of the API is protected, and transient failures become retry attempts instead of user-visible errors.

Why Use Queues?

Without queue (blocking):

app.post('/send-email', async (req, res) => {
  await sendEmail(req.body); // Takes 2-3 seconds
  res.json({ message: 'Email sent' });
});

// User waits 2-3 seconds for response
// If email fails, request fails

With queue (non-blocking):

app.post('/send-email', async (req, res) => {
  await emailQueue.add(req.body); // <10ms
  res.json({ message: 'Email queued' });
});

// User gets instant response
// Email processed in background
// Automatic retries on failure

Installation

npm install bull

Requires Redis:

# Docker
docker run -d -p 6379:6379 redis

# macOS
brew install redis
redis-server

# Ubuntu
sudo apt install redis-server

Basic Queue

const Queue = require('bull');

// Create queue
const emailQueue = new Queue('email', {
  redis: {
    host: '127.0.0.1',
    port: 6379,
  }
});

// Add job to queue
await emailQueue.add({
  to: '[email protected]',
  subject: 'Welcome!',
  body: 'Thanks for signing up',
});

// Process jobs
emailQueue.process(async (job) => {
  console.log('Processing job:', job.id);
  await sendEmail(job.data);
  console.log('Job completed:', job.id);
});

Behind that small .add() / .process() pair sits a fair amount of Redis machinery. Every Bull queue is backed by a set of Redis keys: a wait list holding job IDs ready to be picked up, an active list for jobs currently being processed, completed/failed sets for terminal states, and a hash per job storing its data and metadata. When you call .add(), Bull atomically pushes the job onto the wait list via a Lua script (so the enqueue itself is a single, race-free Redis operation even under heavy concurrency). Workers use a blocking BRPOPLPUSH-style pattern to pull jobs off wait and onto active, which is what gives Bull near-instant pickup latency instead of polling on an interval.

This architecture has a direct consequence you need to design around: Bull gives you at-least-once delivery, not exactly-once. If a worker crashes after it has moved a job to active but before it acknowledges completion, Bull’s stalled-job recovery will eventually put that job back on the queue and another worker will pick it up — which means the same job can run twice. This is not a bug to work around; it is the fundamental trade-off every Redis- or broker-backed queue makes, because true exactly-once delivery across a network boundary is not achievable without cooperation from the consumer. The practical implication is that every job handler must be idempotent — running it twice with the same input must produce the same end state as running it once. We come back to concrete idempotency patterns later in this guide, but keep it in mind from the very first .process() callback you write.

Job Options

await emailQueue.add({
  to: '[email protected]',
  subject: 'Hello',
}, {
  // Job options
  attempts: 3,              // Retry up to 3 times
  backoff: {
    type: 'exponential',    // Exponential backoff
    delay: 2000,            // Start with 2 second delay
  },
  delay: 5000,              // Delay job by 5 seconds
  priority: 1,              // Priority (1 = highest)
  timeout: 30000,           // Timeout after 30 seconds
  removeOnComplete: true,   // Remove from Redis when done
  removeOnFail: false,      // Keep failed jobs for debugging
});

The attempts and backoff options are where most of the reliability engineering in a queue-based system actually lives, and the choice between backoff strategies is not cosmetic. A fixed delay backoff (retry every N seconds regardless of attempt count) is appropriate for jobs that fail due to short, predictable blips — a lock held by another process, a rate limit that resets on a fixed window. An exponential backoff, as used above, is the right default for anything talking to an external service, because it spaces retries out in a way that stops hammering a system that is already struggling (1st retry after 2s, 2nd after 4s, 3rd after 8s, and so on). Retrying too aggressively against a degraded downstream service is a well-known way to turn a brief outage into a self-inflicted denial-of-service against your own dependency — and against yourself, since Bull’s Redis connection and CPU also pay for every retry attempt.

attempts should be set thoughtfully per job type rather than left at a single global default. A payment webhook might warrant 5-8 attempts with a long exponential ceiling, because losing that job silently is expensive. A best-effort analytics ping might warrant 1-2 attempts, because retrying it aggressively costs more in Redis and worker capacity than the data is worth. removeOnFail: false is the right default while you are still building confidence in a job type — it keeps failed jobs in Redis (in the failed set) for inspection — but for high-volume, low-value jobs you will want to bound this (e.g. removeOnFail: { count: 1000 } in later Bull versions, or manual queue.clean() calls as shown further below) so failed jobs don’t grow Redis memory unbounded.

Priority Queues

// Add jobs with different priorities
await queue.add({ task: 'critical' }, { priority: 1 });    // Highest
await queue.add({ task: 'normal' }, { priority: 5 });
await queue.add({ task: 'low' }, { priority: 10 });        // Lowest

// High priority jobs processed first

Priority in Bull is implemented with a Redis sorted set rather than the plain list used by the default queue, and that has a real performance cost: every .add() and every pickup becomes an O(log n) sorted-set operation instead of an O(1) list push/pop. For queues processing a handful of jobs per second this is irrelevant. For queues pushing thousands of jobs per second, mixing priority into every job adds measurable overhead, which is why many production setups reserve priority for a small, genuinely urgent subset of jobs and route everything else through a plain FIFO queue (or, as shown in section 9, a separate named processor with its own concurrency). Treat priority as a tool for “this specific class of job needs to jump the line,” not as a default you attach to every job you enqueue.

Delayed Jobs

// Process job after 1 hour
await queue.add({ task: 'reminder' }, {
  delay: 60 * 60 * 1000, // 1 hour in milliseconds
});

// Process at specific time
const scheduledTime = new Date('2026-12-25T00:00:00Z');
await queue.add({ task: 'holiday-email' }, {
  delay: scheduledTime.getTime() - Date.now(),
});

Delayed jobs also live in a Redis sorted set (scored by the timestamp at which they become eligible to run), and a background process in Bull periodically moves due jobs from that “delayed” set into the active wait list. Two things follow from this. First, delay precision is not real-time-guaranteed — under normal conditions jobs fire within a second or two of their target time, but under heavy Redis load or with many workers polling, you can see drift of a few seconds. Do not use Bull delays for anything that needs sub-second timing accuracy. Second, computing a delay as scheduledTime.getTime() - Date.now() at enqueue time means clock skew between the enqueuing process and Redis (or a long-running process that enqueues far in advance) can shift the effective fire time — for delays measured in weeks or months, prefer storing the target timestamp in your own database and using a shorter periodic re-check job instead of one giant Bull delay.

Repeatable Jobs (Cron)

// Run every 5 minutes
await queue.add({ task: 'cleanup' }, {
  repeat: {
    every: 5 * 60 * 1000, // 5 minutes
  }
});

// Cron syntax
await queue.add({ task: 'daily-report' }, {
  repeat: {
    cron: '0 9 * * *', // Every day at 9 AM
  }
});

// With timezone
await queue.add({ task: 'morning-email' }, {
  repeat: {
    cron: '0 8 * * *',
    tz: 'America/New_York',
  }
});

Repeatable jobs are keyed internally by the combination of job name, cron expression (or every interval), and options — Bull deduplicates by that key, so calling queue.add() with the same repeat configuration on every server restart does not create duplicate schedules. This is convenient, but it also means a subtle bug is easy to introduce: if you run this add() call from every instance in a horizontally-scaled deployment (every app server calling it on boot, for example), that’s fine for the schedule itself, but you must make sure only one worker actually processes each firing — otherwise a “daily report” job fires once but N workers race to pick it up, and depending on your processor’s concurrency you can end up sending the same report N times before idempotency checks (if any) catch it. The safe pattern is to register the repeatable schedule from a single source of truth (a deploy step or a dedicated scheduler process) and let any number of workers safely .process() the resulting jobs, since Bull will only ever enqueue one job per firing regardless of how many workers are listening.

Job Progress

// Report progress
emailQueue.process(async (job) => {
  await job.progress(0);
  
  const emails = job.data.emails;
  
  for (let i = 0; i < emails.length; i++) {
    await sendEmail(emails[i]);
    await job.progress(Math.round(((i + 1) / emails.length) * 100));
  }
  
  return { sent: emails.length };
});

// Listen for progress
queue.on('progress', (job, progress) => {
  console.log(`Job ${job.id} is ${progress}% done`);
});

Job Events

// Job completed
queue.on('completed', (job, result) => {
  console.log(`Job ${job.id} completed with result:`, result);
});

// Job failed
queue.on('failed', (job, err) => {
  console.error(`Job ${job.id} failed:`, err.message);
});

// Job stalled (taking too long)
queue.on('stalled', (job) => {
  console.warn(`Job ${job.id} stalled`);
});

// Job removed
queue.on('removed', (job) => {
  console.log(`Job ${job.id} removed`);
});

// Global completed (all jobs)
queue.on('global:completed', (jobId, result) => {
  console.log(`Job ${jobId} completed globally`);
});

The distinction between the local events (completed, failed, stalled) and the global:* variants matters in any deployment with more than one process. Local events only fire on the exact process instance that ran the job — useful inside the worker process itself, but useless if you want, say, an API server process to push a WebSocket notification to a browser when any worker anywhere finishes a job. The global:* events are published through Redis pub/sub, so every process connected to the same queue name receives them regardless of which process actually did the work. This is the mechanism you’d use to bridge Bull job completion into a real-time UI: an API process subscribes to global:completed, looks up which client is waiting on that jobId, and pushes the result over a WebSocket or SSE connection — without the API process ever needing direct knowledge of which worker handled it.

stalled deserves special attention because it is a symptom, not just an event. A job is marked stalled when Bull’s lock-renewal mechanism fails to hear back from the worker within the expected interval — usually because the event loop was blocked by synchronous, CPU-heavy work, or because the process crashed outright. A stalled job is automatically requeued and retried (subject to maxStalledCount), which is another source of the “job ran twice” scenario idempotency has to cover. If you see stalled jobs in production, the fix is almost always to move CPU-intensive work off the main event loop (worker threads, a child process, or a separate queue with lower concurrency) rather than to just increase timeouts.

Multiple Processors

// Processor 1: High priority
queue.process('high-priority', 5, async (job) => {
  // Process up to 5 high-priority jobs concurrently
  await processHighPriority(job.data);
});

// Processor 2: Low priority
queue.process('low-priority', 2, async (job) => {
  // Process up to 2 low-priority jobs concurrently
  await processLowPriority(job.data);
});

// Add jobs
await queue.add('high-priority', { task: 'urgent' });
await queue.add('low-priority', { task: 'background' });

Named processors are how you get multiple, independently-tuned job types sharing one Redis queue without one slow job type starving another. Note the concurrency values passed as the second argument — 5 for high-priority, 2 for low-priority. This is not thread-based parallelism; Node.js is still single-threaded, so “concurrency 5” means up to 5 jobs can be in flight at once, interleaved on the event loop while each awaits I/O (a database call, an HTTP request, a Redis read). This makes concurrency tuning mostly an I/O-bound-work problem: for jobs that spend most of their time waiting on a network call, high concurrency (10-50) is often safe and beneficial. For jobs that do real CPU work (image resizing, PDF generation, cryptographic hashing), high concurrency does nothing but contend for the same event loop and will make every job on that process slower, including unrelated ones — those job types are better run in a dedicated worker process (via NODE_ENV-driven separate entrypoints, or Node’s worker_threads) rather than sharing a process with I/O-bound jobs.

Real-World Example: Email Service

const Queue = require('bull');
const nodemailer = require('nodemailer');

// Create queue
const emailQueue = new Queue('email', {
  redis: process.env.REDIS_URL,
  defaultJobOptions: {
    attempts: 3,
    backoff: {
      type: 'exponential',
      delay: 2000,
    },
    removeOnComplete: 100, // Keep last 100 completed
    removeOnFail: false,   // Keep all failed for debugging
  }
});

// Email transporter
const transporter = nodemailer.createTransport({ /* config */ });

// Process emails
emailQueue.process(async (job) => {
  const { to, subject, html } = job.data;
  
  console.log(`Sending email to ${to}`);
  
  try {
    await transporter.sendMail({ to, subject, html });
    return { sent: true, to };
  } catch (error) {
    console.error(`Failed to send email to ${to}:`, error);
    throw error; // Will retry
  }
});

// Listen for events
emailQueue.on('completed', (job, result) => {
  console.log(`Email sent to ${result.to}`);
});

emailQueue.on('failed', (job, err) => {
  console.error(`Email to ${job.data.to} failed after ${job.attemptsMade} attempts`);
  // Notify admin or log to monitoring service
});

// API endpoint
app.post('/api/send-email', async (req, res) => {
  const { to, subject, html } = req.body;
  
  const job = await emailQueue.add({
    to,
    subject,
    html,
  });
  
  res.json({
    message: 'Email queued',
    jobId: job.id,
  });
});

// Check job status
app.get('/api/jobs/:id', async (req, res) => {
  const job = await emailQueue.getJob(req.params.id);
  
  if (!job) {
    return res.status(404).json({ error: 'Job not found' });
  }
  
  const state = await job.getState();
  const progress = job.progress();
  
  res.json({
    id: job.id,
    state,
    progress,
    data: job.data,
  });
});

This example is worth pausing on because it shows the two halves of a production queue system working together: the /api/send-email endpoint does almost nothing (a single .add() call, back in under 10ms), while the emailQueue.process() callback does the actual work off the request path, with retries handled automatically by the defaultJobOptions configured on the queue. The /api/jobs/:id polling endpoint is a common pattern for surfacing async job state to a client that needs to know when work finishes — it works well for a handful of concurrent jobs per user, though at higher volume you’d typically replace polling with the global:completed pub/sub event described earlier, pushed over a WebSocket instead.

The failed handler here logs and comments “notify admin,” which points at a gap many teams leave unaddressed until it bites them: what happens to a job after it exhausts every retry attempt? By default, a permanently-failed job just sits in Redis’s failed set — it does not vanish, but it also does not do anything on its own. Treat this the way you would treat a dead-letter queue in SQS or RabbitMQ: after the failed event fires and job.attemptsMade equals the job’s configured attempts, push a record into your own database or alerting system (PagerDuty, Slack, an admin dashboard) so a human actually sees it. Relying on someone to periodically check Bull Board (section 12) for red entries does not scale past a handful of jobs — you want a push notification for anything financially or operationally significant, and Bull Board as the tool you open after being paged, to investigate.

Rate Limiting

const queue = new Queue('api-calls', {
  redis: process.env.REDIS_URL,
  limiter: {
    max: 100,        // Max 100 jobs
    duration: 60000, // Per 60 seconds
  }
});

// Jobs are automatically rate-limited
await queue.add({ url: 'https://api.example.com/data' });

It is worth being precise about how limiter differs from concurrency, because they solve different problems and beginners frequently conflate them. Concurrency controls how many jobs a single worker process runs in parallel at any instant — it is a local resource-management knob for your own CPU/event-loop/connection-pool budget. Rate limiting controls how many jobs the entire queue completes per time window across all workers combined — it exists to protect an external system (a third-party API with a documented quota, a database that chokes above a certain write rate) rather than to protect your own process. You typically want both configured together on any queue that calls a metered external API: concurrency keeps a single worker from being resource-starved by too many parallel in-flight calls, and the rate limiter keeps the aggregate traffic across every worker under the third party’s published quota, since spinning up more worker processes to increase throughput would otherwise silently multiply your effective request rate and trip the provider’s rate limiting or get your API key temporarily blocked.

Bull Board (UI Dashboard)

npm install bull-board
const { createBullBoard } = require('@bull-board/api');
const { BullAdapter } = require('@bull-board/api/bullAdapter');
const { ExpressAdapter } = require('@bull-board/express');

const serverAdapter = new ExpressAdapter();
serverAdapter.setBasePath('/admin/queues');

createBullBoard({
  queues: [
    new BullAdapter(emailQueue),
    new BullAdapter(imageQueue),
    new BullAdapter(reportQueue),
  ],
  serverAdapter: serverAdapter,
});

app.use('/admin/queues', serverAdapter.getRouter());

// Visit http://localhost:3000/admin/queues

Bull Board matters more than “nice-to-have UI” once a system has more than one or two queues, because Redis gives you almost no visibility on its own — everything is opaque keys and sorted sets unless you go inspecting with redis-cli. In production, gate this route behind authentication (it is trivially exposed if mounted on a public Express app without a middleware check) — the dashboard can pause queues, retry jobs, and delete data, so treat /admin/queues with the same access-control seriousness as any other admin panel, not as a read-only status page.

Job Cleanup

// Remove completed jobs older than 1 day
await queue.clean(24 * 3600 * 1000, 'completed');

// Remove failed jobs older than 1 week
await queue.clean(7 * 24 * 3600 * 1000, 'failed');

// Remove all jobs
await queue.empty();

// Automatic cleanup
emailQueue.on('completed', async (job) => {
  await job.remove();
});

Job cleanup is not optional housekeeping — it is a correctness concern for Redis memory. By default Bull keeps every completed and failed job’s full data and result payload in Redis indefinitely unless you set removeOnComplete/removeOnFail or run .clean() periodically. On a low-volume queue this is invisible; on a queue processing tens of thousands of jobs per day, unbounded retention will grow Redis memory usage steadily until you hit maxmemory and Redis starts evicting keys (or refusing writes, depending on your eviction policy) — which can silently corrupt queue state, not just “run out of disk.” The safest default for high-volume, low-value jobs is removeOnComplete: true combined with removeOnFail: false (or a bounded count) as shown in the email service example — keep failures for debugging, discard successes immediately since their audit trail typically belongs in your application database, not in Redis.

Concurrency

// Process 5 jobs concurrently
queue.process(5, async (job) => {
  return await processJob(job.data);
});

// Named processors with different concurrency
queue.process('email', 10, async (job) => {
  // Up to 10 email jobs at once
});

queue.process('image', 3, async (job) => {
  // Up to 3 image jobs at once (CPU intensive)
});

Graceful Shutdown

// Handle shutdown
async function shutdown() {
  console.log('Shutting down gracefully...');
  
  // Close queue (waits for active jobs)
  await emailQueue.close();
  
  process.exit(0);
}

process.on('SIGTERM', shutdown);
process.on('SIGINT', shutdown);

queue.close() matters because container orchestrators (Kubernetes, ECS, systemd) send SIGTERM before killing a process, and give it a bounded grace period (Kubernetes defaults to 30 seconds) before following up with SIGKILL. If you don’t handle SIGTERM and instead let the process die immediately, any job that is mid-processing when the signal arrives gets abandoned — Bull’s stalled-job detection will eventually notice the lock isn’t being renewed and requeue it, so the job doesn’t disappear, but it does mean the work restarts from scratch, and if the handler had already produced a side effect (partially sent a batch of emails, partially written a file) before being killed, that side effect happened without the job being marked complete. queue.close() waits for active jobs to finish (or until its own internal timeout) before releasing the Redis connection, which is what turns an abrupt kill into a clean handoff. In practice you also want to bound this wait — pair it with a hard timeout so a single stuck job doesn’t block your entire deployment rollout from completing within the orchestrator’s grace period.

Redis Connection Loss and Job Data Limits

Two operational realities deserve explicit attention because they rarely show up in tutorials but reliably show up in production incidents.

Redis connection loss. Bull’s underlying ioredis client will attempt to reconnect automatically on a dropped connection, and by default it retries with backoff rather than failing immediately. While disconnected, .add() calls will queue up in memory waiting for the connection (or eventually reject, depending on enableOfflineQueue and your timeout configuration), and workers stop pulling new jobs — they don’t crash, they just idle. The danger case is a managed Redis failover (e.g. AWS ElastiCache or Redis Cloud promoting a replica), where the connection can flip a few times in quick succession; make sure your queue’s redis options include reasonable retryStrategy/maxRetriesPerRequest settings rather than relying on ioredis defaults blindly, and add a queue.on('error', ...) handler — without one, connection errors surface only as silent stalls with no log line telling you why jobs stopped moving.

Job data serialization limits. Job data is serialized to JSON and stored in a Redis hash field. This means anything you pass to .add() must be JSON-serializable — no Buffer, no class instances with methods, no circular references, and no undefined values (they’re silently dropped by JSON.stringify). It also means job payload size is bounded in practice by Redis’s per-value limits and, more importantly, by performance: a 5 MB payload stored per job, multiplied across thousands of queued jobs, becomes a meaningful chunk of your Redis memory budget and slows every read/write of that job. The correct pattern for “large” job data (a file to process, a big JSON export) is to store the reference in the job — an S3 key, a database row ID — and have the worker fetch the actual payload during processing, not to stuff the payload itself into the job.

Making retries safe and spotting a backlog

Idempotent Jobs

// Ensure jobs can be retried safely
queue.process(async (job) => {
  const { userId, emailType } = job.data;
  
  // Check if already sent
  const sent = await db.emailLogs.findOne({ userId, emailType });
  if (sent) {
    console.log('Email already sent, skipping');
    return;
  }
  
  // Send email
  await sendEmail(userId, emailType);
  
  // Log
  await db.emailLogs.create({ userId, emailType, sentAt: new Date() });
});

The check-before-write pattern above works, but it has a race window: if two workers pick up the same job simultaneously (the stalled-job double-processing scenario from section 2), both can pass the findOne check before either writes the log, and you end up sending the email twice anyway. For jobs where that race matters, use a database-level uniqueness constraint instead of an application-level check — a unique index on (userId, emailType) in the emailLogs table turns the second insert into a constraint violation that you catch and treat as “already handled,” which is race-safe in a way that read-then-write never is. The general principle: idempotency guarantees are only as strong as the mechanism enforcing them, and only a database constraint, a Redis SETNX, or an idempotency key checked atomically actually closes the race that concurrent retries can hit.

Queue health signals

setInterval(async () => {
  const waiting = await queue.getWaitingCount();
  const active = await queue.getActiveCount();
  const failed = await queue.getFailedCount();
  
  console.log(`Queue status - Waiting: ${waiting}, Active: ${active}, Failed: ${failed}`);
  
  // Alert if too many jobs waiting
  if (waiting > 1000) {
    console.warn('Queue backlog too high!');
    // Send alert
  }
}, 60000); // Every minute

The specific metrics worth alerting on go beyond a raw waiting-count threshold. waiting growing steadily over time (rather than bursting and draining) means your consumers can’t keep up with producers — either raise concurrency, add more worker processes, or find out why processing got slower. A rising failed count relative to completed is a leading indicator of a downstream dependency degrading before it fully breaks — catching that trend early, before every job in the queue starts failing, is the difference between a quiet fix and a paging incident. And active staying elevated without completed increasing is the signature of stuck or slow-running jobs (often the CPU-blocking-the-event-loop problem from section 8) rather than a throughput problem — each of these calls for a different fix, which is why a single “queue backlog too high” alert is a reasonable starting point but not sufficient on its own for a queue carrying real production traffic.

The failure that runs a job twice: stalled jobs

Bull holds a lock on each active job and renews it periodically from the worker process. If the processor blocks the Node.js event loop, for example with CPU-heavy synchronous work such as image processing or large JSON parsing, the lock is not renewed in time. Bull then treats the job as stalled and moves it back to waiting, and another worker picks it up while the first one is still running it. The symptom is duplicate emails or double charges with nothing in the logs that looks like an error.

Two changes prevent it. Run CPU-heavy processors as sandboxed processors, by passing the path of a processor file to queue.process instead of a function, so the work happens in a separate process and the lock keeps renewing. And make every job idempotent, for example by recording a job or business key before performing the side effect, because a job can legitimately run more than once after a crash anyway.

Next Steps:

Resources:


Frequently Asked Questions (FAQ)

Q. When would I reach for Bull instead of just calling a function directly?

A. Whenever the work is slow, unreliable, or has a failure mode you want to retry automatically — sending email, calling a third-party API, generating a report, resizing an image. If the request handler can afford to fail the whole request when that single step fails, you may not need a queue; if you want the API response to be fast and the work to survive transient failures, that’s exactly what Bull is for.

Q. My jobs are being processed twice — is that a bug?

A. No — it’s expected behavior. Bull (like most Redis- or broker-backed queues) guarantees at-least-once delivery, not exactly-once. A worker crash, a stalled-job timeout, or a network blip during acknowledgment can all cause a job to run more than once. The fix is to make every job handler idempotent (see the Idempotent Jobs section above), not to try to eliminate duplicate delivery entirely.

Q. Should I use Bull or BullMQ for a new project?

A. Use BullMQ. It’s the actively maintained, TypeScript-first successor with better Redis Cluster support and a cleaner API; Bull is effectively in maintenance mode. The concepts in this guide — queues, jobs, retries, backoff, concurrency, rate limiting — carry over directly, so this guide is still useful context even if you write the code against BullMQ.

Q. How do I know if my queue is falling behind in production?

A. Watch trends, not just snapshots: a waiting count that climbs steadily rather than draining means consumers can’t keep pace with producers; a rising failed-to-completed ratio signals a downstream dependency degrading; active staying high without completed increasing points to stuck or slow jobs. Wire these into your existing monitoring rather than relying on manually checking Bull Board.