Microservices Architecture: Service Boundaries, Messaging, Sagas, and Resilience
Key takeaways
Microservices solve specific scaling and team-autonomy problems — but introduce distributed systems complexity. This guide covers architecture patterns, communication strategies, and the honest trade-offs you need to weigh before committing.
Microservices vs Monolith: The Honest Trade-off
Before choosing microservices, understand what you’re trading:
| Monolith | Microservices | |
|---|---|---|
| Development speed | Fast initially | Slower initially |
| Deployment | Simple, one unit | Complex, many units |
| Scaling | Scale everything | Scale independently |
| Team autonomy | Low | High |
| Debugging | Easy (local) | Hard (distributed) |
| Data consistency | Easy (single DB) | Hard (distributed transactions) |
| Operational overhead | Low | High |
Rule of thumb: Start with a monolith. Extract services when you have a specific, validated reason (team bottleneck, scaling requirement, technology mismatch) — not because microservices sound modern.
The row that surprises teams most is data consistency. In a monolith, “create the order and reserve the stock” is one database transaction: either both happen or neither does. Once orders and inventory live in different services with different databases, that guarantee is gone, and every multi-step operation needs an explicit plan for partial failure (section 5). The second surprise is that a function call which never failed becomes a network call that can time out, be retried, arrive twice, or succeed on the server while the client sees an error. Most of the patterns in this guide exist to deal with those two facts.
A useful middle ground is the modular monolith: one deployable, but with modules that own their tables and talk through explicit interfaces. If the module boundaries hold up for a while, extracting one into a service later is mostly mechanical. If they keep leaking, you have learned that cheaply, before paying for a network in between.
Service Decomposition
Domain-Driven Design (DDD) approach
Split services along bounded contexts — areas of the domain with their own language and models:
E-commerce system
├── user-service — registration, auth, profiles
├── product-service — catalog, inventory, pricing
├── order-service — cart, orders, order history
├── payment-service — payment processing, refunds
├── notification-service — email, SMS, push
└── shipping-service — shipping, tracking, returns
Signs a service boundary is wrong:
- Services constantly need to call each other to complete one operation
- You always deploy multiple services together
- A change in one service always requires a change in another
Size heuristics
- Too small: a service that only wraps a database table with CRUD
- Too large: a service that a team can’t understand in a day
- Right: a service owned by one team, deployable independently
The failure mode I see most often in decompositions is splitting by technical layer or by entity instead of by business capability: a “customer-data service” that every other service calls for every request. It looks tidy on a diagram and behaves like a shared database with extra latency. The test that catches it is to take one real user action, such as “place an order”, and count how many services must be up and respond synchronously for it to succeed. If the answer is most of them, the boundaries are drawn around nouns rather than around work that can happen independently.
Communication Patterns
Synchronous (REST / gRPC)
Client → API Gateway → Order Service ──▶ [sync call] → Inventory Service
└─▶ [sync call] → Payment Service
Good for: queries, reads, user-facing requests that need immediate response.
// Order service calling payment service
async function processOrder(order) {
// Synchronous HTTP call
const paymentResult = await fetch('http://payment-service/charge', {
method: 'POST',
body: JSON.stringify({ amount: order.total, userId: order.userId }),
}).then(r => r.json())
if (!paymentResult.success) throw new Error('Payment failed')
return updateOrderStatus(order.id, 'confirmed')
}
Problem: if Payment Service is down, Order Service fails too — cascading failures.
The slow case is worse than the down case. fetch has no default timeout, so if Payment Service accepts connections but responds slowly, every in-flight order request waits, the Order Service’s connection pool and memory fill up, and it starts failing for requests that never touch payment. Always set a deadline, for example fetch(url, { signal: AbortSignal.timeout(2000) }) in Node 18+, and keep it shorter than the caller’s own deadline. The example also omits the Content-Type: application/json header, which many servers need before they parse the body.
There is also an ambiguity you cannot remove: if the call times out, you do not know whether the card was charged. Payment endpoints therefore usually accept an idempotency key (for example the order ID) so that a retry of the same charge is recognized and not executed twice.
Asynchronous (Events / Message Queue)
Order Service → [event: OrderPlaced] → Kafka/RabbitMQ
├──▶ Payment Service (consumes)
├──▶ Inventory Service (consumes)
└──▶ Notification Service (consumes)
Good for: writes, workflows, processes that don’t need an immediate response.
// Order service publishes an event
await kafka.publish('orders', {
event: 'OrderPlaced',
orderId: order.id,
userId: order.userId,
items: order.items,
total: order.total,
})
// Returns immediately — doesn't wait for payment/inventory
// Payment service subscribes
kafka.subscribe('orders', async (message) => {
if (message.event === 'OrderPlaced') {
const result = await chargeCustomer(message.userId, message.total)
await kafka.publish('payments', {
event: result.success ? 'PaymentSucceeded' : 'PaymentFailed',
orderId: message.orderId,
})
}
})
(kafka.publish/kafka.subscribe are simplified pseudo-APIs; with kafkajs these would be producer.send and consumer.run.)
Asynchronous messaging removes the temporal coupling, since Payment Service can be down for a minute and catch up afterwards, but it introduces three problems you have to design for:
- Duplicates. Kafka and RabbitMQ deliver at-least-once in their normal configurations. If the payment consumer crashes after charging but before committing the offset, it will see
OrderPlacedagain. Consumers must be idempotent, typically by recording processed event IDs or order IDs in their own database. - The dual write. The order service must save the order and publish the event. If it writes to its database and then crashes before publishing, the order exists but nobody ever charges for it; publish first and the reverse happens. The standard fix is the transactional outbox: write the event to an
outboxtable in the same database transaction as the order, and have a separate relay (or change-data-capture tool such as Debezium) publish rows from that table. - Ordering. Kafka only orders messages within a partition. Using
orderIdas the message key keeps all events for one order in order; events for different orders can interleave.
The dual-write bug is the one that tends to reach production, because it only shows up when a process dies at exactly the wrong moment, and the symptom is a handful of “stuck” orders that no log explains.
API Gateway
The gateway is the single entry point — handles cross-cutting concerns so services don’t have to:
Client
└─▶ API Gateway
├── Auth (JWT validation)
├── Rate limiting
├── SSL termination
├── Request routing
├── Load balancing
├── Request/response transformation
│ ├── /api/users/* → user-service
│ ├── /api/products/* → product-service
│ └── /api/orders/* → order-service
Popular options: Kong, Nginx, AWS API Gateway, Traefik, Envoy.
# Kong route example
services:
- name: user-service
url: http://user-service:3001
routes:
- name: users-route
paths: ["/api/users"]
methods: ["GET", "POST", "PUT", "DELETE"]
plugins:
- name: jwt # Auth
- name: rate-limiting
config:
minute: 100
This is a fragment of Kong’s declarative (DB-less) configuration; a complete file also needs _format_version: "3.0" at the top, and the jwt plugin needs consumers with credentials before any request will pass.
Keep the gateway thin. It is tempting to put business logic there (merging responses, applying discounts, checking order state) because it is the one place every request passes through. Once that happens, every feature change requires a gateway deploy, and the gateway becomes the shared component that all teams queue behind, which is the monolith problem again. If different clients need differently shaped responses, a per-client Backend-for-Frontend service is usually a cleaner place for aggregation. Also remember that validating a JWT at the gateway does not remove the need for services to authorize the request: the gateway knows who the caller is, but only the order service knows whether that user may see order 123.
Service Discovery
Services need to find each other without hardcoded IPs. In Kubernetes, this is built-in via DNS:
# Kubernetes Service — DNS: user-service.default.svc.cluster.local
apiVersion: v1
kind: Service
metadata:
name: user-service
spec:
selector:
app: user-service
ports:
- port: 80
targetPort: 3001
Services call each other by name:
const user = await fetch('http://user-service/users/123').then(r => r.json())
Outside Kubernetes: use Consul or AWS Cloud Map for service registry.
The short name user-service resolves because the pod’s DNS search path includes its own namespace; from another namespace you need user-service.<namespace>. The Service gives a stable virtual IP that kube-proxy load-balances across ready pods at the connection level. That detail matters for gRPC and HTTP/2: a client keeps one long-lived connection, so all its requests go to the same pod and the load balancing you expected does not happen. Client-side load balancing over a headless Service, or a service mesh, fixes this.
The Saga Pattern (Distributed Transactions)
When an operation spans multiple services, you can’t use a database transaction. Use the Saga pattern:
Choreography (event-driven)
OrderService → OrderCreated event
→ InventoryService reserves stock → StockReserved event
→ PaymentService charges card → PaymentProcessed event
→ ShippingService creates shipment
If payment fails: PaymentFailed event → InventoryService releases stock → OrderService marks order failed.
// Each service listens and reacts
eventBus.on('PaymentFailed', async ({ orderId }) => {
await releaseReservedStock(orderId)
await eventBus.emit('StockReleased', { orderId })
})
Orchestration (central coordinator)
A saga orchestrator directs each step:
class OrderSaga {
async execute(order) {
try {
await this.reserveStock(order)
await this.processPayment(order)
await this.scheduleShipping(order)
await this.completeOrder(order)
} catch (error) {
await this.compensate(order, error.failedStep)
}
}
async compensate(order, failedStep) {
if (failedStep === 'payment') await this.releaseStock(order)
if (failedStep === 'shipping') {
await this.refundPayment(order)
await this.releaseStock(order)
}
await this.cancelOrder(order)
}
}
The sketch assumes each step throws an error tagged with failedStep; in real code you record which steps completed, and compensate those in reverse order. Three properties make sagas work in practice:
- Compensations are not rollbacks. A refund is a new business action, visible to the customer, and it can fail too. Compensation steps must be retried until they succeed, so they need to be idempotent.
- The saga’s progress must be persisted. If the orchestrator process dies between
processPaymentandscheduleShipping, an in-memorytry/catchloses track of the order. Production orchestrators store saga state in a database, or use a workflow engine such as Temporal or AWS Step Functions that does this for you. - Intermediate states are visible. Between “stock reserved” and “payment processed”, other requests can see the reservation. That is why orders typically carry explicit states like
PENDINGand the UI shows them, instead of pretending the operation is atomic.
Resilience Patterns
Circuit Breaker
Stop cascading failures when a downstream service is unhealthy:
import CircuitBreaker from 'opossum'
const options = {
timeout: 3000, // fail if takes > 3s
errorThresholdPercentage: 50, // open circuit if 50% fail
resetTimeout: 30000, // try again after 30s
}
const breaker = new CircuitBreaker(callPaymentService, options)
breaker.fallback(() => ({ status: 'payment-pending', retry: true }))
const result = await breaker.fire(paymentRequest)
States: Closed (normal) → Open (failing, use fallback) → Half-Open (testing recovery).
The point of an open circuit is to fail fast: instead of every request waiting 3 seconds for a timeout from a service that is already struggling, calls return the fallback immediately, and the downstream service gets breathing room to recover. The fallback here returns payment-pending, which is only honest if the rest of the system really will retry the payment later; a fallback that pretends success is worse than an error. Also note that the percentage threshold needs enough traffic to mean anything: with opossum’s defaults it is computed over a rolling window, and a low-traffic endpoint can open the circuit after one or two failures.
Retry with exponential backoff
async function fetchWithRetry(url, maxRetries = 3) {
for (let attempt = 1; attempt <= maxRetries; attempt++) {
try {
return await fetch(url)
} catch (err) {
if (attempt === maxRetries) throw err
const delay = Math.pow(2, attempt) * 100 + Math.random() * 100
await new Promise(r => setTimeout(r, delay))
}
}
}
Two details in this function matter more than the backoff formula. First, fetch only rejects on network errors; an HTTP 503 resolves normally, so this loop never retries server errors. Real code checks response.ok or the status and retries only on 502/503/504 and timeouts, never on 4xx. Second, only retry operations that are safe to repeat: GETs, or writes protected by an idempotency key. The random jitter is there so that many clients that failed at the same moment do not all retry at the same moment too.
Retries multiply across layers. If the gateway, the order service and its HTTP client each retry three times, one user request can become 27 calls to a struggling payment service, which is how a brief slowdown becomes an outage. Pick one layer to own retries, and let the circuit breaker stop them when the downstream is clearly unhealthy.
Distributed Tracing
With many services, debugging requires tracing a request across all of them:
// OpenTelemetry setup (OTLP works with Jaeger, Zipkin via collector, Tempo)
import { NodeSDK } from '@opentelemetry/sdk-node'
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http'
import { trace, SpanStatusCode } from '@opentelemetry/api'
const sdk = new NodeSDK({
serviceName: 'order-service',
traceExporter: new OTLPTraceExporter({ url: 'http://jaeger:4318/v1/traces' }),
})
sdk.start()
// Create spans
const tracer = trace.getTracer('order-service')
const span = tracer.startSpan('process-order')
span.setAttribute('order.id', orderId)
try {
await processOrder(orderId)
span.setStatus({ code: SpanStatusCode.OK })
} catch (err) {
span.setStatus({ code: SpanStatusCode.ERROR, message: err.message })
} finally {
span.end()
}
An earlier version of this snippet used @opentelemetry/exporter-jaeger, which is deprecated: Jaeger now accepts OTLP directly (port 4318 for HTTP, 4317 for gRPC), so the OTLP exporter works with Jaeger, Grafana Tempo and most commercial backends.
A trace only spans services if the trace context travels with each request. OpenTelemetry’s auto-instrumentation (@opentelemetry/auto-instrumentations-node) injects the W3C traceparent header into outgoing HTTP calls and reads it on incoming ones. It does not magically cross a message queue unless the Kafka or RabbitMQ client is instrumented too, and that is where traces usually break: the synchronous part of a request is visible, and the asynchronous half starts a new, unconnected trace. Check that before you rely on tracing to debug a saga.
Database Per Service
Each service owns its own database — never share databases between services:
user-service → PostgreSQL (users DB)
product-service → MongoDB (products DB)
order-service → PostgreSQL (orders DB)
session-service → Redis (sessions)
search-service → Elasticsearch (search index)
Benefits: services can use the best database for their needs, schema changes don’t affect other services, independent scaling.
Challenge: cross-service queries require API calls or event-driven data synchronization.
“Order history with product names” is the classic example. The order service cannot JOIN the products table. Either it calls product-service for every order it displays (slow, and coupled to product-service’s availability), or it keeps its own copy of the few product fields it needs, updated from ProductUpdated events. The copy is eventually consistent, and that is usually fine for display: a product renamed a few seconds ago showing the old name briefly is acceptable, while failing to show order history because the catalog is down is not.
The rule is about ownership, not hardware. Several services can use the same PostgreSQL server with separate schemas and credentials; what breaks independence is one service reading or writing another service’s tables, because then neither can change its schema without coordinating.
Health Checks and Observability
Every service should expose:
// Health check endpoint
app.get('/health', (req, res) => {
res.json({
status: 'healthy',
version: process.env.APP_VERSION,
uptime: process.uptime(),
timestamp: new Date().toISOString(),
})
})
// Readiness check (dependencies OK?)
app.get('/ready', async (req, res) => {
try {
await db.ping()
res.json({ status: 'ready' })
} catch {
res.status(503).json({ status: 'not ready', reason: 'db unreachable' })
}
})
Use Prometheus + Grafana for metrics, ELK stack or Loki for logs, Jaeger for traces.
In Kubernetes, map /health to the liveness probe and /ready to the readiness probe, and keep them different. A failing readiness probe removes the pod from the Service’s endpoints, which is the right reaction to “my database is unreachable”. A failing liveness probe restarts the container. If the liveness probe also checks the database, a short database outage restarts every pod of every service that uses it at once, and they all reconnect at the same moment when the database returns. Liveness should only answer “is this process stuck?”.
Before you split the monolith
Microservices are a team scaling solution, not a technical one. Before adopting them: make sure your monolith is well-structured, your team is large enough to own services independently, and you have the operational maturity to run distributed systems (monitoring, tracing, deployment pipelines). Done right, microservices enable organizational agility. Done wrong, they’re a distributed monolith with network calls instead of function calls.
A quick way to tell which one you are building: if shipping one feature regularly requires coordinated deploys of several services, or two services read and write the same tables, the boundaries are wrong, and adding more services will make it worse rather than better.
Frequently Asked Questions (FAQ)
Q. Should I implement a saga with choreography or orchestration?
A. Choreography (services reacting to each other’s events) works well for short flows with two or three steps, because there is no central component to build or operate. As the number of steps and compensation paths grows, the flow becomes hard to follow because no single place describes it. Orchestration puts the sequence and the compensation logic in one coordinator, which is easier to reason about, test and monitor, at the cost of an extra component that the flow depends on.
Related Articles
- Kubernetes in Practice: Pods, Deployments, Services, Ingress, HPA and kubectl Troubleshooting
- Docker Compose for Multi-Container Apps
- Event Streaming with Kafka