gRPC Services with Protocol Buffers: Unary and Streaming RPC in Node.js and Go
Key takeaways
A complete guide to building high-performance APIs with gRPC—Protocol Buffers, service definitions, Unary and Streaming RPC, and hands-on Node.js and Go examples including errors, JWT metadata, and Docker deployment.
What this post covers
This is a guide to building high-performance APIs with gRPC. It covers Protocol Buffers, service definitions, Unary and Streaming RPC, and hands-on Node.js and Go examples.
From the field: the gains from moving internal calls to gRPC tend to come less from raw speed than from the contract — generated clients in every language stop the “field renamed, consumer broke” class of bug. The surprise that usually follows is operational: load balancing, debugging with
curl, and browser access all work differently than with REST.
Introduction: “Our REST API feels slow”
Real-world scenarios
Scenario 1: JSON parsing is expensive
Large payloads feel sluggish. gRPC uses efficient binary serialization.
Scenario 2: Weak type safety
The API contract is vague. gRPC gives you strong, generated types.
Scenario 3: You need streaming
Classic request/response REST is not enough. gRPC supports bidirectional streaming.
What is gRPC?
Key characteristics
gRPC is a high-performance RPC framework originally developed by Google.
Main benefits:
- Performance: binary protocol over HTTP/2
- Type safety: Protocol Buffers as the interface definition language
- Streaming: server, client, and bidirectional streaming
- Polyglot: first-class support for many languages
- HTTP/2: multiplexing, header compression, and efficient connections
Rough performance comparison (illustrative, not a benchmark):
- REST (JSON): ~100 ms, ~10 KB payload
- gRPC (Protobuf): ~20 ms, ~2 KB payload
Where the difference comes from matters more than the numbers. Protobuf encodes field numbers and compact varints instead of repeating field names as text, so messages are smaller and parsing does not involve scanning strings. HTTP/2 multiplexes many concurrent calls over one connection, avoiding connection setup per request and head-of-line blocking at the HTTP layer. For a service that spends most of its time in database queries, those savings are a small fraction of the total; for chatty internal traffic with many small calls, they add up. The real trade-offs are elsewhere: payloads are not human-readable, browsers cannot speak native gRPC, and every client needs generated code from the same .proto files.
Protocol Buffers
.proto files
// user.proto
syntax = "proto3";
package user;
message User {
int32 id = 1;
string name = 2;
string email = 3;
int32 age = 4;
}
message GetUserRequest {
int32 id = 1;
}
message GetUserResponse {
User user = 1;
}
message ListUsersRequest {
int32 page = 1;
int32 page_size = 2;
}
message ListUsersResponse {
repeated User users = 1;
int32 total = 2;
}
service UserService {
rpc GetUser(GetUserRequest) returns (GetUserResponse);
rpc ListUsers(ListUsersRequest) returns (ListUsersResponse);
rpc CreateUser(User) returns (User);
}
The numbers after each field (= 1, = 2) are the field’s identity on the wire, and they are what makes schema evolution safe: you can rename a field or add new ones freely, and old clients simply ignore numbers they don’t know. The rules that follow from this are strict. Never change the number or type of an existing field, and never reuse the number of a deleted field — mark it reserved 4; instead — because an old client would decode the new data into the wrong meaning. In proto3, scalar fields have no “unset” state: a missing age reads as 0, and an empty name as "". If you need to distinguish “not provided” from zero, declare the field optional int32 age = 4; (supported since protoc 3.15) or use a wrapper type.
Two design choices in this file are worth reconsidering in a real API. CreateUser(User) reuses the resource message as the request, which lets clients send an id the server must then ignore; a dedicated CreateUserRequest without id makes the contract clearer and can evolve independently. And id as int32 caps you at about two billion users, while int64 arrives in JavaScript as a Long object or string rather than a number — a decision best made up front.
Node.js implementation
Install dependencies
npm install @grpc/grpc-js @grpc/proto-loader
Server
// server.ts
import * as grpc from '@grpc/grpc-js';
import * as protoLoader from '@grpc/proto-loader';
const PROTO_PATH = './user.proto';
const packageDefinition = protoLoader.loadSync(PROTO_PATH);
const userProto = grpc.loadPackageDefinition(packageDefinition).user as any;
const users = [
{ id: 1, name: 'John', email: '[email protected]', age: 30 },
{ id: 2, name: 'Jane', email: '[email protected]', age: 25 },
];
const server = new grpc.Server();
server.addService(userProto.UserService.service, {
getUser: (call: any, callback: any) => {
const user = users.find(u => u.id === call.request.id);
if (user) {
callback(null, { user });
} else {
callback({
code: grpc.status.NOT_FOUND,
message: 'User not found',
});
}
},
listUsers: (call: any, callback: any) => {
callback(null, { users, total: users.length });
},
createUser: (call: any, callback: any) => {
const newUser = {
...call.request,
id: users.length + 1, // after the spread, so a client-sent id cannot override it
};
users.push(newUser);
callback(null, newUser);
},
});
server.bindAsync(
'0.0.0.0:50051',
grpc.ServerCredentials.createInsecure(),
(err, port) => {
if (err) throw err; // e.g. EADDRINUSE
console.log(`gRPC server running on :${port}`);
// server.start() is no longer required in recent @grpc/grpc-js versions
}
);
This server loads the .proto file at runtime with @grpc/proto-loader, which is the quickest way to start but leaves every handler typed as any. The loader’s defaults also cause the most common first bug: without options, loadSync converts field names to camelCase, so the server sees pageSize while a client sends { page_size: 10 } — the field is silently dropped and arrives as 0. Pass options explicitly, typically protoLoader.loadSync(PROTO_PATH, { keepCase: true, longs: String, enums: String, defaults: true, oneofs: true }): keepCase preserves the names from the .proto file, longs: String turns 64-bit integers into strings instead of Long objects, and defaults: true fills unset fields with their zero values so handlers do not have to check for undefined. For production code, generating TypeScript types (with protoc-gen-ts, ts-proto, or Buf’s tooling) gives the type safety that is supposed to be gRPC’s selling point.
Handlers follow a callback convention: call callback(null, response) exactly once for success, or callback(error) with a gRPC status code. A handler that throws or never calls back leaves the client waiting until its deadline. The bindAsync callback reports failures such as a port already in use; ignoring its error argument — as many examples do — turns that into a server that silently never starts. createInsecure() means plaintext HTTP/2, fine on localhost; anything crossing a network should use ServerCredentials.createSsl(...) or run behind a proxy or service mesh that terminates TLS.
Client
// client.ts
import * as grpc from '@grpc/grpc-js';
import * as protoLoader from '@grpc/proto-loader';
const PROTO_PATH = './user.proto';
const packageDefinition = protoLoader.loadSync(PROTO_PATH);
const userProto = grpc.loadPackageDefinition(packageDefinition).user as any;
const client = new userProto.UserService(
'localhost:50051',
grpc.credentials.createInsecure()
);
// GetUser
client.getUser({ id: 1 }, (error: any, response: any) => {
if (error) {
console.error('Error:', error);
} else {
console.log('User:', response.user);
}
});
// ListUsers
client.listUsers({ page: 1, page_size: 10 }, (error: any, response: any) => {
if (error) {
console.error('Error:', error);
} else {
console.log('Users:', response.users);
}
});
// CreateUser
client.createUser(
{ name: 'Bob', email: '[email protected]', age: 35 },
(error: any, response: any) => {
if (error) {
console.error('Error:', error);
} else {
console.log('Created:', response);
}
}
);
The client object is long-lived: it owns an HTTP/2 channel that connects lazily on the first call and is reused for every call after that, so create one per target service at startup rather than per request. The three calls above run concurrently on that one connection. If the server is down, a call fails with 14 UNAVAILABLE: No connection established; this code is the one most worth retrying (with backoff), while INVALID_ARGUMENT or NOT_FOUND should not be retried. Every call should also carry a deadline, as the FAQ below explains — without one, a hung server keeps the call open forever.
Streaming
Server streaming
service LogService {
rpc StreamLogs(StreamLogsRequest) returns (stream LogEntry);
}
// Server
streamLogs: (call: any) => {
const logs = [
{ message: 'Log 1', timestamp: Date.now() },
{ message: 'Log 2', timestamp: Date.now() },
{ message: 'Log 3', timestamp: Date.now() },
];
logs.forEach((log) => {
call.write(log);
});
call.end();
},
// Client
const call = client.streamLogs({});
call.on('data', (log: any) => {
console.log('Log:', log);
});
call.on('end', () => {
console.log('Stream ended');
});
A server-streaming call is a Node stream on both sides: the server writes messages and ends the stream; the client receives data events followed by end, or an error event if the stream fails with a status code — and a client that does not listen for error crashes on an unhandled error event. Two production concerns are absent from this example. call.write() returns false when the client is reading slower than the server writes, and a server that ignores it buffers without limit; for large streams, wait for the drain event. And when the client cancels or disconnects, the server’s call emits cancelled, which is the signal to stop producing — otherwise a server streaming a large export keeps working for a client that is gone.
Bidirectional streaming
service ChatService {
rpc Chat(stream ChatMessage) returns (stream ChatMessage);
}
// Server
chat: (call: any) => {
call.on('data', (message: any) => {
console.log('Received:', message);
// Echo back to the same client (example)
call.write({
user: message.user,
text: message.text,
timestamp: Date.now(),
});
});
call.on('end', () => {
call.end();
});
},
// Client
const call = client.chat();
call.on('data', (message: any) => {
console.log('Message:', message);
});
call.write({ user: 'John', text: 'Hello!' });
In a bidirectional stream, each side reads and writes independently, and the stream finishes only when both have ended. The server here ends its side when the client ends, so the client must eventually call call.end(); this snippet never does, which is fine for a demo but leaves the call open. A real chat service would also broadcast to other connected clients, which means keeping a set of active calls, removing them on end, error, and cancelled, and handling slow readers. Long-lived streams also run into infrastructure limits: proxies and load balancers with idle timeouts close quiet streams, so configure keepalive pings (grpc.keepalive_time_ms on the client, with a matching server policy — servers reject overly frequent pings with ENHANCE_YOUR_CALM) and design clients to reconnect and resume.
Go implementation
Server
// server.go
package main
import (
"context"
"log"
"net"
pb "myapp/proto"
"google.golang.org/grpc"
)
type server struct {
pb.UnimplementedUserServiceServer
}
func (s *server) GetUser(ctx context.Context, req *pb.GetUserRequest) (*pb.GetUserResponse, error) {
user := &pb.User{
Id: req.Id,
Name: "John",
Email: "[email protected]",
Age: 30,
}
return &pb.GetUserResponse{User: user}, nil
}
func main() {
lis, err := net.Listen("tcp", ":50051")
if err != nil {
log.Fatalf("Failed to listen: %v", err)
}
s := grpc.NewServer()
pb.RegisterUserServiceServer(s, &server{})
log.Println("gRPC server running on :50051")
if err := s.Serve(lis); err != nil {
log.Fatalf("Failed to serve: %v", err)
}
}
The pb package is generated, not written: protoc --go_out=. --go-grpc_out=. user.proto (with the protoc-gen-go and protoc-gen-go-grpc plugins installed) produces the message types and the UserServiceServer interface, and the .proto file needs an option go_package = "myapp/proto"; line for the import path to work. Embedding pb.UnimplementedUserServiceServer is required by default and serves forward compatibility: when a new RPC is added to the .proto, existing servers still compile and return UNIMPLEMENTED for it instead of failing to build. This server implements only GetUser, so ListUsers and CreateUser answer with that code.
Go’s context is where gRPC’s deadline and cancellation model shows most clearly: ctx carries the client’s deadline across the network, so passing it to database calls (db.QueryContext(ctx, ...)) makes them stop when the client gives up. For production, add s.GracefulStop() on SIGTERM so in-flight RPCs finish before shutdown, register the standard health service (google.golang.org/grpc/health) for Kubernetes probes, and enable server reflection so tools like grpcurl can call the service without the .proto file.
Error handling
import * as grpc from '@grpc/grpc-js';
// Server
getUser: (call: any, callback: any) => {
const user = findUser(call.request.id);
if (!user) {
return callback({
code: grpc.status.NOT_FOUND,
message: 'User not found',
details: 'No user with the given ID exists',
});
}
callback(null, { user });
},
// Client
client.getUser({ id: 999 }, (error: any, response: any) => {
if (error) {
if (error.code === grpc.status.NOT_FOUND) {
console.error('User not found');
} else {
console.error('Error:', error.message);
}
} else {
console.log('User:', response.user);
}
});
gRPC errors are a status code plus a message, not an HTTP status: every response travels as HTTP 200 at the transport level, with the real outcome in the grpc-status trailer. The standard codes are few and meaningful, so choose them carefully — NOT_FOUND and INVALID_ARGUMENT tell clients not to retry, UNAVAILABLE tells them a retry may help, FAILED_PRECONDITION signals a state problem the client must fix, and INTERNAL should be reserved for genuine server bugs. On the client, error.code is the number to branch on and error.details the message string. For structured error data (field violations, retry hints), the “richer error model” attaches google.rpc.Status details in the metadata; plain codes are enough for most internal services.
Authentication
JWT metadata
// Client
import * as grpc from '@grpc/grpc-js';
const metadata = new grpc.Metadata();
metadata.add('authorization', `Bearer ${token}`);
client.getUser({ id: 1 }, metadata, (error: any, response: any) => {
// ...
});
// Server
getUser: (call: any, callback: any) => {
const metadata = call.metadata;
const auth = metadata.get('authorization')[0];
if (!auth || !verifyToken(auth)) {
return callback({
code: grpc.status.UNAUTHENTICATED,
message: 'Invalid token',
});
}
// ...
},
Metadata is gRPC’s equivalent of HTTP headers, and keys must be lowercase (the library normalizes them). Checking the token inside every handler, as here, works but repeats itself; server interceptors (available in recent @grpc/grpc-js versions, and standard in Go via grpc.UnaryInterceptor and grpc.StreamInterceptor) apply authentication once to every method, including streaming ones. On the client side, grpc.credentials.createFromMetadataGenerator attaches the token automatically to every call, and combining it with SSL channel credentials is deliberately required: the library refuses to send call credentials over an insecure channel, which is the right default — a bearer token in plaintext is visible to anyone on the network path. Inside a service mesh with mutual TLS, identity often comes from the client certificate instead, and tokens carry end-user identity only.
Deployment
Docker
FROM node:20-alpine
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
EXPOSE 50051
CMD ["node", "server.js"]
Since the examples are TypeScript, the image needs a compile step (or a runtime such as tsx) before node server.js can run, and user.proto must be copied into the image because proto-loader reads it at startup — a missing proto file fails with ENOENT on boot.
The deployment problem that surprises most teams is load balancing. A Kubernetes Service or any layer-4 load balancer balances connections, and a gRPC client keeps one long-lived HTTP/2 connection per target. The result is that all traffic from a client goes to whichever pod it connected to first, new pods receive nothing after a scale-up, and one replica can run hot while the others idle. The fixes are to balance per request instead: a layer-7 proxy or service mesh that understands HTTP/2 (Envoy, Linkerd, Istio), or client-side load balancing across all pod addresses (a headless Service plus the round_robin policy). Setting a maximum connection age on the server (grpc.max_connection_age_ms) makes clients reconnect periodically, which spreads load more evenly as a simpler mitigation.
Browsers are the other boundary. They cannot open raw HTTP/2 gRPC calls, so browser clients use gRPC-Web through a translating proxy (Envoy has a filter for it), or the Connect protocol, which serves gRPC, gRPC-Web, and plain JSON over HTTP from the same service. Many teams simply expose REST or GraphQL at the edge and keep gRPC internal.
The load-balancing surprise, and two other things to settle early
The gRPC problem that most often reaches production unnoticed is load balancing. Because a client keeps one long-lived HTTP/2 connection and multiplexes every call over it, a connection-level (L4) balancer — including a plain Kubernetes ClusterIP Service — picks a backend once and then sends all of that client’s calls to the same pod. Adding replicas does nothing for a busy client, and after a rolling deploy traffic stays pinned to whichever pods happened to be up first. The fix is to balance per request: an L7 proxy or service mesh that understands HTTP/2 (Envoy, Linkerd, Istio), or client-side balancing, pointing the client at a headless Service and enabling the round_robin policy.
The other two are cheaper to get right at the start than to retrofit. Put a deadline on every call, so a stuck downstream service fails fast instead of holding requests open. And treat field numbers in .proto files as permanent: when you remove a field, mark its number (and name) as reserved so nobody reuses it, because old clients will otherwise decode the new field’s bytes as the old one’s.
Frequently asked questions (FAQ)
Q. gRPC vs REST—which should I choose?
A. gRPC is usually much faster for internal RPC. REST is simpler and better suited to browsers and public HTTP APIs. Use gRPC between services; expose REST or GraphQL at the edge when needed.
Q. Can I use gRPC in the browser?
A. Use gRPC-Web for limited browser support. For most web apps, REST or GraphQL remains the pragmatic choice.
Q. Do I have to learn Protocol Buffers?
A. Yes, but the basics are small. Think of .proto files as typed, versioned API contracts—similar in spirit to JSON Schema, but compiled to efficient code.
Q. Is gRPC safe for production?
A. Yes. Large organizations run gRPC at scale. Pair it with TLS, robust load balancing, health checks, and observability (metrics, tracing, structured logs).
Q. Why should I set a deadline on every gRPC call?
A. gRPC calls have no deadline by default, so a slow or stuck server can leave a client call waiting indefinitely while it holds a connection and memory. In @grpc/grpc-js you pass it in the call options, for example client.getUser({ id: 1 }, metadata, { deadline: Date.now() + 2000 }, callback), and the client receives grpc.status.DEADLINE_EXCEEDED when it expires. Handle that code separately from NOT_FOUND or UNAUTHENTICATED in the error handling shown above, because it usually means retry or fail fast rather than a bad request.