C++ Observability: Prometheus and Grafana for Server
Introduction: see why it got slow—with data
You need metrics to respond
Parts 43-1 and 43-2 covered RPC and security. In operations, metrics (request counts, latency, CPU usage, etc.) are essential. Prometheus uses a pull model: the server scrapes targets over HTTP and stores time series. Grafana visualizes them on dashboards. To expose Prometheus text format from C++, define Counter, Gauge, and Histogram, and serve them as text on a path such as /metrics. You can use prometheus-cpp or implement a minimal exporter yourself.
Real-world scenarios
Scenario 1: API suddenly slow, root cause unknown
Situation: C++ gRPC server latency spikes to 10s at 2 AM
Problem: Logs alone do not show *where* it blocks
Result: Without Prometheus metrics you cannot see RPS, latency distribution, or error-rate trends
→ Hours to triage, incident response drags
Scenario 2: Memory usage creeps up
Situation: C++ server memory rises for 3 days straight
Problem: No time series for heap, connections, or queue depth
Result: Without Gauge metrics you only suspect leaks, with no evidence
→ Restart as a band-aid, root cause unfixed
Scenario 3: One endpoint has high errors
Situation: Overall error rate 1%, but /api/payment is 30%
Problem: Without per-path metrics you cannot pinpoint the route
Result: Only a global Counter → no fine-grained analysis
→ You optimize the wrong thing or miss the bad path
Scenario 4: Regression after deploy
Situation: Users feel slowness after a new release
Problem: No p99 or RPS to compare before/after
Result: Without Histograms you cannot compare percentiles
→ Roll back blindly or leave the outage running
This article shows how to prevent those issues with a Prometheus + Grafana pipeline and complete examples.
Prometheus metrics
Counter, Gauge, Histogram
- Counter: monotonically increasing (requests, bytes). Use rate() for per-second increase.
- Gauge: goes up and down (connections, queue length, memory).
- Histogram: distributions (latency). Expose buckets plus sum and count; use histogram_quantile in Prometheus for percentiles.
- Labels: attach labels (e.g. method, path, status) for filtering and grouping. Keep cardinality bounded.
The distinction between the three types is not cosmetic; it decides which PromQL functions give meaningful answers. A counter’s raw value is almost useless on a dashboard because it only reflects how long the process has been up. What you actually chart is rate() or increase(), and those functions assume the value only goes up. When the process restarts and the counter drops back to zero, Prometheus treats the drop as a counter reset and compensates for it, which is exactly why you must never decrement a counter or reuse a counter for something that can go down. If you expose “active connections” as a counter and call fetch_sub on it, rate() will interpret every decrease as a restart and produce spikes that never happened.
Histograms deserve special attention because the exposition format is cumulative: the le="0.1" bucket counts every observation less than or equal to 0.1 s, including the ones already counted in le="0.05". The +Inf bucket therefore always equals _count. histogram_quantile() relies on that property and linearly interpolates inside the bucket where the requested rank falls, so the percentile it returns is only as precise as your bucket boundaries. If your SLO is “p99 under 300 ms” and you have buckets at 0.1 and 0.5, every p99 between those two values is an interpolation guess. Put a boundary right at the SLO threshold so the question “are we under 300 ms?” has an exact answer.
The alternative Prometheus type, the summary, computes quantiles inside the process. It is precise for one instance but cannot be aggregated: averaging the p99 of ten pods does not give you the fleet’s p99. That is why this article uses histograms everywhere; the server only counts, and Prometheus does the math across instances.
Histogram bucket hints (seconds)
| Service type | Suggested buckets (s) | Notes |
|---|---|---|
| Low-latency API | 0.001, 0.005, 0.01, 0.025, 0.05, 0.1 | ms-scale latency |
| Typical API | 0.005, 0.025, 0.1, 0.5, 1.0, 2.5 | REST/gRPC |
| Batch | 1, 5, 10, 30, 60, 120 | Long jobs |
Prometheus text format example
# HELP http_requests_total Total number of HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",path="/api"} 1234
http_requests_total{method="POST",path="/api"} 567
# HELP http_request_duration_seconds Request duration in seconds
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{le="0.05"} 50
http_request_duration_seconds_bucket{le="0.1"} 100
http_request_duration_seconds_bucket{le="0.5"} 200
http_request_duration_seconds_bucket{le="1.0"} 250
http_request_duration_seconds_bucket{le="+Inf"} 300
http_request_duration_seconds_sum 45.2
http_request_duration_seconds_count 300
Collection architecture
flowchart LR
subgraph Cpp[C++ server]
M["/metrics endpoint"]
end
subgraph Prom[Prometheus]
S[Scrape]
TS[Time series DB]
end
subgraph Graf[Grafana]
D[Dashboards]
A[Alerts]
end
M -->|HTTP GET| S
S --> TS
TS -->|PromQL| D
TS -->|Alert rules| A
Scrape sequence
sequenceDiagram
participant P as Prometheus
participant C as C++ server
loop scrape_interval (e.g. 15s)
P->>C: GET /metrics
C->>C: call export_metrics()
C->>P: 200 OK, text/plain
P->>P: parse and store in TSDB
end
Exposing metrics from C++
Library vs manual
- prometheus-cpp: register Counter/Gauge/Histogram and serialize to text. Under multi-threaded access, protect with atomics or locks.
- Manual: std::atomic counters and per-bucket counts, assemble strings in the /metrics handler. Set Content-Type: text/plain; charset=utf-8 as expected.
- Placement: put metrics on an admin port or separate path, use auth and network isolation so they are not public.
Minimal manual example
Increment request_count with fetch_add(1, memory_order_relaxed) on each request; export_metrics() returns a Prometheus line (name value\n). The /metrics handler returns that body. memory_order_relaxed is enough for a simple counter; use seq_cst if you need ordering across metrics.
#include <atomic>
#include <string>
// Conceptual single Counter
std::atomic<uint64_t> request_count{0};
void on_request() {
request_count.fetch_add(1, std::memory_order_relaxed);
}
std::string export_metrics() {
return "http_requests_total " + std::to_string(request_count.load()) + "\n";
}
Manual Counter, Gauge, Histogram
The struct below keeps every metric in its own std::atomic, so request threads never take a lock on the hot path. Two details matter. First, the bucket update uses independent if statements rather than an else if chain: a 3 ms request must increment the 5 ms, 25 ms, 100 ms, 500 ms, 1 s and +Inf buckets, because the format is cumulative. An else if chain that only bumps the first matching bucket is a common bug in hand-written exporters; Prometheus accepts the output without complaint, but histogram_quantile() then reports nonsense because higher buckets hold fewer samples than lower ones. Second, std::atomic<double> only gained fetch_add in C++20, so the sum is updated with a compare-exchange loop that works in C++11 and later.
#include <atomic>
#include <string>
#include <sstream>
#include <mutex>
#include <chrono>
// Per-label counters (path, method) — simplified global example
struct Metrics {
std::atomic<uint64_t> requests_total{0};
std::atomic<uint64_t> errors_total{0};
std::atomic<uint64_t> active_connections{0};
std::atomic<uint64_t> queue_length{0};
// Histogram buckets: 5ms, 25ms, 100ms, 500ms, 1s, +Inf
static constexpr double buckets[] = {0.005, 0.025, 0.1, 0.5, 1.0, -1}; // -1 = +Inf
std::atomic<uint64_t> duration_bucket_5ms{0};
std::atomic<uint64_t> duration_bucket_25ms{0};
std::atomic<uint64_t> duration_bucket_100ms{0};
std::atomic<uint64_t> duration_bucket_500ms{0};
std::atomic<uint64_t> duration_bucket_1s{0};
std::atomic<uint64_t> duration_bucket_inf{0};
std::atomic<double> duration_sum{0};
std::atomic<uint64_t> duration_count{0};
void record_request(bool error, double duration_sec) {
requests_total.fetch_add(1, std::memory_order_relaxed);
if (error) errors_total.fetch_add(1, std::memory_order_relaxed);
auto add_bucket = [this](std::atomic<uint64_t>& b) {
b.fetch_add(1, std::memory_order_relaxed);
};
// Buckets are cumulative: increment every bucket whose bound >= duration
if (duration_sec <= 0.005) add_bucket(duration_bucket_5ms);
if (duration_sec <= 0.025) add_bucket(duration_bucket_25ms);
if (duration_sec <= 0.1) add_bucket(duration_bucket_100ms);
if (duration_sec <= 0.5) add_bucket(duration_bucket_500ms);
if (duration_sec <= 1.0) add_bucket(duration_bucket_1s);
add_bucket(duration_bucket_inf);
double expected;
do {
expected = duration_sum.load(std::memory_order_relaxed);
} while (!duration_sum.compare_exchange_weak(
expected, expected + duration_sec, std::memory_order_relaxed));
duration_count.fetch_add(1, std::memory_order_relaxed);
}
void connection_opened() {
active_connections.fetch_add(1, std::memory_order_relaxed);
}
void connection_closed() {
active_connections.fetch_sub(1, std::memory_order_relaxed);
}
void queue_inc() { queue_length.fetch_add(1, std::memory_order_relaxed); }
void queue_dec() { queue_length.fetch_sub(1, std::memory_order_relaxed); }
std::string export_prometheus() const {
std::ostringstream out;
out << "# HELP http_requests_total Total HTTP requests\n";
out << "# TYPE http_requests_total counter\n";
out << "http_requests_total " << requests_total.load() << "\n";
out << "# HELP http_errors_total Total HTTP errors\n";
out << "# TYPE http_errors_total counter\n";
out << "http_errors_total " << errors_total.load() << "\n";
out << "# HELP http_active_connections Active connections\n";
out << "# TYPE http_active_connections gauge\n";
out << "http_active_connections " << active_connections.load() << "\n";
out << "# HELP http_queue_length Current queue length\n";
out << "# TYPE http_queue_length gauge\n";
out << "http_queue_length " << queue_length.load() << "\n";
out << "# HELP http_request_duration_seconds Request duration\n";
out << "# TYPE http_request_duration_seconds histogram\n";
out << "http_request_duration_seconds_bucket{le=\"0.005\"} " << duration_bucket_5ms.load() << "\n";
out << "http_request_duration_seconds_bucket{le=\"0.025\"} " << duration_bucket_25ms.load() << "\n";
out << "http_request_duration_seconds_bucket{le=\"0.1\"} " << duration_bucket_100ms.load() << "\n";
out << "http_request_duration_seconds_bucket{le=\"0.5\"} " << duration_bucket_500ms.load() << "\n";
out << "http_request_duration_seconds_bucket{le=\"1\"} " << duration_bucket_1s.load() << "\n";
out << "http_request_duration_seconds_bucket{le=\"+Inf\"} " << duration_bucket_inf.load() << "\n";
out << "http_request_duration_seconds_sum " << duration_sum.load() << "\n";
out << "http_request_duration_seconds_count " << duration_count.load() << "\n";
return out.str();
}
};
Because each atomic is read separately during export, a scrape that races with a request can see, for example, the 5 ms bucket already incremented but +Inf not yet. The skew is at most a handful of in-flight requests and Prometheus’s quantile function tolerates small non-monotonic buckets, so for most services this is an acceptable price for a lock-free hot path. If you need the snapshot to be exactly consistent (for example, because you also export derived values), guard record_request and export_prometheus with one mutex instead; at a few thousand requests per second the contention is rarely measurable.
The hand-rolled version also has no labels. The moment you want method or status dimensions, you need a map from label set to metric object, with a lock around insertion and careful escaping on output. That is roughly the point where writing it yourself stops being cheaper than taking the library dependency.
prometheus-cpp example
prometheus-cpp splits the API into a family (one metric name plus help text, created with BuildCounter() and friends) and children (one time series per label combination, created with family.Add({...})). The Exposer runs a small embedded HTTP server (civetweb) on its own thread and serves every registered registry on /metrics, so you do not have to wire the endpoint into your own HTTP stack. Histogram bucket boundaries are passed per child in Add(), not on the builder.
#include <prometheus/counter.h>
#include <prometheus/gauge.h>
#include <prometheus/histogram.h>
#include <prometheus/registry.h>
#include <prometheus/exposer.h>
#include <memory>
int main() {
// Expose /metrics on port 8080
prometheus::Exposer exposer{"127.0.0.1:8080"};
auto registry = std::make_shared<prometheus::Registry>();
// Counter with labels for path and method
auto& request_counter = prometheus::BuildCounter()
.Name("http_requests_total")
.Help("Total HTTP requests")
.Labels({{"service", "cpp-server"}})
.Register(*registry);
auto& get_requests = request_counter.Add({{"method", "GET"}, {"path", "/api"}});
auto& post_requests = request_counter.Add({{"method", "POST"}, {"path", "/api"}});
// Gauge: active connections
auto& conn_gauge = prometheus::BuildGauge()
.Name("http_active_connections")
.Help("Active connections")
.Register(*registry);
// Histogram: latency (5ms, 25ms, 100ms, 500ms, 1s)
auto& duration_hist = prometheus::BuildHistogram()
.Name("http_request_duration_seconds")
.Help("Request duration")
.Register(*registry);
// Bucket boundaries are given per child in Add()
auto& get_duration = duration_hist.Add({{"method", "GET"}}, prometheus::Histogram::BucketBoundaries{0.005, 0.025, 0.1, 0.5, 1.0});
exposer.RegisterCollectable(registry);
// During request handling
get_requests.Increment();
conn_gauge.Increment();
auto start = std::chrono::steady_clock::now();
// ... handle request ...
auto elapsed = std::chrono::duration<double>(std::chrono::steady_clock::now() - start).count();
get_duration.Observe(elapsed);
conn_gauge.Decrement();
return 0;
}
Keep the references returned by Add() and reuse them. Add() takes a lock and looks up the label set in a map, so calling request_counter.Add({{"method", m}, {"path", p}}).Increment() on every request works but moves a map lookup and a mutex into the hot path. A common layout is to pre-create the children for the known routes at startup and store the references in a struct. Note also that main in this sketch returns immediately, which destroys the Exposer; in a real server the exposer must live as long as the process, typically as a member of your server object.
Building prometheus-cpp
# vcpkg (recommended)
vcpkg install prometheus-cpp
# CMakeLists.txt
find_package(prometheus-cpp CONFIG REQUIRED)
target_link_libraries(my_server prometheus-cpp::pull) # pull = core + Exposer
# Or FetchContent (clones the civetweb submodule as well)
include(FetchContent)
FetchContent_Declare(
prometheus-cpp
GIT_REPOSITORY https://github.com/jupp0r/prometheus-cpp.git
GIT_TAG v1.2.2
)
FetchContent_MakeAvailable(prometheus-cpp)
target_link_libraries(my_server prometheus-cpp::pull)
The library is split into core (registry and metric types), pull (the Exposer HTTP server) and push (Pushgateway client, which needs libcurl). A frequent link error when people first add it is an undefined reference to prometheus::Exposer::Exposer(...): they linked only prometheus-cpp::core, which does not contain the HTTP side. If you already run an HTTP server and only want the text serializer, link core and call prometheus::TextSerializer on registry->Collect() from your own /metrics handler; that avoids a second listening socket and the civetweb dependency.
Prometheus configuration and scraping
Basic prometheus.yml
global:
scrape_interval: 15s # default scrape interval
evaluation_interval: 15s # alert rule evaluation
alerting:
alertmanagers:
- static_configs:
- targets: []
rule_files: []
scrape_configs:
- job_name: 'cpp-server'
scrape_interval: 10s # scrape C++ server every 10s
scrape_timeout: 5s
static_configs:
- targets: ['localhost:8080']
labels:
env: 'production'
service: 'cpp-api'
scrape_timeout must not exceed scrape_interval; Prometheus refuses to load a config where it does. The interval is also the resolution floor for everything you query: with a 10 s interval, rate(x[15s]) often contains only one sample and returns nothing, which is why the examples use a [5m] window. A practical rule is to make the range at least four times the scrape interval so that one missed scrape does not produce gaps in the graph.
Prometheus stores the time of the scrape, not a timestamp from your process, so the C++ side does not need a synchronized clock. It does need /metrics to answer quickly: if serialization takes longer than scrape_timeout, the whole scrape fails and the up metric for that target drops to 0, which looks exactly like the server being down.
Dynamic targets (service discovery)
# Scrape C++ pods in Kubernetes
scrape_configs:
- job_name: 'cpp-pods'
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path]
action: replace
target_label: __metrics_path__
regex: (.+)
- source_labels: [__address__, __meta_kubernetes_pod_annotation_prometheus_io_port]
action: replace
regex: ([^:]+)(?::\d+)?;(\d+)
replacement: ${1}:${2}
target_label: __address__
Grafana integration
Data source
- Add Prometheus as a Grafana data source and query with PromQL.
- URL:
http://prometheus:9090(Docker/K8s) orhttp://localhost:9090
Useful PromQL examples
# Requests per second
rate(http_requests_total[5m])
# p99 latency (seconds)
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
# Error rate (%)
100 * sum(rate(http_errors_total[5m])) / sum(rate(http_requests_total[5m]))
# Active connections (Gauge — no rate)
http_active_connections
# Queue length
http_queue_length
Two of these queries hide aggregation choices worth understanding. histogram_quantile(0.99, rate(..._bucket[5m])) without an aggregation returns one p99 per series, so with three instances you get three lines. To get the fleet-wide p99 you must sum the buckets first while keeping the le label: histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))). Summing the buckets is valid; averaging three per-instance p99 values is not. The error-rate query already wraps both sides in sum(), which is what you want; if you forget it on one side, the division tries to match series label by label and returns an empty result whenever the label sets differ.
The first time I wired a dashboard like this, the error-rate panel showed “No data” for hours even though errors were clearly being counted. The cause is a classic one: a counter that has never been incremented does not exist as a series if it carries labels that are only created on first use, and rate() of a missing series is empty rather than zero. Creating the child series at startup, so they are exported as 0, fixes both the empty panel and alerts that silently never fire.
Dashboard panels
- Graphs: RPS, latency percentiles (p50, p95, p99), error rate over time
- Single stat: current connections, queue length
- Table: requests by path, errors by method
- Alerts: e.g. p99 > 1s, error rate > 5% → Slack/email
More PromQL
# p50, p95, p99
histogram_quantile(0.50, rate(http_request_duration_seconds_bucket[5m]))
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
# Average latency (sum/count)
rate(http_request_duration_seconds_sum[5m]) / rate(http_request_duration_seconds_count[5m])
# RPS per instance
sum by (instance) (rate(http_requests_total[5m]))
# Errors in 5 minutes
increase(http_errors_total[5m])
Grafana alert channel (Slack example)
# Configuration → Alerting → Contact points → New contact point
# Type: Slack
# Webhook URL: https://hooks.slack.com/services/xxx/yyy/zzz
# Channel: #alerts-cpp-server
Dashboard variables (filter by instance)
# Dashboard Settings → Variables → New variable
# Name: instance
# Type: Query
# Data source: Prometheus
# Query: label_values(http_requests_total, instance)
# Multi-value: Yes
# In panel queries: {instance=~"$instance"}
End-to-end Prometheus + Grafana examples
Full stack with Docker Compose
# docker-compose.yml
version: '3.8'
services:
cpp-server:
build: .
ports:
- "8080:8080"
environment:
- METRICS_PORT=8080
prometheus:
image: prom/prometheus:v2.47.0
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
ports:
- "9090:9090"
command:
- '--config.file=/etc/prometheus/prometheus.yml'
- '--storage.tsdb.retention.time=15d'
grafana:
image: grafana/grafana:10.2.0
ports:
- "3000:3000"
environment:
- GF_SECURITY_ADMIN_PASSWORD=admin
- GF_USERS_ALLOW_SIGN_UP=false
volumes:
- grafana-data:/var/lib/grafana
depends_on:
- prometheus
volumes:
grafana-data:
C++ server + /metrics (Boost.Beast sketch)
#include <boost/beast/core.hpp>
#include <boost/beast/http.hpp>
#include <boost/asio.hpp>
#include <atomic>
#include <chrono>
#include <string>
#include <thread>
namespace beast = boost::beast;
namespace http = beast::http;
namespace net = boost::asio;
// Global metrics (prefer singleton or DI in production)
std::atomic<uint64_t> g_requests_total{0};
std::atomic<uint64_t> g_errors_total{0};
std::atomic<uint64_t> g_active_connections{0};
void handle_metrics(http::request<http::string_body> const& req,
http::response<http::string_body>& res) {
res.set(http::field::content_type, "text/plain; charset=utf-8");
res.body() = "# HELP http_requests_total Total requests\n"
"# TYPE http_requests_total counter\n"
"http_requests_total " + std::to_string(g_requests_total.load()) + "\n"
"# HELP http_errors_total Total errors\n"
"# TYPE http_errors_total counter\n"
"http_errors_total " + std::to_string(g_errors_total.load()) + "\n"
"# HELP http_active_connections Active connections\n"
"# TYPE http_active_connections gauge\n"
"http_active_connections " + std::to_string(g_active_connections.load()) + "\n";
res.prepare_payload();
}
// For /metrics call handle_metrics; other paths run business logic
The Compose file pins image versions on purpose: Grafana dashboards and Prometheus flags change between major versions, and latest turns a routine docker compose pull into a debugging session. The top-level version: key is ignored by current Docker Compose and can be dropped. Inside the Compose network, Prometheus reaches the server as cpp-server:8080, not localhost:8080; localhost inside the Prometheus container is the Prometheus container itself, which is the single most common reason the target shows as down in this setup.
In the Beast handler, prepare_payload() sets Content-Length from the body; forgetting it makes some clients wait for the connection to close. Keep /metrics on the same io_context only if serialization is cheap. With thousands of labeled series, building the text can take milliseconds, and doing that on the thread that also serves user requests adds latency to exactly the requests you are trying to measure.
Grafana dashboard JSON (core panels)
{
"panels": [
{
"title": "RPS",
"type": "timeseries",
"targets": [{
"expr": "rate(http_requests_total[5m])",
"legendFormat": "{{instance}}"
}]
},
{
"title": "p99 latency (s)",
"type": "timeseries",
"targets": [{
"expr": "histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))",
"legendFormat": "p99"
}]
},
{
"title": "Error rate (%)",
"type": "timeseries",
"targets": [{
"expr": "100 * sum(rate(http_errors_total[5m])) / sum(rate(http_requests_total[5m]))",
"legendFormat": "error_rate"
}]
},
{
"title": "Active connections",
"type": "stat",
"targets": [{
"expr": "http_active_connections",
"legendFormat": "connections"
}]
}
]
}
Common errors and fixes
Prometheus: “connection refused” or “context deadline exceeded”
Cause: /metrics port closed, firewall, or network isolation. Fix:
curl -v http://localhost:8080/metrics
docker exec prometheus wget -qO- http://cpp-server:8080/metrics
# Use Docker service names or K8s Service names
scrape_configs:
- job_name: 'cpp-server'
static_configs:
- targets: ['cpp-server:8080']
“parse error” or “invalid character”
Cause: Text format does not match Prometheus exposition format. Fix:
# Bad: commas, spaces, bad escaping
http_requests_total 1234,567
# Good
# HELP http_requests_total Total requests
# TYPE http_requests_total counter
http_requests_total 1234
http_requests_total{path="/api"} 100
- Include HELP and TYPE where appropriate
- Escape
"and\inside label values - One sample per line:
name{labels} valueorname value
Grafana: “No data”
Cause: PromQL typo, time range, or metric name mismatch. Fix:
{__name__=~"http_.*"}
rate(http_requests_total[5m])
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
Label cardinality explosion
Cause: Using raw paths or user IDs as labels. Every distinct label combination is a separate time series with its own memory in both your process and the Prometheus head block. A path label fed from the raw URL turns every scanner hitting /wp-login.php?x=... into a new series; the process memory grows steadily and looks like a leak, and Prometheus itself eventually slows down or runs out of memory. In my experience this is the problem that bites hand-written C++ exporters most often, because nothing fails loudly: the metrics keep working until the day memory graphs of the monitoring system itself go vertical.
Fix:
// Risky: thousands of paths → thousands of series
request_counter.Add({{"path", user_provided_path}});
// Safer: normalize templates
std::string normalize_path(const std::string& path) {
if (path.find("/api/users/") == 0) return "/api/users/:id";
if (path.find("/api/orders/") == 0) return "/api/orders/:id";
return path;
}
Histogram races with atomics
Cause: Bucket updates and sum/count not consistent as a group. Fix: use atomics per bucket and CAS loop for sum, or one mutex around record_request.
Slow /metrics (hundreds of ms)
Cause: Heavy allocation or lock contention during export, usually proportional to the number of series.
Fix: reduce cardinality first; then reserve the output buffer, avoid std::ostringstream locale overhead on very large outputs, and make sure export does not hold the same lock that request threads need.
”Out of order” or duplicate samples
Cause: The exporter writes explicit timestamps (the optional third field on a sample line) from a clock that jumped, or two targets end up producing identical label sets (for example, the same server listed in two jobs with honor_labels: true).
Fix: do not emit timestamps from a normal exporter; let Prometheus assign the scrape time. Ensure each target appears in one job.
Best practices
Naming
- Counter:
_totalsuffix (e.g.http_requests_total) - Units:
_seconds,_bytes, etc. - Lowercase snake_case
Labels
- Bound cardinality (hundreds of combinations, not millions)
- Static labels: env, service, region
- Avoid high-cardinality dynamic values (user_id, request_id)
Scrape intervals
- App: 10–15s
- Infra: 30s–1m
- Expensive metrics: 1–5m
Securing /metrics
- Internal networks only
- Basic Auth or mTLS
- Separate port from user-facing API (e.g. 8080 API, 9090 metrics)
Performance
| Topic | Recommendation |
|---|---|
| Atomics | memory_order_relaxed when order across metrics does not matter |
| Histogram | Prefer atomic buckets on hot paths over a global mutex |
| Export | Minimize string building on each scrape |
| Labels | Keep label count small; cardinality < 100 for typical setups |
Production patterns
Pattern 1: separate metrics port
// API :8080, metrics :9090 (bind to internal IP)
void run_metrics_server(const std::string& bind_addr, uint16_t port) {
tcp::acceptor acceptor(ctx, {net::ip::make_address(bind_addr), port});
}
Binding the metrics listener to an internal address keeps it off the public internet without any auth code in the C++ server. Pick a port that does not collide with your monitoring stack: 9090 is Prometheus’s own default, so running the metrics endpoint on 9090 on the same host as Prometheus fails with “address already in use”.
Pattern 2: initialize metrics at startup
The point is not zeroing plain atomics (they are already zero), but making every labeled series exist before the first event, so dashboards and absent()-style alerts see a 0 instead of a missing series.
void init_metrics() {
g_requests_total.store(0);
g_errors_total.store(0);
}
Pattern 3: RAII request scope
struct ScopedRequestMetrics {
Metrics& m;
std::chrono::steady_clock::time_point start;
bool error = false;
ScopedRequestMetrics(Metrics& metrics) : m(metrics), start(std::chrono::steady_clock::now()) {
m.connection_opened();
}
~ScopedRequestMetrics() {
auto dur = std::chrono::duration<double>(
std::chrono::steady_clock::now() - start).count();
m.record_request(error, dur);
m.connection_closed();
}
};
The RAII guard guarantees that a request is recorded even when the handler throws or returns early, which is where manual record() calls usually get skipped. Set error = true on the failure path before the guard goes out of scope. Note that it tracks in-flight requests, not TCP connections; with keep-alive, one connection carries many requests, so name the gauge accordingly (http_requests_in_flight) if you adopt this pattern.
Pattern 4: Prometheus alert rules
The for: clause is what keeps these alerts from paging on a single bad scrape: the condition must hold continuously for that long before the alert fires. The error-rate expression divides by the request rate, so it returns nothing when there is no traffic at all; if a dead server should also page, add a separate up == 0 alert rather than relying on this one.
groups:
- name: cpp-server
rules:
- alert: HighErrorRate
expr: 100 * sum(rate(http_errors_total[5m])) / sum(rate(http_requests_total[5m])) > 5
for: 2m
labels:
severity: critical
annotations:
summary: "C++ server error rate {{ $value | humanize }}% exceeded"
- alert: HighLatency
expr: histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])) > 1
for: 5m
labels:
severity: warning
annotations:
summary: "p99 latency exceeded 1s"
Implementation checklist
- Implement /metrics on the C++ server
- Expose Counters (requests, errors)
- Expose Gauges (connections, queue length)
- Expose Histogram (latency) with sensible buckets
- Set
Content-Type: text/plain; charset=utf-8 - Add scrape_config in prometheus.yml
- Add Prometheus data source in Grafana
- Panels for RPS, p99, error rate, connections
- Alert rules (errors, latency)
- Secure metrics port/path
- Review label cardinality
Summary
| Topic | Summary |
|---|---|
| Prometheus | Counter/Gauge/Histogram, labels, pull, text exposition |
| C++ | Atomics, library or manual serialization, /metrics handler |
| Grafana | PromQL, dashboards, alerts |
| Production | Split ports, alert rules, bounded labels |
Series 43 covered gRPC/Protobuf → secure coding/OpenSSL → Observability (Prometheus + Grafana) for large distributed systems.
FAQ
When is this useful in production?
A. Whenever you run C++ services in production and need scrapeable metrics, dashboards, and alerts. Use the examples above as templates.
prometheus-cpp vs manual?
A. prometheus-cpp: full feature set and labels when you can take the dependency. Manual: minimal deps, embedded systems, or very small metric sets.
Where to go deeper?
A. Prometheus docs, Grafana docs, prometheus-cpp.
Prometheus scrapes your C++ /metrics; Grafana turns them into dashboards and alerts.