C++ 실시간 모니터링 대시보드: Grafana·Prometheus 통합
C++ REST API 서버나 게임 서버를 운영하다 보면 “어제 새벽 3시쯤 API가 5분 동안 응답하지 않았다”는 보고를 받고 로그를 뒤지게 되는 순간이 옵니다. 로그는 개별 사건을 기록하지만, 언제부터 어떤 엔드포인트의 지연이 늘었는지, 실패한 요청이 전체의 몇 퍼센트였는지 같은 추세와 비율은 로그에서 재구성하기 어렵습니다. 응답 시간이 며칠에 걸쳐 서서히 나빠지는 경우라면 더더욱 그렇습니다.
메트릭은 이런 질문에 답하도록 만든 데이터입니다. 서버가 요청 수, 지연 분포, 연결 수 같은 숫자를 계속 노출하고, Prometheus가 이를 주기적으로 수집해 시계열로 저장하며, Grafana가 그래프와 알람으로 보여 줍니다. 이 글은 C++17 서버에 작은 메트릭 레지스트리를 직접 구현해 /metrics로 노출하고, Prometheus 수집 설정, Grafana 패널, 알람 규칙까지 연결합니다. 그리고 카디널리티 폭증이나 히스토그램 버킷 설정처럼 운영에서 실제로 문제가 되는 지점을 짚습니다.
전체 구조
flowchart LR
subgraph Cpp[C++ 서버]
M1["/metrics 엔드포인트"]
M2["Counter / Gauge / Histogram"]
M2 --> M1
end
subgraph Prom[Prometheus]
P1["주기적 스크레이프"]
P2["TSDB"]
P3["알람 규칙 평가"]
P1 --> P2 --> P3
end
AM[Alertmanager]
G[Grafana]
Cpp -->|HTTP GET, 텍스트 포맷| P1
P2 -->|PromQL| G
P3 -->|발생한 알람| AM
AM -->|Slack, PagerDuty| U[담당자]
Prometheus는 pull 방식입니다. 서버가 메트릭을 보내는 것이 아니라, Prometheus가 설정된 간격마다 서버의 /metrics를 HTTP로 가져갑니다. 그래서 서버 쪽 구현은 “현재 값을 텍스트로 출력하는 엔드포인트” 하나면 되고, 수집 간격이나 대상 변경은 Prometheus 설정에서 합니다. 알람은 두 경로가 있습니다. Prometheus가 규칙을 평가해 Alertmanager로 보내는 방식과, Grafana가 직접 쿼리를 평가해 알림을 보내는 Grafana Alerting 방식입니다. 이 글은 설정을 코드로 관리하기 쉬운 앞쪽을 기준으로 합니다.
sequenceDiagram
participant Cpp as C++ 서버
participant Prom as Prometheus
participant Graf as Grafana
participant User as 사용자
loop scrape_interval마다
Prom->>Cpp: GET /metrics
Cpp-->>Prom: 200 OK 텍스트 메트릭
Prom->>Prom: TSDB에 저장
end
User->>Graf: 대시보드 조회
Graf->>Prom: PromQL 쿼리
Prom-->>Graf: 시계열 데이터
Graf-->>User: 패널 렌더링
메트릭 타입과 텍스트 포맷
| 타입 | 의미 | 예 |
|---|---|---|
| Counter | 단조 증가만 하는 누적값 | http_requests_total |
| Gauge | 오르내리는 현재값 | active_connections |
| Histogram | 관측값을 버킷별로 누적 집계 | http_request_duration_seconds |
Counter는 누적값 자체보다 rate()로 계산한 초당 증가율로 씁니다. 서버가 재시작해 값이 0으로 돌아가도 rate()가 이를 리셋으로 인식해 보정하므로, Counter를 Gauge처럼 줄이거나 임의로 초기화하면 안 됩니다.
Histogram에서 가장 헷갈리는 부분은 버킷이 누적이라는 점입니다. le="0.1" 버킷은 0.1초 이하인 모든 관측의 개수이고, le="0.25" 버킷은 0.25초 이하인 모든 관측의 개수이므로 0.1초 이하도 포함합니다. 마지막 le="+Inf" 버킷은 전체 개수와 같아야 하고, _sum과 _count가 함께 나갑니다.
# HELP http_request_duration_seconds HTTP request latency
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{path="/api/users",le="0.05"} 120
http_request_duration_seconds_bucket{path="/api/users",le="0.1"} 180
http_request_duration_seconds_bucket{path="/api/users",le="+Inf"} 200
http_request_duration_seconds_sum{path="/api/users"} 13.7
http_request_duration_seconds_count{path="/api/users"} 200
C++ 메트릭 레지스트리
아래 구현은 요청 경로에서 호출되는 inc()와 observe()를 잠금 없이 원자 연산만으로 처리하고, 레지스트리 조회와 직렬화에만 mutex를 씁니다.
// metrics_registry.hpp
#pragma once
#include <algorithm>
#include <atomic>
#include <locale>
#include <map>
#include <memory>
#include <mutex>
#include <sstream>
#include <string>
#include <vector>
namespace monitoring {
using Labels = std::map<std::string, std::string>;
// 라벨 값의 \, ", 줄바꿈은 이스케이프해야 텍스트 포맷이 깨지지 않음
inline std::string escape_label_value(const std::string& v) {
std::string out;
for (char c : v) {
if (c == '\\') out += "\\\\";
else if (c == '"') out += "\\\"";
else if (c == '\n') out += "\\n";
else out += c;
}
return out;
}
inline std::string format_double(double v) {
std::ostringstream os;
os.imbue(std::locale::classic()); // 로캘과 무관하게 소수점은 '.'
os << v; // 0.005 → "0.005"
return os.str();
}
inline std::string format_labels(const Labels& labels,
const std::string& extra_key = {},
const std::string& extra_value = {}) {
if (labels.empty() && extra_key.empty()) return "";
std::string out = "{";
bool first = true;
auto append = [&](const std::string& k, const std::string& v) {
if (!first) out += ',';
out += k + "=\"" + escape_label_value(v) + "\"";
first = false;
};
for (const auto& [k, v] : labels) append(k, v);
if (!extra_key.empty()) append(extra_key, extra_value);
return out + "}";
}
// 원자적으로 double에 더하기 (C++20 이전에는 atomic<double>::fetch_add가 없음)
inline void atomic_add(std::atomic<double>& target, double delta) {
double cur = target.load(std::memory_order_relaxed);
while (!target.compare_exchange_weak(cur, cur + delta, std::memory_order_relaxed)) {
}
}
class Counter {
public:
void inc(std::uint64_t n = 1) { value_.fetch_add(n, std::memory_order_relaxed); }
std::uint64_t get() const { return value_.load(std::memory_order_relaxed); }
private:
std::atomic<std::uint64_t> value_{0};
};
class Gauge {
public:
void set(double v) { value_.store(v, std::memory_order_relaxed); }
void inc(double d = 1.0) { atomic_add(value_, d); }
void dec(double d = 1.0) { atomic_add(value_, -d); }
double get() const { return value_.load(std::memory_order_relaxed); }
private:
std::atomic<double> value_{0.0};
};
class Histogram {
public:
explicit Histogram(std::vector<double> bounds)
: bounds_(std::move(bounds)), counts_(bounds_.size() + 1) { // 마지막 칸은 +Inf
std::sort(bounds_.begin(), bounds_.end());
}
void observe(double v) {
// v 이상인 첫 경계 = v가 속하는 가장 작은 le 버킷
auto idx = std::lower_bound(bounds_.begin(), bounds_.end(), v) - bounds_.begin();
counts_[idx].fetch_add(1, std::memory_order_relaxed);
atomic_add(sum_, v);
}
void write(std::ostream& os, const std::string& name, const Labels& labels) const {
std::uint64_t cumulative = 0; // 버킷은 누적으로 출력해야 함
for (size_t i = 0; i < bounds_.size(); ++i) {
cumulative += counts_[i].load(std::memory_order_relaxed);
os << name << "_bucket" << format_labels(labels, "le", format_double(bounds_[i]))
<< ' ' << cumulative << '\n';
}
cumulative += counts_.back().load(std::memory_order_relaxed);
os << name << "_bucket" << format_labels(labels, "le", "+Inf") << ' ' << cumulative << '\n';
os << name << "_sum" << format_labels(labels) << ' '
<< format_double(sum_.load(std::memory_order_relaxed)) << '\n';
os << name << "_count" << format_labels(labels) << ' ' << cumulative << '\n';
}
private:
std::vector<double> bounds_;
std::vector<std::atomic<std::uint64_t>> counts_; // 버킷별 비누적 개수
std::atomic<double> sum_{0.0};
};
class MetricsRegistry {
public:
Counter& counter(const std::string& name, const std::string& help, const Labels& labels = {}) {
return get_or_create(counters_, name, help, labels, [] { return std::make_unique<Counter>(); });
}
Gauge& gauge(const std::string& name, const std::string& help, const Labels& labels = {}) {
return get_or_create(gauges_, name, help, labels, [] { return std::make_unique<Gauge>(); });
}
Histogram& histogram(const std::string& name, const std::string& help,
const std::vector<double>& bounds, const Labels& labels = {}) {
return get_or_create(histograms_, name, help, labels,
[&] { return std::make_unique<Histogram>(bounds); });
}
std::string export_text() const {
std::lock_guard<std::mutex> lock(mutex_);
std::ostringstream os;
os.imbue(std::locale::classic());
for (const auto& [name, fam] : counters_) {
os << "# HELP " << name << ' ' << fam.help << "\n# TYPE " << name << " counter\n";
for (const auto& [labels, c] : fam.series)
os << name << format_labels(labels) << ' ' << c->get() << '\n';
}
for (const auto& [name, fam] : gauges_) {
os << "# HELP " << name << ' ' << fam.help << "\n# TYPE " << name << " gauge\n";
for (const auto& [labels, g] : fam.series)
os << name << format_labels(labels) << ' ' << format_double(g->get()) << '\n';
}
for (const auto& [name, fam] : histograms_) {
os << "# HELP " << name << ' ' << fam.help << "\n# TYPE " << name << " histogram\n";
for (const auto& [labels, h] : fam.series) h->write(os, name, labels);
}
return os.str();
}
private:
template <class M>
struct Family {
std::string help;
std::map<Labels, std::unique_ptr<M>> series; // 라벨 조합 하나 = 시계열 하나
};
template <class M, class Make>
M& get_or_create(std::map<std::string, Family<M>>& families, const std::string& name,
const std::string& help, const Labels& labels, Make make) {
std::lock_guard<std::mutex> lock(mutex_);
auto& fam = families[name];
if (fam.help.empty()) fam.help = help;
auto& ptr = fam.series[labels];
if (!ptr) ptr = make();
return *ptr; // unique_ptr이라 map이 커져도 참조는 유효
}
mutable std::mutex mutex_;
std::map<std::string, Family<Counter>> counters_;
std::map<std::string, Family<Gauge>> gauges_;
std::map<std::string, Family<Histogram>> histograms_;
};
} // namespace monitoring
몇 가지 설계 선택을 짚어 둡니다. 텍스트 포맷은 같은 이름의 시계열이 # TYPE 줄 아래에 모여 있어야 하므로, 이름별로 묶은 Family 구조로 저장합니다. Histogram은 버킷별로 비누적 개수를 원자 변수에 세고, 출력할 때 누적합으로 바꿉니다. 관측할 때마다 누적 버킷 여러 개를 갱신하는 방식보다 요청 경로의 원자 연산이 적습니다. 스크레이프는 여러 원자 변수를 차례로 읽으므로 _sum과 _count가 같은 순간의 스냅샷은 아니지만, 수집 간격에 비하면 무시할 만한 차이이고 Prometheus도 이런 미세한 어긋남을 전제로 계산합니다.
레지스트리 조회(counter(...))는 전역 mutex와 map 탐색을 거칩니다. 라벨 조합이 고정되어 있다면 시작할 때 참조를 받아 두고 요청마다 inc()만 호출하는 편이 빠릅니다. 아래 예제처럼 경로별로 동적으로 조회한다면, 부하가 큰 서버에서는 이 mutex가 경합 지점이 될 수 있다는 점을 염두에 둡니다.
/metrics 엔드포인트와 요청 계측
// main.cpp 일부 - Boost.Beast
#include "metrics_registry.hpp"
#include <boost/beast.hpp>
namespace beast = boost::beast;
namespace http = beast::http;
inline monitoring::MetricsRegistry& metrics() {
static monitoring::MetricsRegistry reg;
return reg;
}
http::response<http::string_body> handle_metrics(const http::request<http::string_body>& req) {
http::response<http::string_body> res{http::status::ok, req.version()};
res.set(http::field::content_type, "text/plain; version=0.0.4; charset=utf-8");
res.body() = metrics().export_text();
res.prepare_payload();
return res;
}
version=0.0.4는 Prometheus 텍스트 포맷 버전을 나타냅니다. /metrics에는 내부 구조가 드러나므로 외부에 공개하지 않는 것이 좋습니다. 뒤에서 설명하듯 메트릭을 별도 포트로 분리하면 네트워크 정책으로 막기 쉽습니다.
// 요청 하나가 끝날 때 호출
std::string normalize_path(const std::string& path); // 아래 카디널리티 절 참고
void record_request(const std::string& method, const std::string& path,
int status, double duration_sec) {
const std::string route = normalize_path(path);
metrics().counter("http_requests_total", "Total HTTP requests",
{{"method", method}, {"path", route}, {"status", std::to_string(status)}})
.inc();
metrics().histogram("http_request_duration_seconds", "HTTP request latency",
{0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0},
{{"path", route}})
.observe(duration_sec);
}
// 연결이 살아 있는 동안 active_connections를 1 올려 두는 RAII 객체
class ConnectionTracker {
public:
ConnectionTracker() { gauge().inc(); }
~ConnectionTracker() { gauge().dec(); }
ConnectionTracker(const ConnectionTracker&) = delete;
ConnectionTracker& operator=(const ConnectionTracker&) = delete;
private:
static monitoring::Gauge& gauge() {
static auto& g = metrics().gauge("active_connections", "Open client connections");
return g;
}
};
Gauge를 RAII로 관리하면 예외나 조기 반환으로 연결 처리가 끝나도 감소가 빠지지 않습니다. 증가와 감소를 서로 다른 코드 경로에서 직접 호출하면, 에러 경로 하나에서 dec()가 빠져 값이 계속 올라가는 버그가 생기기 쉽습니다. 저는 Gauge가 며칠에 걸쳐 조금씩 올라가는 그래프를 보면 실제 누수보다 이런 계측 누락부터 의심합니다.
메모리 사용량은 Linux라면 /proc/self/status의 VmRSS를 주기적으로 읽어 Gauge에 넣을 수 있습니다.
#include <fstream>
void update_memory_metric() {
#ifdef __linux__
std::ifstream status("/proc/self/status");
std::string line;
while (std::getline(status, line)) {
if (line.rfind("VmRSS:", 0) == 0) { // "VmRSS: 123456 kB"
double kb = std::stod(line.substr(6));
static auto& g = metrics().gauge("process_resident_memory_bytes",
"Resident memory size in bytes");
g.set(kb * 1024);
break;
}
}
#endif
}
Prometheus 설정과 실행
# docker-compose.monitoring.yml
services:
cpp-api:
build: .
ports:
- "8080:8080" # API
expose:
- "8081" # /metrics, 컨테이너 네트워크 안에서만
prometheus:
image: prom/prometheus:v2.47.0
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
- ./alerts.yml:/etc/prometheus/alerts.yml
ports:
- "9090:9090"
command:
- '--config.file=/etc/prometheus/prometheus.yml'
- '--storage.tsdb.retention.time=15d'
- '--storage.tsdb.retention.size=50GB'
alertmanager:
image: prom/alertmanager:v0.26.0
volumes:
- ./alertmanager.yml:/etc/alertmanager/alertmanager.yml
grafana:
image: grafana/grafana:10.2.0
ports:
- "3000:3000"
environment:
- GF_SECURITY_ADMIN_PASSWORD__FILE=/run/secrets/grafana_admin
- GF_USERS_ALLOW_SIGN_UP=false
volumes:
- grafana-data:/var/lib/grafana
depends_on:
- prometheus
volumes:
grafana-data:
# prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s
rule_files:
- /etc/prometheus/alerts.yml
alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093']
scrape_configs:
- job_name: 'cpp-api'
metrics_path: /metrics
scrape_interval: 10s
scrape_timeout: 5s
static_configs:
- targets: ['cpp-api:8081']
보존 기간과 용량은 설정 파일이 아니라 --storage.tsdb.retention.time, --storage.tsdb.retention.size 명령행 플래그로 정합니다. 설정 파일에 storage.tsdb.retention 같은 항목을 적어도 적용되지 않습니다. scrape_timeout은 scrape_interval보다 길 수 없습니다. Grafana 관리자 비밀번호는 예제처럼 파일이나 비밀 관리 도구로 주입하고, 기본값 admin을 그대로 두지 않습니다.
Kubernetes에서는 정적 대상 대신 서비스 디스커버리로 파드를 찾고, 파드 이름을 라벨로 붙여 특정 파드만 이상한 경우를 구분할 수 있게 합니다.
scrape_configs:
- job_name: 'cpp-api'
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_label_app]
regex: cpp-api
action: keep
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
Grafana 패널과 PromQL
Grafana에서 Prometheus 데이터 소스를 추가할 때 URL은 Grafana 서버 기준으로 Prometheus에 닿는 주소여야 합니다. Docker Compose 안이라면 http://prometheus:9090입니다.
# 경로별 초당 요청 수
sum by (path) (rate(http_requests_total[5m]))
# 5xx 비율(%)
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# 경로별 p50 / p95 / p99 지연
histogram_quantile(0.99, sum by (le, path) (rate(http_request_duration_seconds_bucket[5m])))
# 현재 연결 수
sum(active_connections)
에러율 식에서 분자와 분모 모두 sum으로 라벨을 없애는 것이 중요합니다. rate(...{status=~"5.."}) / rate(...)처럼 그대로 나누면 PromQL은 라벨이 완전히 같은 시계열끼리만 나누는데, 분자에는 status="500", 분모에는 status="200" 같은 시계열이 섞여 있어 짝이 맞는 것끼리만 계산되고 결과가 엉뚱하거나 비어 버립니다. 마찬가지로 histogram_quantile에 여러 인스턴스를 합칠 때는 le를 남긴 채 sum by (le)로 버킷을 합쳐야 합니다. 백분위수를 인스턴스별로 구한 뒤 평균 내는 것은 통계적으로 의미가 없습니다.
rate()의 범위는 스크레이프 간격의 최소 네 배 정도로 잡는 것이 무난합니다. 범위가 너무 짧으면 구간 안의 샘플이 두 개가 안 되어 값이 비거나 들쭉날쭉합니다. Grafana의 $__rate_interval 변수는 이 기준에 맞춰 범위를 자동으로 정해 줍니다. 범위를 길게 잡으면 곡선이 매끄러워지는 대신 짧은 스파이크가 묻힙니다.
패널 JSON의 핵심은 쿼리와 단위입니다.
{
"title": "p99 지연 시간",
"type": "timeseries",
"targets": [
{
"expr": "histogram_quantile(0.99, sum by (le, path) (rate(http_request_duration_seconds_bucket[$__rate_interval])))",
"legendFormat": "{{path}}"
}
],
"fieldConfig": { "defaults": { "unit": "s", "min": 0 } }
}
알람 규칙
# alerts.yml
groups:
- name: cpp-api-alerts
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "5xx 비율 5% 초과"
- alert: HighLatencyP99
expr: |
histogram_quantile(0.99,
sum by (le) (rate(http_request_duration_seconds_bucket[5m]))) > 1
for: 5m
labels:
severity: warning
annotations:
summary: "p99 지연 1초 초과"
- alert: TargetDown
expr: up{job="cpp-api"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "C++ API 스크레이프 실패"
for: 5m은 조건이 5분 동안 연속으로 참이어야 알람을 보낸다는 뜻이라, 순간적인 스파이크로 밤중에 호출되는 일을 줄입니다. up == 0은 Prometheus가 대상을 알고 있는데 스크레이프에 실패한 경우만 잡습니다. 서비스 디스커버리에서 대상 자체가 사라지면 up 시계열이 없어져 이 규칙이 아예 평가되지 않으므로, 그런 구성이라면 absent(up{job="cpp-api"}) 규칙을 함께 둡니다.
# alertmanager.yml
route:
receiver: slack-warning
group_by: ['alertname']
routes:
- matchers: ['severity="critical"']
receiver: pagerduty
receivers:
- name: slack-warning
slack_configs:
- api_url: https://hooks.slack.com/services/xxx/yyy/zzz
channel: '#alerts'
- name: pagerduty
pagerduty_configs:
- routing_key: <PagerDuty 통합 키>
라우팅 조건은 예전 문서의 match: 대신 matchers: 문법을 쓰는 것이 현재 권장 방식입니다.
자주 겪는 문제
Prometheus가 파싱에 실패함
스크레이프가 실패하고 Targets 화면에 파싱 에러가 보인다면 텍스트 포맷 위반입니다. 흔한 원인은 세 가지입니다. 라벨 값에 따옴표나 역슬래시가 이스케이프되지 않은 채 들어간 경우, 메트릭 이름이 [a-zA-Z_:][a-zA-Z0-9_:]* 규칙을 어긴 경우(http-requests처럼 하이픈을 쓰면 안 됩니다), 그리고 로캘 때문에 소수점이 쉼표로 출력된 경우입니다. 마지막 것은 서버의 전역 로캘을 바꾸는 코드가 있을 때만 나타나서 개발 환경에서는 재현되지 않는 경우가 많습니다. 앞의 구현이 std::locale::classic()을 명시적으로 쓰는 이유입니다. promtool check metrics에 /metrics 출력을 넣으면 포맷 문제를 미리 확인할 수 있습니다.
curl -s http://localhost:8081/metrics | promtool check metrics
Grafana에 No data
# 타겟 상태와 마지막 에러
curl -s http://localhost:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health, lastError}'
# 서버 출력 직접 확인
curl -s http://localhost:8081/metrics | head
타겟이 DOWN이면 호스트 이름과 포트, 컨테이너 네트워크를 확인합니다. UP인데 값이 없으면 쿼리의 메트릭 이름·라벨이 실제 출력과 다른 경우가 대부분이고, Prometheus UI에서는 나오는데 Grafana에서만 비면 데이터 소스 URL과 대시보드 시간 범위를 봅니다.
카디널리티 폭증
Prometheus는 라벨 조합 하나마다 별도의 시계열을 메모리에 유지합니다. path 라벨에 /users/12345 같은 원래 경로를 그대로 넣으면 사용자 수만큼 시계열이 생기고, 이것이 status와 method 조합과 곱해져 Prometheus 메모리가 급격히 늘어납니다. 서버 쪽 레지스트리의 map도 같은 속도로 커지므로 C++ 프로세스 메모리도 함께 늘어납니다.
#include <set>
// 라우터가 알고 있는 경로 템플릿을 쓰는 것이 가장 정확함
std::string normalize_path(const std::string& path) {
if (path.rfind("/users/", 0) == 0) return "/users/:id";
if (path.rfind("/orders/", 0) == 0) return "/orders/:id";
static const std::set<std::string> known = {"/", "/health", "/api/users"};
return known.count(path) ? path : "other"; // 모르는 경로는 한 버킷으로
}
모르는 경로를 그대로 두지 않고 other로 묶는 것이 중요합니다. 스캐너가 임의 경로로 요청을 퍼붓는 것만으로도 시계열이 무한히 늘어날 수 있기 때문입니다. 사용자 ID, 요청 ID, 전체 URL, 에러 메시지 원문은 라벨에 넣지 않고, 필요하면 로그나 트레이스로 보냅니다.
히스토그램 버킷이 실제 분포와 맞지 않음
histogram_quantile은 버킷 경계 사이를 선형 보간해 백분위수를 추정합니다. 그래서 버킷이 {1, 2, 5}초뿐인데 실제 요청이 대부분 10ms라면, 모든 관측이 첫 버킷(0~1초)에 몰리고 p99가 0.99초 근처로 계산됩니다. 실제보다 100배 가까이 부풀려진 값입니다. 반대로 관측값이 가장 큰 유한 버킷보다 크면 결과는 그 버킷의 경계값으로 잘립니다. 버킷은 예상 지연 범위를 촘촘하게 덮도록, 특히 SLO 기준값(예: 300ms) 근처에 경계를 두도록 정합니다. 버킷마다 시계열이 하나씩 늘어나므로 무작정 많이 두는 것도 카디널리티 비용입니다.
/metrics 응답이 느림
시계열이 수만 개로 늘면 직렬화 자체가 수백 ms 걸릴 수 있고, 그동안 레지스트리 mutex를 잡고 있어 요청 경로의 레지스트리 조회까지 막힙니다. 근본 해결은 카디널리티를 줄이는 것입니다. 그래도 필요하다면 직렬화 결과를 짧게 캐시할 수 있는데, 캐시 문자열을 잠금 없이 읽고 쓰면 데이터 레이스가 되므로 공유 포인터를 교체하는 방식이 안전합니다.
std::string cached_metrics() {
static std::mutex m;
static std::shared_ptr<const std::string> cached;
static std::chrono::steady_clock::time_point last;
std::lock_guard<std::mutex> lock(m);
auto now = std::chrono::steady_clock::now();
if (!cached || now - last >= std::chrono::seconds(1)) {
cached = std::make_shared<const std::string>(metrics().export_text());
last = now;
}
return *cached;
}
SLO 추적
“30일 동안 요청의 99.9%가 성공” 같은 목표를 추적할 때 rate(...[30d])를 대시보드에서 직접 계산하면 쿼리마다 30일치 샘플을 읽어 무겁습니다. 짧은 구간의 비율을 recording rule로 미리 계산해 두고, 그 결과를 길게 평균 내는 방식이 일반적입니다.
# alerts.yml에 함께 둘 수 있음
groups:
- name: cpp-api-slo
interval: 30s
rules:
- record: job:http_requests:rate5m
expr: sum by (job) (rate(http_requests_total[5m]))
- record: job:http_errors:rate5m
expr: sum by (job) (rate(http_requests_total{status=~"5.."}[5m]))
# 30일 가용성
1 - (sum_over_time(job:http_errors:rate5m[30d]) / sum_over_time(job:http_requests:rate5m[30d]))
# 남은 에러 버짓 비율 (목표 99.9% → 허용 실패 비율 0.001)
1 - (
(sum_over_time(job:http_errors:rate5m[30d]) / sum_over_time(job:http_requests:rate5m[30d]))
/ 0.001
)
두 번째 식은 허용된 실패 비율(0.1%) 대비 실제로 쓴 실패 비율을 빼서 남은 에러 버짓을 구합니다. 1이면 버짓을 전혀 쓰지 않은 상태이고, 0 이하이면 이번 기간의 목표를 이미 어긴 상태입니다. 지연 SLO도 같은 방식으로, 기준값을 버킷 경계로 두고 le="0.3" 버킷 비율을 recording rule로 계산하면 백분위수보다 정확하게 “300ms 이내 요청 비율”을 추적할 수 있습니다.
같이 보면 좋은 글
- C++ 시리즈 전체 보기
- C++ Observability: Prometheus와 Grafana로 C++ 서버 모니터링 구축하기
- C++ 서버 배포: Docker 멀티스테이지, systemd, Kubernetes 무중단 배포, Prometheus 모니터링
- C++에서 PostgreSQL 연동
- C++ Python과 C++의 만남 | pybind11으로 고성능 엔진 만들기 [#35-1]
다음 글: [C++ 실전 가이드 #51-1] C++ 프로파일러 비교: perf, gprof, Valgrind, VTune, Tracy 이전 글: [C++ 실전 가이드 #50-5] 프로덕션 배포 자동화