C++ 멀티스레드 네트워크 서버: io_context 스레드 풀, strand로 순서 보장, 세션 수명과 data race
들어가며: io_context를 여러 스레드에서 돌리면 생기는 일
Boost.Asio 서버를 멀티코어로 확장하는 가장 간단한 방법은 하나의 io_context에 대해 여러 스레드에서 run()을 호출하는 것입니다.
boost::asio::io_context io;
tcp::acceptor acceptor(io, tcp::endpoint(tcp::v4(), 8080));
std::vector<std::thread> threads;
for (int i = 0; i < 4; ++i) {
threads.emplace_back([&io]() { io.run(); });
}
이렇게 하면 완료된 비동기 연산의 핸들러는 run()을 호출한 스레드 중 그 순간 비어 있는 아무 스레드에서나 실행됩니다. 한 연결에 대해 비동기 연산이 하나씩만 걸려 있는 단순한 에코 서버라면(읽기가 끝나야 쓰기를 시작하고, 쓰기가 끝나야 다시 읽기를 시작) 같은 연결의 핸들러가 동시에 실행될 일이 없습니다. Asio 문서는 이를 “암묵적 strand”라고 부릅니다.
문제는 한 연결에 여러 연산이 동시에 걸릴 때 생깁니다. 읽기를 기다리는 중에 다른 연결이 보낸 채팅 메시지를 이 소켓에 써야 하거나, 타임아웃 타이머가 함께 돌거나, 여러 메시지를 연달아 쓰는 경우입니다. 그러면 같은 세션의 버퍼와 상태를 두 스레드가 동시에 건드리는 data race가 생기고, 같은 소켓에 async_write 두 개가 겹쳐 메시지 바이트가 섞일 수도 있습니다. 여기에 여러 연결이 함께 쓰는 연결 목록 같은 공유 상태, 그리고 비동기 연산이 끝나기 전에 세션 객체가 사라지는 수명 문제가 더해집니다.
이 글은 이 세 가지를 차례로 다룹니다. strand로 한 연결의 핸들러를 직렬화하고, shared_ptr로 세션 수명을 비동기 연산에 묶고, 공유 상태는 뮤텍스나 별도 strand로 보호합니다. 예제는 Boost 1.74 이상(any_io_executor 도입 이후)을 기준으로 합니다.
스레딩 모델
flowchart TB
subgraph Clients[클라이언트들]
C1[클라이언트 1]
C2[클라이언트 2]
C3[클라이언트 N]
end
subgraph Server[멀티스레드 서버]
Acceptor[acceptor]
IO[io_context]
subgraph ThreadPool[run을 호출하는 스레드들]
T1[스레드 1]
T2[스레드 2]
T3[스레드 3]
T4[스레드 4]
end
subgraph Sessions[세션들]
S1[세션 A + strand]
S2[세션 B + strand]
S3[세션 C + strand]
end
Shared["공유 상태(연결 목록 등)"]
end
C1 --> Acceptor
C2 --> Acceptor
C3 --> Acceptor
Acceptor --> S1
Acceptor --> S2
Acceptor --> S3
IO --> T1
IO --> T2
IO --> T3
IO --> T4
S1 --> IO
S2 --> IO
S3 --> IO
S1 -.->|뮤텍스 또는 전용 strand| Shared
S2 -.->|뮤텍스 또는 전용 strand| Shared
S3 -.->|뮤텍스 또는 전용 strand| Shared
io_context는 완료된 핸들러들의 큐를 하나 갖고 있고, run()을 호출한 스레드들이 이 큐에서 핸들러를 하나씩 가져가 실행합니다. 어느 스레드가 어떤 핸들러를 실행할지는 정해져 있지 않습니다. strand는 이 위에 “같은 strand에 속한 핸들러는 동시에 실행되지 않고, 들어온 순서대로 실행된다”는 보장을 더하는 executor입니다.
여러 스레드에서 run 호출하기
#include <boost/asio.hpp>
#include <algorithm>
#include <iostream>
#include <thread>
#include <vector>
class ThreadPoolServer {
boost::asio::io_context io_;
boost::asio::executor_work_guard<boost::asio::io_context::executor_type> work_;
std::vector<std::thread> threads_;
public:
explicit ThreadPoolServer(unsigned num_threads =
std::max(1u, std::thread::hardware_concurrency()))
: work_(boost::asio::make_work_guard(io_)) {
threads_.reserve(num_threads);
for (unsigned i = 0; i < num_threads; ++i) {
threads_.emplace_back([this]() { io_.run(); });
}
}
~ThreadPoolServer() { stop(); }
void stop() {
work_.reset();
io_.stop();
for (auto& t : threads_) {
if (t.joinable()) t.join();
}
threads_.clear();
}
template <typename F>
void post(F&& f) { boost::asio::post(io_, std::forward<F>(f)); }
boost::asio::io_context& context() { return io_; }
};
work guard는 io_context에 “아직 할 일이 남아 있다”는 표시를 남겨, 처리할 비동기 연산이 하나도 없을 때 run()이 바로 반환되지 않게 합니다. 서버가 시작 직후 아직 accept를 등록하기 전에 스레드가 먼저 run()에 들어가는 경우를 생각하면 필요한 이유가 분명합니다. hardware_concurrency()는 값을 알 수 없을 때 0을 반환할 수 있으므로 최소 1로 맞춥니다.
strand로 한 연결의 핸들러 직렬화하기
strand를 쓰는 방법은 두 가지입니다. 하나는 핸들러마다 bind_executor(strand, handler)로 감싸는 것이고, 다른 하나는 처음부터 소켓을 strand executor 위에 만드는 것입니다. 두 번째 방식이 빠뜨릴 여지가 적습니다. 소켓의 기본 executor가 strand이면, 그 소켓에서 시작한 비동기 연산의 완료 핸들러는 따로 감싸지 않아도 그 strand에서 실행됩니다.
// accept할 때 새 소켓을 연결마다 새 strand 위에 만든다
acceptor_.async_accept(
boost::asio::make_strand(io_),
[this](boost::system::error_code ec, tcp::socket socket) {
if (!ec) std::make_shared<Session>(std::move(socket))->start();
start_accept();
});
이때 socket.get_executor()는 그 연결 전용 strand를 돌려주므로, 세션 안에서 타이머를 만들거나 post할 때도 같은 executor를 쓰면 연결 하나의 모든 작업이 자동으로 직렬화됩니다.
strand가 순서를 보장하는 원리는 내부 큐입니다. strand에 들어온 핸들러는 strand별 큐에 쌓이고, 한 번에 하나만 io_context로 넘어가 실행됩니다. 구현 내부에서는 이 큐를 보호하는 데 짧은 뮤텍스를 쓰므로 비용이 0은 아니지만, 사용자 코드가 락을 잡고 기다리는 일은 없습니다. 같은 연결의 핸들러끼리는 서로 기다리지 않고 순서대로 처리되며, 다른 연결의 핸들러는 다른 스레드에서 동시에 실행됩니다.
뮤텍스와 비교하면, 연결 하나의 상태(버퍼, 쓰기 큐, 타이머)를 보호하는 데는 strand가 맞습니다. 핸들러 안에서 락을 잡고 기다리지 않으므로 스레드가 막히지 않고, “읽기 핸들러 실행 중에는 쓰기 핸들러가 끼어들지 않는다”는 순서까지 함께 보장되기 때문입니다. 반대로 여러 연결이 아주 잠깐씩 건드리는 카운터나 맵처럼 작은 공유 상태는 뮤텍스가 더 단순합니다.
세션과 수명 관리
연결 하나를 세션 객체 하나로 표현하고, 소켓과 버퍼를 멤버로 둡니다. 비동기 연산을 시작할 때마다 shared_from_this()로 얻은 shared_ptr을 완료 핸들러에 캡처하면, 연산이 걸려 있는 동안 세션이 살아 있고 마지막 핸들러가 끝나면서 참조가 사라지면 자동으로 소멸합니다.
#include <boost/asio.hpp>
#include <array>
#include <iostream>
#include <memory>
using boost::asio::ip::tcp;
using boost::system::error_code;
class Session : public std::enable_shared_from_this<Session> {
tcp::socket socket_; // 연결별 strand 위에서 생성된 소켓
std::array<char, 4096> buffer_;
public:
explicit Session(tcp::socket socket) : socket_(std::move(socket)) {}
void start() { do_read(); }
private:
void do_read() {
socket_.async_read_some(
boost::asio::buffer(buffer_),
[self = shared_from_this()](error_code ec, std::size_t bytes) {
if (ec) {
if (ec != boost::asio::error::eof &&
ec != boost::asio::error::operation_aborted) {
std::cerr << "Read error: " << ec.message() << "\n";
}
return; // 더 이상 연산이 없으므로 self가 사라지며 세션 소멸
}
self->do_write(bytes);
});
}
void do_write(std::size_t bytes) {
boost::asio::async_write(
socket_, boost::asio::buffer(buffer_, bytes),
[self = shared_from_this()](error_code ec, std::size_t) {
if (!ec) self->do_read();
});
}
};
shared_from_this()는 객체가 이미 shared_ptr로 관리되고 있을 때만 쓸 수 있으므로, 세션은 반드시 std::make_shared<Session>(...)으로 만든 뒤 start()를 호출해야 합니다. 생성자 안에서 shared_from_this()를 부르면 C++17부터 std::bad_weak_ptr 예외가 납니다. 핸들러에 this만 캡처하면, 연결이 끊겨 세션이 해제된 뒤 남은 핸들러가 실행될 때 use-after-free가 생깁니다.
bind_executor로 핸들러를 감싸는 방식을 쓴다면 strand 멤버의 타입에 주의하세요. Boost 1.74부터 socket.get_executor()는 io_context::executor_type이 아니라 any_io_executor를 돌려주므로, boost::asio::strand<boost::asio::io_context::executor_type> strand_{make_strand(socket_.get_executor())}처럼 쓰면 타입이 맞지 않아 컴파일되지 않습니다. boost::asio::strand<tcp::socket::executor_type>로 선언하거나, 앞에서처럼 소켓 자체를 strand 위에 만드는 편이 간단합니다.
공유 상태 동기화
여러 연결이 공유하는 데이터(연결 목록, 채팅방, 카운터)는 연결별 strand로 보호되지 않습니다. 서로 다른 strand의 핸들러는 동시에 실행되기 때문입니다. 뮤텍스나, 그 공유 상태 전용 strand가 필요합니다.
방법 1: 뮤텍스
#include <mutex>
#include <string>
#include <unordered_set>
class ConnectionManager {
std::unordered_set<std::string> connections_;
mutable std::mutex mutex_;
public:
void add(const std::string& id) {
std::lock_guard<std::mutex> lock(mutex_);
connections_.insert(id);
}
void remove(const std::string& id) {
std::lock_guard<std::mutex> lock(mutex_);
connections_.erase(id);
}
std::size_t count() const {
std::lock_guard<std::mutex> lock(mutex_);
return connections_.size();
}
};
count()가 const 멤버 함수이므로 뮤텍스를 mutable로 선언해야 락을 잡을 수 있습니다. 락을 잡은 채로 소켓 I/O나 오래 걸리는 작업을 하지 않는 것이 원칙입니다.
방법 2: 전용 strand로 직렬화
공유 상태를 그 상태 전용 strand에서만 건드리면 뮤텍스 없이 단일 스레드처럼 다룰 수 있습니다. 대신 결과를 바로 반환할 수 없고 콜백으로 받아야 합니다.
class StrandConnectionManager {
boost::asio::strand<boost::asio::io_context::executor_type> strand_;
std::unordered_set<std::string> connections_; // strand_ 안에서만 접근
public:
explicit StrandConnectionManager(boost::asio::io_context& io)
: strand_(boost::asio::make_strand(io)) {}
void add(std::string id) {
boost::asio::post(strand_, [this, id = std::move(id)]() { connections_.insert(id); });
}
void remove(std::string id, std::function<void(std::size_t)> on_done = nullptr) {
boost::asio::post(strand_, [this, id = std::move(id), on_done]() {
connections_.erase(id);
if (on_done) on_done(connections_.size());
});
}
};
여기서 make_strand(io)는 io_context::executor_type 기반 strand를 만들므로 멤버 타입과 맞습니다. 람다가 this를 캡처하므로 StrandConnectionManager는 걸려 있는 핸들러가 모두 끝날 때까지 살아 있어야 합니다.
전체 에코 서버
// 앞의 Session 클래스와 함께 컴파일
class Server {
boost::asio::io_context& io_;
tcp::acceptor acceptor_;
public:
Server(boost::asio::io_context& io, unsigned short port)
: io_(io), acceptor_(boost::asio::make_strand(io), tcp::endpoint(tcp::v4(), port)) {
start_accept();
}
void close() { // acceptor 전용 strand에서 닫아 accept 핸들러와 겹치지 않게 함
boost::asio::post(acceptor_.get_executor(), [this]() { acceptor_.close(); });
}
private:
void start_accept() {
acceptor_.async_accept(
boost::asio::make_strand(io_), // 연결마다 새 strand
[this](error_code ec, tcp::socket socket) {
if (ec == boost::asio::error::operation_aborted) return; // acceptor 닫힘
if (!ec) std::make_shared<Session>(std::move(socket))->start();
start_accept();
});
}
};
int main() {
boost::asio::io_context io;
Server server(io, 8080);
boost::asio::signal_set signals(io, SIGINT, SIGTERM);
signals.async_wait([&](error_code, int) {
server.close(); // 새 연결을 받지 않고, 기존 세션은 자연스럽게 끝나게 둠
});
std::vector<std::thread> threads;
unsigned n = std::max(1u, std::thread::hardware_concurrency());
for (unsigned i = 0; i < n; ++i) threads.emplace_back([&io]() { io.run(); });
for (auto& t : threads) t.join();
}
이 서버는 work guard를 쓰지 않습니다. acceptor가 항상 accept를 하나 걸어 두므로 할 일이 떨어지지 않고, 시그널을 받아 acceptor를 닫으면 남은 세션들이 끝나는 대로 할 일이 0이 되어 run()이 스스로 반환됩니다. 즉시 끊고 싶다면 io.stop()을 호출하면 되지만, 이 경우 진행 중이던 쓰기가 중간에 끊길 수 있습니다. signal_set은 main의 지역 변수로 두어 서버가 도는 동안 살아 있게 했습니다. 함수 안에서 만들고 바로 반환하면 소멸하면서 대기가 취소됩니다.
채팅 서버: 쓰기 큐와 브로드캐스트
채팅 서버는 앞의 에코 서버와 달리 한 연결에 연산이 동시에 걸립니다. 읽기를 기다리는 동안 다른 사용자의 메시지를 이 소켓에 써야 하고, 메시지가 연달아 오면 쓰기도 여러 개가 겹칩니다. 같은 소켓에 async_write를 겹쳐 호출하면 두 메시지의 바이트가 섞일 수 있으므로, 연결마다 쓰기 큐를 두고 한 번에 하나씩만 쓰는 것이 핵심입니다. 큐와 소켓은 그 연결의 strand에서만 건드립니다.
#include <boost/asio.hpp>
#include <array>
#include <deque>
#include <memory>
#include <mutex>
#include <set>
#include <string>
using boost::asio::ip::tcp;
using boost::system::error_code;
class ChatSession;
class ChatRoom { // 여러 연결이 공유: 뮤텍스로 보호
std::mutex mutex_;
std::set<std::shared_ptr<ChatSession>> members_;
public:
void join(std::shared_ptr<ChatSession> s) {
std::lock_guard<std::mutex> lock(mutex_);
members_.insert(std::move(s));
}
void leave(const std::shared_ptr<ChatSession>& s) {
std::lock_guard<std::mutex> lock(mutex_);
members_.erase(s);
}
void broadcast(const std::string& msg, const ChatSession* from);
};
class ChatSession : public std::enable_shared_from_this<ChatSession> {
tcp::socket socket_; // 연결별 strand 위의 소켓
ChatRoom& room_;
std::array<char, 4096> buffer_;
std::deque<std::string> write_queue_; // strand 안에서만 접근
std::string nickname_;
public:
ChatSession(tcp::socket socket, ChatRoom& room, std::string nickname)
: socket_(std::move(socket)), room_(room), nickname_(std::move(nickname)) {}
void start() {
room_.join(shared_from_this());
do_read();
}
// 다른 스레드에서 호출될 수 있음 → 이 세션의 strand로 넘겨서 처리
void deliver(std::string msg) {
boost::asio::post(socket_.get_executor(),
[self = shared_from_this(), msg = std::move(msg)]() mutable {
bool idle = self->write_queue_.empty();
self->write_queue_.push_back(std::move(msg));
if (idle) self->do_write(); // 진행 중인 쓰기가 없을 때만 시작
});
}
private:
void do_read() {
socket_.async_read_some(boost::asio::buffer(buffer_),
[self = shared_from_this()](error_code ec, std::size_t n) {
if (ec) { self->room_.leave(self); return; }
self->room_.broadcast("[" + self->nickname_ + "] " +
std::string(self->buffer_.data(), n), self.get());
self->do_read();
});
}
void do_write() {
boost::asio::async_write(socket_, boost::asio::buffer(write_queue_.front()),
[self = shared_from_this()](error_code ec, std::size_t) {
if (ec) { self->room_.leave(self); return; }
self->write_queue_.pop_front();
if (!self->write_queue_.empty()) self->do_write();
});
}
};
void ChatRoom::broadcast(const std::string& msg, const ChatSession* from) {
std::vector<std::shared_ptr<ChatSession>> targets;
{
std::lock_guard<std::mutex> lock(mutex_);
for (auto& s : members_) if (s.get() != from) targets.push_back(s);
}
for (auto& s : targets) s->deliver(msg); // 락을 푼 뒤 각 세션의 strand로 전달
}
async_write에 넘긴 버퍼는 쓰기가 끝날 때까지 유효해야 하는데, 여기서는 큐의 맨 앞 문자열을 쓰고 완료 핸들러에서야 pop_front하므로 그 조건이 지켜집니다. std::deque는 양 끝에 원소를 넣고 빼도 다른 원소의 참조가 무효화되지 않으므로, 쓰기 도중 뒤에 메시지가 추가되어도 안전합니다. ChatRoom은 members_의 shared_ptr로 세션을 붙잡고 있으므로, 연결이 끊길 때 leave로 빼 주지 않으면 세션이 영원히 해제되지 않습니다. 이 예제는 TCP 스트림에서 읽은 바이트를 그대로 메시지로 취급하는데, 실제 프로토콜에서는 줄바꿈이나 길이 접두사로 메시지 경계를 나눠야 합니다.
줄 단위 요청-응답 서버
줄 하나를 요청으로 받아 응답을 돌려주는 단순한 프로토콜은 async_read_until로 구현합니다. 요청과 응답이 번갈아 하나씩만 걸리므로 쓰기 큐가 필요 없습니다.
class LineSession : public std::enable_shared_from_this<LineSession> {
tcp::socket socket_;
boost::asio::streambuf buffer_;
std::string response_;
public:
explicit LineSession(tcp::socket socket) : socket_(std::move(socket)) {}
void start() { do_read_line(); }
private:
void do_read_line() {
boost::asio::async_read_until(socket_, buffer_, '\n',
[self = shared_from_this()](error_code ec, std::size_t) {
if (ec) return;
std::istream is(&self->buffer_);
std::string line;
std::getline(is, line);
if (!line.empty() && line.back() == '\r') line.pop_back();
self->response_ = self->process(line) + "\n";
self->do_write();
});
}
std::string process(const std::string& req) {
if (req == "PING") return "PONG";
if (req == "STATUS") return "OK";
return "UNKNOWN: " + req;
}
void do_write() {
boost::asio::async_write(socket_, boost::asio::buffer(response_),
[self = shared_from_this()](error_code ec, std::size_t) {
if (!ec) self->do_read_line();
});
}
};
async_read_until은 구분자 뒤의 데이터까지 streambuf에 미리 읽어 둘 수 있습니다. 그래서 getline으로 한 줄만 꺼내고 나머지는 버퍼에 남겨 두어야 다음 요청이 사라지지 않습니다.
스레드마다 io_context를 두는 방식
strand 대신, 스레드 수만큼 io_context를 만들고 각자 한 스레드에서만 run()하게 한 뒤, 새 연결을 이 중 하나에 돌아가며 배정하는 방식도 있습니다. 한 연결의 모든 핸들러가 항상 같은 스레드에서 실행되므로 연결별 strand가 필요 없고 스레드 간 경합도 적습니다.
#include <atomic>
#include <memory>
#include <thread>
#include <vector>
class IoContextPool {
std::vector<std::unique_ptr<boost::asio::io_context>> contexts_;
std::vector<boost::asio::executor_work_guard<boost::asio::io_context::executor_type>> guards_;
std::vector<std::thread> threads_;
std::atomic<std::size_t> next_{0};
public:
explicit IoContextPool(std::size_t n) {
for (std::size_t i = 0; i < n; ++i) {
contexts_.push_back(std::make_unique<boost::asio::io_context>(1)); // 단일 스레드 힌트
guards_.push_back(boost::asio::make_work_guard(*contexts_.back()));
}
for (auto& c : contexts_) threads_.emplace_back([ctx = c.get()]() { ctx->run(); });
}
boost::asio::io_context& next() { return *contexts_[next_++ % contexts_.size()]; }
void stop() {
guards_.clear();
for (auto& c : contexts_) c->stop();
for (auto& t : threads_) if (t.joinable()) t.join();
}
};
// accept: acceptor_.async_accept(pool.next(), handler);
단점은 부하 분산이 accept 시점의 배정으로 끝난다는 것입니다. 특정 연결이 무거우면 그 연결이 배정된 스레드만 바쁘고 다른 스레드는 놀 수 있습니다. 또 여러 연결이 공유하는 상태는 이 방식에서도 여전히 뮤텍스 등으로 보호해야 합니다. 단일 io_context + strand 방식은 핸들러 단위로 부하가 고르게 퍼지는 대신 스레드 간 큐 경합이 있습니다. 어느 쪽이 나은지는 연결 수와 핸들러 비용에 따라 달라지므로 측정으로 정합니다.
자주 만나는 문제
같은 소켓에 쓰기를 겹쳐 호출
async_write는 내부적으로 async_write_some을 여러 번 호출해 데이터를 나눠 보낼 수 있는 합성 연산입니다. 앞의 쓰기가 끝나기 전에 같은 소켓에 다음 async_write를 시작하면 두 데이터의 조각이 섞여 전송될 수 있습니다. 앞의 채팅 예제처럼 쓰기 큐를 두고 완료 핸들러에서 다음 쓰기를 시작해야 합니다.
세션을 지역 변수로 만들기
// 잘못된 예
void on_accept(tcp::socket socket) {
Session session(std::move(socket));
session.start(); // 비동기 읽기 등록
} // 여기서 session 소멸 → 나중에 실행되는 핸들러가 해제된 객체를 건드림
// 올바른 예
void on_accept(tcp::socket socket) {
std::make_shared<Session>(std::move(socket))->start(); // 핸들러가 shared_ptr를 붙잡음
}
strand 핸들러 안에서 기다리기
post는 핸들러를 큐에 넣고 바로 반환하므로, 뮤텍스를 잡은 채로 post해도 그 자체로 데드락이 생기지는 않습니다. 위험한 것은 핸들러 안에서 무언가를 블로킹으로 기다리는 경우입니다. 예를 들어 strand 핸들러 안에서 같은 strand에 post한 작업의 future.get()을 기다리면, 그 작업은 지금 핸들러가 끝나야 실행되므로 영원히 끝나지 않습니다. 또 dispatch는 이미 그 executor 안에 있으면 핸들러를 즉시 그 자리에서 실행하므로, 뮤텍스를 잡은 채 dispatch한 핸들러가 같은 뮤텍스를 잡으려 하면 데드락이 됩니다. 핸들러 안에서는 기다리지 말고, 다음 일을 다시 비동기로 이어 붙이는 것이 원칙입니다.
스레드를 너무 많이 만들기
Asio의 비동기 I/O는 스레드를 막지 않으므로, 핸들러가 짧다면 CPU 코어 수 이상의 스레드는 대개 이득이 없고 컨텍스트 스위칭만 늘어납니다. 핸들러 안에서 블로킹 DB 호출 같은 작업을 해야 한다면 스레드를 늘리기보다 그 작업을 별도 스레드 풀로 넘기고 결과를 post로 돌려받는 편이 낫습니다.
핸들러에서 예외가 빠져나감
핸들러에서 던진 예외는 그 핸들러를 실행한 run() 밖으로 전파됩니다. 스레드 함수가 이 예외를 잡지 않으면 std::terminate로 프로세스 전체가 종료됩니다. 핸들러 안에서 파싱처럼 예외가 날 수 있는 코드는 try/catch로 감싸 해당 연결만 정리하거나, 스레드 함수에서 run()을 루프로 감싸 예외를 로깅한 뒤 다시 run()을 호출하는 방식을 씁니다.
쓰기 버퍼의 수명
// 잘못된 예: msg는 함수가 끝나면 사라지지만 쓰기는 계속 진행 중
void send(const std::string& msg) {
boost::asio::async_write(socket_, boost::asio::buffer(msg), handler);
}
boost::asio::buffer는 데이터를 복사하지 않고 가리키기만 합니다. 버퍼를 세션 멤버(쓰기 큐 등)에 두거나, shared_ptr<std::string>을 핸들러에 캡처해 쓰기가 끝날 때까지 살려 둬야 합니다.
성능 측정
절대 수치는 CPU, 커널 설정, 메시지 크기, 클라이언트 위치에 따라 몇 배씩 달라지므로 자기 환경에서 직접 재야 합니다. 단일 스레드 서버는 연결이 적을 때 지연이 가장 낮을 수 있지만 한 코어가 포화되면 처리량이 더 늘지 않습니다. 여러 스레드 + strand 구성은 코어 수까지는 처리량이 늘다가, 연결이 아주 많아지면 스레드 간 작업 분배 비용 때문에 증가 폭이 줄어듭니다.
wrk는 HTTP 부하 도구라 HTTP를 말하지 않는 에코 서버에는 쓸 수 없습니다. 에코 서버는 tcpkali 같은 TCP 부하 도구를 쓰거나 간단한 클라이언트를 직접 만듭니다.
// 동기식 에코 클라이언트: 연결 하나로 왕복 횟수를 잼 (여러 개를 병렬로 띄워 사용)
void echo_client(const char* host, unsigned short port, int requests) {
boost::asio::io_context io;
tcp::socket socket(io);
socket.connect(tcp::endpoint(boost::asio::ip::make_address(host), port));
socket.set_option(tcp::no_delay(true));
const std::string msg = "ping\n";
std::array<char, 64> buf;
auto start = std::chrono::steady_clock::now();
for (int i = 0; i < requests; ++i) {
boost::asio::write(socket, boost::asio::buffer(msg));
boost::asio::read(socket, boost::asio::buffer(buf.data(), msg.size())); // 보낸 만큼만 읽음
}
double ms = std::chrono::duration<double, std::milli>(
std::chrono::steady_clock::now() - start).count();
std::cout << "round trips/s: " << requests * 1000.0 / ms << "\n";
}
boost::asio::read는 버퍼가 가득 찰 때까지 읽으므로, 버퍼 전체 크기를 넘기면 서버가 보낸 5바이트 이후 영원히 기다리게 됩니다. 보낸 크기만큼만 읽도록 버퍼를 잘라 넘깁니다. 측정할 때는 클라이언트가 서버와 같은 코어를 두고 경쟁하지 않도록 다른 머신이나 분리된 코어에서 돌리고, 평균보다 p99 지연을 함께 봅니다. 작은 메시지를 주고받는 프로토콜이라면 tcp::no_delay(true)로 Nagle 알고리즘을 끄는 것이 지연에 큰 영향을 줍니다.
운영 패턴
연결 수 제한
class ConnectionLimiter {
std::atomic<std::size_t> count_{0};
std::size_t max_;
public:
explicit ConnectionLimiter(std::size_t max) : max_(max) {}
bool try_acquire() {
std::size_t c = count_.load(std::memory_order_relaxed);
while (c < max_ && !count_.compare_exchange_weak(c, c + 1)) {}
return c < max_;
}
void release() { count_.fetch_sub(1, std::memory_order_relaxed); }
};
// accept 핸들러: if (!limiter.try_acquire()) { socket.close(); } else { ... }
// 세션 소멸자에서 limiter.release()
compare_exchange_weak가 실패하면 c에 현재 값이 다시 채워지므로, 루프는 한도에 닿거나 증가에 성공할 때까지 반복합니다.
헬스체크 엔드포인트
로드밸런서용 헬스체크는 별도 포트에서 고정 응답을 돌려주면 됩니다. acceptor와 응답 버퍼가 비동기 연산이 끝날 때까지 살아 있어야 하므로 shared_ptr로 관리합니다.
void start_health_check(std::shared_ptr<tcp::acceptor> acceptor) {
acceptor->async_accept([acceptor](error_code ec, tcp::socket socket) {
if (ec == boost::asio::error::operation_aborted) return;
if (!ec) {
auto sock = std::make_shared<tcp::socket>(std::move(socket));
static const std::string reply =
"HTTP/1.1 200 OK\r\nContent-Length: 2\r\nConnection: close\r\n\r\nOK";
boost::asio::async_write(*sock, boost::asio::buffer(reply),
[sock](error_code, std::size_t) {}); // sock이 쓰기 완료까지 소켓을 살림
}
start_health_check(acceptor);
});
}
// 사용: start_health_check(std::make_shared<tcp::acceptor>(io, tcp::endpoint(tcp::v4(), 8081)));
메트릭
활성 연결 수, 누적 요청 수, 에러 수 정도는 std::atomic 카운터로 모아 두고 주기적으로 내보내면 됩니다. 세션 생성자와 소멸자에서 활성 연결 수를 올리고 내리면 수명 버그(세션이 해제되지 않는 누수)도 이 지표로 드러납니다.
struct ServerMetrics {
std::atomic<std::uint64_t> total_connections{0};
std::atomic<std::uint64_t> active_connections{0};
std::atomic<std::uint64_t> total_errors{0};
};
이전 글: C++ 실전 가이드 #29-2: 비동기 이벤트 루프 다음 글: [C++ 실전 가이드 #30-1] WebSocket 구현: 핸드셰이크와 프레임