C++에서 split·join·trim·replace 구현하기: getline·string_view·std::regex 비교

표준에는 split 함수가 없습니다. 구분자가 하나면 getline, 복사를 피하고 싶으면 string_view와 find를 씁니다. 반복문 안에서 stringstream이나 regex를 매번 새로 만들지 않도록 주의하세요. 11-1 문자열 기초를 먼저 보면 연결이 잘 됩니다.

들어가며: “split 함수가 없어요”

Python에서는 "a,b,c".split(",") 한 줄로 문자열을 나눌 수 있지만, C++ 표준 라이브러리의 std::string에는 split, join, trim이 없습니다. 제공되는 것은 find, substr, replace, erase 같은 저수준 연산뿐이고, 고수준 알고리즘은 직접 조합하거나 Boost.StringAlgo, Abseil 같은 라이브러리를 써야 합니다.

직접 구현하다 보면 비슷한 곳에서 걸립니다. CSV 한 줄을 나누다가 끝의 빈 필드가 사라지고, 공백뿐인 입력을 trim하다가 std::out_of_range 예외가 나고, string::replace가 첫 번째 매칭만 바꾼다는 사실을 늦게 알아차리고, 이메일 검증 루프 안에서 std::regex를 매번 생성해서 느려집니다. 이 글에서는 split, join, trim, replace를 각각 몇 가지 방식으로 구현하면서 이런 함정이 어디서 생기는지 짚고, 정규식은 언제 쓰고 어떻게 재사용하는지 정리합니다.


split: getline, find, string_view, 정규식

stringstream과 getline (구분자 1개)

,나 | 같은 단일 문자 구분자로 나눌 때 가장 흔히 쓰는 방법입니다.

#include <sstream>
#include <string>
#include <vector>

std::vector<std::string> split(const std::string& str, char delimiter) {
    std::vector<std::string> result;
    std::istringstream iss(str);
    std::string token;
    while (std::getline(iss, token, delimiter)) {
        result.push_back(token);
    }
    return result;
}

// split("Alice,25,Engineer", ',') → {"Alice", "25", "Engineer"}

getline(iss, token, delimiter)는 구분자 직전까지를 token에 넣고 구분자는 버립니다. 중간의 빈 토큰은 유지되어 "a,,b"는 {"a", "", "b"}가 됩니다. 하지만 마지막 구분자 뒤의 빈 토큰은 사라집니다. "a,b,"는 {"a", "b"}가 되는데, 마지막 getline이 아무것도 읽지 못한 채 스트림 끝에 도달해 실패하기 때문입니다. CSV에서 마지막 열이 비어 있을 수 있다면 이 차이 때문에 열 개수가 줄어듭니다. 빈 문자열도 {}가 되며, Python의 "".split(",")이 ['']를 돌려주는 것과 다릅니다.

find와 substr (스트림 없이)

스트림 객체 생성과 로캘 처리 비용을 피하고, 빈 토큰을 어떻게 다룰지 직접 정하고 싶을 때 씁니다.

#include <string>
#include <vector>

// 빈 토큰을 모두 유지: "a,b," → {"a", "b", ""}, "" → {""}
std::vector<std::string> split_find(const std::string& str, char delimiter) {
    std::vector<std::string> result;
    std::size_t start = 0;
    while (true) {
        std::size_t pos = str.find(delimiter, start);
        if (pos == std::string::npos) {
            result.push_back(str.substr(start));
            break;
        }
        result.push_back(str.substr(start, pos - start));
        start = pos + 1;
    }
    return result;
}

루프 조건을 while (start < str.size())로 쓰면 끝의 빈 토큰이 빠지므로, 위처럼 구분자를 더 찾을 수 없을 때 남은 부분을 넣고 끝내는 구조가 Python의 동작과 같습니다. substr은 토큰마다 새 std::string을 만들지만, 짧은 문자열은 대부분의 구현에서 SSO(Small String Optimization) 덕분에 힙 할당 없이 객체 안에 저장됩니다.

string_view로 복사 없이 split (C++17)

토큰을 잠깐 검사만 하고 버린다면 복사할 필요가 없습니다. string_view::substr은 새 문자열을 만들지 않고 원본의 구간만 가리킵니다.

#include <string_view>
#include <vector>

std::vector<std::string_view> split_sv(std::string_view str, char delimiter) {
    std::vector<std::string_view> result;
    std::size_t start = 0;
    while (true) {
        std::size_t pos = str.find(delimiter, start);
        if (pos == std::string_view::npos) {
            result.push_back(str.substr(start));
            break;
        }
        result.push_back(str.substr(start, pos - start));
        start = pos + 1;
    }
    return result;
}

대가는 수명 관리입니다. 결과 벡터는 원본 문자열이 살아 있고 변경되지 않는 동안에만 유효합니다.

std::vector<std::string_view> get_tokens() {
    std::string line = read_line();
    return split_sv(line, ',');   // line이 소멸하면서 모든 뷰가 댕글링
}

이 코드는 컴파일도 되고 경고도 없는 경우가 많아서 더 위험합니다. 반환된 뷰를 읽으면 해제된 메모리를 읽게 되고, 테스트에서는 그 메모리가 아직 덮어써지지 않아 우연히 맞는 값이 나오기도 합니다. 결과를 원본보다 오래 보관해야 한다면 std::string으로 복사해서 반환해야 합니다.

공백류로 나누기

공백, 탭, 개행이 섞인 입력을 단어 단위로 나눌 때는 >> 연산자가 가장 간단합니다.

std::vector<std::string> split_whitespace(const std::string& str) {
    std::vector<std::string> result;
    std::istringstream iss(str);
    std::string token;
    while (iss >> token) result.push_back(token);
    return result;
}
// split_whitespace("  a b\tc\n") → {"a", "b", "c"}

>>는 연속된 공백을 하나로 취급하고 앞뒤 공백도 건너뛰므로 빈 토큰이 생기지 않습니다. 쉼표처럼 공백이 아닌 구분자는 지정할 수 없습니다.

정규식으로 split

\s+나 [,;]처럼 구분자가 패턴일 때는 std::sregex_token_iterator를 씁니다.

#include <regex>

std::vector<std::string> split_regex(const std::string& str, const std::regex& re) {
    return {std::sregex_token_iterator(str.begin(), str.end(), re, -1),
            std::sregex_token_iterator()};
}

// static const std::regex ws(R"(\s+)");
// split_regex("a  b   c", ws) → {"a", "b", "c"}

네 번째 인자 -1은 “매칭된 부분이 아니라 매칭 사이의 부분”을 토큰으로 달라는 뜻입니다. 정규식을 인자로 받도록 한 이유는 호출할 때마다 패턴을 컴파일하지 않게 하기 위해서입니다. 문자열 앞에 구분자가 있으면(" a b") 첫 토큰으로 빈 문자열이 나옵니다.

C++20 std::views::split

C++20부터는 범위 라이브러리의 std::views::split으로 지연 평가 방식의 분할을 할 수 있습니다. 초기 C++20 명세에서는 결과 요소가 연속 범위로 취급되지 않아 string_view로 바꾸기 불편했는데, P2210이 결함 보고(DR)로 반영되면서 GCC 12, 최신 MSVC 등에서는 각 요소로 std::string_view를 바로 만들 수 있습니다. 예전 동작은 std::views::lazy_split으로 이름이 바뀌었습니다. 컴파일러 버전에 따라 동작이 다르므로, 여러 환경을 지원해야 한다면 위의 find 기반 구현이 여전히 무난합니다.


join: 이어 붙이기

ostringstream

#include <sstream>

template <typename Container>
std::string join(const Container& items, const std::string& delimiter) {
    std::ostringstream oss;
    bool first = true;
    for (const auto& item : items) {
        if (!first) oss << delimiter;
        oss << item;
        first = false;
    }
    return oss.str();
}
// join(std::vector<std::string>{"Hello", "World", "C++"}, ", ") → "Hello, World, C++"

<<를 쓰므로 std::vector<int>처럼 문자열이 아닌 요소도 그대로 이어 붙일 수 있다는 것이 장점입니다. 속도 면에서는 스트림의 서식 처리 비용이 있어서, 문자열만 이어 붙이는 경우 아래의 reserve 방식보다 느린 것이 보통입니다.

std::accumulate를 쓰면 안 되는 이유

#include <numeric>

std::string join_accumulate(const std::vector<std::string>& items,
                            const std::string& delimiter) {
    if (items.empty()) return "";
    return std::accumulate(std::next(items.begin()), items.end(), items[0],
        [&delimiter](const std::string& a, const std::string& b) {
            return a + delimiter + b;
        });
}

한 줄로 끝나서 보기엔 깔끔하지만, 람다가 매번 지금까지의 결과 a 전체를 복사한 새 문자열을 만듭니다. 결과 길이가 L이고 항목이 n개면 복사량이 대략 O(nL)로 늘어나, 항목이 많을수록 급격히 느려집니다. C++20부터 accumulate가 누적값을 std::move로 넘기긴 하지만, 람다가 const std::string&로 받는 한 복사는 그대로 일어납니다.

reserve 후 +=

std::string join_reserve(const std::vector<std::string>& items,
                         const std::string& delimiter) {
    if (items.empty()) return "";
    std::size_t total = delimiter.size() * (items.size() - 1);
    for (const auto& s : items) total += s.size();

    std::string result;
    result.reserve(total);
    result += items[0];
    for (std::size_t i = 1; i < items.size(); ++i) {
        result += delimiter;
        result += items[i];
    }
    return result;
}

최종 길이를 먼저 계산해 한 번만 할당하므로 재할당이 전혀 없습니다. 문자열 벡터를 합치는 경우라면 이 방식이 가장 예측 가능한 성능을 냅니다.


trim: 앞뒤 공백 제거

#include <string>

std::string trim(const std::string& str, const char* chars = " \t\n\r\f\v") {
    std::size_t first = str.find_first_not_of(chars);
    if (first == std::string::npos) return "";
    std::size_t last = str.find_last_not_of(chars);
    return str.substr(first, last - first + 1);
}

std::string ltrim(const std::string& str, const char* chars = " \t\n\r\f\v") {
    std::size_t first = str.find_first_not_of(chars);
    return first == std::string::npos ? "" : str.substr(first);
}

std::string rtrim(const std::string& str, const char* chars = " \t\n\r\f\v") {
    std::size_t last = str.find_last_not_of(chars);
    return last == std::string::npos ? "" : str.substr(0, last + 1);
}

// trim("  hello  \n") → "hello"
// trim(",,,hello,,,", ",") → "hello"

npos 검사를 빠뜨리면 문제가 생깁니다. 문자열이 비어 있거나 전부 공백이면 first와 last가 모두 npos이고, last - first + 1은 1이 됩니다. 그러면 str.substr(npos, 1)을 호출하게 되는데, 시작 위치가 문자열 길이보다 크므로 std::out_of_range 예외가 던져집니다. 사용자 입력을 trim하는 코드에서 빈 입력으로 테스트하지 않으면 운영 중에야 드러나는 전형적인 버그입니다. first가 npos가 아니면 공백이 아닌 문자가 적어도 하나 있다는 뜻이므로 last도 반드시 유효합니다.

복사가 필요 없다면 string_view 버전이 더 가볍습니다.

#include <string_view>

std::string_view trim_sv(std::string_view str, std::string_view chars = " \t\n\r\f\v") {
    std::size_t first = str.find_first_not_of(chars);
    if (first == std::string_view::npos) return {};
    std::size_t last = str.find_last_not_of(chars);
    return str.substr(first, last - first + 1);
}

replace: 첫 매칭, 전체 매칭, 정규식 치환

std::string::replace(pos, len, str)는 지정한 위치의 구간 하나를 바꾸는 함수입니다. 이름만 보고 Python의 str.replace처럼 모든 매칭을 바꾼다고 생각하기 쉽지만, 위치를 찾는 일도 직접 해야 하고 한 번에 하나만 바꿉니다.

std::string s = "foo bar foo";
std::size_t pos = s.find("foo");
if (pos != std::string::npos) s.replace(pos, 3, "bar");
// s: "bar bar foo"

모든 매칭을 바꾸려면 find 루프를 돌립니다.

void replace_all(std::string& str, const std::string& from, const std::string& to) {
    if (from.empty()) return;          // 빈 패턴은 모든 위치에 매칭되어 끝나지 않는다
    std::size_t pos = 0;
    while ((pos = str.find(from, pos)) != std::string::npos) {
        str.replace(pos, from.size(), to);
        pos += to.size();              // 방금 넣은 문자열은 다시 검사하지 않는다
    }
}
// replace_all(s = "foofoofoo", "foo", "bar") → "barbarbar"
// replace_all(s = "aaa", "a", "aa") → "aaaaaa"

검색을 재개할 위치를 pos += to.size()로 옮기는 부분이 핵심입니다. pos를 그대로 두거나 pos + 1로 옮기면, "a"를 "aa"로 바꿀 때 방금 넣은 "aa" 안의 "a"를 또 찾아 끝없이 늘어납니다. 위 구현은 삽입한 문자열 뒤에서 검색을 이어가므로 to가 from을 포함해도 무한 루프에 빠지지 않습니다. 진짜 위험한 경우는 from이 빈 문자열일 때입니다. find("")는 어느 위치에서든 매칭되므로 루프가 끝나지 않습니다.

이 구현은 매칭마다 replace가 뒤쪽 문자를 밀거나 당기므로, 매칭이 아주 많은 긴 문자열에서는 O(n·매칭 수)가 됩니다. 그런 경우에는 새 문자열을 만들어 매칭 사이의 구간과 to를 차례로 이어 붙이는 방식이 O(n)입니다.

std::regex_replace

패턴 기반 치환은 std::regex_replace가 기본적으로 모든 매칭을 바꿉니다.

#include <regex>

std::string text = "Price: $100, Tax: $10";
static const std::regex money(R"(\$\d+)");
std::string redacted = std::regex_replace(text, money, "[REDACTED]");
// "Price: [REDACTED], Tax: [REDACTED]"

std::string email = "[email protected]";
static const std::regex email_mask(R"((.{2}).*@(.*))");
std::string masked = std::regex_replace(email, email_mask, "$1***@$2");
// "us***@example.com"

치환 문자열에서 $1, $2는 캡처 그룹, $&는 매칭 전체를 뜻합니다. 반대로 말하면 치환 문자열에 사용자 데이터를 넣을 때 $가 포함되어 있으면 의도치 않게 해석됩니다. 문자 그대로의 $는 $$로 써야 합니다.


regex_match, regex_search와 정규식 재사용

regex_match는 문자열 전체가 패턴과 일치해야 참이고, regex_search는 문자열 어딘가에 패턴이 있으면 참입니다. 입력 검증에는 앞의 것을, 로그에서 정보를 뽑아낼 때는 뒤의 것을 주로 씁니다.

#include <iostream>
#include <regex>

int main() {
    static const std::regex email_re(R"([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})");
    std::cout << std::regex_match("[email protected]", email_re) << "\n";   // 1

    std::string log = "2026-03-10 14:30:00 [ERROR] Connection failed";
    static const std::regex level_re(R"(\[(\w+)\]\s+(.+))");
    std::smatch m;
    if (std::regex_search(log, m, level_re)) {
        std::cout << m[1].str() << "\n";   // ERROR
        std::cout << m[2].str() << "\n";   // Connection failed
    }
}

m[0]은 매칭 전체, m[1]부터는 괄호로 묶은 캡처 그룹입니다. std::smatch는 원본 문자열의 반복자를 들고 있으므로, 임시 문자열에 대해 regex_search를 호출하는 오버로드는 C++14부터 삭제되어 컴파일 에러가 납니다. 매칭 결과를 쓰는 동안 원본이 살아 있어야 한다는 점은 string_view와 같습니다.

std::regex 객체를 만들 때마다 패턴을 파싱해 내부 오토마톤을 구성하므로, 루프 안에서 생성하면 이 비용을 반복마다 치릅니다.

// 느림: 반복마다 정규식 생성
for (const auto& email : emails) {
    std::regex re(R"([a-z]+@[a-z]+\.[a-z]+)");
    if (std::regex_match(email, re)) { /* ... */ }
}

// 한 번만 생성해 재사용
static const std::regex email_re(R"([a-z]+@[a-z]+\.[a-z]+)");
for (const auto& email : emails) {
    if (std::regex_match(email, email_re)) { /* ... */ }
}

함수 안의 static const 지역 변수는 C++11부터 초기화가 스레드 안전하게 한 번만 일어나고, 이후 regex_match처럼 const 객체를 읽기만 하는 호출은 여러 스레드에서 동시에 해도 됩니다. 패턴이 런타임에 정해진다면 캐시를 둘 수 있는데, 아래 구현은 뮤텍스가 없으므로 스레드 하나에서만 써야 합니다.

#include <unordered_map>

class RegexCache {
public:
    const std::regex& get(const std::string& pattern) {
        auto it = cache_.find(pattern);
        if (it == cache_.end()) it = cache_.emplace(pattern, std::regex(pattern)).first;
        return it->second;   // unordered_map의 요소 참조는 재해시 후에도 유효
    }
private:
    std::unordered_map<std::string, std::regex> cache_;
};

정규식 재사용과 별개로, 표준 std::regex 구현은 전반적으로 느린 편이라는 평가가 많습니다. libstdc++ 구현은 역추적 방식이라 긴 입력에서 재귀가 깊어져 스택 오버플로가 난 사례도 보고되어 있습니다. 단순 구분자 분할처럼 find로 되는 일에는 정규식을 쓰지 않고, 정규식 처리량이 중요한 곳에서는 RE2, PCRE2, 또는 컴파일 타임 정규식 라이브러리(CTRE)를 검토하는 편이 좋습니다.


예제: CSV, 로그, URL 쿼리, 템플릿

숫자 CSV 한 줄

#include <charconv>
#include <optional>
#include <string_view>
#include <vector>

std::optional<int> to_int(std::string_view s) {
    int value{};
    auto [ptr, ec] = std::from_chars(s.data(), s.data() + s.size(), value);
    if (ec != std::errc{} || ptr != s.data() + s.size()) return std::nullopt;
    return value;
}

std::optional<std::vector<int>> parse_int_csv(std::string_view line) {
    std::vector<int> out;
    for (std::string_view field : split_sv(line, ',')) {
        auto v = to_int(trim_sv(field));
        if (!v) return std::nullopt;       // 숫자가 아닌 필드가 있으면 실패
        out.push_back(*v);
    }
    return out;
}
// parse_int_csv(" 1 , 2 , 3 ") → {1, 2, 3}

std::stoi는 잘못된 입력에 예외를 던지고, "12abc"처럼 앞부분만 숫자여도 12를 돌려줍니다. std::from_chars는 예외도 로캘도 쓰지 않고, 반환된 ptr로 입력을 끝까지 소비했는지 확인할 수 있어 검증에 적합합니다. 반면 from_chars는 앞의 공백이나 + 부호를 받아들이지 않으므로 먼저 trim해야 합니다. 따옴표로 감싼 필드 안의 쉼표("Seoul, Korea")까지 처리해야 하는 실제 CSV라면, 이런 단순 분할로는 부족하고 상태를 가진 파서가 필요합니다.

로그 라인

#include <optional>
#include <regex>

struct LogEntry {
    std::string timestamp;
    std::string level;
    std::string message;
};

std::optional<LogEntry> parse_log(const std::string& line) {
    static const std::regex re(R"((\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}) \[(\w+)\] (.+))");
    std::smatch m;
    if (!std::regex_match(line, m, re)) return std::nullopt;
    return LogEntry{m[1].str(), m[2].str(), m[3].str()};
}

URL 쿼리 문자열

#include <map>

std::map<std::string, std::string> parse_query(std::string_view query) {
    std::map<std::string, std::string> params;
    for (std::string_view pair : split_sv(query, '&')) {
        if (pair.empty()) continue;
        std::size_t eq = pair.find('=');
        std::string_view key = pair.substr(0, eq);
        std::string_view value = eq == std::string_view::npos ? "" : pair.substr(eq + 1);
        params[std::string(key)] = std::string(value);
    }
    return params;
}
// parse_query("name=alice&age=30&flag") → {age: "30", flag: "", name: "alice"}

이 함수는 %20이나 + 같은 퍼센트 인코딩을 해제하지 않습니다. 실제 URL을 다룬다면 키와 값을 저장하기 전에 디코딩 단계가 필요합니다. 결과는 std::string으로 복사해 저장하므로 원본 query가 사라져도 안전합니다.

템플릿 문자열 치환

std::string render(std::string tmpl, const std::map<std::string, std::string>& vars) {
    for (const auto& [key, value] : vars) {
        replace_all(tmpl, "{{" + key + "}}", value);
    }
    return tmpl;
}
// render("Hello {{name}}!", {{"name", "World"}}) → "Hello World!"

이 작업에 std::regex_replace를 쓰면 두 가지 문제가 생깁니다. 키에 .이나 + 같은 정규식 메타문자가 들어 있으면 패턴으로 해석되고, 값에 $가 들어 있으면 캡처 그룹 참조로 해석됩니다. 고정 문자열을 찾아 바꾸는 일이라면 find 기반 치환이 더 빠르고 안전합니다. 다만 위 구현은 앞에서 치환한 값 안에 {{other}}가 들어 있으면 그것도 다음 키에서 치환되므로, 사용자 입력을 값으로 넣는다면 템플릿을 한 번만 훑으며 치환하는 방식으로 바꿔야 합니다.


성능 관련 정리

반복문 안에서 std::istringstream을 매번 생성하면 스트림 객체 생성과 로캘 초기화 비용을 반복해서 치릅니다. 스트림이 꼭 필요하다면 루프 밖에서 하나를 만들고 clear()로 상태 플래그를 지운 뒤 str(line)으로 내용을 바꿔 재사용할 수 있습니다. clear()를 빠뜨리면 이전 줄을 끝까지 읽으며 켜진 eof 플래그 때문에 다음 줄에서 아무것도 읽히지 않습니다.

std::istringstream iss;
for (const auto& line : lines) {
    iss.clear();
    iss.str(line);
    // 파싱...
}

std::cin >> n 다음에 std::getline을 호출하면 빈 줄이 읽히는 것도 스트림에서 자주 겪는 문제입니다. >>가 숫자 뒤의 개행을 버퍼에 남겨 두기 때문으로, std::cin.ignore(std::numeric_limits<std::streamsize>::max(), '\n')로 남은 줄을 버린 뒤 getline을 호출합니다.

그 외에는 토큰 수를 대략 알면 결과 벡터에 reserve를 하고, 토큰을 잠깐만 쓰면 string_view로 복사를 없애고, 숫자 변환은 std::from_chars를 쓰는 것이 효과가 큽니다. 어느 쪽이 병목인지는 입력에 따라 다르므로, 실제 데이터로 프로파일링해 확인하는 것이 가장 확실합니다.


자주 묻는 질문 (FAQ)

Q. stringstream과 find/substr 중 뭐가 더 빠른가요?

A. 단순 구분자 분할에서는 스트림 객체 생성과 로캘 처리가 없는 find/substr이 보통 더 빠르고, string_view로 복사까지 없애면 더 가볍습니다. stringstream은 >>로 타입 변환을 함께 할 수 있다는 편의성이 있으므로, 입력 크기가 작다면 읽기 쉬운 쪽을 고르면 됩니다.

Q. UTF-8 한글이 섞인 문자열도 이 방법으로 나눠도 되나요?

A. 구분자가 ,나 공백 같은 ASCII 문자라면 그대로 써도 안전합니다. UTF-8에서 멀티바이트 문자를 이루는 바이트는 모두 0x80 이상이라서, ASCII 구분자와 절대 겹치지 않기 때문입니다. 반면 “앞에서 세 글자”처럼 문자 단위로 자르거나 한글 문자를 구분자로 쓰려면 바이트가 아니라 코드 포인트 단위로 처리해야 하므로 ICU 같은 라이브러리를 쓰는 편이 안전합니다. trim에서 std::isspace를 쓴다면 unsigned char로 변환해서 넘겨야 음수 바이트로 인한 미정의 동작을 피할 수 있습니다.


이전 글: C++ 실전 가이드 #11-1: 파일 I/O 기초 다음 글: C++ 실전 가이드 #11-3: stringstream과 포맷팅


같이 보면 좋은 글