C++ 문자열 다루기: std::string 연산, C 문자열 변환, string_view와 SSO 성능

💡 핵심 개념: 두 const char*를 ==로 비교하면 내용이 아니라 주소를 비교합니다. 내용 비교는 std::string의 ==나 strcmp로 하고, 읽기만 하는 함수 인자에는 string_view를 먼저 고려합니다. 11-1 파일 I/O와 순서를 바꿔 읽어도 됩니다.

들어가며: 문자열 비교에 ==를 썼는데 결과가 이상할 때

C 스타일 문자열(char* 또는 const char*)을 사용하는 레거시 API와 연동할 때, ==로 비교하면 포인터 주소가 비교됩니다. 문자열 내용이 아니라 “같은 메모리 주소를 가리키는지”만 확인하게 됩니다.

자주 만나는 C 문자열 함정

strcmp는 “같으면 0, 작으면 음수, 크면 양수”를 반환하는데, if (strcmp(a, b))처럼 쓰면 “같을 때” false가 됩니다. C++에서 if (x)는 x가 0이 아니면 true이므로, strcmp가 0(같음)을 반환할 때 false가 되어 의도와 반대입니다.

for 루프 안에서 result += str를 반복하면 용량이 찰 때마다 재할당과 전체 복사가 일어납니다. 용량은 배수로 늘어나므로 재할당 횟수 자체는 많지 않지만, 최종 길이를 안다면 reserve로 한 번에 잡아 두는 편이 낫습니다. result = result + str처럼 쓰면 매번 새 문자열을 만들어 훨씬 느려집니다.

std::string_view로 받은 문자열의 일부를 가리킬 때, 원본이 수정되거나 삭제되면 댕글링 참조가 됩니다. string_view는 복사하지 않고 “보기만” 하므로, 참조하는 메모리가 유효한 동안만 사용해야 합니다.

printf("%s", str)처럼 가변 인자 C 함수에 std::string을 그대로 넘기면 GCC·Clang에서는 컴파일 에러이고, 통과시키는 컴파일러에서도 정의되지 않은 동작입니다. C API는 const char*를 기대하므로 str.c_str()을 넘겨야 합니다.

std::string은 크기가 자동으로 늘어나고 메모리를 스스로 관리하는 표준 문자열 클래스이고, string_view(C++17)는 문자열을 소유하지 않고 가리키기만 하는 가벼운 뷰입니다. 처음 예로 든 포인터 비교 문제는 다음과 같습니다.

const char* a = "hello";
const char* b = "hello";
if (a == b) {  // ❌ 포인터 비교! 내용은 같아도 주소가 다를 수 있음
    std::cout << "Same\n";
}

a와 b는 서로 다른 메모리 주소를 가리킬 수 있습니다. 컴파일러가 문자열 리터럴을 “같은 주소”로 합치는 최적화(string pooling)를 할 수도 있지만, 런타임에 만든 문자열이나 다른 번역 단위에선 보장되지 않습니다. 내용 비교가 목적이면 strcmp를 써야 합니다. 올바른 비교:

#include <cstring>
// 변수 선언 및 초기화
const char* a = "hello";
const char* b = "hello";
if (strcmp(a, b) == 0) {  // ✅ 내용 비교
    std::cout << "Same\n";
}

문자열 타입 선택을 요약하면 아래와 같습니다.

flowchart TB
  subgraph choice[문자열 타입 선택]
    A[문자열이 필요할 때] --> B{용도}
    B -->|소유·수정 필요| C["std string"]
    B -->|읽기만·함수 인자| D["std string_view"]
    B -->|C API 연동| E["const char*"]
  end
  subgraph caution[주의]
    C --> F["reserve로 재할당 최소화"]
    D --> G["원본 수명 확인"]
    E --> H["strcmp로 비교"]
  end

대소문자, null 문자열, 경로 구분자에서 생기는 문제

시나리오 1: 로그 파싱에서 “user”와 “User”가 같은 것으로 처리됨

대소문자를 구분하지 않는 비교가 필요할 때, ==는 대소문자를 구분합니다. "user" == "User"는 false입니다. std::tolower로 한 글자씩 소문자로 바꿔 비교하거나, POSIX의 strcasecmp(Windows는 _stricmp)를 씁니다. 아래 방식은 ASCII 범위에서만 정확하며, 한글이나 독일어 ß 같은 문자의 대소문자 처리는 ICU 같은 라이브러리가 필요합니다.

// 복사해 붙여넣은 뒤: g++ -std=c++17 -o case_compare case_compare.cpp && ./case_compare
#include <iostream>
#include <string>
#include <algorithm>
#include <cctype>
bool equalsIgnoreCase(const std::string& a, const std::string& b) {
    if (a.size() != b.size()) return false;
    return std::equal(a.begin(), a.end(), b.begin(),
        [](char ca, char cb) {
            return std::tolower(static_cast<unsigned char>(ca)) ==
                   std::tolower(static_cast<unsigned char>(cb));
        });
}
int main() {
    std::string a = "user";
    std::string b = "User";
    std::cout << (equalsIgnoreCase(a, b) ? "Same" : "Different") << "\n";
    return 0;
}

실행 결과:

Same

시나리오 2: JSON 키에서 “null” 문자열과 실제 null 구분 실패

파싱된 JSON에서 "null"(문자열)과 null(JSON null 값)을 구분해야 할 때, 문자열 비교를 잘못하면 혼동됩니다.

std::string value = getJsonString(key);
if (value == "null") {  // JSON의 null 문자열
    // ...
}
// value가 비어 있거나 파싱 실패와 구분 필요

시나리오 3: HTTP 헤더에서 “Content-Type” vs “content-type”

HTTP 헤더는 대소문자를 구분하지 않는 경우가 많습니다. ==로 비교하면 “Content-Type”과 “content-type”이 다르게 처리됩니다.

시나리오 4: 파일 경로에서 슬래시/백슬래시 혼용

Windows 경로 "C:\\Users\\file.txt"와 Linux 스타일 "C:/Users/file.txt"를 같은 것으로 처리하려면 정규화가 필요합니다.


std::string 생성, 연결, 검색, 숫자 변환

기본 생성과 초기화

// 복사해 붙여넣은 뒤: g++ -std=c++17 -o string_basic string_basic.cpp && ./string_basic
#include <iostream>
#include <string>
int main() {
    std::string s1;                    // 빈 문자열
    std::string s2("hello");           // C 문자열로 초기화
    std::string s3 = "world";          // 복사 초기화
    std::string s4(5, 'a');            // 'a' 5개: "aaaaa"
    std::string s5(s2, 1, 3);         // s2의 1번 인덱스부터 3글자: "ell"
    std::string s6(s2.begin(), s2.end());  // 반복자 범위
    std::cout << s1 << "|" << s2 << "|" << s4 << "\n";
    return 0;
}

실행 결과:

|hello|aaaaa

s1은 빈 문자열, s2는 C 문자열 리터럴로 초기화, s4는 (개수, 문자)로 반복 문자, s5는 (문자열, 시작위치, 길이)로 부분 문자열을 만듭니다.

주요 연산: 연결, 추가, 삽입

#include <string>
#include <iostream>
int main() {
    std::string a = "Hello";
    std::string b = "World";
    // 연결: + 연산자 (새 문자열 반환)
    std::string c = a + " " + b;       // "Hello World"
    std::string d = a + std::string("!");  // "Hello!"
    // 추가: += (기존 문자열 수정)
    a += " ";      // a = "Hello "
    a += b;        // a = "Hello World"
    // append
    std::string e = "Hi";
    e.append(" there");     // "Hi there"
    e.append(3, '!');       // "Hi there!!!"
    // push_back: 문자 하나 추가
    e.push_back('?');       // "Hi there!!!?"
    std::cout << c << "\n" << a << "\n" << e << "\n";
    return 0;
}

실행 결과:

Hello World
Hello World
Hi there!!!?

주의: + 연산자는 새 문자열을 반환하므로 임시 객체가 생성됩니다. 루프 안에서 result = result + piece를 반복하면 매번 새 문자열이 만들어져 비효율적입니다. += 또는 append를 사용하면 기존 버퍼에 추가합니다.

검색과 치환

#include <string>
#include <iostream>
int main() {
    std::string s = "Hello World, Hello C++";
    // find: 처음 발견 위치 (없으면 npos)
    size_t pos = s.find("Hello");
    std::cout << "First 'Hello' at: " << pos << "\n";  // 0
    pos = s.find("Hello", 1);  // 1번 인덱스부터 검색
    std::cout << "Second 'Hello' at: " << pos << "\n";  // 13
    // rfind: 끝에서부터 역방향 검색 (마지막 발견 위치)
    pos = s.rfind("Hello");
    std::cout << "Last 'Hello' at: " << pos << "\n";  // 13
    // find_first_of: 문자 집합 중 하나라도 처음
    pos = s.find_first_of("aeiou");
    std::cout << "First vowel at: " << pos << "\n";  // 1 (e)
    // replace: 치환
    s.replace(0, 5, "Hi");  // 0번부터 5글자를 "Hi"로
    std::cout << s << "\n";  // "Hi World, Hello C++"
    return 0;
}

실행 결과:

First 'Hello' at: 0
Second 'Hello' at: 13
Last 'Hello' at: 13
First vowel at: 1
Hi World, Hello C++

substr: 부분 문자열 추출

#include <string>
#include <iostream>
int main() {
    std::string s = "Hello World";
    // substr(시작): 시작부터 끝까지
    std::string sub1 = s.substr(6);   // "World"
    // substr(시작, 길이): 시작부터 지정 길이
    std::string sub2 = s.substr(0, 5);  // "Hello"
    // substr은 새 문자열을 반환 (복사)
    std::cout << sub1 << " " << sub2 << "\n";
    return 0;
}

실행 결과:

World Hello

비교 연산

#include <string>
#include <iostream>
int main() {
    std::string a = "apple";
    std::string b = "banana";
    // ==, !=, <, <=, >, >= 모두 지원
    std::cout << (a == b) << "\n";   // 0 (false)
    std::cout << (a < b) << "\n";    // 1 (true, 사전순)
    // compare: 0=같음, <0=a가 작음, >0=a가 큼
    int cmp = a.compare(b);
    std::cout << "compare: " << cmp << "\n";
    // C 문자열과도 비교 가능
    std::cout << (a == "apple") << "\n";  // 1 (true)
    return 0;
}

실행 결과:

0
1
compare: -1
1

compare는 음수·0·양수만 보장하므로 -1이 아닌 다른 음수가 나올 수도 있습니다. 결과는 부호로만 판단합니다.

숫자 변환 (C++11)

#include <string>
#include <iostream>
int main() {
    // 문자열 → 숫자
    std::string s1 = "42";
    int i = std::stoi(s1);
    std::string s2 = "3.14";
    double d = std::stod(s2);
    // 숫자 → 문자열 (C++11)
    std::string s3 = std::to_string(42);
    std::string s4 = std::to_string(3.14);
    std::cout << i << " " << d << " " << s3 << " " << s4 << "\n";
    return 0;
}

실행 결과:

42 3.14 42 3.140000

주의: stoi는 변환 실패 시 std::invalid_argument 또는 std::out_of_range 예외를 던집니다. 사용자 입력처럼 잘못된 값이 자주 들어오는 경로에서는 std::from_chars(C++17) 또는 strtol+에러 검사가 안전합니다.


strcmp, strncmp와 C 문자열 변환

strcmp 사용법

// 복사해 붙여넣은 뒤: g++ -std=c++17 -o strcmp_demo strcmp_demo.cpp && ./strcmp_demo
#include <iostream>
#include <cstring>
int main() {
    const char* a = "apple";
    const char* b = "banana";
    const char* c = "apple";
    // strcmp(a, b): a < b 이면 음수, a == b 이면 0, a > b 이면 양수
    std::cout << "strcmp(a,b): " << strcmp(a, b) << "\n";   // 음수 (정확한 값은 구현마다 다름)
    std::cout << "strcmp(a,c): " << strcmp(a, c) << "\n";   // 0
    // 올바른 비교: == 0일 때 "같음"
    if (strcmp(a, c) == 0) {
        std::cout << "a and c are equal\n";
    }
    // ❌ 잘못된 사용: strcmp(a,c)가 0이면 false
    // if (strcmp(a, c)) { ... }  // "다를 때" 실행됨!
    return 0;
}

실행 결과:

strcmp(a,b): -1
strcmp(a,c): 0
a and c are equal

strncmp: 길이 제한 비교

#include <cstring>
#include <iostream>
int main() {
    const char* a = "apple";
    const char* b = "application";
    // 앞 3글자만 비교
    if (strncmp(a, b, 3) == 0) {
        std::cout << "First 3 chars match\n";
    }
    // strncmp는 null 종료 전까지만 비교
    // strncmp("app", "apple", 5) → "app"은 3글자만 있으므로 3글자까지만 비교
    return 0;
}

strcasecmp: 대소문자 무시 비교 (POSIX)

#include <iostream>
#if defined(_WIN32)
    #include <string.h>   // _stricmp
    #define strcasecmp _stricmp
#else
    #include <strings.h>  // strcasecmp (POSIX)
#endif
int main() {
    const char* a = "Hello";
    const char* b = "hello";
    if (strcasecmp(a, b) == 0) {
        std::cout << "Equal (case-insensitive)\n";
    }
    return 0;
}

std::string ↔ C 문자열 변환

#include <string>
#include <iostream>
#include <cstdio>
int main() {
    // std::string → const char*
    std::string s = "hello";
    const char* cstr = s.c_str();   // null 종료 보장
    // C API에 넘길 때
    printf("%s\n", s.c_str());
    // 주의: c_str() 반환값은 s가 수정되면 무효화됨
    s += " world";  // 재할당 가능
    // printf("%s", cstr);  // ❌ 댕글링 가능
    // C 문자열 → std::string
    const char* input = "from C";
    std::string s2(input);
    // std::string은 중간에 null 문자를 포함할 수 있음 (c_str()로 넘기면 그 앞에서 잘림)
    std::string s3 = "hello";
    s3 += '\0';
    s3 += "world";
    std::cout << s3.size() << "\n";  // 11 (null 포함)
    return 0;
}

복사 없이 문자열을 보는 std::string_view

string_view란?

std::string_view는 문자열을 소유하지 않고 “보기만” 하는 경량 타입입니다. 복사 없이 std::string, const char*, 리터럴을 받을 수 있어, 함수 인자로 전달할 때 효율적입니다.

// 복사해 붙여넣은 뒤: g++ -std=c++17 -o string_view_basic string_view_basic.cpp && ./string_view_basic
#include <iostream>
#include <string>
#include <string_view>
void print(std::string_view sv) {
    std::cout << sv << " (size=" << sv.size() << ")\n";
}
int main() {
    std::string s = "Hello World";
    const char* cstr = "C string";
    const char* literal = "Literal";
    print(s);        // std::string 복사 없음
    print(cstr);     // C 문자열
    print(literal);  // 리터럴
    print("Inline"); // inline 리터럴
    // substring: 복사 없음
    print(std::string_view(s).substr(0, 5));  // "Hello"
    return 0;
}

실행 결과:

Hello World (size=11)
C string (size=8)
Literal (size=7)
Inline (size=6)
Hello (size=5)

string_view 주의사항: 수명

#include <string>
#include <string_view>
#include <iostream>
std::string_view getBadView() {
    std::string s = "temporary";
    return s;  // ❌ 위험! s가 파괴되면 댕글링
}
std::string_view getGoodView() {
    static std::string s = "static";
    return s;  // ✅ s가 프로그램 종료까지 유효
}
int main() {
    // std::string_view bad = getBadView();  // ❌ 댕글링
    std::string_view good = getGoodView();
    std::cout << good << "\n";
    return 0;
}

핵심: string_view가 가리키는 원본이 string_view보다 먼저 파괴되면 안 됩니다. 함수 반환값으로 string_view를 반환할 때, 그 안에서 만든 지역 std::string을 반환하면 댕글링입니다.

string_view 주요 연산

#include <string_view>
#include <iostream>
int main() {
    std::string_view sv = "Hello World";
    // substr: 복사 없음, 뷰만 생성
    std::string_view sub = sv.substr(0, 5);  // "Hello"
    // find, rfind, find_first_of 등 std::string과 유사
    size_t pos = sv.find(' ');
    std::cout << "Space at: " << pos << "\n";
    // remove_prefix, remove_suffix: 뷰 범위 조정 (C++17)
    sv.remove_prefix(6);  // "World"
    std::cout << sv << "\n";
    return 0;
}

CSV 파싱, trim, 토큰 분리 예제

예제 1: CSV 한 줄 파싱

// 복사해 붙여넣은 뒤: g++ -std=c++17 -o csv_parse csv_parse.cpp && ./csv_parse
#include <sstream>
#include <string>
#include <vector>
#include <iostream>
std::vector<std::string> splitCSV(const std::string& line) {
    std::vector<std::string> result;
    std::stringstream ss(line);
    std::string cell;
    while (std::getline(ss, cell, ',')) {
        result.push_back(cell);
    }
    return result;
}
int main() {
    std::string line = "apple,banana,cherry";
    auto cells = splitCSV(line);
    for (const auto& c : cells) {
        std::cout << "[" << c << "] ";
    }
    std::cout << "\n";
    return 0;
}

실행 결과:

[apple] [banana] [cherry]

이 방식은 따옴표 안의 쉼표("a,b",c)나 줄바꿈이 든 필드를 처리하지 못합니다. 실제 CSV 파일을 다룬다면 RFC 4180을 따르는 파서 라이브러리를 쓰는 편이 안전합니다.

예제 2: trim (앞뒤 공백 제거)

#include <string>
#include <iostream>
std::string trim(const std::string& s) {
    size_t start = s.find_first_not_of(" \t\n\r");
    if (start == std::string::npos) return "";
    size_t end = s.find_last_not_of(" \t\n\r");
    return s.substr(start, end - start + 1);
}
int main() {
    std::string s = "  hello world  ";
    std::cout << "[" << trim(s) << "]\n";
    return 0;
}

실행 결과:

[hello world]

예제 3: 문자열 연결 (reserve 활용)

#include <string>
#include <vector>
#include <iostream>
std::string join(const std::vector<std::string>& parts, const std::string& sep) {
    if (parts.empty()) return "";
    size_t total = 0;
    for (const auto& p : parts) total += p.size();
    total += sep.size() * (parts.size() - 1);
    std::string result;
    result.reserve(total);  // 한 번에 공간 확보
    result += parts[0];
    for (size_t i = 1; i < parts.size(); ++i) {
        result += sep;
        result += parts[i];
    }
    return result;
}
int main() {
    std::vector<std::string> v = {"a", "bb", "ccc"};
    std::cout << join(v, ", ") << "\n";
    return 0;
}

실행 결과:

a, bb, ccc

예제 4: string_view로 토큰 파싱 (복사 없음)

#include <string>
#include <string_view>
#include <vector>
#include <iostream>
std::vector<std::string_view> splitView(std::string_view sv, char delim) {
    std::vector<std::string_view> result;
    while (!sv.empty()) {
        size_t pos = sv.find(delim);
        if (pos == std::string_view::npos) {
            result.push_back(sv);
            break;
        }
        result.push_back(sv.substr(0, pos));
        sv.remove_prefix(pos + 1);
    }
    return result;
}
int main() {
    std::string s = "one,two,three";
    auto tokens = splitView(s, ',');
    for (auto t : tokens) {
        std::cout << "[" << t << "] ";
    }
    std::cout << "\n";
    return 0;
}

remove_prefix는 C++17 std::string_view 멤버로, 뷰의 시작 위치를 앞으로 당깁니다. 복사 없이 토큰을 나눌 수 있어 대용량 파싱에 유리합니다.


C 문자열 == 비교, c_str() 보관, string_view 댕글링

에러 1: C 문자열을 ==로 비교

const char*끼리의 ==는 포인터 주소 비교라서, 같은 내용이어도 거짓이 되거나 리터럴 병합 여부에 따라 결과가 바뀝니다.

// ❌ 잘못된 코드
const char* a = "hello";
const char* b = "hello";
if (a == b) { /* ... */ }
// ✅ 올바른 코드
if (strcmp(a, b) == 0) { /* ... */ }
// 또는 std::string 사용
std::string sa(a); std::string sb(b);
if (sa == sb) { /* ... */ }

에러 2: strcmp 반환값 잘못 해석

if (strcmp(a, b))는 “다를 때” 참이므로, 같은 문자열인데 조건이 거짓으로 처리됩니다.

// ❌ 잘못된 코드
if (strcmp(a, b)) {
    std::cout << "Same\n";  // 반대! 다를 때 실행됨
}
// ✅ 올바른 코드
if (strcmp(a, b) == 0) {
    std::cout << "Same\n";
}

에러 3: c_str() 반환값을 오래 보관

std::string을 수정하면 내부 버퍼가 재할당될 수 있고, 그 전에 받아 둔 c_str() 포인터는 무효가 됩니다. 크래시나 쓰레기 값으로 나타납니다.

// ❌ 잘못된 코드
const char* p = s.c_str();
s += " more";  // 재할당 가능
use(p);        // ❌ 댕글링
// ✅ 올바른 코드
s += " more";
use(s.c_str());  // 수정 직후 호출

에러 4: string_view가 댕글링

string_view가 가리키는 원본이 먼저 파괴되면 크래시나 쓰레기 값으로 나타납니다.

// ❌ 잘못된 코드
std::string_view getView() {
    std::string s = "temp";
    return s;  // s 파괴 후 댕글링
}
// ✅ 올바른 코드: std::string 반환 또는 원본 수명 보장
std::string getString() {
    return "temp";
}

에러 5: stoi 예외 미처리

stoi는 변환에 실패하면 예외를 던지므로, "abc" 같은 입력에서 예외를 잡지 않으면 프로그램이 종료됩니다.

// ❌ 잘못된 코드
int n = std::stoi(user_input);  // "abc" → 예외
// ✅ 올바른 코드
try {
    int n = std::stoi(user_input);
} catch (const std::invalid_argument&) {
    // 잘못된 입력 처리
} catch (const std::out_of_range&) {
    // 오버플로우 처리
}
// 또는 std::from_chars (C++17, 예외 없음)
int n;
auto [p, ec] = std::from_chars(user_input.data(),
                               user_input.data() + user_input.size(), n);
if (ec != std::errc{}) {
    // 에러 처리
}

에러 6: 인덱스 범위 초과

s[i]는 범위를 검사하지 않으므로 범위 밖 인덱스는 미정의 동작입니다. substr(start, len)은 start가 size()보다 크면 std::out_of_range를 던집니다.

// ❌ 잘못된 코드
std::string s = "hi";
char c = s[10];  // undefined behavior
// ✅ 올바른 코드
if (i < s.size()) {
    char c = s[i];
}
// 또는 at() 사용 (범위 검사, 예외)
char c = s.at(i);  // 범위 초과 시 std::out_of_range

reserve, SSO, from_chars로 문자열 비용 줄이기

tip 1: reserve로 재할당 최소화

// ❌ 용량이 찰 때마다 재할당·복사
std::string result;
for (const auto& piece : parts) {
    result += piece;
}
// ✅ 효율: 한 번에 공간 확보
std::string result;
size_t total = 0;
for (const auto& piece : parts) total += piece.size();
result.reserve(total);
for (const auto& piece : parts) {
    result += piece;
}

tip 2: 읽기 전용 인자는 string_view

// ❌ 리터럴이나 const char*를 넘기면 임시 std::string이 만들어짐
void process(const std::string& s);
// ✅ std::string, const char*, 리터럴 모두 복사 없이 받음
void process(std::string_view s);

tip 3: 작은 문자열은 SSO 활용

주요 표준 라이브러리 구현은 모두 SSO(Small String Optimization)를 씁니다. 64비트 기준으로 libstdc++와 MSVC STL은 15자, libc++는 22자 이하 문자열을 힙 할당 없이 객체 내부에 저장합니다. 그래서 짧은 문자열은 std::string을 그대로 써도 할당 비용이 없습니다.

tip 4: from_chars로 숫자 변환 (C++17)

std::from_chars는 예외를 던지지 않고 로케일도 참조하지 않아, 일반적으로 stoi/stod보다 빠릅니다. 부동소수점 버전은 GCC 11, MSVC 2019 16.4 이후부터 지원됩니다.

#include <charconv>
#include <iostream>
int main() {
    std::string s = "12345";
    int value;
    auto [ptr, ec] = std::from_chars(s.data(), s.data() + s.size(), value);
    if (ec == std::errc{}) {
        std::cout << "Parsed: " << value << "\n";
    } else {
        std::cout << "Parse failed\n";
    }
    return 0;
}

string_view 인자, 메시지 조립, C API 래퍼 패턴

패턴 1: 문자열 인자로 string_view 사용

// 함수가 문자열을 읽기만 할 때
std::string findAndReplace(std::string_view text,
                           std::string_view find,
                           std::string_view replace) {
    std::string result;
    result.reserve(text.size());
    size_t pos = 0;
    while (true) {
        size_t found = text.find(find, pos);
        if (found == std::string_view::npos) {
            result.append(text.substr(pos));
            break;
        }
        result.append(text.substr(pos, found - pos));
        result.append(replace);
        pos = found + find.size();
    }
    return result;
}

패턴 2: 로그/에러 메시지 조립

#include <sstream>
#include <string>
std::string formatError(const std::string& filename, int line, const std::string& msg) {
    std::ostringstream oss;
    oss << filename << ":" << line << ": " << msg;
    return oss.str();
}
// C++20: std::format
// return std::format("{}:{}: {}", filename, line, msg);

패턴 3: 문자열 풀 (공통 문자열 캐싱)

#include <string>
#include <unordered_set>
class StringPool {
    std::unordered_set<std::string> pool;
public:
    // unordered_set은 노드 기반이라 재해시해도 원소 주소가 바뀌지 않으므로
    // 반환한 string_view는 풀이 살아 있는 동안 유효함
    std::string_view intern(const std::string& s) {
        auto it = pool.find(s);
        if (it != pool.end()) {
            return *it;
        }
        auto [inserted, _] = pool.insert(s);
        return *inserted;
    }
};

패턴 4: 안전한 C API 래퍼

void legacyApi(const char* str);
void safeCall(const std::string& s) {
    legacyApi(s.c_str());  // 호출 시점에만 c_str() 사용
}
// string_view는 null 종료를 보장하지 않으므로 C API에 넘기려면 복사가 필요
// (같은 이름으로 오버로드하면 리터럴 인자에서 호출이 모호해지므로 이름을 나눔)
void safeCallView(std::string_view sv) {
    std::string temp(sv);
    legacyApi(temp.c_str());
}

패턴 5: 환경 변수/설정 파싱

#include <string>
#include <iostream>
#include <cstdlib>
std::string getEnvOrDefault(const std::string& key, const std::string& def) {
    const char* val = std::getenv(key.c_str());
    return val ? std::string(val) : def;
}

같이 보면 좋은 글


참고 자료

한글(UTF-8) 문자열에 대해

std::string은 문자가 아니라 바이트를 저장합니다. UTF-8에서 한글 한 글자는 3바이트이므로 "한글"의 size()는 6이고, s[0]은 글자가 아니라 첫 바이트입니다. 글자 단위로 자르거나 세려면 UTF-8 디코딩이 필요하며, 대소문자 변환이나 정규화까지 다룬다면 ICU 같은 라이브러리를 씁니다.

다음 글: C++ 실전 가이드 #11-2: 바이너리 직렬화 이전 글: C++ 실전 가이드 #10-3: STL 알고리즘