C++ Regular Expression Basics: std::regex Matching, Capture Groups and Tokenizing

Matching a pattern with std::regex

This post is the beginner-focused entry point in the series — if you already know regular expressions from another language (Python, JavaScript, PCRE) and just want the C++-specific API mapped out end to end, the companion complete std::regex guide covers the same ground with more depth on performance and backtracking behavior. Here, the goal is just to get the core mental model straight: a regex object holds a compiled pattern, and functions like regex_search and regex_match (the difference is explained just below) take that compiled pattern and a string and tell you whether/where it matches.

#include <regex>
#include <iostream>
using namespace std;

int main() {
    regex pattern("\\d+");  // digit pattern
    
    string text = "abc123def456";
    
    // search
    if (regex_search(text, pattern)) {
        cout << "found digits" << endl;
    }
}

This is the single most common beginner mistake with C++ regex, so it’s worth memorizing early: regex_match only returns true if the pattern accounts for the entire string from the first character to the last, as if you’d wrapped it in ^ and $. regex_search just needs to find the pattern somewhere inside the string. If you want to know “does this string contain a number anywhere,” reach for regex_search; if you want to know “is this string, in its entirety, a valid number,” use regex_match. Picking the wrong one is a quiet bug — the code compiles and runs, it just silently rejects or accepts the wrong inputs.

regex pattern("\\d+");

string s1 = "123";
string s2 = "abc123";

// regex_match: whole string must match
cout << regex_match(s1, pattern) << endl;  // 1 (true)
cout << regex_match(s2, pattern) << endl;  // 0 (false)

// regex_search: substring match
cout << regex_search(s1, pattern) << endl;  // 1
cout << regex_search(s2, pattern) << endl;  // 1

Capture groups

Parentheses in a pattern do two things at once: they group parts of the pattern together for quantifiers/alternation, and they “capture” the text that part matched so you can retrieve it afterward. match[0] is always the whole match, not the first group — the first parenthesized group is match[1], counted left to right by where the opening parenthesis appears in the pattern, which is worth double-checking whenever a pattern has nested or nested-looking groups.

regex pattern("(\\d{3})-(\\d{4})-(\\d{4})");
string phone = "010-1234-5678";

smatch match;
if (regex_match(phone, match, pattern)) {
    cout << "full: " << match[0] << endl;   // 010-1234-5678
    cout << "part1: " << match[1] << endl;  // 010
    cout << "part2: " << match[2] << endl;  // 1234
    cout << "part3: " << match[3] << endl;  // 5678
}

Email checks, URL parsing, replacement, and log lines

A loose email check

Treat this pattern as “good enough to catch obviously wrong input,” not as a spec-accurate email validator — the real email grammar (RFC 5322) allows for things this pattern doesn’t handle (quoted local parts, comments) and most production systems don’t bother implementing it fully. A regex check followed by an actual verification email (or an API call to a validation service) is the more reliable combination in practice.

bool isValidEmail(const string& email) {
    regex pattern(R"(^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$)");
    return regex_match(email, pattern);
}

int main() {
    cout << isValidEmail("[email protected]") << endl;  // 1
    cout << isValidEmail("invalid.email") << endl;     // 0
}

Splitting a URL into scheme, host, port, and path

The ? after (?::(\d+)) makes that whole group optional, which is why match[3] (the port) can legitimately come back empty for a URL like https://example.com/path with no port specified — check match[3].matched before treating an empty string as meaningful, since “the group didn’t participate at all” and “the group matched an empty string” are technically different outcomes that str() alone doesn’t distinguish.

struct URL {
    string protocol;
    string host;
    string port;
    string path;
};

URL parseURL(const string& url) {
    regex pattern(R"(^(\w+)://([^:/]+)(?::(\d+))?(/.*)?$)");
    smatch match;
    
    if (regex_match(url, match, pattern)) {
        return {
            match[1],  // protocol
            match[2],  // host
            match[3],  // port
            match[4]   // path
        };
    }
    
    return {};
}

int main() {
    auto url = parseURL("https://example.com:8080/path/to/page");
    
    cout << "protocol: " << url.protocol << endl;
    cout << "host: " << url.host << endl;
    cout << "port: " << url.port << endl;
    cout << "path: " << url.path << endl;
}

Replacing matches with regex_replace

By default regex_replace replaces every match in the string, which surprises people coming from tools where “replace” defaults to the first occurrence only. Pass regex_constants::format_first_only explicitly when you only want the first match touched — worth remembering as a flag rather than something you have to hand-roll with a loop and regex_search.

#include <regex>

int main() {
    string text = "Hello World, Hello C++";
    regex pattern("Hello");
    
    // replace all
    string result = regex_replace(text, pattern, "Hi");
    cout << result << endl;  // Hi World, Hi C++
    
    // first occurrence only
    result = regex_replace(text, pattern, "Hi", regex_constants::format_first_only);
    cout << result << endl;  // Hi World, Hello C++
}

Parsing structured log lines

Because this uses regex_match (not regex_search), the whole line has to fit the pattern exactly — a line with an extra trailing space, a different timestamp format, or a multi-line stack trace continuation simply won’t produce a LogEntry and is silently skipped, with no indication anything went wrong. For anything beyond a quick script, it’s worth logging or counting the lines that fail to match so malformed input doesn’t just vanish unnoticed.

struct LogEntry {
    string timestamp;
    string level;
    string message;
};

vector<LogEntry> parseLog(const string& log) {
    vector<LogEntry> entries;
    
    regex pattern(R"(\[([\d\-: ]+)\] \[(\w+)\] (.+))");
    
    istringstream iss(log);
    string line;
    
    while (getline(iss, line)) {
        smatch match;
        if (regex_match(line, match, pattern)) {
            entries.push_back({
                match[1],  // timestamp
                match[2],  // level
                match[3]   // message
            });
        }
    }
    
    return entries;
}

int main() {
    string log = R"([2026-03-11 10:30:00] [INFO] server started
[2026-03-11 10:30:05] [ERROR] connection failed
[2026-03-11 10:30:10] [WARN] retrying)";
    
    auto entries = parseLog(log);
    
    for (const auto& entry : entries) {
        cout << entry.timestamp << " | " 
             << entry.level << " | " 
             << entry.message << endl;
    }
}

Finding every match with sregex_iterator

Use sregex_iterator any time you need every match in a string rather than just the first one — regex_search alone only ever finds one match per call. The default-constructed sregex_iterator end; might look odd if you’re used to STL container iterators that come from .end(), but it’s the same pattern: a value-initialized iterator of this type acts as the sentinel meaning “no more matches.”

// declare and initialize
string text = "abc123def456ghi789";
regex pattern("\\d+");

// find every match
sregex_iterator it(text.begin(), text.end(), pattern);
sregex_iterator end;

while (it != end) {
    cout << it->str() << endl;  // 123, 456, 789
    ++it;
}

Splitting strings with sregex_token_iterator

The -1 argument here is easy to gloss over but changes everything: it tells sregex_token_iterator to give you the text between matches of the delimiter, rather than the matches themselves. Change it to 0 and you’d iterate over the commas instead of the words — a small detail, but the one that makes this useful as a splitting mechanism at all.

string text = "apple,banana,cherry";
regex delimiter(",");

// token iterator
sregex_token_iterator it(text.begin(), text.end(), delimiter, -1);
sregex_token_iterator end;

while (it != end) {
    cout << *it << endl;  // apple, banana, cherry
    ++it;
}

Escaping, recompilation, greed, and invalid patterns

Backslashes and raw string literals

This trips up nearly everyone learning C++ regex for the first time: \d isn’t a valid escape in an ordinary C++ string literal, so "\d+" typically just becomes the two characters \ and d passed through as-is (some compilers will warn, many won’t), which is not the digit pattern you intended. Getting into the habit of writing regex patterns as raw string literals — R"(\d+)" — from day one sidesteps the whole class of double-escaping bugs, since backslashes inside R"(...)" are treated completely literally.

// wrong: insufficient escaping
regex pattern("\d+");  // \d is not a regex escape as intended

// correct: double backslash
regex pattern("\\d+");

// raw string (recommended)
regex pattern(R"(\d+)");

Rebuilding the regex on every call

Constructing a regex object isn’t free — the pattern string gets parsed and compiled into an internal matching engine at that point, not lazily on first use. Rebuilding the same unchanging pattern inside a loop pays that compilation cost on every single iteration for no reason; move the construction outside the loop and the regex object can be reused freely, including across multiple calls to regex_search/regex_match, since matching doesn’t modify it.

// bad: construct regex every iteration
for (const string& text : texts) {
    regex pattern("\\d+");  // wasteful
    regex_search(text, pattern);
}

// good: reuse one regex
regex pattern("\\d+");
for (const string& text : texts) {
    regex_search(text, pattern);
}

Greedy quantifiers matching too much

By default, *, +, and similar quantifiers grab as much text as they possibly can while still letting the rest of the pattern succeed — that’s what “greedy” means. Against <div>content</div>, a greedy .* between < and > doesn’t stop at the first > it sees; it keeps consuming characters all the way to the last > in the string, since that still lets the pattern match overall. Adding ? after the quantifier (.*?) makes it lazy instead, matching as little as possible — which is almost always what people actually want when parsing something tag-like.

string html = "<div>content</div>";

// greedy
regex greedy("<.*>");
// matches: <div>content</div> (whole string)

// non-greedy
regex nonGreedy("<.*?>");
// matches: <div>, </div> separately

Invalid patterns throw std::regex_error

A pattern is compiled when the std::regex object is constructed, and a syntax error throws std::regex_error at that point, at runtime, not at compile time. That matters whenever the pattern comes from configuration or user input:

try {
    std::regex re(user_pattern);          // may throw
} catch (const std::regex_error& e) {
    std::cerr << "bad pattern: " << e.what() << " (code " << e.code() << ")\n";
}

e.code() returns a std::regex_constants::error_type such as error_paren or error_brack. Some implementations can also throw error_complexity or error_stack while matching, when a pattern with heavy backtracking meets a long input, so matching untrusted text needs the same try/catch.

Performance: construct once, and know the limits of std::regex

Constructing a std::regex is expensive, since it parses and compiles the pattern into a matcher. Build it once (a static const std::regex or a member) rather than inside a loop. Even then, the standard library implementations (libstdc++, libc++, MSVC) are known to be much slower than dedicated engines, and they use backtracking, so some patterns (nested quantifiers such as (a+)+$) take exponential time on crafted input. Some implementations also recurse deeply and can overflow the stack on long inputs. For hot paths or untrusted input, consider RE2 (linear-time, no backreferences), PCRE2, or compile-time regex libraries such as CTRE. For fixed-format parsing (numbers, simple delimiters), std::from_chars and hand-written scanning are usually both faster and clearer.

Regex syntax (ECMAScript-style overview)

std::regex defaults to the ECMAScript grammar (the same family JavaScript’s regex uses), which is why most patterns you’d find in a JavaScript or general regex tutorial translate directly. This is a quick-reference cheat sheet of the pieces used throughout this post — worth bookmarking rather than memorizing up front, since the constructs make more sense once you’ve written a few patterns of your own.

// character classes
\d  // digit [0-9]
\w  // word [a-zA-Z0-9_]
\s  // whitespace
.   // any character (except newline, depending on flags)

// quantifiers
*   // zero or more
+   // one or more
?   // zero or one
{n} // exactly n
{n,m}  // between n and m

// anchors
^   // start
$   // end
\b  // word boundary

// groups
()  // capturing group
(?:)  // non-capturing group

FAQ

Q1: When should I use regular expressions?

A:

  • String validation
  • Parsing
  • Search and replace
  • Data extraction

Q2: What about performance?

A: For a beginner project or one-off script, don’t worry about it — std::regex is fine. If you find yourself running the same regex over huge amounts of text in a hot loop, make sure you’re not reconstructing the regex object each time (see “Rebuilding the regex on every call” above), and if it’s still too slow, simple string member functions like find/substr are usually faster for basic checks that don’t need real pattern matching.

Q3: Why use raw string literals?

A: R"(...)" avoids manual backslash escaping in the pattern text.

Q4: ECMAScript vs POSIX?

A: The default is ECMAScript; you can change the grammar with regex_constants when constructing std::regex.

Q5: How do I debug regex?

A:

Q6: Where can I learn more?

A:


Posts that connect well with this topic: