C++ sregex_iterator | Regex iterators for all matches
Key takeaways
Use sregex_iterator and sregex_token_iterator to enumerate matches, split tokens, avoid dangling iterators, and handle empty matches and UTF-8 limits with std::regex.
What is regex_iterator?
Traverse all matches (C++11)
#include <regex>
std::string text = "C++ 11, C++ 14, C++ 17";
std::regex pattern{R"(\d+)"};
auto begin = std::sregex_iterator(text.begin(), text.end(), pattern);
auto end = std::sregex_iterator();
for (auto it = begin; it != end; ++it) {
std::cout << it->str() << std::endl;
}
// 11
// 14
// 17
std::regex_search on its own only ever finds the first match in a string — to get “11”, “14”, and “17” without an iterator, you’d need to manually re-search starting after each match’s end position, tracking that offset yourself. sregex_iterator packages exactly that loop into an iterator interface: each ++it internally re-invokes regex_search starting right after the previous match, so you get the same begin/end/range-for ergonomics as iterating a container, even though there’s no actual container of matches sitting in memory — each match is computed lazily as you advance.
Iterating over every match
#include <regex>
#include <string>
std::string text = "abc 123 def 456";
std::regex pattern{R"(\d+)"};
// Create the iterators
auto begin = std::sregex_iterator(text.begin(), text.end(), pattern);
auto end = std::sregex_iterator();
// Iterate
for (auto it = begin; it != end; ++it) {
std::smatch match = *it;
std::cout << match.str() << std::endl;
}
The default-constructed std::sregex_iterator() as end looks unusual the first time you see it — there’s no obvious “one past the last match” position to construct, since matches are found lazily. The standard sidesteps this by defining the default-constructed sregex_iterator as a special universal end-of-sequence marker: every sregex_iterator eventually compares equal to it once regex_search fails to find any further match, regardless of which string or pattern the iterator started from.
Words, capture groups, URLs, and tokens
Extracting words
#include <regex>
#include <vector>
std::vector<std::string> extractWords(const std::string& text) {
std::regex pattern{R"(\b\w+\b)"};
auto begin = std::sregex_iterator(text.begin(), text.end(), pattern);
auto end = std::sregex_iterator();
std::vector<std::string> words;
for (auto it = begin; it != end; ++it) {
words.push_back(it->str());
}
return words;
}
int main() {
auto words = extractWords("Hello, World! C++ 2026");
for (const auto& word : words) {
std::cout << word << std::endl;
}
// Hello
// World
// C
// 2026
}
Notice "C++" comes out as just "C" — \w only matches word characters (letters, digits, underscore), and + and + themselves aren’t word characters, so the \b\w+\b pattern stops at the first +. This is a genuinely common surprise when running general-purpose word-extraction regexes over text that contains programming language names, versioned identifiers, or other tokens with punctuation baked in — the fix is a pattern tailored to the actual token shape you need, not a generic \w+.
Reading capture groups from each match
#include <regex>
int main() {
std::string text = "[email protected], [email protected]";
std::regex pattern{R"((\w+)@(\w+\.\w+))"};
auto begin = std::sregex_iterator(text.begin(), text.end(), pattern);
auto end = std::sregex_iterator();
for (auto it = begin; it != end; ++it) {
std::smatch match = *it;
std::cout << "Email: " << match[0] << std::endl;
std::cout << "User: " << match[1] << std::endl;
std::cout << "Domain: " << match[2] << std::endl;
std::cout << std::endl;
}
}
match[0] is easy to forget about since it’s not one of the parenthesized groups you wrote — it’s always the entire matched substring, with match[1], match[2], etc. corresponding to the first, second, and so on parenthesized capture group in the pattern, in the order their opening parenthesis appears. This numbering is consistent across every std::regex API (regex_match, regex_search, iterators), so it’s worth memorizing once rather than re-deriving each time.
Pulling hosts and paths out of URLs
#include <regex>
struct URL {
std::string protocol;
std::string host;
std::string path;
};
std::vector<URL> extractURLs(const std::string& text) {
std::regex pattern{R"((https?)://([^/]+)(/[^\s]*))"};
auto begin = std::sregex_iterator(text.begin(), text.end(), pattern);
auto end = std::sregex_iterator();
std::vector<URL> urls;
for (auto it = begin; it != end; ++it) {
std::smatch match = *it;
urls.push_back({
match[1].str(), // protocol
match[2].str(), // host
match[3].str() // path
});
}
return urls;
}
Treat this as a demonstration of iterator + capture-group mechanics, not a production-ready URL parser — real URLs have query strings, fragments, userinfo, IPv6 host literals, and percent-encoding that a three-group regex like this doesn’t attempt to handle. For anything beyond quick text scraping, a dedicated URL-parsing library will handle the edge cases a hand-written pattern like this one will quietly get wrong.
Splitting into tokens
#include <regex>
std::vector<std::string> tokenize(const std::string& text) {
std::regex pattern{R"(\s+)"}; // whitespace
std::sregex_token_iterator begin(text.begin(), text.end(), pattern, -1);
std::sregex_token_iterator end;
return {begin, end};
}
int main() {
auto tokens = tokenize("Hello World C++");
for (const auto& token : tokens) {
std::cout << "[" << token << "]" << std::endl;
}
// [Hello]
// [World]
// [C++]
}
regex_token_iterator
std::string text = "a,b,c,d";
std::regex pattern{","};
// -1 means "everything between matches" (i.e. the delimiter is excluded)
std::sregex_token_iterator begin(text.begin(), text.end(), pattern, -1);
std::sregex_token_iterator end;
for (auto it = begin; it != end; ++it) {
std::cout << *it << std::endl;
}
// a
// b
// c
// d
The -1 argument is the detail that makes sregex_token_iterator behave like a splitter rather than a matcher — passed a submatch index of -1, it yields the text between matches (exactly what a “split” operation needs), whereas passing 0 (the default) would yield the matched delimiters themselves. This is std::regex’s equivalent of str.split(",") in languages with a built-in split method, just expressed through the iterator abstraction instead of a dedicated function.
Dangling iterators, recompiled patterns, empty matches, and group indexes
Iterating over a temporary string
// ❌ Dangling iterator
auto getIterator() {
std::string text = "hello 123";
std::regex pattern{R"(\d+)"};
return std::sregex_iterator(text.begin(), text.end(), pattern);
// text is destroyed here, but the returned iterator still refers to it
}
// ✅ Ensure the string outlives the iterator
std::string text = "hello 123";
auto it = std::sregex_iterator(text.begin(), text.end(), pattern);
sregex_iterator stores iterators into the original string, not a copy of its contents — it never owns the text it’s searching. Returning an iterator constructed from a local std::string hands the caller something that refers to memory that’s about to be freed the moment getIterator() returns, which is undefined behavior the instant the caller dereferences it. This is the same category of bug as returning a reference or pointer to a local variable, just less obvious because sregex_iterator doesn’t look like a reference at the call site.
Compiling the regex in a loop
// ❌ Recompiling every time
for (const auto& text : texts) {
std::regex pattern{R"(\d+)"}; // compiled on every iteration
std::regex_search(text, pattern);
}
// ✅ Compile once, reuse
std::regex pattern{R"(\d+)"};
for (const auto& text : texts) {
std::regex_search(text, pattern);
}
Constructing a std::regex compiles the pattern into an internal automaton, which is genuinely expensive relative to a single search — measured in real codebases, recompiling the same pattern inside a loop can easily dominate the total runtime, turning what should be a cheap scan into the slowest part of the program. Hoisting the std::regex construction outside the loop (or making it static, as the log-parsing example further down does) is one of the most impactful, easiest-to-apply regex performance fixes there is.
Patterns that match the empty string
std::string text = "hello";
std::regex pattern{R"(\d*)"}; // zero or more digits — can match an empty string
// Empty matches are possible here
auto begin = std::sregex_iterator(text.begin(), text.end(), pattern);
auto end = std::sregex_iterator();
for (auto it = begin; it != end; ++it) {
if (!it->str().empty()) {
std::cout << it->str() << std::endl;
}
}
A pattern like \d* (zero-or-more) can match a length-zero string at every single character position, since “zero digits” is trivially satisfied everywhere. sregex_iterator handles the resulting risk of an infinite loop correctly by internally advancing by at least one character after an empty match — but the output still includes those empty matches unless you filter them out yourself, as the !it->str().empty() check here does. Forgetting that check with a zero-or-more pattern is a common source of “why do I have all these blank lines in my output” bugs.
Off-by-one group indexes
std::string text = "[email protected]";
std::regex pattern{R"((\w+)@(\w+)\.(\w+))"};
std::smatch matches;
if (std::regex_match(text, matches, pattern)) {
// matches[0]: the entire match
// matches[1]: first capture group
// matches[2]: second capture group
// ...
}
Which regex function to reach for
// 1. Validate
bool isValid = std::regex_match(text, pattern);
// 2. Search
std::smatch matches;
std::regex_search(text, matches, pattern);
// 3. Replace
auto result = std::regex_replace(text, pattern, replacement);
// 4. All matches
auto it = std::sregex_iterator(begin, end, pattern);
These four functions cover the four things you’d ever want to do with a regex, and picking the wrong one for the job is a common mistake: regex_match requires the entire string to match the pattern end to end (good for validation, like “is this a valid email address”), while regex_search only needs to find the pattern somewhere in the string (good for “does this text contain an email address anywhere”). Using regex_search where you meant regex_match is a real bug pattern — a validation function using regex_search would accept "garbage [email protected] garbage" as a “valid” email string, since the pattern matches somewhere inside it even though the whole string clearly isn’t a valid email.
Behavior of regular expression iterative matching
std::regex_search searches only one section at a time. To iterate over all non-nested matches in a string, use std::regex_iterator (for character sequences) / std::sregex_iterator (for std::string iterators). The iterator is internally implemented as a pattern that calls regex_search again starting from the position after the end of the previous match.
- Not a global match: the default iterator lists partial matches. Use
regex_matchto see if an entire string matches a pattern exactly. - Nested/Overlapping: Standard iterators generally move the next search start point to the end of the match, so overlapping patterns (e.g.
aaovera) require separate design depending on requirements.
sregex_iterator type family
sregex_iterator: works overstd::string::const_iteratorranges, initialized with astd::regex.cregex_iterator: forconst char*ranges.wsregex_iterator: forstd::wstring.
The widely used idiom is to leave the terminal iterator as the default-constructed std::sregex_iterator().
auto begin = std::sregex_iterator(text.begin(), text.end(), pattern);
auto end = std::sregex_iterator();
for (auto it = begin; it != end; ++it) {
const std::smatch& m = *it;
// m.ready(), m.size(), m.str(n)
}
Parsing timestamp, level, and message from log lines
Below is an example of extracting Timestamp, Level, Message from one line (pattern adjusted to suit log format).
#include <regex>
#include <string>
#include <iostream>
#include <vector>
struct LogLine {
std::string timestamp;
std::string level;
std::string message;
};
// Example: "2026-03-30 12:00:00 ERROR something failed"
bool parseLogLine(const std::string& line, LogLine& out) {
static const std::regex re(
R"((\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}) (\w+) (.*))");
std::smatch m;
if (!std::regex_match(line, m, re)) {
return false;
}
out.timestamp = m[1].str();
out.level = m[2].str();
out.message = m[3].str();
return true;
}
// Collect only lines at a specific level, across multiple lines
std::vector<std::string> extractErrors(const std::string& text) {
std::regex levelLine(R"(\b(ERROR|CRITICAL)\b.*)");
auto begin = std::sregex_iterator(text.begin(), text.end(), levelLine);
auto end = std::sregex_iterator();
std::vector<std::string> errors;
for (auto it = begin; it != end; ++it) {
errors.push_back(it->str());
}
return errors;
}
The static const std::regex re(...) inside parseLogLine is doing exactly the fix for compiling the regex in a loop — static ensures the pattern is compiled exactly once, the first time parseLogLine is called, and every subsequent call (potentially millions, over a large log file) reuses that same compiled automaton instead of recompiling the pattern per line.
Tip: If you load a whole file into a string and then split it into lines, sregex_iterator works well for “multiple tokens within a line,” while a line-level loop (as extractErrors does over multi-line text) works well for “record boundaries.”
Keeping std::regex fast enough
- Compilation cost: The
std::regexconstructor compiles the pattern. Create and reuse it only once outside the loop. - Engine: GCC/LLVM’s
std::regexmay be slower than expected on very large inputs or complex patterns. If it’s a hot path, look at Boost.Regex, RE2-style libraries, or a manual parser after profiling. - Allocation: Using
smatch/sregex_iteratorcan internally allocate substrings. For bulk logs, a custom scanner based onstring_viewor a fixed buffer parser may be better. std::regex_constants::optimize: May hint at optimization depending on the implementation, but not guaranteed to always be faster — measurement takes precedence.
FAQ
Q1: What about Regex?
A: Regular expression support (C++11).
Q2: Iterator?
A: regex_iterator enumerates every match.
Q3: Capture group?
A: Use (). Access via matches[N].
Q4: Performance?
A: Compilation is relatively slow — compile once and reuse.
Q5: Token split?
A: regex_token_iterator.
Q6: What are the learning resources?
A:
- “Mastering Regular Expressions”
- cppreference.com
- “C++ Primer”
Related Articles
- C++ Regular Expressions Basics — Complete std::regex Guide
- std::async and Launch Policies
- C++ std::atomic
- C++ Attributes