Walking Directories with std::filesystem: directory_iterator, Recursion, Symlinks and error_code
Key takeaways
directory_iterator vs recursive_directory_iterator, directory_options, symlinks, error_code overloads, filtering, disk usage, and performance tips for C++17 filesystem.
What is directory_iterator?
Iterate directory entries (C++17).
#include <filesystem>
namespace fs = std::filesystem;
for (const auto& entry : fs::directory_iterator(".")) {
std::cout << entry.path() << std::endl;
}
Before C++17, listing a directory meant opendir/readdir on POSIX and FindFirstFile/FindNextFile on Windows, usually behind a hand-written wrapper or Boost.Filesystem. std::filesystem standardized Boost’s design, so the same loop now works on every major platform. Each entry is a directory_entry, which holds the path and may cache some metadata (file type, sometimes size) that the operating system returned while reading the directory.
A few details are worth knowing from the start. The order of entries is unspecified: it is whatever order the file system returns, which differs between NTFS, ext4 and APFS, so sort the results if order matters (for example, in tests or reproducible builds). The special entries . and .. are never returned. And printing a path with << puts it in quotes ("./main.cpp"), because operator<< for paths uses std::quoted; use entry.path().string() if you want the raw text. On older GCC (before 9) and Clang (before 9), you also needed to link -lstdc++fs or -lc++fs, and forgetting it produced undefined reference to std::filesystem::... linker errors.
Recursive traversal
for (const auto& entry : fs::recursive_directory_iterator(".")) {
std::cout << entry.path() << std::endl;
}
recursive_directory_iterator visits every entry in the subtree, depth-first: when it reaches a directory, it yields that directory and then descends into it before continuing with the directory’s siblings. It keeps a stack of open directory handles internally, one per level, so very deep trees hold many handles open at once. By default it does not follow directory symlinks, which is the safe choice; you opt in with directory_options::follow_directory_symlink.
The simple range-for form has one sharp edge: any error during the walk, such as a subdirectory you are not allowed to read, throws std::filesystem::filesystem_error and ends the whole traversal. On a developer machine that is rarely noticed; run the same code over / or a user’s home directory with protected folders and it stops halfway with filesystem error: directory iterator cannot open directory: Permission denied. The skip_permission_denied option and the error_code sections below are the fixes.
directory_iterator vs recursive_directory_iterator
directory_iterator | recursive_directory_iterator | |
|---|---|---|
| Scope | Immediate children only | Full subtree |
| Default | Does not descend | Descends automatically |
| Options | directory_options | Same family + disable_recursion_pending |
| Typical use | List one level | Full-tree search, size totals |
There is no standard max-depth flag. Implement depth limits with manual recursion/queues, or call disable_recursion_pending() to skip entering specific directories (e.g. build).
for (fs::recursive_directory_iterator it(root), end; it != end; ++it) {
if (it->is_directory() && it->path().filename() == "build") {
it.disable_recursion_pending();
}
}
disable_recursion_pending() must be called while the iterator points at the directory, before the next increment; it tells the iterator “do not descend into this one”. This explicit-iterator loop is necessary because a range-for hides the iterator. The same loop can implement a depth limit using it.depth(), which returns 0 for entries directly under root: if (it->is_directory() && it.depth() >= maxDepth) it.disable_recursion_pending();. There is also it.pop(), which abandons the rest of the current directory and moves back up a level, useful when you find what you were looking for and the siblings do not matter. Skipping .git, node_modules or build output directories this way can make a scan of a source tree dramatically faster, since those are often the largest parts of it.
Listing, filtering, and sizing a directory
Listing the entries of one directory
void listFiles(const fs::path& dir) {
for (const auto& entry : fs::directory_iterator(dir)) {
if (entry.is_regular_file()) {
std::cout << entry.path().filename() << std::endl;
}
}
}
Finding files by extension
std::vector<fs::path> findFiles(const fs::path& dir,
const std::string& ext) {
std::vector<fs::path> result;
for (const auto& entry : fs::recursive_directory_iterator(dir)) {
if (entry.is_regular_file() &&
entry.path().extension() == ext) {
result.push_back(entry.path());
}
}
return result;
}
extension() returns the part from the last dot, including the dot, so the caller must pass ".txt", not "txt"; passing "txt" silently finds nothing. The comparison is exact and case-sensitive, which matters on Windows and macOS, where the file system itself is usually case-insensitive and photo.JPG and photo.jpg are equally common. Normalize both sides to lower case if you want a case-insensitive match. Note also that archive.tar.gz has extension .gz, and a dotfile such as .gitignore has an empty extension (its stem is .gitignore), which surprises people writing “skip hidden files” filters.
entry.is_regular_file() follows symlinks, so a symlink to a regular file counts as one here. That is usually what you want for “find all source files”, but it means the same file can appear twice if it is reachable both directly and through a link.
Summing the size of a directory tree
uintmax_t calculateDirSize(const fs::path& dir) {
uintmax_t size = 0;
for (const auto& entry : fs::recursive_directory_iterator(dir)) {
if (entry.is_regular_file()) {
size += entry.file_size();
}
}
return size;
}
This sums the logical size of each file, which is what ls -l shows, not the disk space used, which is what du shows. The two differ for sparse files, compressed or deduplicated file systems, and small files that still occupy a whole allocation block, so do not expect this number to match your file manager’s “size on disk”. Every file_size() call can also throw: a file deleted between reading the directory and asking for its size, or a symlink whose target is missing, raises filesystem_error and aborts the total. The error-tolerant tree_bytes below handles that.
Filtering files by size
void filterFiles(const fs::path& dir) {
for (const auto& entry : fs::directory_iterator(dir)) {
if (entry.is_regular_file()) {
auto size = entry.file_size();
if (size > 1024 * 1024) {
std::cout << entry.path() << ": "
<< size << " bytes" << std::endl;
}
}
}
}
Filtering patterns
The standard has no glob/regex filter—combine:
- Extension:
path::extension(),path::stem(). - Type:
is_regular_file(),is_directory(). - Exclude dirs: skip paths containing
.git,node_modules, etc. - Regex: apply to
path::string()if needed—mind Windows encodings. With C++20 you can chainstd::rangesadapters for cleaner filters.
For example, for (auto& e : fs::directory_iterator(dir) | std::views::filter([](const fs::directory_entry& e) { return e.is_regular_file(); })) works because directory_iterator models an input range. Be aware that these are single-pass input iterators: you cannot iterate the same directory_iterator twice, and range adaptors that need to go back (such as views::reverse) will not compile on it. Collect into a std::vector<fs::path> first if you need to sort, reverse or iterate more than once.
When comparing paths against exclusion lists, compare path components, not substrings. A check like path.string().find(".git") != npos also skips my.github.io and .gitignore. Iterating over path itself yields components, so std::ranges::find(entry.path(), fs::path(".git")) != entry.path().end() tests whether .git is one of the directory names.
Iterator options
fs::directory_iterator it(dir);
fs::directory_iterator it(dir,
fs::directory_options::follow_directory_symlink);
fs::directory_iterator it(dir,
fs::directory_options::skip_permission_denied);
(The three declarations are alternatives; in one scope, reusing the name it would not compile.) The options are bit flags, so they combine with |: fs::directory_options::follow_directory_symlink | fs::directory_options::skip_permission_denied. follow_directory_symlink has no effect on the non-recursive directory_iterator, because it never descends anyway; it matters for recursive_directory_iterator. skip_permission_denied makes the iterator silently skip directories it cannot open instead of reporting an error. It only covers “permission denied”; other errors, such as a directory removed mid-walk or an I/O error on a failing disk, are still reported.
Symlinks
- Directory symlinks: by default,
recursive_directory_iteratorlists a directory symlink as an entry but does not descend into it. Recursion follows them only when you passfollow_directory_symlink—which can mean unexpectedly large trees or cycles. follow_directory_symlink: explicitly follow; still no automatic cycle detection—track visited canonical paths on sensitive workloads.- Link vs target:
is_symlink(),symlink_statusvsstatusdistinguish link metadata from final type. - Hard links: the same inode through multiple paths can double-count size—deduplicate if required.
A symlink cycle is easy to create by accident: ln -s .. parent inside a directory makes dir/parent/parent/parent/... infinitely deep. With follow_directory_symlink, a walk over such a tree runs until the path exceeds the operating system’s length limit and then fails with an error such as File name too long, after producing a huge number of bogus entries. Package managers and build tools create symlinks liberally (node_modules/.bin, virtualenvs, Nix stores), so this is a real risk when scanning developer machines. If you must follow links, keep a set of fs::canonical(dir) results for directories already entered and call disable_recursion_pending() when you see one again; on POSIX systems, comparing device and inode numbers is the more robust identity check.
To inspect the link itself rather than its target, use entry.is_symlink() and entry.symlink_status(). entry.status() and entry.is_directory() follow the link, so a dangling symlink reports file_type::not_found through status() while symlink_status() still says it is a symlink. On Windows, directory junctions and symlinks behave slightly differently from POSIX symlinks, and creating symlinks may require developer mode or elevated privileges, so test traversal code on the platforms you ship to.
Searching a tree for a file name
#include <filesystem>
#include <string>
#include <vector>
namespace fs = std::filesystem;
std::vector<fs::path> find_by_name(const fs::path& root, std::string_view key) {
std::vector<fs::path> out;
auto opts = fs::directory_options::skip_permission_denied;
std::error_code ec;
for (const auto& e : fs::recursive_directory_iterator(root, opts, ec)) {
if (ec) break;
if (!e.is_regular_file()) continue;
auto fn = e.path().filename().string();
if (fn.find(key) != std::string::npos)
out.push_back(e.path());
}
return out;
}
This function looks error-tolerant, but it has a gap that catches almost everyone: in a range-for, the error_code passed to the constructor is only filled in by the constructor. Every later step of the loop calls the throwing operator++, so an error deep in the tree still throws filesystem_error, and if (ec) break; inside the loop only ever sees the result of opening root. To handle errors at every step without exceptions, drive the iterator by hand with increment(ec):
std::error_code ec;
fs::recursive_directory_iterator it(root, fs::directory_options::skip_permission_denied, ec);
for (fs::recursive_directory_iterator end; !ec && it != end; it.increment(ec)) {
// use *it
}
if (ec) std::cerr << "walk stopped: " << ec.message() << '\n';
filename().string() is also a portability trap on Windows, where paths are stored as UTF-16. Converting a name containing characters outside the current code page to std::string can throw or produce mojibake; path::u8string() (which returns std::u8string in C++20) or comparing path objects directly avoids the lossy conversion.
Byte count of a tree with error_code
std::uintmax_t tree_bytes(const fs::path& p, std::error_code& ec) {
std::uintmax_t total = 0;
auto opts = fs::directory_options::skip_permission_denied;
for (const auto& e : fs::recursive_directory_iterator(p, opts, ec)) {
if (ec) break;
if (!e.is_regular_file()) continue;
std::error_code fe;
auto sz = fs::file_size(e.path(), fe);
if (!fe) total += sz;
}
return total;
}
Here the per-file error_code fe is used correctly: a file that vanished or cannot be stat’ed is simply skipped, and the total reflects everything that could be measured. The directory-level error handling has the same range-for limitation as find_by_name, so for truly unattended use, combine this per-file pattern with the manual increment(ec) loop above. A second subtle point is e.is_regular_file() itself: the no-argument version can throw if the entry’s status cannot be determined, and the is_regular_file(ec) overload avoids that. Deciding in advance what “best effort” means (skip and continue, skip and report a count of failures, or stop) makes these functions far easier to reason about than sprinkling try/catch around the loop.
error_code overloads
std::error_code ec;
for (const auto& entry : fs::directory_iterator("maybe_missing", ec)) {
if (ec) {
std::cerr << ec.message() << '\n';
break;
}
std::cout << entry.path() << '\n';
}
Combine with skip_permission_denied so permission errors on single entries do not kill the walk. Still check ec around file_size/status on critical paths.
As with the recursive examples, this range-for only uses ec for the constructor. If "maybe_missing" does not exist, the constructor sets ec and produces an end iterator, so the loop body never runs and the if (ec) inside it never executes. The error check belongs before the loop: construct the iterator, test ec, then iterate. In my experience this is the single most common misunderstanding of the std::filesystem error-code API: code reviews show loops with an if (ec) inside that can never trigger, and the program quietly reports “no files” for a mistyped path.
The choice between exceptions and error_code is about what errors mean in your program. For a command-line tool that should stop and report on the first problem, the throwing versions with one try/catch around the walk are simpler and less error-prone. For a long-running service or an indexer that must survive unreadable folders, the error_code overloads with manual increment give per-entry control. Mixing the two in one function is where bugs creep in.
Traversal cost is system calls, not C++
- Avoid unnecessary deep recursion when you only need shallow scans.
- Avoid copying
entry.path(): it returns aconst path&, so bind it to a reference (const auto& p = entry.path();) rather than creating a newpathfor each use. - Syscalls:
directory_entrymay cache metadata—do not assume, but avoid redundant queries. - Parallelism: split by subtrees + thread pools—watch disk contention.
- Network drives: high latency—batch and tune cache policy.
Directory traversal is almost always I/O-bound, and the dominant cost is the number of system calls, not C++ overhead. The biggest win is usually asking directory_entry for metadata instead of calling the free functions: entry.is_directory() and entry.file_size() can use values the OS already returned while reading the directory (Windows’ FindNextFile returns type, size and timestamps; Linux readdir returns the type on most file systems), whereas fs::file_size(entry.path()) always performs a fresh stat. On a warm cache over a local SSD the difference is modest; over a network share, where each stat is a round trip, it can be the difference between seconds and minutes. Parallelizing by subtree helps on SSDs and network storage but can make things worse on a single spinning disk, where concurrent walks cause seeking.
Traps in real directory walks
One unreadable folder aborts the whole walk
Use skip_permission_denied and error_code when you cannot afford exceptions on every unreadable folder.
Symlink cycles in recursive traversal
Document traversal policy; add a visited set for risky trees.
Deleting files while the iterator is live
If files are added or removed in a directory after its iterator was created, the standard leaves it unspecified whether the iterator will see the change—so a delete-as-you-go loop may skip entries or report ones that no longer exist. Deleting the directory the iterator is currently inside can make the next increment fail. Collect paths first, then delete (or use fs::remove_all for whole trees).
FAQ
Q1: What is directory_iterator?
A: C++17 facility to enumerate directory entries.
Q2: Recursive walks?
A: recursive_directory_iterator.
Q3: Errors?
A: Exceptions or error_code overloads; skip_permission_denied helps. In a range-for, the error_code only covers construction—use increment(ec) in a manual loop to handle errors at every step.
Q4: Performance?
A: Reuse path, minimize syscalls, consider parallel subtree walks.
Q5: Modify during iteration?
A: Avoid—collect then mutate.
Related Articles
- std::filesystem Across Platforms: Paths, Directory Iteration, File Operations and Permissions
- C++ File Status
- C++ Filesystem overview
- C++ path