C++17 std::filesystem in Practice: Directory Iteration, Copy and Remove, Symlinks and TOCTOU Races

Key takeaways

std::filesystem (C++17) is the standard way to work with directories, bulk file operations, and disk space in portable C++. This guide focuses on directory iteration, copy/remove semantics, symlink handling, and the race conditions and error-handling trade-offs that trip people up in production code.

What is std::filesystem, and what this article covers

std::filesystem, introduced in C++17 and modeled closely on Boost.Filesystem, gives you a portable way to enumerate directories, copy and remove files in bulk, query disk space, and inspect symlinks without shelling out to system() or writing #ifdef _WIN32 branches around dirent.h versus FindFirstFile. Before it existed, “list every file under this directory recursively” was a different fifteen lines of code on Windows than it was on Linux, and getting the edge cases right (trailing separators, UNC paths, symlink cycles) was a recurring source of bugs.

This article deliberately narrows its scope to directory iteration, copy/remove operations, disk space queries, and symlink handling — the operations you reach for when writing installers, backup tools, cache cleanup logic, or build systems. The fs::path type itself (how paths are decomposed, compared, and normalized) is covered in more depth in the companion post on std::filesystem::path, and stream-based file I/O (ifstream/ofstream, buffering, binary vs text mode) is covered separately in the file I/O guide. If you came here expecting a deep dive on path decomposition or <fstream> usage, those two posts are the better starting point; this one assumes you already have a fs::path in hand and want to know what you can safely do with it against a real filesystem.

#include <filesystem>

namespace fs = std::filesystem;

fs::path p = "/home/user/file.txt";
bool exists = fs::exists(p);

Basic operations

#include <filesystem>

namespace fs = std::filesystem;

// Check existence
if (fs::exists("file.txt")) {}

// Create directory
fs::create_directory("mydir");

// Delete file
fs::remove("file.txt");

// Copy
fs::copy("src.txt", "dst.txt");

These four calls look trivial, and for a one-off script they are. The trouble starts once this code runs unattended — in a CI job, a background service, or an installer — where the assumptions that were true on your development machine (the file exists, the directory is writable, nothing else touches this path concurrently) stop holding. The rest of this article is essentially an extended answer to “what happens when those assumptions break.”

Directory iteration: what actually happens under the hood

directory_iterator and recursive_directory_iterator are the two workhorses for walking a filesystem tree. They look like ordinary forward iterators you can drop into a range-based for, but their failure modes are not obvious from the interface alone.

void listFiles(const fs::path& dir) {
    for (const auto& entry : fs::directory_iterator(dir)) {
        std::cout << entry.path() << std::endl;
    }
}

void listFilesRecursive(const fs::path& dir) {
    for (const auto& entry : fs::recursive_directory_iterator(dir)) {
        std::cout << entry.path() << std::endl;
    }
}

A few things about this that are easy to get wrong the first time you write a directory-scanning tool:

Symlink loops. By default, recursive_directory_iterator does not follow symlinks into directories — it treats a symlink to a directory as a leaf entry unless you explicitly pass fs::directory_options::follow_directory_symlink. This is a deliberate safety default: without it, a symlink that points back up its own ancestor chain (a common outcome of a careless deployment script or a user creating ln -s .. loop) would send the iterator into infinite recursion. If you do pass follow_directory_symlink, you take on the responsibility of detecting cycles yourself — the standard library will not do it for you, and there is no built-in cycle guard once you opt in.

Permission-denied subdirectories. When the iterator hits a directory it cannot read (common in system directories, mounted volumes with restrictive ACLs, or directories another user owns), the default behavior is to throw fs::filesystem_error from the constructor or from operator++. If you want the scan to skip the unreadable directory and keep going — usually what you want for a best-effort backup or search tool — you need fs::directory_options::skip_permission_denied:

std::error_code ec;
for (auto it = fs::recursive_directory_iterator(
         dir, fs::directory_options::skip_permission_denied, ec);
     it != fs::recursive_directory_iterator();
     it.increment(ec)) {
    if (ec) {
        // log and continue, rather than letting the whole scan die
        std::cerr << "Skipping entry: " << ec.message() << std::endl;
        ec.clear();
        continue;
    }
    std::cout << it->path() << std::endl;
}

Notice that even with skip_permission_denied, you still pass an error_code and check it after every increment. The flag changes what the iterator tolerates when it descends into a directory; it does not guarantee every individual stat()-like call inside the loop body will succeed, especially if a file disappears mid-scan (see the TOCTOU section below). A directory tree walker written for a network share or a user-writable directory needs both the option flag and per-step error checking — I have seen “working” scan code that only had one of the two and would either throw and abort the whole scan on the first locked-down subdirectory, or silently swallow errors elsewhere and produce an incomplete file list without any indication that it was incomplete.

flowchart TD
    A["Start recursive_directory_iterator"] --> B{Entry readable?}
    B -- yes --> C{Is directory symlink?}
    C -- "no, or follow_directory_symlink not set" --> D[Descend / list entry]
    C -- "yes, and follow_directory_symlink set" --> E{Cycle back to ancestor?}
    E -- yes --> F[Infinite recursion risk — not guarded by the library]
    E -- no --> D
    B -- "no (permission denied)" --> G{skip_permission_denied set?}
    G -- yes --> H[Skip subtree, continue scan]
    G -- no --> I[Throw filesystem_error]

The exceptions vs. error_code trade-off

Almost every function in std::filesystem ships in two forms: a throwing overload and an overload that takes a trailing std::error_code& and reports failure by setting it instead of throwing.

// Throwing form
fs::remove("file.txt");

// error_code form — never throws for filesystem-related failures
std::error_code ec;
fs::remove("file.txt", ec);
if (ec) {
    std::cout << "Error: " << ec.message() << std::endl;
}

The existence of two parallel APIs for the same operation is not an accident or a legacy wart — it reflects a real design tension, and picking the wrong one for the situation produces code that is either needlessly fragile or silently wrong.

Use the throwing overload when failure means “something is fundamentally broken and continuing is pointless.” If your application creates its own working directory at startup and that create_directories call fails, there usually isn’t a sensible fallback — you want the exception to propagate up and abort startup with a clear message. Wrapping every one-off filesystem call in a local try/catch just to log and rethrow adds noise without adding safety.

Use the error_code overload when failure is an expected, routine outcome of the operation, not an exceptional one. Scanning a directory tree that includes entries you don’t control — user uploads, a mounted network share, another process’s temp directory — means you will hit permission-denied files, files that vanish between listing and processing, and broken symlinks. Throwing on each of these turns “skip and log one bad entry” into “unwind the entire scan,” which is both slower (stack unwinding has real cost when it happens thousands of times during a large scan) and harder to reason about, since you now need try/catch around every iteration step anyway. The error_code form lets the surrounding loop stay in control of what “one entry failed” means.

A rule of thumb I’ve settled on: reach for error_code overloads by default for anything running inside a loop over externally-supplied paths (uploads, scans, bulk operations), and reserve the throwing form for single, load-bearing operations at a well-defined point in your program (application startup/shutdown, a user explicitly clicking “delete this one file”). Mixing styles inconsistently within the same function is the actual anti-pattern, not choosing one style over the other.

TOCTOU: the race condition built into every filesystem check

TOCTOU (time-of-check-to-time-of-use) is not a corner case in filesystem programming — it’s the default behavior of every filesystem API, because the filesystem is shared, mutable state that other processes, other threads, and the OS itself can change at any moment between your check and your action.

// This is not atomic. The file can be deleted, replaced,
// or have its permissions changed between the two calls.
if (fs::exists(path)) {
    fs::remove(path);   // may still fail, or worse, remove a *different* file
                         // that now happens to occupy that path
}

fs::exists() answers “was this true at the moment I asked” — nothing more. There is no way to turn exists() + remove() into a single atomic “remove if it exists” operation using std::filesystem alone, because the standard filesystem library is a thin, portable wrapper over OS calls (stat, unlink, CreateFile, DeleteFile) that were never atomic across a check-then-act pair to begin with. Some POSIX-specific fallbacks exist outside the standard (O_EXCL, renameat2 with RENAME_NOREPLACE, opening a file descriptor and doing subsequent operations relative to that descriptor instead of the path), but the standard library gives you none of them.

The practical fix is almost always to invert the logic: attempt the operation and handle its failure, instead of checking first and hoping nothing changes before you act.

// Better: attempt, then interpret the failure
std::error_code ec;
fs::remove(path, ec);
if (ec && ec != std::errc::no_such_file_or_directory) {
    // a real error, not "someone already removed it"
    std::cerr << "Failed to remove " << path << ": " << ec.message() << std::endl;
}

This version treats “the file was already gone” as a non-error, which is usually the correct semantics for cleanup code — you wanted the file gone, and it’s gone, regardless of who removed it.

I ran into this the hard way in a log-rotation tool: it listed files older than a retention window with directory_iterator, checked fs::exists() on each one to skip anything a concurrent process might have already rotated, and then called fs::remove(). Under normal load this worked fine for months. Under heavier load, with a second rotation job kicked off by a retry after a transient timeout, the two processes would occasionally race on the same file: the exists() check would pass in both processes almost simultaneously, one process would remove the file, and the second process’s subsequent remove() call would throw filesystem_error because the file was already gone — which, since the code treated any exception from remove() as a hard failure, marked the entire rotation batch as failed and paged someone at 3 a.m. for a “failure” that was actually the system behaving exactly as intended. The fix was exactly the pattern above: drop the exists() check entirely, call remove() with an error_code, and only treat codes other than “no such file” as real errors.

Copying and removing: options, defaults, and the remove_all trap

fs::copy and fs::copy_file take a copy_options bitmask that controls what happens on conflicts and how symlinks and directories are treated. The default options are conservative — plain fs::copy("src.txt", "dst.txt") throws if dst.txt already exists, which surprises people who expect it to behave like cp.

void backupFiles(const fs::path& src, const fs::path& dst) {
    if (!fs::exists(dst)) {
        fs::create_directories(dst);
    }

    for (const auto& entry : fs::directory_iterator(src)) {
        fs::path dstPath = dst / entry.path().filename();

        if (fs::is_regular_file(entry)) {
            fs::copy_file(entry.path(), dstPath,
                          fs::copy_options::overwrite_existing);
        }
    }
}

Note that this function only copies regular files one directory level deep — it does not recurse, and it silently skips symlinks and subdirectories. That’s a reasonable choice for a flat backup of one directory’s files, but it’s worth stating explicitly in a comment at the call site, because “backup” reads as if it should be comprehensive. For a full recursive copy, fs::copy accepts copy_options::recursive combined with overwrite_existing, and separately copy_options::copy_symlinks or skip_symlinks to control whether a symlink inside the source tree gets followed (copying the target’s contents) or replicated as a symlink in the destination.

fs::remove and fs::remove_all look like a matched pair — one for files, one for directories — but the asymmetry in what they do to symlinks is where the real danger lives:

// remove() never follows a symlink: it removes the link itself.
// remove_all() recurses into what a path *resolves to*, so if any
// component of that path is a symlink pointing somewhere unexpected,
// you can delete far more than you intended.
fs::remove_all("build/output");

remove_all resolves the path you give it and then deletes everything under the resulting directory tree. If build/output is itself a symlink — left behind by a previous build configuration, a Docker volume mount, or a developer’s local convenience symlink into a shared folder — remove_all deletes the contents of whatever that symlink points to, not just the symlink. Combined with relative path resolution (the current working directory of the process calling remove_all might not be what you assume, particularly in a script invoked from a Makefile, a CI runner, or a cron job with an unexpected cwd), this is a well-known way to lose data that has nothing to do with your project.

I had exactly this happen with a “clean” target in an internal build script: fs::remove_all(project_root / "build") was meant to wipe a scratch build directory before a fresh compile. On one developer’s machine, build had at some point been replaced with a symlink to a much larger shared cache directory outside the repository — done deliberately, to avoid duplicating gigabytes of intermediate artifacts across several checkouts of the same repo. The clean target had no way to know that, followed the symlink, and recursively deleted the shared cache directory that three other checkouts were also pointing at. Nothing in the C++ code was wrong per se — remove_all did exactly what it’s documented to do — but the assumption that “the build directory is just a build directory” turned out to be false in one environment. The lesson that stuck: before calling remove_all on any path that wasn’t just created by the same process in the same run, check fs::is_symlink() on that specific path first (not exists(), which follows symlinks and tells you nothing about whether the entry itself is a link) and refuse to proceed — or resolve with fs::read_symlink() and log exactly what’s about to be deleted — rather than trusting a directory name to mean what it usually means.

// Defensive remove_all: refuse to recurse through an unexpected symlink
std::error_code ec;
if (fs::is_symlink(target)) {
    std::cerr << "Refusing to remove_all through a symlink: " << target
              << " -> " << fs::read_symlink(target, ec) << std::endl;
    return;
}
auto removedCount = fs::remove_all(target, ec);
if (ec) {
    std::cerr << "remove_all failed: " << ec.message() << std::endl;
}

remove_all also has a return value that’s easy to overlook: it returns the number of files and directories removed (as std::uintmax_t), and on the error_code overload it returns static_cast<std::uintmax_t>(-1) on failure rather than throwing. Checking only the error_code and ignoring the count, or vice versa, will occasionally miss a partial failure where some entries were removed before an error occurred partway through.

Disk space queries

fs::space() reports capacity, free space, and available space for the filesystem that contains a given path, returning an fs::space_info struct with capacity, free, and available members (all in bytes).

#include <filesystem>
#include <iostream>

namespace fs = std::filesystem;

void checkDiskSpace(const fs::path& p) {
    std::error_code ec;
    fs::space_info info = fs::space(p, ec);
    if (ec) {
        std::cerr << "Could not query space for " << p << ": " << ec.message() << std::endl;
        return;
    }

    std::cout << "Capacity:  " << info.capacity  << " bytes\n";
    std::cout << "Free:      " << info.free      << " bytes\n";
    std::cout << "Available: " << info.available << " bytes\n";
}

The distinction between free and available matters and is frequently ignored: free is the total unused space on the filesystem, while available is what’s actually usable by an unprivileged process — on POSIX systems, this accounts for space reserved for the root user (typically 5% on ext4 by default), so available can be meaningfully smaller than free on a nearly-full disk. If you’re writing a pre-flight check before a large write (“do we have enough space to extract this archive?”), compare against available, not free — otherwise your check can pass while the actual write later fails with ENOSPC because the remaining space was reserved and off-limits to your process.

Also worth knowing: fs::space() reports figures for the filesystem mounted at the given path, which is not necessarily the filesystem your own working directory lives on. On systems with multiple mounted volumes, a network share, or a container with bind-mounted volumes, checking space on / tells you nothing about the space available at /mnt/data if that’s a separate mount. Always query space on the actual target path you intend to write to, and be aware that on network filesystems the reported numbers can be stale or approximate depending on the underlying protocol.

A recurring source of confusion is that most type-checking functions (fs::is_directory, fs::is_regular_file, fs::exists) follow symlinks by default — they report on what the symlink points to, not the symlink entry itself.

fs::path p = "file.txt";

if (fs::is_regular_file(p)) {
    std::cout << "Regular file" << std::endl;
}

if (fs::is_directory(p)) {
    std::cout << "Directory" << std::endl;
}

if (fs::is_symlink(p)) {
    std::cout << "Symbolic link" << std::endl;
}

Note that fs::is_symlink(p) and fs::is_regular_file(p) are not mutually exclusive in the way the code above implies: if p is a symlink pointing at a regular file, both is_symlink(p) and is_regular_file(p) return true (the latter because it follows the link). If you want to distinguish “this exact path entry is a symlink” from “this path, after resolving symlinks, refers to a regular file,” you need fs::symlink_status(p) instead of fs::status(p) — symlink_status does not follow the final symlink component, so is_regular_file(fs::symlink_status(p)) correctly returns false for a symlink even if its target is a regular file. This distinction matters most in exactly the security-sensitive code paths discussed above: anywhere you’re deciding whether to recurse into, copy, or delete something, checking the resolved status when you meant to check the link itself (or vice versa) is how symlink-based surprises happen.

A dangling symlink (pointing at a target that no longer exists) makes this sharper still: fs::exists(p) returns false for a dangling symlink (because exists follows the link and the target isn’t there), while fs::is_symlink(p) still correctly returns true. Code that does if (!fs::exists(p)) { create it } without first checking is_symlink can end up creating a brand-new file at a location a stale symlink was already occupying, leaving the original symlink target-less and the new file sitting somewhere the symlink no longer even points to logically.

Four filesystem mistakes that reach production

Ignoring exceptions from calls you expect to sometimes fail

// Fragile: any transient failure aborts whatever loop this is in
fs::remove("file.txt");

// Better for expected failures: error_code and explicit handling
std::error_code ec;
fs::remove("file.txt", ec);
if (ec) {
    std::cout << "Error: " << ec.message() << std::endl;
}

Hard-coded path separators

// Platform-specific and fragile on Windows vs POSIX
fs::path p = "dir\\file.txt";

// Portable: let path's operator/ build the separator for the current platform
fs::path p2 = fs::path("dir") / "file.txt";

Assuming a relative path resolves the way you expect

fs::path rel = "file.txt";
fs::path abs = fs::absolute(rel);

fs::absolute() resolves against the process’s current working directory at the time of the call — not the directory the executable lives in, not the directory a config file was loaded from. In a long-running service where the working directory can be set by whatever launched the process (a systemd unit, a Docker WORKDIR, a parent shell), relying on relative paths for anything security- or correctness-sensitive is asking for the wrong file to be touched. Resolve important paths once, explicitly, against a known base directory rather than trusting cwd.

fs::remove() on a non-empty directory

// Fails: remove() only removes empty directories or single files
fs::remove("dir");

// Recursively removes the directory and everything under it —
// see the remove_all warnings above before using this on any
// path you didn't just create yourself
fs::remove_all("dir");

Symptoms and their usual causes

  • “filesystem_error: Permission denied” during a recursive scan — add fs::directory_options::skip_permission_denied and check error_code after every increment; the option alone does not suppress every possible error.
  • A cleanup job intermittently throws on files that “should” exist — this is almost always a TOCTOU race with a concurrent writer or another instance of the same job; switch to attempt-then-check-error-code instead of check-then-act.
  • remove_all deleted more than expected — check whether any path component was a symlink; remove_all follows the resolved path, and is_symlink() before deletion is the cheapest safeguard.
  • A “sufficient disk space” pre-check passes but the write still fails with ENOSPC — you likely compared against space_info::free instead of space_info::available, or checked space on the wrong mount point.
  • is_regular_file() returns true for something you expected to be treated as a link — you called fs::status() (follows symlinks) where you needed fs::symlink_status() (does not).

FAQ

Q1: Which standard added filesystem?

A: C++17. It is the standard library for filesystem operations, modeled on the earlier Boost.Filesystem library.

Q2: Is it portable?

A: Yes. Windows, Linux, and macOS are supported by mainstream toolchains, though some behaviors around permissions, symlinks, and reserved disk space differ by platform since they ultimately map onto different OS primitives.

Q3: How do I handle errors?

A: Use try/catch for filesystem_error on single, load-bearing operations where failure should stop the program. Use the std::error_code overloads inside loops over externally-supplied or untrusted paths, where failure of one entry is expected and shouldn’t unwind the whole operation.

Q4: Performance?

A: Every filesystem call is a system call, and recursive_directory_iterator in particular can generate a large number of them on deep trees. Avoid redundant exists()/is_directory() calls right before an operation that will report the same failure itself — the extra syscalls add latency without eliminating the underlying race.

Q5: Path separators?

A: Use the / operator on path for portability instead of hard-coding \ or /.

Q6: Is fs::exists() followed by an action safe?

A: No — see the TOCTOU section above. Prefer attempting the operation and inspecting its error_code over checking first and acting second.

Q7: Where can I learn more?

A: