std::call_once and once_flag: Thread-Safe One-Time Initialization vs static Locals

Key takeaways

std::call_once runs a callable exactly once across all threads, blocks the others until it finishes, and lets a later call retry if it throws. How it compares with function-local statics (which cover most cases), and the pitfalls: recursive calls on the same flag, ignored arguments on later calls, and platform bugs with throwing callables.

What call_once guarantees

std::call_once(flag, f, args...) runs f(args...) exactly once for a given std::once_flag, no matter how many threads call it at the same time. The precise guarantees from the standard are worth knowing, because they are what make it safe:

  • Exactly one call is active: if several threads arrive together, one runs f and the others block until it finishes.
  • When that call returns normally, the flag is done. Every later call_once on the flag returns immediately, and all writes made by f are visible to threads that return from call_once (there is a happens-before relationship, so no extra synchronization is needed to read the initialized data).
  • If f throws, the exception propagates to that caller, the flag is not marked done, and the next call on the flag picks a new thread to try again.
#include <iostream>
#include <mutex>

std::once_flag flag;

void init() {
    std::cout << "Initialization" << std::endl;
}

void func() {
    std::call_once(flag, init);  // Only once
}

std::once_flag has a constexpr default constructor, so a namespace-scope flag is initialized at compile time and is never subject to the static initialization order problem. It is neither copyable nor movable, since two copies of “has this happened yet” would defeat the purpose.

std::once_flag flag1;
// std::once_flag flag2 = flag1;              // error: copy constructor is deleted
// std::once_flag flag3 = std::move(flag1);   // error: no move either

Why not a mutex and a bool?

// Correct, but locks on every call
std::mutex mtx;
bool initialized = false;

void init() {
    std::lock_guard<std::mutex> lock(mtx);
    if (!initialized) {
        // Initialize
        initialized = true;
    }
}

// call_once: same guarantee, cheap after the first call
std::once_flag flag;

void init() {
    std::call_once(flag, []() {
        // Initialize (only once)
    });
}

The mutex version is correct. Its cost is that every call, forever, takes and releases a lock, and under contention every thread serializes on it just to discover that nothing needs doing. The tempting “optimization” is double-checked locking: test initialized without the lock, and only lock if it is false. With a plain bool that is a data race and undefined behavior. It can appear to work for years, then fail on a weaker memory model (ARM) or after the compiler reorders the flag write before the initialization writes, so another thread sees initialized == true and reads a half-built object. Doing it correctly needs std::atomic<bool> with acquire/release ordering, which is exactly what call_once implements for you.

Conceptually, the implementation looks like this:

// Conceptual operation, not real code
void call_once(std::once_flag& flag, Callable&& func) {
    if (flag.done_acquire_load()) {
        return;                      // fast path: one atomic load
    }
    lock();
    if (!flag.done()) {
        func();                      // if this throws, flag stays "not done"
        flag.mark_done_release();
    }
    unlock();
}

Singletons, Config Loading, and Passing Arguments

A lazily created singleton

class Singleton {
    static std::once_flag initFlag;
    static Singleton* instance;

    Singleton() {
        std::cout << "Singleton created" << std::endl;
    }

public:
    static Singleton& getInstance() {
        std::call_once(initFlag, []() {
            instance = new Singleton();
        });
        return *instance;
    }
};

std::once_flag Singleton::initFlag;
Singleton* Singleton::instance = nullptr;

The object is deliberately never deleted. That leak is a common choice for singletons, because it means the instance is still usable from other static destructors during shutdown. A singleton held by value would be destroyed at exit in reverse order of construction, and any code that touches it from another static object’s destructor afterwards uses a dead object. If the class holds resources that must be released (files that must be flushed), give it an explicit shutdown function instead of relying on the destructor.

Loading configuration once

std::once_flag configFlag;
Config config;

void loadConfig() {
    std::cout << "Loading config" << std::endl;
    config = Config::load("config.json");
}

Config& getConfig() {
    std::call_once(configFlag, loadConfig);
    return config;
}

Here the flag and the storage are separate globals. That works, but config is default-constructed during static initialization before anyone asks for it, so Config must have a cheap default constructor. The function-local static version shown later avoids that.

Passing arguments and member functions

std::once_flag flag;
int value = 0;

void func() {
    std::call_once(flag, []() {        // globals need no capture
        value = expensiveComputation();
    });
    std::cout << "Value: " << value << std::endl;
}

void init(int x, const std::string& s) {
    std::cout << x << ", " << s << std::endl;
}

std::once_flag flag2;
void func2() {
    std::call_once(flag2, init, 42, "Hello");   // extra arguments are forwarded
}

class MyClass {
    std::once_flag flag_;
    void init() { std::cout << "Initialization" << std::endl; }
public:
    void process() {
        std::call_once(flag_, &MyClass::init, this);   // like std::invoke
    }
};

The callable is invoked as if by std::invoke, so member function pointers take the object as the first argument. A once_flag member is the case where call_once has no simple static-local equivalent: each MyClass object initializes itself once, on first use, independently of other objects. Note that the flag makes MyClass non-copyable and non-movable unless you write those operations yourself.

Exceptions and retry

std::once_flag flag;
int attempt = 0;

void init() {
    attempt++;
    if (attempt < 3) {
        throw std::runtime_error("Retry");
    }
    std::cout << "Success" << std::endl;
}

void func() {
    try {
        std::call_once(flag, init);
    } catch (...) {
        // flag is still "not done": the next call_once runs init again
    }
}

A throwing call leaves the flag untouched, so retrying is automatic: call call_once again. This is useful for initialization that depends on something external, such as opening a device or reading a file that may not exist yet. attempt in this example is only safe because call_once runs one attempt at a time; it needs no atomic even with many threads, since each attempt happens-before the next.

The exceptional path is the least-tested part of many implementations. For years, libstdc++ implemented call_once on top of pthread_once, which knows nothing about C++ exceptions, and GCC bug 66146 documents targets on which a throwing callable left the flag in a state where the next call hung forever instead of retrying. If your design depends on retry after an exception, write a test that exercises it on the exact compiler, standard library and platform you ship; or avoid the question by catching inside the callable and recording failure in a separate variable.

A retry loop around it looks like this:

class RetryableInit {
    std::once_flag flag_;
    int attempts_ = 0;

    void tryInit() {
        attempts_++;
        if (attempts_ < 3) {
            throw std::runtime_error("Init failed, retrying");
        }
        std::cout << "Init successful\n";
    }

public:
    bool initialize() {
        try {
            std::call_once(flag_, [this]() { tryInit(); });
            return true;
        } catch (const std::exception& e) {
            std::cerr << e.what() << '\n';
            return false;
        }
    }
};

RetryableInit init;
while (!init.initialize()) {
    std::this_thread::sleep_for(std::chrono::seconds(1));
}

call_once vs function-local statics

Since C++11, initialization of a block-scope static variable is thread-safe. The compiler emits a guard variable and calls into the runtime (__cxa_guard_acquire / __cxa_guard_release on the Itanium ABI) that does what call_once does: one thread initializes, the others wait, and a throwing initializer is retried on the next call.

// call_once
std::once_flag flag;
Resource* resource = nullptr;

Resource& getResource() {
    std::call_once(flag, []() {
        resource = new Resource();
    });
    return *resource;
}

// Function-local static: same guarantees, less code
Resource& getResource() {
    static Resource resource;
    return resource;
}

That is why “retry after an exception” is not a reason to prefer call_once; both have it. The real differences are structural:

SituationBetter fit
One object, created on first useFunction-local static
Per-object lazy initialization (flag as a class member)call_once
A one-time action that produces no object (register a codec, call a C library’s global init)call_once
Arguments known only at the first call sitecall_once
Code compiled with thread-safe statics disabled (-fno-threadsafe-statics, or MSVC /Zc:threadSafeInit-)call_once

With MSVC, thread-safe statics have been the default since Visual Studio 2015; older compilers generated unguarded initialization, which is where much of the older advice to avoid local statics in multithreaded code comes from.

Recursion Deadlocks and Other call_once Traps

Recursion on the same flag deadlocks. If f (directly or through other code) calls call_once on the flag it is currently running under, the inner call waits for the outer call to finish, which never happens. The standard makes it undefined; in practice the thread hangs. This tends to show up with layered lazy initialization, where getLogger() initializes a config and the config loader logs a warning through getLogger(). The same cycle with function-local statics also deadlocks or, on some implementations, throws __gnu_cxx::recursive_init_error.

Arguments on later calls are ignored. Only the arguments of the call that actually ran matter. A lazy wrapper like the one below looks general, but get("other-host", 1) after the first call silently returns the object built with the first arguments:

template<typename T>
class LazyInit {
    std::once_flag flag_;
    std::unique_ptr<T> instance_;

public:
    template<typename... Args>
    T& get(Args&&... args) {
        std::call_once(flag_, [this, &args...]() {
            instance_ = std::make_unique<T>(std::forward<Args>(args)...);
        });
        return *instance_;
    }
};

LazyInit<Database> db;
db.get("localhost", 5432).query("SELECT * FROM users");
db.get("elsewhere", 1);   // returns the localhost instance, arguments unused

If different call sites might pass different values, split the API into a configure(args) step and an argument-free get(), or assert that later arguments match.

A flag cannot be reset. This makes code built on global once_flags awkward to unit test: the first test that triggers initialization fixes the state for every other test in the same process. Where tests need a fresh start, make the flag a member of an object the test creates, rather than a global.

Blocking is real. Threads that arrive while f runs wait for it. If f does a slow network call, every thread that needs the resource stalls for that long, and if f waits on something that one of those blocked threads was supposed to provide, the program deadlocks.

The mistake I have run into most with this API is the first one, in a form where the recursion was not obvious: a callback registered during initialization, invoked synchronously by the library being initialized, which called back into the same accessor. The hang shows up as all threads waiting inside call_once or __cxa_guard_acquire, and a stack trace of the stuck thread shows the accessor twice on the same stack, which is the quickest way to confirm it.

Cost After the First Call

After the flag is done, call_once is an acquire load and a branch, comparable to the guard check of a function-local static. It is not free in the sense of being optimized away, and in a tight loop it still prevents some optimizations because the compiler cannot prove the initialized data does not change. Call it (or the accessor) once before the loop and keep a reference.