std::locale in C++: Number, Money and Date Formatting, imbue vs Global, and UTF-8 Caveats
Key takeaways
A std::locale is a bundle of facets (numpunct, moneypunct, time_put, ctype, ...) that streams consult when they format and parse. Named locales like en_US.UTF-8 depend on the OS and can throw; custom facets do not. Prefer imbue on one stream over locale::global, keep machine-readable formats in the classic locale, and do not expect std::locale to understand UTF-8 characters.
What a locale actually is
std::locale is C++‘s mechanism for culture-dependent formatting and parsing: the character used as decimal point, how digits are grouped, the currency symbol, how dates are written, and which characters count as letters. Internally, a locale is an immutable, reference-counted bundle of facets, one object per category:
| Facet | Used for |
|---|---|
std::numpunct<char> | Decimal point, thousands separator, digit grouping |
std::num_put / std::num_get | Writing and reading numbers in streams |
std::moneypunct, std::money_put | Currency symbol and money layout (std::put_money) |
std::time_put, std::time_get | Dates and times (std::put_time, std::get_time) |
std::ctype<char> | Classification (isalpha, isupper) and case conversion |
std::codecvt | Conversion between internal and external character encodings |
Every stream holds a locale. When you write std::cout << 1234.5, the stream’s num_put facet asks its numpunct facet which decimal point to use. Change the stream’s locale, and the same line prints something else. That is both the feature and the source of most locale bugs: formatting that looks fixed in the code actually depends on state held elsewhere.
There are two ways to change which locale applies, and they have very different reach:
stream.imbue(loc)changes the locale of one stream.std::locale::global(loc)changes the process-wide default that every later default-constructedstd::localecopies. Iflochas a name, it also callsstd::setlocale(LC_ALL, name), which changes the C library’s behavior forprintf,strtod,isalphafrom<cctype>, and so on.
A detail that surprises many people: locale::global does not change streams that already exist. std::cout is constructed before main with the classic locale, and it keeps it. That is why code that sets a global locale usually follows it with std::cout.imbue(std::locale());.
#include <iostream>
#include <locale>
int main() {
std::cout << std::locale().name() << '\n'; // "C" at program start
std::locale::global(std::locale("")); // user's environment locale, may throw
std::cout.imbue(std::locale()); // cout does not pick it up by itself
std::cout << 1234567.5 << '\n';
}
std::locale("") means “whatever the environment says” (the LANG and LC_* variables on POSIX systems). It is the right choice for a command-line tool that formats output for the user running it, and the wrong one for anything that writes files other programs will read.
Named locales depend on the operating system
A named locale such as "en_US.UTF-8" or "de_DE.UTF-8" is not part of the C++ library; it is loaded from the operating system, and constructing one that does not exist throws std::runtime_error. With libstdc++ the message is not very descriptive:
try {
std::locale loc("en_US.UTF-8");
} catch (const std::runtime_error& e) {
std::cerr << e.what() << '\n';
// locale::facet::_S_create_c_locale name not valid
}
This is the portability problem I hit first with std::locale. Code that works on a developer’s Linux laptop throws in a minimal Docker image, where only the C and POSIX locales are installed until you generate others (locale -a lists what is available). On MinGW, libstdc++ uses its “generic” locale model, which supports only "C" and "POSIX", so the constructor above throws even on a fully configured Windows machine; the message shown is from MinGW g++ 10.3. MSVC’s standard library uses Windows locale names, which follow a different naming scheme. Treat every named locale as optional: catch the exception, log it, and fall back to std::locale::classic().
Custom facets: formatting without OS locales
When what you want is a specific format rather than “whatever the user’s system does”, you do not need a named locale at all. Derive from std::numpunct<char>, override the virtual do_ functions, and combine the facet with an existing locale:
#include <iomanip>
#include <iostream>
#include <locale>
#include <sstream>
#include <string>
struct GroupedThousands : std::numpunct<char> {
protected:
char do_thousands_sep() const override { return ','; }
std::string do_grouping() const override { return "\3"; } // groups of three digits
};
struct CommaDecimal : std::numpunct<char> {
protected:
char do_decimal_point() const override { return ','; }
char do_thousands_sep() const override { return '.'; }
std::string do_grouping() const override { return "\3"; }
};
int main() {
std::cout << "default: " << 1234567 << ' ' << 1234.5 << '\n';
std::cout.imbue(std::locale(std::cout.getloc(), new GroupedThousands));
std::cout << "grouped: " << 1234567 << ' ' << 1234.5 << '\n';
std::ostringstream de;
de.imbue(std::locale(std::locale::classic(), new CommaDecimal));
de << std::fixed << std::setprecision(2) << 1234567.25;
std::cout << "comma decimal: " << de.str() << '\n';
}
default: 1234567 1234.5
grouped: 1,234,567 1,234.5
comma decimal: 1.234.567,25
This runs identically on every platform because nothing is loaded from the OS. The new is not a leak: the locale takes ownership of the facet and deletes it when the last locale referring to it goes away (facets start with a reference count of zero unless you pass a nonzero value to the constructor). The grouping string is a sequence of small integers, not digit characters: "\3" means “groups of three”, which is why writing "3" (the character with value 51) produces no grouping you would recognize.
Parsing numbers with a locale
Locales apply in both directions. The same text parses differently depending on the stream’s locale, and that is where silent data corruption comes from:
std::istringstream classic("1,234.56");
double v = 0;
classic >> v; // v == 1, stops at ','
std::istringstream grouped("1,234.56");
grouped.imbue(std::locale(std::locale::classic(), new GroupedThousands));
double w = 0;
grouped >> w; // w == 1234.56
std::istringstream bad("12,34.5");
bad.imbue(std::locale(std::locale::classic(), new GroupedThousands));
double z = 0;
bad >> z; // failbit set: grouping does not match "\3"
The classic-locale read is the dangerous one. It does not fail; it reads 1, leaves ,234.56 in the stream, and the next extraction reads garbage or fails somewhere unrelated. Always check the stream state after extraction, and check that the whole input was consumed (stream >> std::ws followed by stream.eof()) if you require that.
The mistake that shows up most often in practice runs the other way. A program writes a CSV or configuration file with the user’s locale active, and in a German locale 3.14 becomes 3,14. A comma-separated file now has an extra column, and a reader running in the classic locale parses 3. The rule I follow is that anything a program writes for another program uses the classic locale (or std::to_chars / std::from_chars, which ignore locales entirely and are also much faster), and only text shown to a human uses the user’s locale.
Money
std::cout.imbue(std::locale("en_US.UTF-8")); // may throw, see above
std::cout << std::showbase << std::put_money(123456) << '\n'; // $1,234.56 with glibc
std::put_money takes a long double or a string of digits in the smallest currency unit. 123456 means 123456 cents, which is why it prints as 1,234.56. Without std::showbase, the currency symbol is omitted. For currencies without minor units the same value would print as 123,456, which is a good reason to keep amounts as integers of minor units in your program and never as floating-point major units.
Dates and times
#include <chrono>
#include <ctime>
#include <iomanip>
#include <iostream>
int main() {
std::time_t t = std::chrono::system_clock::to_time_t(std::chrono::system_clock::now());
std::tm local{};
#ifdef _WIN32
localtime_s(&local, &t);
#else
localtime_r(&t, &local);
#endif
std::cout << std::put_time(&local, "%Y-%m-%d %H:%M:%S") << '\n'; // fixed format
std::cout << std::put_time(&local, "%c") << '\n'; // locale's preferred format
}
Only some conversion specifiers depend on the locale: %c, %x, %X, %a/%A (weekday names), %b/%B (month names) and %p. Numeric ones like %Y-%m-%d are the same everywhere, which makes them the right choice for logs. Note the use of localtime_r / localtime_s: plain std::localtime returns a pointer to a shared static buffer and is not safe to call from several threads at once. In C++20, std::chrono gains time zones and std::format support for dates, which avoid both problems where available.
Character classification and UTF-8
std::locale loc = std::locale::classic();
char c = 'A';
bool alpha = std::isalpha(c, loc); // true
char lower = std::tolower(c, loc); // 'a'
The <locale> overloads take the locale explicitly and accept any char, unlike the <cctype> versions, which have undefined behavior for negative values other than EOF (a real concern for bytes above 127 when char is signed).
What neither version can do is understand UTF-8. std::ctype<char> classifies one byte at a time, and a character like é is two bytes in UTF-8, neither of which is a letter on its own. toupper on each byte of a UTF-8 string will either do nothing or corrupt it. The standard’s UTF-8 conversion facets (std::codecvt_utf8, std::wstring_convert) were deprecated in C++17 because they were underspecified and error-prone. For case mapping, collation, word boundaries or normalization, use ICU or a library built on it; std::locale is about formatting numbers and dates, not about Unicode text.
Performance
Constructing a named locale is not free: the library has to find and load the locale data from the system. Build locale objects once and reuse them:
// Construct once, share: locales are cheap to copy (reference counted)
const std::locale userLocale = [] {
try { return std::locale(""); } catch (const std::runtime_error&) { return std::locale::classic(); }
}();
for (int i = 0; i < 1000; ++i) {
std::ostringstream oss;
oss.imbue(userLocale);
oss << i;
}
Copying and imbuing a locale is a reference-count operation. The stream itself is the more expensive part of that loop; if you are formatting many numbers for machines, std::to_chars into a buffer avoids both the stream and the locale.
Global locale in multi-threaded servers
std::locale::global is shared, mutable, process-wide state. In a server that handles users from different regions, switching the global locale per request is a real bug: a request on another thread that constructs a default locale or calls printf in the middle of the switch picks up someone else’s settings. Since a named global locale also calls setlocale, C library functions used by third-party code change too. Keep the global locale fixed (usually classic) for the lifetime of the process, and give each request its own std::locale object that you imbue into the streams formatting that request’s output.