C++ Random Numbers in Practice: Seeding, Reproducibility, and Common Pitfalls
Key takeaways
The <random> API is easy to use and easy to use slightly wrong. This guide focuses on the practical problems: bias from rand() % n, under-seeded engines, results that change between compilers, broken benchmarks, and using a simulation-grade generator where security matters.
For an overview of the engines and distributions themselves, see C++ random and C++ distributions. This article is about using them correctly: the problems that show up in real programs even after you have moved away from rand(). The code samples assume #include <random>, the usual standard headers, and using namespace std; to keep them short.
Problems with rand()
// ❌ Legacy style
srand(time(0));
int x = rand() % 100; // 0-99
// Issues:
// 1. Not a uniform distribution
// 2. Low statistical quality
// 3. Not thread-safe
Each of these problems is concrete. Bias: rand() returns values from 0 to RAND_MAX. If RAND_MAX + 1 is not a multiple of 100, the lowest remainders occur one extra time. With MSVC, RAND_MAX is only 32767, so rand() % 100 makes 0–67 about 0.3% more likely than 68–99, and rand() % 50000 can never return anything above 32767 at all. Quality: the standard does not specify the algorithm, and many implementations use a simple linear congruential generator whose low bits follow short, visible patterns. Threads: rand() has one hidden global state. It is not required to be thread-safe, and even where it is, every thread contends for the same state.
srand(time(0)) adds its own problem: time has one-second resolution, so two processes started in the same second, such as parallel test workers or a set of containers launched together, produce identical sequences.
Modern random numbers (C++11)
#include <random>
int main() {
// Seed
random_device rd;
// Engine
mt19937 gen(rd());
// Distribution
uniform_int_distribution<> dis(1, 100);
// Generate
for (int i = 0; i < 10; i++) {
cout << dis(gen) << " ";
}
}
The three objects have separate jobs, and keeping them apart is what makes the library flexible. random_device is a source of entropy, usually the operating system’s randomness. It is slow and meant to be used rarely, typically just once for seeding. The engine (mt19937) is a fast, deterministic algorithm: given the same seed, it produces exactly the same sequence of raw 32-bit numbers. The distribution turns those raw numbers into values with the shape you need (an integer between 1 and 100, a normal value around a mean) without bias. You can swap any of the three independently, for example a fixed seed in tests, a different engine for speed, or a different distribution.
Random engines
// Mersenne Twister (recommended)
mt19937 gen32; // 32-bit
mt19937_64 gen64; // 64-bit
// Linear congruential (fast, lower quality)
minstd_rand gen;
// Subtract-with-carry
ranlux24 gen;
mt19937 is a reasonable default, but it is not free. Its state is 624 32-bit words, about 2.5 KB, so constructing one is noticeably more expensive than constructing a small LCG, and a million of them (for example, one per object) cost real memory. That is the reason behind the engine-per-call pitfall below. For simulation-heavy code, many developers today prefer small modern generators outside the standard library (such as PCG or xoshiro), which are faster and have a tiny state, but mt19937 is the portable choice that ships with every compiler.
Distributions
Uniform
// Integers
uniform_int_distribution<> intDis(1, 6); // dice
int dice = intDis(gen);
// Reals
uniform_real_distribution<> realDis(0.0, 1.0);
double x = realDis(gen);
Normal
normal_distribution<> normalDis(100.0, 15.0); // mean 100, stddev 15
double iq = normalDis(gen);
Other distributions
// Bernoulli (true/false)
bernoulli_distribution coinFlip(0.5); // 50%
bool result = coinFlip(gen);
// Binomial
binomial_distribution<> binDis(10, 0.5);
int heads = binDis(gen);
// Poisson
poisson_distribution<> poisDis(4.0);
int events = poisDis(gen);
// Exponential
exponential_distribution<> expDis(1.0);
double time = expDis(gen);
Note that uniform_int_distribution<> dis(1, 6) includes both ends (1 through 6), while uniform_real_distribution<> dis(0.0, 1.0) produces values in the half-open range [0, 1). Mixing up those conventions is a common off-by-one source when converting code from other languages.
The same seed does not give the same numbers everywhere
This is the most important portability fact about <random>, and it surprises many people. The standard specifies the engines exactly: mt19937 seeded with the default seed must produce 4123659995 as its 10,000th output on every conforming implementation. It does not specify the algorithms of the distributions. uniform_int_distribution, normal_distribution, and the others are implemented differently in libstdc++ (GCC), libc++ (Clang), and the Microsoft STL. The same seed with the same engine gives the same raw numbers, but different dice rolls, normals, and shuffles (std::shuffle is also implementation-specific).
In practice, this breaks anything that treats “same seed” as “same results” across platforms: procedural game worlds that must be identical on every client, replays, simulation results you want to reproduce on another machine, and unit tests with golden values. A typical symptom is a test with golden values that passes on the Linux CI and fails on Windows developer machines (or the other way around), with nothing wrong in the code except this assumption. It is a confusing failure the first time, because every developer checks the seed first and finds it identical. If you need cross-platform reproducibility, use the engine’s raw output and write the distribution yourself, or use a library that documents its algorithm.
Dice, random strings, shuffling, and weighted picks
A dice-roll histogram
#include <random>
#include <map>
int main() {
random_device rd;
mt19937 gen(rd());
uniform_int_distribution<> dice(1, 6);
map<int, int> histogram;
// Roll 10000 times
for (int i = 0; i < 10000; i++) {
int roll = dice(gen);
histogram[roll]++;
}
// Print histogram
for (const auto& [value, count] : histogram) {
cout << value << ": " << string(count / 100, '*') << endl;
}
}
With 10,000 rolls, each face should come up about 1,667 times, but you will see differences of a few dozen between faces in every run. That is expected: the standard deviation for each count is about 37. A histogram is a good sanity check for gross errors (a face that never appears, one that appears twice as often), but it cannot show subtle bias. That requires a statistical test and far more samples.
Random strings (not for secrets)
string generateRandomString(size_t length) {
static random_device rd;
static mt19937 gen(rd());
static uniform_int_distribution<> dis(0, 61);
const string chars =
"0123456789"
"ABCDEFGHIJKLMNOPQRSTUVWXYZ"
"abcdefghijklmnopqrstuvwxyz";
string result;
for (size_t i = 0; i < length; i++) {
result += chars[dis(gen)];
}
return result;
}
int main() {
cout << generateRandomString(10) << endl;
cout << generateRandomString(20) << endl;
}
This function is fine for test data or temporary file names. It is not safe for anything an attacker should not be able to guess: session IDs, password reset tokens, API keys. mt19937 is not a cryptographic generator. Its output is a linear function of its state, and after observing 624 consecutive 32-bit outputs, an attacker can reconstruct the full state and predict every future value. There are public tools that do exactly this. For security-relevant values, use the operating system’s CSPRNG (getrandom on Linux, BCryptGenRandom on Windows, arc4random_buf on BSD and macOS) or a library such as libsodium (randombytes_buf).
The static variables also make the function unsafe to call from several threads at once: they share one engine and one distribution without a lock. Use a thread_local engine instead (see Q5 below).
Shuffling a container
#include <algorithm>
int main() {
vector<int> v = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10};
random_device rd;
mt19937 gen(rd());
shuffle(v.begin(), v.end(), gen);
for (int x : v) {
cout << x << " ";
}
}
Weighted discrete choice
#include <random>
int main() {
random_device rd;
mt19937 gen(rd());
// Weights: 10%, 30%, 60%
discrete_distribution<> dis({10, 30, 60});
map<int, int> histogram;
for (int i = 0; i < 10000; i++) {
int choice = dis(gen);
histogram[choice]++;
}
for (const auto& [choice, count] : histogram) {
cout << "choice " << choice << ": " << count << " times" << endl;
}
}
discrete_distribution normalizes the weights, so {10, 30, 60} and {1, 3, 6} behave the same. It returns an index, not a value, so you use it to pick from an array of items, such as loot table entries or A/B test variants. Building the distribution is O(n) and each draw is fast, so construct it once when the weights are known, not before every draw.
Seeding
// Time-based (not reproducible, and identical for processes started in the same second)
mt19937 gen1(time(0));
// random_device (recommended for non-fixed seeds)
random_device rd;
mt19937 gen2(rd());
// Fixed seed (reproducible)
mt19937 gen3(12345);
// Seed sequence
seed_seq seq{1, 2, 3, 4, 5};
mt19937 gen4(seq);
// Seed sequence filled from random_device (more entropy)
seed_seq seq2{rd(), rd(), rd(), rd(), rd(), rd(), rd(), rd()};
mt19937 gen5(seq2);
mt19937 gen(rd()) is the most common seeding idiom, and it has a limitation worth knowing. rd() returns 32 bits, so this engine can start in at most 2³² ≈ 4.3 billion different states, even though its state space is 19,937 bits. For a dice game that is irrelevant. For a large Monte Carlo simulation that runs many instances, it means some instances will start from the same state and produce correlated results, and it means an attacker could simply try every possible seed. Filling a seed_seq with several rd() values, as in gen5, gives the engine far more starting states. seed_seq has known statistical weaknesses of its own, but it is much better than a single 32-bit seed.
For reproducibility, the usual approach is to generate a seed once, log it, and use it for that run. When a test fails or a simulation produces a strange result, you can re-run it with the logged seed and get the same sequence (on the same platform, see the portability note above).
Also be careful with random_device on unusual platforms. It is allowed to be deterministic if no entropy source is available. A well-known case: MinGW-w64 builds of GCC before version 9.2 had a random_device that returned the same sequence on every run, so programs “seeded from hardware” produced identical output every time they started. If your program runs on embedded targets or old toolchains, check that two runs really differ.
Engine construction, reseeding, and modulo bias
Creating an engine on every call
// ❌ Inefficient
int getRandom() {
random_device rd;
mt19937 gen(rd()); // created every time (slow)
uniform_int_distribution<> dis(1, 100);
return dis(gen);
}
// ✅ static storage
int getRandom() {
static random_device rd;
static mt19937 gen(rd());
static uniform_int_distribution<> dis(1, 100);
return dis(gen);
}
The first version is not only slow. Constructing an mt19937 initializes 2.5 KB of state, and each random_device call may be a system call, so a hot loop that calls getRandom() a million times does a million system calls. Worse, on platforms where random_device is weak or deterministic, every call returns the same “random” number. The static version fixes both, but it is not thread-safe (see the generateRandomString example). For multithreaded code, replace static with thread_local, which gives each thread its own engine and still initializes it only once per thread.
Reseeding inside a loop
// ❌ Same short sequence each iteration
for (int i = 0; i < 10; i++) {
mt19937 gen(12345); // always the same seed
cout << gen() << endl;
}
// ✅ Reuse one engine
mt19937 gen(12345);
for (int i = 0; i < 10; i++) {
cout << gen() << endl;
}
Biased ranges from rand() % n
// ❌ Biased
int x = rand() % 100; // not uniform
// ✅ Uniform distribution
uniform_int_distribution<> dis(0, 99);
int x = dis(gen);
Performance comparison
#include <chrono>
int main() {
const int N = 10000000;
// rand()
srand(time(0));
unsigned long long sum = 0; // use the results, or the loop may be optimized away
auto start = chrono::steady_clock::now();
for (int i = 0; i < N; i++) {
sum += rand();
}
auto end = chrono::steady_clock::now();
cout << "rand(): " << chrono::duration_cast<chrono::milliseconds>(end - start).count() << "ms" << endl;
// mt19937
random_device rd;
mt19937 gen(rd());
start = chrono::steady_clock::now();
for (int i = 0; i < N; i++) {
sum += gen();
}
end = chrono::steady_clock::now();
cout << "mt19937: " << chrono::duration_cast<chrono::milliseconds>(end - start).count() << "ms" << endl;
cout << "(checksum " << sum << ")" << endl; // prevents dead-code elimination
}
The sum variable is essential. A loop like int x = gen(); whose result is never used can be removed entirely by an optimizing compiler (with mt19937, whose code is fully visible in the header, this is quite likely), and the benchmark then reports close to 0 ms and proves nothing. Printing a value derived from every result forces the work to happen. steady_clock is used because high_resolution_clock can be an alias for the system clock, which may jump if the time is adjusted during the measurement.
Expect mt19937 to be at least as fast as rand() in a real measurement, and often faster, since rand() is an out-of-line library call that may take a lock. Performance is almost never a reason to stay with rand().
FAQ
Q1: rand() vs the random library?
A:
- rand(): legacy, poor quality.
- random: modern, higher quality, more flexible.
Q2: Must I always use random_device?
A: Use it for seeding. Generate numbers with an engine.
Q3: Which engine should I use?
A: In most cases mt19937 is a solid default.
Q4: How do I get reproducible sequences?
A: Use a fixed or logged seed. Remember that distributions differ between standard library implementations, so results are only reproducible on the same toolchain.
Q5: Is it thread-safe?
A: No. Give each thread its own engine, for example thread_local mt19937 gen{random_device{}()};. Distribution objects are cheap and can be created where needed.