C++ Random Distributions: Choosing and Using uniform, normal, Poisson, and discrete

Key takeaways

A distribution turns an engine's raw bits into values with a specific shape. Choosing the right one is a modeling question: uniform for equal chances, normal for measurements around a mean, Poisson and exponential for arrivals, discrete for weighted picks. This guide covers what each models and the usage details that trip people up.

What is a distribution?

An engine such as std::mt19937 produces raw, uniformly distributed 32-bit integers. A distribution turns those raw numbers into values with the shape you need: an integer between 1 and 6, a real number with a bell curve around 100, true 70% of the time.

#include <random>
std::mt19937 gen{std::random_device{}()};
// Uniform distribution
std::uniform_int_distribution<> uniform{1, 6};
int dice = uniform(gen);
// Normal distribution
std::normal_distribution<> normal{0.0, 1.0};
double value = normal(gen);

Why not just compute gen() % 6 + 1? Because turning raw bits into an unbiased range is surprisingly subtle. The modulo creates a small bias whenever the engine’s range is not a multiple of 6, and producing a correct normal or Poisson value requires specific algorithms. Distributions encapsulate those algorithms, so code states what it wants rather than how to compute it.

The practical skill is choosing the right one. That is a modeling decision, not a C++ one, and a wrong choice gives results that look plausible but are systematically off. The sections below are organized around what each distribution models.

Uniform distributions: every value equally likely

#include <random>
std::mt19937 gen{std::random_device{}()};
// Integers: both ends included
std::uniform_int_distribution<> intDist{0, 99};
int randomInt = intDist(gen);
// Reals: [0.0, 1.0), upper end excluded
std::uniform_real_distribution<> realDist{0.0, 1.0};
double randomDouble = realDist(gen);

Use uniform when there is genuinely no reason for any value to be preferred: dice, a random index, a random position on a map, jitter added to a retry delay. The integer version includes both bounds, and the real version excludes the upper bound. That asymmetry is deliberate (a closed real interval would make hitting exactly 1.0 a special case), but it causes off-by-one bugs when people translate between the two.

Two restrictions are easy to trip over. First, uniform_int_distribution<char>, <uint8_t>, and <bool> are not allowed by the standard (only short, int, long, long long and their unsigned versions are). Some compilers accept them, some reject them. Generate an int and cast. Second, the constructor requires a <= b. Passing bounds in the wrong order, for example from a range computed at runtime that turned out empty, is undefined behavior rather than an exception.

Normal, Bernoulli, binomial, Poisson, and weighted choice in practice

Normal distribution: a histogram and the unbounded tails

#include <cmath>
#include <iostream>
#include <map>
#include <random>
#include <string>
int main() {
    std::mt19937 gen{std::random_device{}()};
    std::normal_distribution<> dist{100.0, 15.0};  // mean 100, stddev 15
    
    std::map<int, int> histogram;
    
    for (int i = 0; i < 10000; ++i) {
        double value = dist(gen);
        ++histogram[static_cast<int>(std::round(value / 10) * 10)];
    }
    
    for (const auto& [value, count] : histogram) {
        std::cout << value << ": " << std::string(count / 100, '*') << std::endl;
    }
}

The normal (Gaussian) distribution models quantities that are the sum of many small independent effects: measurement errors, heights, test scores. About 68% of values fall within one standard deviation of the mean (85 to 115 here), and about 95% within two.

The modeling trap is the tails. A normal distribution is unbounded. With mean 100 and standard deviation 15, a value below 0 is extremely rare, but with a mean of 5 seconds and a standard deviation of 3, about 5% of generated “durations” are negative. Code that then clamps negative values to 0 shifts the average and creates a spike at zero. If the quantity must be positive and is skewed (response times, file sizes, incomes), std::lognormal_distribution or std::gamma_distribution usually models it better than a clamped normal.

Bernoulli: one yes/no trial

#include <iostream>
#include <random>
int main() {
    std::mt19937 gen{std::random_device{}()};
    std::bernoulli_distribution dist{0.7};  // 70% probability
    
    int success = 0;
    for (int i = 0; i < 100; ++i) {
        if (dist(gen)) {
            ++success;
        }
    }
    
    std::cout << "Successes: " << success << "/100" << std::endl;
}

bernoulli_distribution is a single yes/no trial, and it returns bool directly. It is clearer than writing realDist(gen) < 0.7 by hand, and it is the natural choice for “does this event happen?”, such as a critical hit chance, a simulated packet loss, or a feature flag rollout percentage. Don’t expect exactly 70 successes out of 100. The count itself varies from run to run with a standard deviation of about 4.6, which is the next distribution.

Binomial: counting successes in a single draw

#include <iostream>
#include <random>
int main() {
    std::mt19937 gen{std::random_device{}()};
    std::binomial_distribution<> dist{10, 0.5};  // 10 trials, 50% each
    
    for (int i = 0; i < 10; ++i) {
        int successes = dist(gen);
        std::cout << "Success count: " << successes << std::endl;
    }
}

Binomial gives the number of successes in n independent Bernoulli trials, in one call. If you only need the count (how many of 10,000 users clicked, how many of 1,000 packets were lost), a single binomial draw replaces a loop of n Bernoulli draws and is much faster for large n.

Arrivals: Poisson and exponential

#include <random>
std::mt19937 gen{std::random_device{}()};
// How many requests arrive in one second, at an average of 4 per second?
std::poisson_distribution<> arrivals{4.0};
int count = arrivals(gen);
// How long until the next request? (rate 4 per second -> mean 0.25 s)
std::exponential_distribution<> gap{4.0};
double seconds = gap(gen);

These two describe the same process from two angles, and they are the standard tools for simulating traffic, queues, and failures. Poisson answers “how many events in a fixed interval?”. Exponential answers “how long until the next event?”. Note the parameter of exponential_distribution is the rate (events per unit time), not the mean interval. Passing 0.25 where you meant “one request every 0.25 seconds” gives a mean gap of 4 seconds, 16 times slower than intended. This mix-up is common when porting from libraries that parameterize by the mean, such as NumPy’s exponential(scale).

The model assumes events are independent and the rate is constant. Real traffic is often burstier than Poisson (many requests arrive together after a deploy or a push notification). A load test that uses Poisson arrivals is therefore usually a little optimistic about peak load.

Weighted discrete choice for loot tables

#include <iostream>
#include <map>
#include <random>
#include <string>
#include <vector>
int main() {
    std::vector<std::string> items = {"common", "uncommon", "rare", "epic", "legendary"};
    std::vector<double> weights = {0.5, 0.3, 0.15, 0.04, 0.01};
    
    std::mt19937 gen{std::random_device{}()};
    std::discrete_distribution<> dist{weights.begin(), weights.end()};
    
    std::map<std::string, int> drops;
    
    for (int i = 0; i < 1000; ++i) {
        int index = dist(gen);
        ++drops[items[index]];
    }
    
    for (const auto& [item, count] : drops) {
        std::cout << item << ": " << count << std::endl;
    }
}

discrete_distribution returns an index i with probability proportional to weights[i]. The weights do not need to sum to 1: they are normalized, so {50, 30, 15, 4, 1} gives the same result. This is the right tool for loot tables, weighted A/B assignment, and picking from a list with priorities.

Look at the “legendary” row after 1,000 draws: the expected count is 10, but individual runs will easily show 5 or 16. Rare outcomes need many more samples before their observed frequency is close to the target. When a game designer or product manager says “the drop rate feels wrong” based on a short test session, this is usually what they are seeing. Before changing weights, compute how many samples the test actually needs. For a 1% event, a few thousand draws is only a rough check.

This is why, when I review code that uses weighted randomness, I look for a test that runs the distribution many times and checks the observed frequencies against the weights with a tolerance derived from the sample size, rather than a fixed “within 1%”. A tolerance that is too tight produces a flaky test that fails every few hundred runs. One that is too loose passes even when the weights are wired to the wrong items. Both failures are common, and both come from forgetting that the observed frequency is itself random.

Distribution families

// Uniform
std::uniform_int_distribution<>      // integers
std::uniform_real_distribution<>     // reals
// Bernoulli family
std::bernoulli_distribution          // bool
std::binomial_distribution<>         // successes in n trials
std::geometric_distribution<>        // failures before the first success
// Normal family
std::normal_distribution<>           // Gaussian
std::lognormal_distribution<>        // positive, right-skewed
// Poisson family
std::poisson_distribution<>          // events per interval
std::exponential_distribution<>      // time between events (rate parameter)
// Sampling
std::discrete_distribution<>         // weighted index
std::piecewise_constant_distribution<>

Range, parameter, and reset() mistakes

Closed ranges: {0, v.size() - 1}, not v.size()

// uniform_int_distribution: [a, b] inclusive
std::uniform_int_distribution<> intDist{0, 9};  // 0..9
// uniform_real_distribution: [a, b) half-open
std::uniform_real_distribution<> realDist{0.0, 1.0};  // 0.0 <= x < 1.0

Picking a random element from a vector is the classic place this goes wrong: the distribution must be {0, v.size() - 1}, not {0, v.size()}. With an empty vector, v.size() - 1 wraps around to a huge value, so check for empty first.

Standard deviation, not variance

// Normal: mean, standard deviation (not variance)
std::normal_distribution<> dist{100.0, 15.0};
// Invalid (negative stddev) — undefined behavior, not an exception
// std::normal_distribution<> bad{100.0, -15.0};
auto params = dist.param();
std::cout << "mean: " << params.mean() << std::endl;
std::cout << "stddev: " << params.stddev() << std::endl;

The second argument is the standard deviation, not the variance. If your model says “variance 225”, pass 15. Invalid parameters are a precondition violation, not something the library checks for you, so validate values that come from configuration files or user input before constructing a distribution.

What reset() really does

std::mt19937 gen{std::random_device{}()};
std::normal_distribution<> dist{0.0, 1.0};
double a = dist(gen);
dist.reset();  // discard any cached value; does NOT change parameters or reseed
// To change the parameters, assign a new distribution:
dist = std::normal_distribution<>{10.0, 2.0};

reset() does not reset the engine, and it does not restore anything visible. Some distributions keep internal state between calls. The common algorithm for normal_distribution produces values in pairs and caches the second one. reset() discards that cache, so the next value depends only on the engine’s future output. This matters for reproducibility. If you reseed the engine to replay a sequence but keep using the same distribution object, the first value may come from the cache of the previous run. Call reset() together with reseeding, or create a fresh distribution.

Construct once, or pass parameters per call

std::mt19937 gen{std::random_device{}()};
// Fine for uniform: construction is trivial
for (int i = 0; i < 1000; ++i) {
    std::uniform_int_distribution<> dist{0, 99};
    int r = dist(gen);
}
// Different range on each call, one object
std::uniform_int_distribution<> dist;
using param_t = std::uniform_int_distribution<>::param_type;
for (int i = 1; i <= 1000; ++i) {
    int r = dist(gen, param_t{0, i});  // range grows with i
}

For uniform_int_distribution, constructing an object is cheap (it just stores two numbers), so creating it inside a loop costs almost nothing. The cost matters for discrete_distribution and piecewise_* distributions, which build a table in the constructor, and for normal_distribution, where a new object throws away the cached second value and wastes half the work. The param_type overload shown above is the clean way to draw from a different range on each call, which is exactly what a Fisher–Yates shuffle needs, without reconstructing anything.

FAQ

Q1: What is a distribution here?

A: An object that maps an engine’s uniform raw output to values following a chosen probability law.

Q2: Which one should I use?

A: Uniform for equal chances, normal for symmetric measurements around a mean, lognormal for positive skewed values, Bernoulli/binomial for yes/no trials, Poisson/exponential for arrivals, discrete for weighted picks.

Q3: Are results the same on every compiler?

A: No. Engines are specified exactly, but distribution algorithms are not, so the same seed can produce different values with GCC, Clang, and MSVC. See C++ random pitfalls.

Q4: Should I reuse the object?

A: Reuse matters for efficiency and reproducibility rather than correctness: a fresh normal_distribution per call still produces normally distributed values but discards its cached second value, and table-based ones (discrete, piecewise) rebuild their tables. For uniform it makes little difference.

Q5: Integer vs real ranges?

A: uniform_int_distribution includes both ends; uniform_real_distribution excludes the upper end.