std::pmr Memory Resources: monotonic_buffer_resource, Pools and Custom Allocators

Introduction: When malloc/new Becomes the Bottleneck

When small objects are frequently allocated and deallocated, heap fragmentation (scattered small free spaces making large contiguous allocations difficult) and allocator overhead accumulate. Profilers often show malloc/free consuming 20-30% of execution time. Memory pools pre-allocate large blocks and distribute slices from them, reducing allocation count. C++17 std::pmr (polymorphic memory resources) provides injectable allocators based on std::memory_resource, allowing containers of the same type to use different pools. What this guide covers:

  • Problem scenarios: When malloc dominates your profile
  • std::pmr complete examples: monotonic_buffer_resource, pool_resource, custom memory_resource, pmr containers
  • Common errors and solutions: Lifetime management, resource mixing, alignment
  • Best practices: Pool selection guide, incremental adoption
  • Production patterns: Frame pools, request-scoped arenas, game entities

Conceptual analogy

Memory pools and PMR are like pre-dividing warehouse bins and taking items only when needed. When allocation/deallocation patterns are predictable, this approach is more cache-friendly than global new.


Problem Scenarios: When malloc is the Bottleneck

Real-world situations

"Profiler shows malloc/free taking 30% of total execution time."
"Allocating tens of thousands of small objects causes heap fragmentation and OOM."
"Creating/deleting entities every game frame causes severe frame drops."
"Parsing HTTP requests with many allocations causes latency spikes under load."
"Parser repeatedly creates std::vector, std::map causing allocation explosion."

Root causes

  1. Excessive allocation count: Repeated alloc/free of small objects → malloc overhead accumulation
  2. Heap fragmentation: Scattered small free spaces → large contiguous allocation failures
  3. Cache misses: Physically scattered allocations → poor cache efficiency
  4. Lock contention: Global heap lock contention in multithreaded scenarios Memory pools mitigate issues 1-3; thread-local pools mitigate issue 4.

Solution by scenario

ScenarioCharacteristicsRecommended Resource
Game frameCreate/destroy per frame, short lifetimemonotonic_buffer_resource + release()
HTTP requestAllocate per request, free at request endmonotonic_buffer_resource
Node-based structuresRepeated fixed-size node alloc/freesynchronized_pool_resource
Long-lived objectsLong lifetime, varying sizesDefault heap or pool_resource
Parsing/serializationMany temporary buffers, free at scope endmonotonic_buffer_resource

Before/After: HTTP Parser Example

Before (default allocation): Every request allocates vector, map, string from heap repeatedly.

// ❌ Default allocation — many malloc calls
void handleRequest(const std::string& raw) {
    std::vector<std::string> path_segments;
    std::map<std::string, std::string> headers;
    std::string body;
    parsePath(raw, path_segments);
    parseHeaders(raw, headers);
    parseBody(raw, body);
    process(path_segments, headers, body);
}

After (pmr applied): Allocate from request-scoped pool, deallocate all at once when function exits.

// ✅ pmr — reduced allocation count & fragmentation
void handleRequest(const std::string& raw) {
    std::pmr::monotonic_buffer_resource request_pool;
    std::pmr::vector<std::pmr::string> path_segments(&request_pool);
    std::pmr::map<std::pmr::string, std::pmr::string, std::less<>>
        headers(&request_pool);
    std::pmr::string body(&request_pool);
    parsePath(raw, path_segments);
    parseHeaders(raw, headers);
    parseBody(raw, body);
    process(path_segments, headers, body);
}  // request_pool destroyed → all allocations freed at once

Memory Pool Concepts

Reducing allocation count

  • Pools allocate large memory blocks once (or a few times), then distribute fixed-size or variable-size slots from them. On deallocation, blocks are returned to the pool, and actual free happens only once when the pool is destroyed.
  • Fixed-size pools: For same-sized objects (node pools, event pools) — simple and minimal fragmentation.
  • Variable-size: Supporting multiple sizes complicates slot management. monotonic_buffer_resource uses a “cut from front only, no individual frees” approach — simple to implement and perfect for frame/request-scoped resets.

Memory pool operation flow

flowchart TB
    subgraph Heap[Global Heap]
        H[Allocate large block once]
    end
    subgraph Pool[Memory Pool]
        B[Block]
        B --> S1[Slot 1]
        B --> S2[Slot 2]
        B --> S3[Slot 3]
    end
    H --> B
    S1 --> A1[Object A]
    S2 --> A2[Object B]
    S3 --> A3[Object C]

monotonic vs pool comparison

flowchart LR
    subgraph Mono[monotonic_buffer_resource]
        M1[Alloc 1] --> M2[Alloc 2] --> M3[Alloc 3]
        M3 -.->|release returns all| M1
    end
    subgraph Pool[pool_resource]
        P1[Slot] <--> P2[Slot]
        P2 <--> P3[Slot]
    end

HTTP request processing sequence (with pmr)

sequenceDiagram
    participant Req as handleRequest
    participant Pool as monotonic_buffer_resource
    participant Vec as pmr::vector
    participant Map as pmr::map
    Req->>Pool: Create (scope entry)
    Req->>Vec: path_segments(&pool)
    Req->>Map: headers(&pool)
    Req->>Vec: push_back, parsePath...
    Req->>Map: insert, parseHeaders...
    Req->>Req: process()
    Req->>Pool: Destroy (scope exit)
    Note over Pool: All allocations freed at once

std::pmr Overview

Why std::pmr?

Traditional std::allocator is a template type fixed at compile time. std::vector<int, MyAllocator<int>> and std::vector<int, OtherAllocator<int>> are different types, making it difficult to pass them to the same function. In contrast, std::pmr::polymorphic_allocator points to a memory_resource* at runtime, so the same std::pmr::vector<int> type can use different pools. This allows APIs to standardize on std::pmr::vector while callers simply switch pools.

Polymorphic Allocator

  • std::pmr::memory_resource is a pure virtual interface defining only: allocate, deallocate, is_equal. Implementations (pool_resource, monotonic_buffer_resource, etc.) can be selected at runtime.
  • std::pmr::polymorphic_allocator<T> is an allocator pointing to this memory_resource. Pass it to containers like std::vector<T, std::pmr::polymorphic_allocator<T>> to allocate from that resource.
  • std::pmr::vector is an alias for std::vector<T, std::pmr::polymorphic_allocator<T>>. Its constructor accepts a memory_resource* to use that pool.

std::pmr architecture diagram

flowchart TB
    subgraph Container[pmr containers]
        V[std::pmr::vector]
        M[std::pmr::map]
        S[std::pmr::string]
    end
    subgraph Alloc[polymorphic_allocator]
        PA[memory_resource*]
    end
    subgraph Resources[memory_resource implementations]
        MONO[monotonic_buffer_resource]
        POOL[synchronized_pool_resource]
        CUSTOM[Custom resource]
    end
    subgraph Backend[Backend]
        HEAP[Global heap]
        BUF[Stack/static buffer]
    end
    V --> PA
    M --> PA
    S --> PA
    PA --> MONO
    PA --> POOL
    PA --> CUSTOM
    MONO --> BUF
    POOL --> HEAP
    CUSTOM --> HEAP

Build and run

# Requires C++17 or later
g++ -std=c++17 -O2 -o pmr_demo pmr_demo.cpp
./pmr_demo

Basic usage example

#include <memory_resource>
#include <vector>
int main() {
    char buffer[1024];
    std::pmr::monotonic_buffer_resource pool{std::data(buffer), std::size(buffer)};
    std::pmr::vector<int> v(&pool);
    v.push_back(1);
    v.push_back(2);
    // v's allocations come from buffer
}

monotonic_buffer_resource Complete Examples

Concept: Forward-only allocation

  • std::pmr::monotonic_buffer_resource: Sequentially slices memory from a given buffer (or upstream resource). deallocate is a no-op; individual frees don’t happen. Use release() to reset everything. Perfect for frame buffers and request scopes.

Example 1: Stack buffer + monotonic (zero heap allocations)

#include <memory_resource>
#include <vector>
#include <array>
void processRequest() {
    // Allocate 64KB buffer on stack — no heap usage
    std::array<std::byte, 65536> stack_buffer;
    std::pmr::monotonic_buffer_resource pool{
        stack_buffer.data(), stack_buffer.size(),
        std::pmr::new_delete_resource()  // Use heap on overflow
    };
    std::pmr::vector<int> ids(&pool);
    std::pmr::vector<std::pmr::string> tokens(&pool);
    for (int i = 0; i < 1000; ++i) {
        ids.push_back(i);
        tokens.push_back("token");
    }
    // Function exit → stack_buffer and pool automatically freed
}

Example 2: Frame pool + release()

#include <memory_resource>
#include <vector>
struct Entity { int id; float x, y; };
struct Component { int type; void* data; };
void gameLoop() {
    std::array<std::byte, 1024*1024> frame_buffer;
    std::pmr::monotonic_buffer_resource frame_pool{
        frame_buffer.data(), frame_buffer.size(),
        std::pmr::new_delete_resource()
    };
    while (running) {
        frame_pool.release();  // Reset previous frame memory
        std::pmr::vector<Entity> entities(&frame_pool);
        std::pmr::vector<Component> components(&frame_pool);
        // ... create entities, frame logic ...
    }
}

Example 3: Multiple pmr containers sharing same pool

#include <memory_resource>
#include <vector>
#include <string>
#include <map>
int main() {
    std::pmr::monotonic_buffer_resource pool;
    std::pmr::vector<int> nums(&pool);
    std::pmr::vector<std::pmr::string> names(&pool);
    std::pmr::map<std::pmr::string, int, std::less<>> scores(&pool);
    nums.push_back(42);
    names.push_back("Alice");
    scores["Bob"] = 100;
    // All allocations from pool
}

pool_resource Complete Examples

Concept: Reusing fixed-size blocks

  • std::pmr::synchronized_pool_resource / unsynchronized_pool_resource: Manage fixed-size block pools and reuse freed blocks. Use pool_options to configure block size, max block count, etc.

Example 1: Configuring block size with pool_options

#include <memory_resource>
int main() {
    std::pmr::pool_options opts;
    opts.max_blocks_per_chunk = 32;   // Max blocks per chunk
    opts.largest_required_pool_block = 256;  // Max block size
    std::pmr::synchronized_pool_resource pool{opts};
    // Suitable for objects ≤256 bytes, thread-safe
}

Example 2: Applying pool to node-based data structures

#include <memory_resource>
#include <list>
struct TreeNode {
    int value;
    TreeNode* left = nullptr;
    TreeNode* right = nullptr;
};
void buildTree(std::pmr::memory_resource* mr) {
    std::pmr::list<TreeNode> nodes(mr);
    for (int i = 0; i < 1000; ++i) {
        nodes.push_back(TreeNode{i});
    }
    // Nodes allocated from pool, returned to pool on list destruction
}
int main() {
    std::pmr::synchronized_pool_resource pool;
    buildTree(&pool);
}

Example 3: unsynchronized_pool (single-thread only)

#include <memory_resource>
void singleThreadWork() {
    // Single thread only — no lock overhead
    std::pmr::unsynchronized_pool_resource pool;
    std::pmr::vector<int> v(&pool);
    for (int i = 0; i < 10000; ++i) {
        v.push_back(i);
    }
}

Custom memory_resource Implementation

Example 1: Logging memory_resource

A debugging resource that logs allocations/deallocations.

#include <memory_resource>
#include <iostream>
#include <cstddef>
class logging_memory_resource : public std::pmr::memory_resource {
public:
    explicit logging_memory_resource(std::pmr::memory_resource* upstream
        = std::pmr::get_default_resource())
        : upstream_(upstream) {}
private:
    void* do_allocate(std::size_t bytes, std::size_t alignment) override {
        void* p = upstream_->allocate(bytes, alignment);
        std::cout << "[alloc] " << bytes << " bytes, align " << alignment
                  << " -> " << p << "\n";
        return p;
    }
    void do_deallocate(void* p, std::size_t bytes, std::size_t alignment) override {
        std::cout << "[dealloc] " << bytes << " bytes @ " << p << "\n";
        upstream_->deallocate(p, bytes, alignment);
    }
    [[nodiscard]] bool do_is_equal(
        const std::pmr::memory_resource& other) const noexcept override {
        return this == &other;
    }
    std::pmr::memory_resource* upstream_;
};

Example 2: Statistics-collecting memory_resource

#include <memory_resource>
#include <atomic>
#include <cstddef>
class stats_memory_resource : public std::pmr::memory_resource {
public:
    explicit stats_memory_resource(std::pmr::memory_resource* upstream
        = std::pmr::get_default_resource())
        : upstream_(upstream) {}
    std::size_t allocation_count() const noexcept { return alloc_count_.load(); }
    std::size_t total_allocated() const noexcept { return total_allocated_.load(); }
    std::size_t peak_allocated() const noexcept { return peak_allocated_.load(); }
private:
    void* do_allocate(std::size_t bytes, std::size_t alignment) override {
        void* p = upstream_->allocate(bytes, alignment);
        alloc_count_.fetch_add(1);
        std::size_t prev = total_allocated_.fetch_add(bytes);
        std::size_t current = prev + bytes;
        for (std::size_t peak = peak_allocated_.load();
             current > peak && !peak_allocated_.compare_exchange_weak(peak, current);
             peak = peak_allocated_.load()) {}
        return p;
    }
    void do_deallocate(void* p, std::size_t bytes, std::size_t alignment) override {
        total_allocated_.fetch_sub(bytes);
        upstream_->deallocate(p, bytes, alignment);
    }
    [[nodiscard]] bool do_is_equal(
        const std::pmr::memory_resource& other) const noexcept override {
        return this == &other;
    }
    std::pmr::memory_resource* upstream_;
    std::atomic<std::size_t> alloc_count_{0};
    std::atomic<std::size_t> total_allocated_{0};
    std::atomic<std::size_t> peak_allocated_{0};
};

Example 3: Thread-local monotonic pool

#include <memory_resource>
#include <vector>
#include <thread>
thread_local std::pmr::monotonic_buffer_resource* tls_pool = nullptr;
void initThreadPool() {
    tls_pool = new std::pmr::monotonic_buffer_resource(
        std::pmr::new_delete_resource());
}
void cleanupThreadPool() {
    delete tls_pool;
    tls_pool = nullptr;
}
std::pmr::memory_resource* getThreadPool() {
    if (!tls_pool) initThreadPool();
    return tls_pool;
}
void worker(int id) {
    std::pmr::vector<int> local_data(getThreadPool());
    for (int i = 0; i < 1000; ++i) {
        local_data.push_back(i * id);
    }
}

Example 4: Fixed-block pool (simple implementation)

#include <memory_resource>
#include <vector>
#include <cstddef>
class fixed_block_pool : public std::pmr::memory_resource {
public:
    explicit fixed_block_pool(std::size_t block_size, std::size_t block_count = 64)
        : block_size_(block_size), blocks_(block_count) {
        storage_.resize(block_size * block_count);
        for (std::size_t i = 0; i < block_count; ++i) {
            free_list_.push_back(storage_.data() + i * block_size);
        }
    }
private:
    void* do_allocate(std::size_t bytes, std::size_t alignment) override {
        if (bytes > block_size_) return nullptr;
        if (free_list_.empty()) return nullptr;
        void* p = free_list_.back();
        free_list_.pop_back();
        return p;
    }
    void do_deallocate(void* p, std::size_t bytes, std::size_t alignment) override {
        if (p >= storage_.data() && p < storage_.data() + storage_.size()) {
            free_list_.push_back(static_cast<std::byte*>(p));
        }
    }
    [[nodiscard]] bool do_is_equal(
        const std::pmr::memory_resource& other) const noexcept override {
        return this == &other;
    }
    std::size_t block_size_;
    std::vector<std::byte> storage_;
    std::vector<std::byte*> blocks_;
    std::vector<std::byte*> free_list_;
};

pmr Containers in Practice

Supported pmr containers

ContainerAliasUse Case
std::pmr::stringvector<char, polymorphic_allocator<char>>Strings
std::pmr::vectorvector<T, polymorphic_allocator<T>>Dynamic arrays
std::pmr::mapmap<K,V,…,polymorphic_allocator<pair<…>>>Sorted maps
std::pmr::setset<T,…,polymorphic_allocator<T>>Sorted sets
std::pmr::unordered_mapunordered_map with pmr allocatorHash maps
std::pmr::listlist<T, polymorphic_allocator<T>>Doubly-linked lists

Nested pmr containers: map<string, vector<string>>

#include <memory_resource>
#include <vector>
#include <string>
#include <map>
void parseConfig(const char* raw) {
    std::pmr::monotonic_buffer_resource pool;
    // Both map keys and values are pmr::string
    // If value is vector<pmr::string>, strings inside also use pool
    std::pmr::map<std::pmr::string, std::pmr::vector<std::pmr::string>, std::less<>>
        config(&pool);
    config["sections"].push_back("a");
    config["sections"].push_back("b");
    config["keys"].push_back("x");
    // All allocations from pool
}

Why nested containers must be pmr too: uses-allocator construction

When a std::pmr::vector<std::pmr::string> creates an element, it does not just call the string’s constructor. It uses uses-allocator construction: because std::pmr::string accepts an allocator, the container passes its own allocator along, so the inner string allocates from the same resource as the outer vector. That is what makes a whole tree of pmr containers live in one arena with no extra code. If the element type is a plain std::string, it does not accept a polymorphic_allocator, so nothing is passed down and every string’s characters go to the global heap, while only the vector’s own array lives in the pool.

std::pmr::monotonic_buffer_resource pool;
std::pmr::vector<std::pmr::vector<int>> outer(&pool);
outer.emplace_back(3, 7);   // inner vector gets &pool automatically
// outer[0].get_allocator().resource() == &pool

Your own types can take part by accepting an allocator_type (declare using allocator_type = std::pmr::polymorphic_allocator<>; and provide constructors that take it). Then pmr containers of your type propagate the resource into it as well.

C++20: polymorphic_allocator<> and null_memory_resource

C++20 gave polymorphic_allocator a default template argument of std::byte, plus helpers allocate_object<T>(), new_object<T>(args...), and delete_object(). std::pmr::polymorphic_allocator<> is therefore the conventional allocator type for your own allocator-aware classes: one type that can allocate any object from the resource.

std::pmr::null_memory_resource() throws std::bad_alloc on every allocation. It is the right upstream when an arena must never fall back to the heap, for example in a real-time loop or a signal-safe path:

std::array<std::byte, 64 * 1024> buf;
std::pmr::monotonic_buffer_resource arena(buf.data(), buf.size(),
                                          std::pmr::null_memory_resource());
// exceeding 64 KB now throws instead of silently calling operator new

That turns “the buffer was too small” (Error 4 below) from a silent performance regression into a test failure you can see.

pmr string caution: Don’t mix with std::string

// ❌ Dangerous: mixing std::string and std::pmr::string
std::pmr::map<std::pmr::string, std::string, std::less<>> m(&pool);
// value is std::string → uses default allocator, not pool
// ✅ Correct: all pmr
std::pmr::map<std::pmr::string, std::pmr::string, std::less<>> m(&pool);

Common Errors and Solutions

Error 1: Container outlives pool (Use-After-Free)

Symptoms: Crash, undefined behavior, heap corruption. Cause: memory_resource destroyed before containers using it.

// ❌ Wrong code
std::pmr::vector<int>* createVector() {
    std::pmr::monotonic_buffer_resource pool;
    return new std::pmr::vector<int>(&pool);  // pool destroyed on function exit!
}
// Returned vector points to freed pool → UB

Solution:

// ✅ Correct: pool lifetime > container lifetime
std::pmr::monotonic_buffer_resource* pool = new std::pmr::monotonic_buffer_resource;
std::pmr::vector<int>* vec = new std::pmr::vector<int>(pool);
// Manage lifetimes to ensure pool outlives vec

Or keep pool and container in same scope:

// ✅ Pool and container lifetimes match
void process() {
    std::pmr::monotonic_buffer_resource pool;
    std::pmr::vector<int> vec(&pool);
    // ... use ...
}  // vec, pool destroyed in order

Error 2: Expecting deallocate from monotonic

Symptoms: Memory not freed, keeps accumulating. Cause: monotonic_buffer_resource’s deallocate is a no-op.

// ❌ Wrong expectation
std::pmr::monotonic_buffer_resource pool;
std::pmr::vector<int> v(&pool);
v.push_back(1);
v.pop_back();  // Internally calls deallocate, but pool doesn't free
// Memory remains in pool until release() called

Solution: Use monotonic only for release() full reset.

// ✅ monotonic: release at scope end
{
    std::pmr::monotonic_buffer_resource pool;
    std::pmr::vector<int> v(&pool);
    v.push_back(1);
    v.pop_back();
    pool.release();  // Full pool reset at this point
}

Error 3: Assuming copies keep the source’s resource

Symptoms: Allocations silently go to the global heap instead of your pool, or a “fast” move is actually a slow copy.

A polymorphic_allocator never propagates on copy assignment, move assignment, or swap (propagate_on_container_* are all false). Each container keeps the resource it was constructed with, and data is copied into that resource. The consequences differ by operation:

std::pmr::monotonic_buffer_resource pool1, pool2;
std::pmr::vector<int> a(&pool1);
a.push_back(42);

std::pmr::vector<int> b(&pool2);
b = a;                         // safe: elements are copied INTO pool2; b keeps pool2

std::pmr::vector<int> c = a;   // surprise: c uses the DEFAULT resource (usually the heap), not pool1
std::pmr::vector<int> c2(a, &pool1);  // what you probably meant: copy into pool1

std::pmr::vector<int> d(std::move(a));  // move construction keeps pool1 (cheap pointer steal)
b = std::move(d);              // different resources: elements are moved one by one into pool2, O(n)

The copy construction case is the trap. Copy construction calls select_on_container_copy_construction(), which for polymorphic_allocator returns a default-constructed allocator, meaning get_default_resource(). Code that copies a pmr container “within” a request arena, or passing it by value, therefore allocates from the global heap without any warning. Pass the resource explicitly when copying (c2 above), and check with get_allocator().resource() in tests.

Move assignment between containers with different resources cannot just steal the buffer (it belongs to another resource), so it degrades to an element-wise move with fresh allocations. And swap of two containers with unequal allocators is undefined behavior. Keep containers that interact on the same resource, or treat cross-resource operations as copies.

In my experience, the copy-construction surprise is the reason many first pmr migrations show no speedup at all: the profile still shows malloc, because the objects that were supposed to live in the arena were copied out of it on the way to where they are used. A statistics-collecting resource (section 6) in a debug build makes this visible immediately.

Error 4: Insufficient stack buffer size

Symptoms: monotonic allocates additional memory from upstream (heap).

// ❌ Buffer may be too small
char buffer[256];
std::pmr::monotonic_buffer_resource pool{buffer, sizeof(buffer)};
std::pmr::vector<int> v(&pool);
for (int i = 0; i < 1000; ++i) v.push_back(i);  // Exceeds 256 bytes → uses heap

Solution: Allocate generous buffer size.

// ✅ Generous buffer
std::array<std::byte, 65536> buffer;
std::pmr::monotonic_buffer_resource pool{
    buffer.data(), buffer.size(),
    std::pmr::new_delete_resource()
};

Error 5: Inappropriate pool_options settings

Symptoms: Memory waste or allocation failures with synchronized_pool_resource. Solution: Configure based on actual maximum block size used.

// ✅ Settings matching usage pattern
std::pmr::pool_options opts;
opts.largest_required_pool_block = 64;   // Mostly ≤64-byte objects
opts.max_blocks_per_chunk = 128;
std::pmr::synchronized_pool_resource pool{opts};

Error 6: Ignoring alignment

Symptoms: Crash or SIGBUS on certain platforms.

// ❌ Wrong implementation (ignoring alignment)
void* do_allocate(std::size_t bytes, std::size_t alignment) override {
    return upstream_->allocate(bytes, 1);  // Ignoring alignment!
}

Solution: Always pass requested alignment.

// ✅ Respecting alignment
void* do_allocate(std::size_t bytes, std::size_t alignment) override {
    return upstream_->allocate(bytes, alignment);
}

Error 7: Thread safety misconceptions

Symptoms: Crash or data corruption in multithreaded scenarios.

// ❌ Dangerous: multiple threads sharing same unsynchronized_pool
std::pmr::unsynchronized_pool_resource pool;
std::thread t1([&] { std::pmr::vector<int> v(&pool); /* ... */ });
std::thread t2([&] { std::pmr::vector<int> v(&pool); /* ... */ });

Solution: Use synchronized_pool_resource for shared pools in multithreaded code.

// ✅ synchronized_pool_resource (thread-safe)
std::pmr::synchronized_pool_resource pool;

Best Practices

Pool selection guide

PatternRecommendedReason
Frame/request unit, bulk freemonotonicNo deallocate, reset with release()
Individual object repeated alloc/freepool_resourceBlock reuse
Thread-local independent allocationThread-local monotonicEliminates lock contention
Debugging/profilingWrap with stats_memory_resourceTrack allocation count & peak

Pool lifetime rules

Always: pool lifetime ≥ all containers using it

Use pmr for nested container elements too

// ✅ Both map keys and values are pmr types
std::pmr::map<std::pmr::string, std::pmr::vector<int>, std::less<>> m(&pool);

Apply after profiling

// Step 1: Identify bottleneck with profiling
// Step 2: Introduce pool only in that function/scope
// Step 3: Measure performance, then expand application

Use statistics resource in debug builds

#ifdef NDEBUG
    std::pmr::memory_resource* resource = std::pmr::get_default_resource();
#else
    static stats_memory_resource stats{std::pmr::get_default_resource()};
    std::pmr::memory_resource* resource = &stats;
#endif
std::pmr::vector<int> data(resource);
// ... processing ...
#ifndef NDEBUG
    std::cout << "Allocations: " << stats.allocation_count()
              << ", Peak: " << stats.peak_allocated() << " bytes\n";
#endif

Performance Benchmarks

Benchmark 1: Measure allocations, not push_back

A benchmark only says something about allocators if the code under test actually allocates many times. A std::vector<int> that calls reserve(N) first allocates exactly once, whatever resource it uses, so timing its push_back loop measures integer stores, not memory management. Node-based containers and containers of strings are where allocation cost shows up:

#include <chrono>
#include <iostream>
#include <list>
#include <memory_resource>
#include <string>

template <class F>
long long time_us(F f) {
    auto t0 = std::chrono::steady_clock::now();
    f();
    auto t1 = std::chrono::steady_clock::now();
    return std::chrono::duration_cast<std::chrono::microseconds>(t1 - t0).count();
}

int main() {
    constexpr int N = 200'000;
    long long sink = 0;  // keep results alive so the work is not optimized away

    auto heap = time_us([&] {
        std::list<std::string> l;
        for (int i = 0; i < N; ++i) l.emplace_back(40, 'x');  // node + string buffer
        sink += l.size();
    });

    auto mono = time_us([&] {
        std::pmr::monotonic_buffer_resource arena;
        std::pmr::list<std::pmr::string> l(&arena);
        for (int i = 0; i < N; ++i) l.emplace_back(40, 'x');
        sink += l.size();
    });  // arena releases everything at once here

    auto pool = time_us([&] {
        std::pmr::unsynchronized_pool_resource p;
        std::pmr::list<std::pmr::string> l(&p);
        for (int i = 0; i < N; ++i) l.emplace_back(40, 'x');
        sink += l.size();
    });

    std::cout << "heap " << heap << " us, monotonic " << mono
              << " us, pool " << pool << " us (" << sink << ")\n";
}

Each element here needs two allocations (the list node and a 40-character string that does not fit in the small-string buffer), so allocation dominates. Build with optimizations (-O2), run it several times, and compare medians.

What to expect, and why numbers vary. The monotonic arena usually wins clearly in this pattern, because allocation is a pointer bump and destruction frees nothing individually: the whole arena is released at once. The pool resource helps less here, and more in workloads that repeatedly free and reallocate same-sized blocks. The size of the gain depends heavily on the platform’s default allocator: glibc’s malloc, jemalloc, mimalloc, and the Windows heap differ a lot for small allocations, and a modern thread-caching malloc narrows the gap considerably. That is why this article does not quote a fixed speedup. Measure on your target platform with your allocation pattern, and use a profiler to confirm that allocation was significant in the first place.

When pmr does not help

  • Few, large allocations: the cost is dominated by touching the memory, not by the allocator call.
  • Long-lived objects freed in random order: a monotonic arena never reuses memory until release, so it grows without bound. Use a pool or the default heap.
  • Cross-thread handoff: objects allocated from a thread’s arena and freed on another thread need a synchronized resource or careful lifetime design.
  • Code that copies pmr containers by copy construction (Error 3): allocations quietly return to the heap.

Production Patterns

Pattern 1: Game frame pool

void gameLoop() {
    std::array<std::byte, 1024*1024> frame_buffer;
    while (running) {
        std::pmr::monotonic_buffer_resource frame_pool{
            frame_buffer.data(), frame_buffer.size(),
            std::pmr::new_delete_resource()
        };
        std::pmr::vector<Entity> entities(&frame_pool);
        std::pmr::vector<Component> components(&frame_pool);
        // ... create entities, frame logic ...
    }  // frame_pool, entities destroyed → same buffer reused next frame
}

Pattern 2: HTTP request scope

void handleRequest(const Request& req) {
    std::pmr::monotonic_buffer_resource request_pool(
        std::pmr::new_delete_resource());
    std::pmr::vector<std::pmr::string> path_segments(&request_pool);
    std::pmr::map<std::pmr::string, std::pmr::string, std::less<>>
        headers(&request_pool);
    parsePath(req.uri, path_segments);
    parseHeaders(req.raw_headers, headers);
    Response response = processRequest(path_segments, headers);
    sendResponse(response);
}  // request_pool destroyed → all allocations freed

Pattern 3: Per-worker pool in thread pool

class Worker {
    std::pmr::monotonic_buffer_resource worker_pool_;
    std::pmr::vector<Task> local_queue_{&worker_pool_};
public:
    Worker() : worker_pool_(std::pmr::new_delete_resource()) {}
    void process(Task t) {
        worker_pool_.release();
        local_queue_.clear();
        // Process t using local_queue_ allocations
    }
};

Pattern 4: Hierarchical resources (upstream)

std::pmr::synchronized_pool_resource global_pool;
std::pmr::monotonic_buffer_resource thread_pool{&global_pool};
std::pmr::monotonic_buffer_resource frame_pool{&thread_pool};
std::pmr::vector<int> frame_data(&frame_pool);
// frame_pool exhausted → thread_pool → global_pool

Pattern 5: Implementation checklist

  • Pool lifetime exceeds all containers using it
  • Clear release() timing for monotonic (frame/request end)
  • No copy/assignment between pmr containers using different pools
  • Generous stack buffer size when used
  • pool_options match actual usage patterns
  • Use synchronized_pool_resource or per-thread pools for multithreading

Summary

TopicSummary
Memory poolsAllocate large block once, distribute slots — reduces allocation count & fragmentation
std::pmrmemory_resource + polymorphic_allocator for injecting pools into containers
monotonicSequential allocation, reset only — perfect for frame/request scopes
pool_resourceFixed-size block reuse — suitable for general object pools
Custom resourceInherit memory_resource for logging, statistics, thread-local pools

Core principles:

  1. Pool lifetime > container lifetime
  2. monotonic frees only via release()
  3. Beware copy/assignment between containers from different pools
  4. Apply after profiling confirms actual benefits

Frequently Asked Questions (FAQ)

Q. When is a custom allocator or std::pmr actually worth it?

A. When profiling shows malloc/free dominating execution time, when allocating many short-lived objects per frame/request (games, HTTP servers), or when node-based data structures repeatedly allocate fixed-size blocks.

Q. monotonic vs pool_resource?

A. Use monotonic for “allocate many, free all at once” patterns (frames/requests). Use pool_resource for repeated individual alloc/free patterns with fixed sizes. std::pmr enables control over memory pools and fragmentation. Next, read SIMD & Intrinsics (#39-3). Previous: High-Performance C++ #39-1: Cache & Data-Oriented Design Next: High-Performance C++ #39-3: SIMD & Parallelization