std::pmr Memory Resources: monotonic_buffer_resource, Pools and Custom Allocators
Introduction: When malloc/new Becomes the Bottleneck
When small objects are frequently allocated and deallocated, heap fragmentation (scattered small free spaces making large contiguous allocations difficult) and allocator overhead accumulate. Profilers often show malloc/free consuming 20-30% of execution time. Memory pools pre-allocate large blocks and distribute slices from them, reducing allocation count. C++17 std::pmr (polymorphic memory resources) provides injectable allocators based on std::memory_resource, allowing containers of the same type to use different pools. What this guide covers:
- Problem scenarios: When malloc dominates your profile
- std::pmr complete examples: monotonic_buffer_resource, pool_resource, custom memory_resource, pmr containers
- Common errors and solutions: Lifetime management, resource mixing, alignment
- Best practices: Pool selection guide, incremental adoption
- Production patterns: Frame pools, request-scoped arenas, game entities
Conceptual analogy
Memory pools and PMR are like pre-dividing warehouse bins and taking items only when needed. When allocation/deallocation patterns are predictable, this approach is more cache-friendly than global new.
Problem Scenarios: When malloc is the Bottleneck
Real-world situations
"Profiler shows malloc/free taking 30% of total execution time."
"Allocating tens of thousands of small objects causes heap fragmentation and OOM."
"Creating/deleting entities every game frame causes severe frame drops."
"Parsing HTTP requests with many allocations causes latency spikes under load."
"Parser repeatedly creates std::vector, std::map causing allocation explosion."
Root causes
- Excessive allocation count: Repeated alloc/free of small objects → malloc overhead accumulation
- Heap fragmentation: Scattered small free spaces → large contiguous allocation failures
- Cache misses: Physically scattered allocations → poor cache efficiency
- Lock contention: Global heap lock contention in multithreaded scenarios Memory pools mitigate issues 1-3; thread-local pools mitigate issue 4.
Solution by scenario
| Scenario | Characteristics | Recommended Resource |
|---|---|---|
| Game frame | Create/destroy per frame, short lifetime | monotonic_buffer_resource + release() |
| HTTP request | Allocate per request, free at request end | monotonic_buffer_resource |
| Node-based structures | Repeated fixed-size node alloc/free | synchronized_pool_resource |
| Long-lived objects | Long lifetime, varying sizes | Default heap or pool_resource |
| Parsing/serialization | Many temporary buffers, free at scope end | monotonic_buffer_resource |
Before/After: HTTP Parser Example
Before (default allocation): Every request allocates vector, map, string from heap repeatedly.
// ❌ Default allocation — many malloc calls
void handleRequest(const std::string& raw) {
std::vector<std::string> path_segments;
std::map<std::string, std::string> headers;
std::string body;
parsePath(raw, path_segments);
parseHeaders(raw, headers);
parseBody(raw, body);
process(path_segments, headers, body);
}
After (pmr applied): Allocate from request-scoped pool, deallocate all at once when function exits.
// ✅ pmr — reduced allocation count & fragmentation
void handleRequest(const std::string& raw) {
std::pmr::monotonic_buffer_resource request_pool;
std::pmr::vector<std::pmr::string> path_segments(&request_pool);
std::pmr::map<std::pmr::string, std::pmr::string, std::less<>>
headers(&request_pool);
std::pmr::string body(&request_pool);
parsePath(raw, path_segments);
parseHeaders(raw, headers);
parseBody(raw, body);
process(path_segments, headers, body);
} // request_pool destroyed → all allocations freed at once
Memory Pool Concepts
Reducing allocation count
- Pools allocate large memory blocks once (or a few times), then distribute fixed-size or variable-size slots from them. On deallocation, blocks are returned to the pool, and actual free happens only once when the pool is destroyed.
- Fixed-size pools: For same-sized objects (node pools, event pools) — simple and minimal fragmentation.
- Variable-size: Supporting multiple sizes complicates slot management. monotonic_buffer_resource uses a “cut from front only, no individual frees” approach — simple to implement and perfect for frame/request-scoped resets.
Memory pool operation flow
flowchart TB
subgraph Heap[Global Heap]
H[Allocate large block once]
end
subgraph Pool[Memory Pool]
B[Block]
B --> S1[Slot 1]
B --> S2[Slot 2]
B --> S3[Slot 3]
end
H --> B
S1 --> A1[Object A]
S2 --> A2[Object B]
S3 --> A3[Object C]
monotonic vs pool comparison
flowchart LR
subgraph Mono[monotonic_buffer_resource]
M1[Alloc 1] --> M2[Alloc 2] --> M3[Alloc 3]
M3 -.->|release returns all| M1
end
subgraph Pool[pool_resource]
P1[Slot] <--> P2[Slot]
P2 <--> P3[Slot]
end
HTTP request processing sequence (with pmr)
sequenceDiagram
participant Req as handleRequest
participant Pool as monotonic_buffer_resource
participant Vec as pmr::vector
participant Map as pmr::map
Req->>Pool: Create (scope entry)
Req->>Vec: path_segments(&pool)
Req->>Map: headers(&pool)
Req->>Vec: push_back, parsePath...
Req->>Map: insert, parseHeaders...
Req->>Req: process()
Req->>Pool: Destroy (scope exit)
Note over Pool: All allocations freed at once
std::pmr Overview
Why std::pmr?
Traditional std::allocator is a template type fixed at compile time. std::vector<int, MyAllocator<int>> and std::vector<int, OtherAllocator<int>> are different types, making it difficult to pass them to the same function. In contrast, std::pmr::polymorphic_allocator points to a memory_resource* at runtime, so the same std::pmr::vector<int> type can use different pools. This allows APIs to standardize on std::pmr::vector while callers simply switch pools.
Polymorphic Allocator
- std::pmr::memory_resource is a pure virtual interface defining only: allocate, deallocate, is_equal. Implementations (pool_resource, monotonic_buffer_resource, etc.) can be selected at runtime.
- std::pmr::polymorphic_allocator<T> is an allocator pointing to this memory_resource. Pass it to containers like std::vector<T, std::pmr::polymorphic_allocator<T>> to allocate from that resource.
- std::pmr::vector is an alias for std::vector<T, std::pmr::polymorphic_allocator<T>>. Its constructor accepts a memory_resource* to use that pool.
std::pmr architecture diagram
flowchart TB
subgraph Container[pmr containers]
V[std::pmr::vector]
M[std::pmr::map]
S[std::pmr::string]
end
subgraph Alloc[polymorphic_allocator]
PA[memory_resource*]
end
subgraph Resources[memory_resource implementations]
MONO[monotonic_buffer_resource]
POOL[synchronized_pool_resource]
CUSTOM[Custom resource]
end
subgraph Backend[Backend]
HEAP[Global heap]
BUF[Stack/static buffer]
end
V --> PA
M --> PA
S --> PA
PA --> MONO
PA --> POOL
PA --> CUSTOM
MONO --> BUF
POOL --> HEAP
CUSTOM --> HEAP
Build and run
# Requires C++17 or later
g++ -std=c++17 -O2 -o pmr_demo pmr_demo.cpp
./pmr_demo
Basic usage example
#include <memory_resource>
#include <vector>
int main() {
char buffer[1024];
std::pmr::monotonic_buffer_resource pool{std::data(buffer), std::size(buffer)};
std::pmr::vector<int> v(&pool);
v.push_back(1);
v.push_back(2);
// v's allocations come from buffer
}
monotonic_buffer_resource Complete Examples
Concept: Forward-only allocation
- std::pmr::monotonic_buffer_resource: Sequentially slices memory from a given buffer (or upstream resource). deallocate is a no-op; individual frees don’t happen. Use release() to reset everything. Perfect for frame buffers and request scopes.
Example 1: Stack buffer + monotonic (zero heap allocations)
#include <memory_resource>
#include <vector>
#include <array>
void processRequest() {
// Allocate 64KB buffer on stack — no heap usage
std::array<std::byte, 65536> stack_buffer;
std::pmr::monotonic_buffer_resource pool{
stack_buffer.data(), stack_buffer.size(),
std::pmr::new_delete_resource() // Use heap on overflow
};
std::pmr::vector<int> ids(&pool);
std::pmr::vector<std::pmr::string> tokens(&pool);
for (int i = 0; i < 1000; ++i) {
ids.push_back(i);
tokens.push_back("token");
}
// Function exit → stack_buffer and pool automatically freed
}
Example 2: Frame pool + release()
#include <memory_resource>
#include <vector>
struct Entity { int id; float x, y; };
struct Component { int type; void* data; };
void gameLoop() {
std::array<std::byte, 1024*1024> frame_buffer;
std::pmr::monotonic_buffer_resource frame_pool{
frame_buffer.data(), frame_buffer.size(),
std::pmr::new_delete_resource()
};
while (running) {
frame_pool.release(); // Reset previous frame memory
std::pmr::vector<Entity> entities(&frame_pool);
std::pmr::vector<Component> components(&frame_pool);
// ... create entities, frame logic ...
}
}
Example 3: Multiple pmr containers sharing same pool
#include <memory_resource>
#include <vector>
#include <string>
#include <map>
int main() {
std::pmr::monotonic_buffer_resource pool;
std::pmr::vector<int> nums(&pool);
std::pmr::vector<std::pmr::string> names(&pool);
std::pmr::map<std::pmr::string, int, std::less<>> scores(&pool);
nums.push_back(42);
names.push_back("Alice");
scores["Bob"] = 100;
// All allocations from pool
}
pool_resource Complete Examples
Concept: Reusing fixed-size blocks
- std::pmr::synchronized_pool_resource / unsynchronized_pool_resource: Manage fixed-size block pools and reuse freed blocks. Use pool_options to configure block size, max block count, etc.
Example 1: Configuring block size with pool_options
#include <memory_resource>
int main() {
std::pmr::pool_options opts;
opts.max_blocks_per_chunk = 32; // Max blocks per chunk
opts.largest_required_pool_block = 256; // Max block size
std::pmr::synchronized_pool_resource pool{opts};
// Suitable for objects ≤256 bytes, thread-safe
}
Example 2: Applying pool to node-based data structures
#include <memory_resource>
#include <list>
struct TreeNode {
int value;
TreeNode* left = nullptr;
TreeNode* right = nullptr;
};
void buildTree(std::pmr::memory_resource* mr) {
std::pmr::list<TreeNode> nodes(mr);
for (int i = 0; i < 1000; ++i) {
nodes.push_back(TreeNode{i});
}
// Nodes allocated from pool, returned to pool on list destruction
}
int main() {
std::pmr::synchronized_pool_resource pool;
buildTree(&pool);
}
Example 3: unsynchronized_pool (single-thread only)
#include <memory_resource>
void singleThreadWork() {
// Single thread only — no lock overhead
std::pmr::unsynchronized_pool_resource pool;
std::pmr::vector<int> v(&pool);
for (int i = 0; i < 10000; ++i) {
v.push_back(i);
}
}
Custom memory_resource Implementation
Example 1: Logging memory_resource
A debugging resource that logs allocations/deallocations.
#include <memory_resource>
#include <iostream>
#include <cstddef>
class logging_memory_resource : public std::pmr::memory_resource {
public:
explicit logging_memory_resource(std::pmr::memory_resource* upstream
= std::pmr::get_default_resource())
: upstream_(upstream) {}
private:
void* do_allocate(std::size_t bytes, std::size_t alignment) override {
void* p = upstream_->allocate(bytes, alignment);
std::cout << "[alloc] " << bytes << " bytes, align " << alignment
<< " -> " << p << "\n";
return p;
}
void do_deallocate(void* p, std::size_t bytes, std::size_t alignment) override {
std::cout << "[dealloc] " << bytes << " bytes @ " << p << "\n";
upstream_->deallocate(p, bytes, alignment);
}
[[nodiscard]] bool do_is_equal(
const std::pmr::memory_resource& other) const noexcept override {
return this == &other;
}
std::pmr::memory_resource* upstream_;
};
Example 2: Statistics-collecting memory_resource
#include <memory_resource>
#include <atomic>
#include <cstddef>
class stats_memory_resource : public std::pmr::memory_resource {
public:
explicit stats_memory_resource(std::pmr::memory_resource* upstream
= std::pmr::get_default_resource())
: upstream_(upstream) {}
std::size_t allocation_count() const noexcept { return alloc_count_.load(); }
std::size_t total_allocated() const noexcept { return total_allocated_.load(); }
std::size_t peak_allocated() const noexcept { return peak_allocated_.load(); }
private:
void* do_allocate(std::size_t bytes, std::size_t alignment) override {
void* p = upstream_->allocate(bytes, alignment);
alloc_count_.fetch_add(1);
std::size_t prev = total_allocated_.fetch_add(bytes);
std::size_t current = prev + bytes;
for (std::size_t peak = peak_allocated_.load();
current > peak && !peak_allocated_.compare_exchange_weak(peak, current);
peak = peak_allocated_.load()) {}
return p;
}
void do_deallocate(void* p, std::size_t bytes, std::size_t alignment) override {
total_allocated_.fetch_sub(bytes);
upstream_->deallocate(p, bytes, alignment);
}
[[nodiscard]] bool do_is_equal(
const std::pmr::memory_resource& other) const noexcept override {
return this == &other;
}
std::pmr::memory_resource* upstream_;
std::atomic<std::size_t> alloc_count_{0};
std::atomic<std::size_t> total_allocated_{0};
std::atomic<std::size_t> peak_allocated_{0};
};
Example 3: Thread-local monotonic pool
#include <memory_resource>
#include <vector>
#include <thread>
thread_local std::pmr::monotonic_buffer_resource* tls_pool = nullptr;
void initThreadPool() {
tls_pool = new std::pmr::monotonic_buffer_resource(
std::pmr::new_delete_resource());
}
void cleanupThreadPool() {
delete tls_pool;
tls_pool = nullptr;
}
std::pmr::memory_resource* getThreadPool() {
if (!tls_pool) initThreadPool();
return tls_pool;
}
void worker(int id) {
std::pmr::vector<int> local_data(getThreadPool());
for (int i = 0; i < 1000; ++i) {
local_data.push_back(i * id);
}
}
Example 4: Fixed-block pool (simple implementation)
#include <memory_resource>
#include <vector>
#include <cstddef>
class fixed_block_pool : public std::pmr::memory_resource {
public:
explicit fixed_block_pool(std::size_t block_size, std::size_t block_count = 64)
: block_size_(block_size), blocks_(block_count) {
storage_.resize(block_size * block_count);
for (std::size_t i = 0; i < block_count; ++i) {
free_list_.push_back(storage_.data() + i * block_size);
}
}
private:
void* do_allocate(std::size_t bytes, std::size_t alignment) override {
if (bytes > block_size_) return nullptr;
if (free_list_.empty()) return nullptr;
void* p = free_list_.back();
free_list_.pop_back();
return p;
}
void do_deallocate(void* p, std::size_t bytes, std::size_t alignment) override {
if (p >= storage_.data() && p < storage_.data() + storage_.size()) {
free_list_.push_back(static_cast<std::byte*>(p));
}
}
[[nodiscard]] bool do_is_equal(
const std::pmr::memory_resource& other) const noexcept override {
return this == &other;
}
std::size_t block_size_;
std::vector<std::byte> storage_;
std::vector<std::byte*> blocks_;
std::vector<std::byte*> free_list_;
};
pmr Containers in Practice
Supported pmr containers
| Container | Alias | Use Case |
|---|---|---|
| std::pmr::string | vector<char, polymorphic_allocator<char>> | Strings |
| std::pmr::vector | vector<T, polymorphic_allocator<T>> | Dynamic arrays |
| std::pmr::map | map<K,V,…,polymorphic_allocator<pair<…>>> | Sorted maps |
| std::pmr::set | set<T,…,polymorphic_allocator<T>> | Sorted sets |
| std::pmr::unordered_map | unordered_map with pmr allocator | Hash maps |
| std::pmr::list | list<T, polymorphic_allocator<T>> | Doubly-linked lists |
Nested pmr containers: map<string, vector<string>>
#include <memory_resource>
#include <vector>
#include <string>
#include <map>
void parseConfig(const char* raw) {
std::pmr::monotonic_buffer_resource pool;
// Both map keys and values are pmr::string
// If value is vector<pmr::string>, strings inside also use pool
std::pmr::map<std::pmr::string, std::pmr::vector<std::pmr::string>, std::less<>>
config(&pool);
config["sections"].push_back("a");
config["sections"].push_back("b");
config["keys"].push_back("x");
// All allocations from pool
}
Why nested containers must be pmr too: uses-allocator construction
When a std::pmr::vector<std::pmr::string> creates an element, it does not just call the string’s constructor. It uses uses-allocator construction: because std::pmr::string accepts an allocator, the container passes its own allocator along, so the inner string allocates from the same resource as the outer vector. That is what makes a whole tree of pmr containers live in one arena with no extra code. If the element type is a plain std::string, it does not accept a polymorphic_allocator, so nothing is passed down and every string’s characters go to the global heap, while only the vector’s own array lives in the pool.
std::pmr::monotonic_buffer_resource pool;
std::pmr::vector<std::pmr::vector<int>> outer(&pool);
outer.emplace_back(3, 7); // inner vector gets &pool automatically
// outer[0].get_allocator().resource() == &pool
Your own types can take part by accepting an allocator_type (declare using allocator_type = std::pmr::polymorphic_allocator<>; and provide constructors that take it). Then pmr containers of your type propagate the resource into it as well.
C++20: polymorphic_allocator<> and null_memory_resource
C++20 gave polymorphic_allocator a default template argument of std::byte, plus helpers allocate_object<T>(), new_object<T>(args...), and delete_object(). std::pmr::polymorphic_allocator<> is therefore the conventional allocator type for your own allocator-aware classes: one type that can allocate any object from the resource.
std::pmr::null_memory_resource() throws std::bad_alloc on every allocation. It is the right upstream when an arena must never fall back to the heap, for example in a real-time loop or a signal-safe path:
std::array<std::byte, 64 * 1024> buf;
std::pmr::monotonic_buffer_resource arena(buf.data(), buf.size(),
std::pmr::null_memory_resource());
// exceeding 64 KB now throws instead of silently calling operator new
That turns “the buffer was too small” (Error 4 below) from a silent performance regression into a test failure you can see.
pmr string caution: Don’t mix with std::string
// ❌ Dangerous: mixing std::string and std::pmr::string
std::pmr::map<std::pmr::string, std::string, std::less<>> m(&pool);
// value is std::string → uses default allocator, not pool
// ✅ Correct: all pmr
std::pmr::map<std::pmr::string, std::pmr::string, std::less<>> m(&pool);
Common Errors and Solutions
Error 1: Container outlives pool (Use-After-Free)
Symptoms: Crash, undefined behavior, heap corruption.
Cause: memory_resource destroyed before containers using it.
// ❌ Wrong code
std::pmr::vector<int>* createVector() {
std::pmr::monotonic_buffer_resource pool;
return new std::pmr::vector<int>(&pool); // pool destroyed on function exit!
}
// Returned vector points to freed pool → UB
Solution:
// ✅ Correct: pool lifetime > container lifetime
std::pmr::monotonic_buffer_resource* pool = new std::pmr::monotonic_buffer_resource;
std::pmr::vector<int>* vec = new std::pmr::vector<int>(pool);
// Manage lifetimes to ensure pool outlives vec
Or keep pool and container in same scope:
// ✅ Pool and container lifetimes match
void process() {
std::pmr::monotonic_buffer_resource pool;
std::pmr::vector<int> vec(&pool);
// ... use ...
} // vec, pool destroyed in order
Error 2: Expecting deallocate from monotonic
Symptoms: Memory not freed, keeps accumulating.
Cause: monotonic_buffer_resource’s deallocate is a no-op.
// ❌ Wrong expectation
std::pmr::monotonic_buffer_resource pool;
std::pmr::vector<int> v(&pool);
v.push_back(1);
v.pop_back(); // Internally calls deallocate, but pool doesn't free
// Memory remains in pool until release() called
Solution: Use monotonic only for release() full reset.
// ✅ monotonic: release at scope end
{
std::pmr::monotonic_buffer_resource pool;
std::pmr::vector<int> v(&pool);
v.push_back(1);
v.pop_back();
pool.release(); // Full pool reset at this point
}
Error 3: Assuming copies keep the source’s resource
Symptoms: Allocations silently go to the global heap instead of your pool, or a “fast” move is actually a slow copy.
A polymorphic_allocator never propagates on copy assignment, move assignment, or swap (propagate_on_container_* are all false). Each container keeps the resource it was constructed with, and data is copied into that resource. The consequences differ by operation:
std::pmr::monotonic_buffer_resource pool1, pool2;
std::pmr::vector<int> a(&pool1);
a.push_back(42);
std::pmr::vector<int> b(&pool2);
b = a; // safe: elements are copied INTO pool2; b keeps pool2
std::pmr::vector<int> c = a; // surprise: c uses the DEFAULT resource (usually the heap), not pool1
std::pmr::vector<int> c2(a, &pool1); // what you probably meant: copy into pool1
std::pmr::vector<int> d(std::move(a)); // move construction keeps pool1 (cheap pointer steal)
b = std::move(d); // different resources: elements are moved one by one into pool2, O(n)
The copy construction case is the trap. Copy construction calls select_on_container_copy_construction(), which for polymorphic_allocator returns a default-constructed allocator, meaning get_default_resource(). Code that copies a pmr container “within” a request arena, or passing it by value, therefore allocates from the global heap without any warning. Pass the resource explicitly when copying (c2 above), and check with get_allocator().resource() in tests.
Move assignment between containers with different resources cannot just steal the buffer (it belongs to another resource), so it degrades to an element-wise move with fresh allocations. And swap of two containers with unequal allocators is undefined behavior. Keep containers that interact on the same resource, or treat cross-resource operations as copies.
In my experience, the copy-construction surprise is the reason many first pmr migrations show no speedup at all: the profile still shows malloc, because the objects that were supposed to live in the arena were copied out of it on the way to where they are used. A statistics-collecting resource (section 6) in a debug build makes this visible immediately.
Error 4: Insufficient stack buffer size
Symptoms: monotonic allocates additional memory from upstream (heap).
// ❌ Buffer may be too small
char buffer[256];
std::pmr::monotonic_buffer_resource pool{buffer, sizeof(buffer)};
std::pmr::vector<int> v(&pool);
for (int i = 0; i < 1000; ++i) v.push_back(i); // Exceeds 256 bytes → uses heap
Solution: Allocate generous buffer size.
// ✅ Generous buffer
std::array<std::byte, 65536> buffer;
std::pmr::monotonic_buffer_resource pool{
buffer.data(), buffer.size(),
std::pmr::new_delete_resource()
};
Error 5: Inappropriate pool_options settings
Symptoms: Memory waste or allocation failures with synchronized_pool_resource.
Solution: Configure based on actual maximum block size used.
// ✅ Settings matching usage pattern
std::pmr::pool_options opts;
opts.largest_required_pool_block = 64; // Mostly ≤64-byte objects
opts.max_blocks_per_chunk = 128;
std::pmr::synchronized_pool_resource pool{opts};
Error 6: Ignoring alignment
Symptoms: Crash or SIGBUS on certain platforms.
// ❌ Wrong implementation (ignoring alignment)
void* do_allocate(std::size_t bytes, std::size_t alignment) override {
return upstream_->allocate(bytes, 1); // Ignoring alignment!
}
Solution: Always pass requested alignment.
// ✅ Respecting alignment
void* do_allocate(std::size_t bytes, std::size_t alignment) override {
return upstream_->allocate(bytes, alignment);
}
Error 7: Thread safety misconceptions
Symptoms: Crash or data corruption in multithreaded scenarios.
// ❌ Dangerous: multiple threads sharing same unsynchronized_pool
std::pmr::unsynchronized_pool_resource pool;
std::thread t1([&] { std::pmr::vector<int> v(&pool); /* ... */ });
std::thread t2([&] { std::pmr::vector<int> v(&pool); /* ... */ });
Solution: Use synchronized_pool_resource for shared pools in multithreaded code.
// ✅ synchronized_pool_resource (thread-safe)
std::pmr::synchronized_pool_resource pool;
Best Practices
Pool selection guide
| Pattern | Recommended | Reason |
|---|---|---|
| Frame/request unit, bulk free | monotonic | No deallocate, reset with release() |
| Individual object repeated alloc/free | pool_resource | Block reuse |
| Thread-local independent allocation | Thread-local monotonic | Eliminates lock contention |
| Debugging/profiling | Wrap with stats_memory_resource | Track allocation count & peak |
Pool lifetime rules
Always: pool lifetime ≥ all containers using it
Use pmr for nested container elements too
// ✅ Both map keys and values are pmr types
std::pmr::map<std::pmr::string, std::pmr::vector<int>, std::less<>> m(&pool);
Apply after profiling
// Step 1: Identify bottleneck with profiling
// Step 2: Introduce pool only in that function/scope
// Step 3: Measure performance, then expand application
Use statistics resource in debug builds
#ifdef NDEBUG
std::pmr::memory_resource* resource = std::pmr::get_default_resource();
#else
static stats_memory_resource stats{std::pmr::get_default_resource()};
std::pmr::memory_resource* resource = &stats;
#endif
std::pmr::vector<int> data(resource);
// ... processing ...
#ifndef NDEBUG
std::cout << "Allocations: " << stats.allocation_count()
<< ", Peak: " << stats.peak_allocated() << " bytes\n";
#endif
Performance Benchmarks
Benchmark 1: Measure allocations, not push_back
A benchmark only says something about allocators if the code under test actually allocates many times. A std::vector<int> that calls reserve(N) first allocates exactly once, whatever resource it uses, so timing its push_back loop measures integer stores, not memory management. Node-based containers and containers of strings are where allocation cost shows up:
#include <chrono>
#include <iostream>
#include <list>
#include <memory_resource>
#include <string>
template <class F>
long long time_us(F f) {
auto t0 = std::chrono::steady_clock::now();
f();
auto t1 = std::chrono::steady_clock::now();
return std::chrono::duration_cast<std::chrono::microseconds>(t1 - t0).count();
}
int main() {
constexpr int N = 200'000;
long long sink = 0; // keep results alive so the work is not optimized away
auto heap = time_us([&] {
std::list<std::string> l;
for (int i = 0; i < N; ++i) l.emplace_back(40, 'x'); // node + string buffer
sink += l.size();
});
auto mono = time_us([&] {
std::pmr::monotonic_buffer_resource arena;
std::pmr::list<std::pmr::string> l(&arena);
for (int i = 0; i < N; ++i) l.emplace_back(40, 'x');
sink += l.size();
}); // arena releases everything at once here
auto pool = time_us([&] {
std::pmr::unsynchronized_pool_resource p;
std::pmr::list<std::pmr::string> l(&p);
for (int i = 0; i < N; ++i) l.emplace_back(40, 'x');
sink += l.size();
});
std::cout << "heap " << heap << " us, monotonic " << mono
<< " us, pool " << pool << " us (" << sink << ")\n";
}
Each element here needs two allocations (the list node and a 40-character string that does not fit in the small-string buffer), so allocation dominates. Build with optimizations (-O2), run it several times, and compare medians.
What to expect, and why numbers vary. The monotonic arena usually wins clearly in this pattern, because allocation is a pointer bump and destruction frees nothing individually: the whole arena is released at once. The pool resource helps less here, and more in workloads that repeatedly free and reallocate same-sized blocks. The size of the gain depends heavily on the platform’s default allocator: glibc’s malloc, jemalloc, mimalloc, and the Windows heap differ a lot for small allocations, and a modern thread-caching malloc narrows the gap considerably. That is why this article does not quote a fixed speedup. Measure on your target platform with your allocation pattern, and use a profiler to confirm that allocation was significant in the first place.
When pmr does not help
- Few, large allocations: the cost is dominated by touching the memory, not by the allocator call.
- Long-lived objects freed in random order: a monotonic arena never reuses memory until release, so it grows without bound. Use a pool or the default heap.
- Cross-thread handoff: objects allocated from a thread’s arena and freed on another thread need a synchronized resource or careful lifetime design.
- Code that copies pmr containers by copy construction (Error 3): allocations quietly return to the heap.
Production Patterns
Pattern 1: Game frame pool
void gameLoop() {
std::array<std::byte, 1024*1024> frame_buffer;
while (running) {
std::pmr::monotonic_buffer_resource frame_pool{
frame_buffer.data(), frame_buffer.size(),
std::pmr::new_delete_resource()
};
std::pmr::vector<Entity> entities(&frame_pool);
std::pmr::vector<Component> components(&frame_pool);
// ... create entities, frame logic ...
} // frame_pool, entities destroyed → same buffer reused next frame
}
Pattern 2: HTTP request scope
void handleRequest(const Request& req) {
std::pmr::monotonic_buffer_resource request_pool(
std::pmr::new_delete_resource());
std::pmr::vector<std::pmr::string> path_segments(&request_pool);
std::pmr::map<std::pmr::string, std::pmr::string, std::less<>>
headers(&request_pool);
parsePath(req.uri, path_segments);
parseHeaders(req.raw_headers, headers);
Response response = processRequest(path_segments, headers);
sendResponse(response);
} // request_pool destroyed → all allocations freed
Pattern 3: Per-worker pool in thread pool
class Worker {
std::pmr::monotonic_buffer_resource worker_pool_;
std::pmr::vector<Task> local_queue_{&worker_pool_};
public:
Worker() : worker_pool_(std::pmr::new_delete_resource()) {}
void process(Task t) {
worker_pool_.release();
local_queue_.clear();
// Process t using local_queue_ allocations
}
};
Pattern 4: Hierarchical resources (upstream)
std::pmr::synchronized_pool_resource global_pool;
std::pmr::monotonic_buffer_resource thread_pool{&global_pool};
std::pmr::monotonic_buffer_resource frame_pool{&thread_pool};
std::pmr::vector<int> frame_data(&frame_pool);
// frame_pool exhausted → thread_pool → global_pool
Pattern 5: Implementation checklist
- Pool lifetime exceeds all containers using it
- Clear release() timing for monotonic (frame/request end)
- No copy/assignment between pmr containers using different pools
- Generous stack buffer size when used
- pool_options match actual usage patterns
- Use
synchronized_pool_resourceor per-thread pools for multithreading
Summary
| Topic | Summary |
|---|---|
| Memory pools | Allocate large block once, distribute slots — reduces allocation count & fragmentation |
| std::pmr | memory_resource + polymorphic_allocator for injecting pools into containers |
| monotonic | Sequential allocation, reset only — perfect for frame/request scopes |
| pool_resource | Fixed-size block reuse — suitable for general object pools |
| Custom resource | Inherit memory_resource for logging, statistics, thread-local pools |
Core principles:
- Pool lifetime > container lifetime
- monotonic frees only via
release() - Beware copy/assignment between containers from different pools
- Apply after profiling confirms actual benefits
Related Articles
- C++ Custom Allocators for STL Containers: Pool, Stack, Tracking Allocators and PMR
- C++ Cache-Efficient Code: Data-Oriented Design Guide
- std::chrono guide
- C++ SIMD & Parallelization: std::execution & Intrinsics Guide
- Building a memory pool in C++
Frequently Asked Questions (FAQ)
Q. When is a custom allocator or std::pmr actually worth it?
A. When profiling shows malloc/free dominating execution time, when allocating many short-lived objects per frame/request (games, HTTP servers), or when node-based data structures repeatedly allocate fixed-size blocks.
Q. monotonic vs pool_resource?
A. Use monotonic for “allocate many, free all at once” patterns (frames/requests). Use pool_resource for repeated individual alloc/free patterns with fixed sizes. std::pmr enables control over memory pools and fragmentation. Next, read SIMD & Intrinsics (#39-3). Previous: High-Performance C++ #39-1: Cache & Data-Oriented Design Next: High-Performance C++ #39-3: SIMD & Parallelization