Custom C++ Memory Pools: Fixed Blocks, TLS, and Benchmarks

Why Memory Pools?

The global heap (new/delete) is a general-purpose allocator that handles any size, any lifetime, from any thread. That flexibility costs performance: lock contention, fragmentation, and bookkeeping overhead.

When a hot path allocates many objects of the same size with the same lifetime pattern, a custom pool can be much cheaper than new/delete for that specific case: allocating is popping a node off a free list, freeing is pushing it back, and there is no size-class search, no locking (with a thread-local pool), and no per-block header to maintain. Game engines, network servers, and database implementations all use pools for exactly this reason.

The tradeoff: pools are less flexible. You must know the object size at pool creation time, and you must carefully track which objects belong to which pool.

It is worth being precise about what the pool beats, because modern general-purpose allocators are much better than their reputation. glibc malloc, jemalloc, tcmalloc and mimalloc all keep per-thread caches of small size classes, so a small new on the fast path is already a few dozen instructions without a lock. What a pool still removes is the size-class lookup, the per-allocation header, and above all the unpredictability: no occasional slow path that returns memory to the OS or consolidates free chunks, and a bounded, pre-reserved memory footprint. In latency-sensitive code, that predictability is often the real reason to use a pool, more than the average speed. It also means that before writing one, it is worth trying a faster drop-in allocator, which requires no code changes at all.


Fixed-Size Block Pool

The simplest pool: a slab of memory divided into fixed-size blocks, linked together as a free list. Allocation pops from the list, deallocation pushes back.

#include <algorithm>
#include <cassert>
#include <cstddef>
#include <cstdlib>
#include <new>
#include <iostream>

class FixedBlockPool {
    struct Block {
        Block* next;  // intrusive free list — stored in the block itself
    };

    char*  slab_;        // raw memory
    Block* free_head_;   // head of free list
    size_t block_size_;  // must be >= sizeof(Block*)
    size_t capacity_;    // total number of blocks
    size_t allocated_;   // currently in use

public:
    FixedBlockPool(size_t block_size, size_t capacity)
        : block_size_(std::max(block_size, sizeof(Block*)))
        , capacity_(capacity)
        , allocated_(0)
    {
        // Round block size up to max_align_t alignment
        block_size_ = (block_size_ + alignof(std::max_align_t) - 1)
                    & ~(alignof(std::max_align_t) - 1);

        slab_ = static_cast<char*>(
            std::aligned_alloc(alignof(std::max_align_t),
                               block_size_ * capacity_));
        if (!slab_) throw std::bad_alloc{};

        // Build the free list by linking all blocks
        free_head_ = nullptr;
        for (size_t i = capacity_; i-- > 0; ) {
            Block* b = reinterpret_cast<Block*>(slab_ + i * block_size_);
            b->next = free_head_;
            free_head_ = b;
        }
    }

    ~FixedBlockPool() {
        std::free(slab_);
    }

    // Non-copyable, non-movable
    FixedBlockPool(const FixedBlockPool&) = delete;
    FixedBlockPool& operator=(const FixedBlockPool&) = delete;

    void* allocate() {
        if (!free_head_) return nullptr;  // pool exhausted
        Block* b = free_head_;
        free_head_ = b->next;
        ++allocated_;
        return b;
    }

    void deallocate(void* ptr) {
        if (!ptr) return;
        Block* b = static_cast<Block*>(ptr);
        b->next = free_head_;
        free_head_ = b;
        --allocated_;
    }

    size_t allocated() const { return allocated_; }
    size_t available() const { return capacity_ - allocated_; }
};

The key idea is the intrusive free list: a free block is not being used for anything, so the pool stores the “next free block” pointer inside the block itself. That is why block_size_ must be at least sizeof(Block*), and why the pool needs no side table and no per-block header. The constructor links the blocks in address order (iterating backwards so the first block ends up at the head), which means early allocations are contiguous in memory, a small cache bonus.

Rounding each block up to alignof(std::max_align_t) (usually 16 bytes) guarantees that any ordinary type placed in a block is correctly aligned, at the cost of some internal waste for small objects: a 20-byte object occupies 32 bytes. Types with stricter alignment, such as a struct declared alignas(64) to avoid false sharing, need a pool that takes the alignment as a parameter. Two portability notes: std::aligned_alloc requires the size to be a multiple of the alignment (satisfied here because the block size is rounded), and it is not available in MSVC’s standard library, where _aligned_malloc/_aligned_free or C++17 ::operator new(size, std::align_val_t{...}) are the alternatives.

Exhaustion returns nullptr rather than growing. That is a deliberate design choice: a fixed-capacity pool has a predictable memory ceiling, and the caller decides whether running out is an error, a reason to fall back to new, or a signal to allocate another slab. A growable pool keeps a list of slabs and allocates a new one when the free list is empty; it is not much more code, but it gives up the guaranteed ceiling.

Using the Pool

struct Packet {
    uint32_t id;
    uint32_t length;
    char data[56];  // 64 bytes total
};

int main() {
    FixedBlockPool pool(sizeof(Packet), 1024);

    // Allocate and construct
    void* mem = pool.allocate();
    Packet* p = new(mem) Packet{};  // placement new — construct in pool memory
    p->id = 1;
    p->length = 10;

    // Destroy and return to pool
    p->~Packet();          // explicit destructor — no delete
    pool.deallocate(mem);

    std::cout << "Available: " << pool.available() << '\n';  // 1024
}

The pool hands out raw memory, not objects. Placement new constructs a Packet in that memory, and the explicit destructor call ends its lifetime before the memory goes back. Forgetting the destructor call is harmless for a trivial type like Packet, but for a type holding a std::string or a file descriptor it leaks those resources silently; calling delete p instead is undefined behavior, since the memory did not come from new. Because this pairing is so easy to get wrong, the next section wraps it.


Object Pool (RAII Wrapper)

A cleaner interface that handles construction/destruction automatically:

#include <memory>
#include <functional>

template<typename T>
class ObjectPool {
    FixedBlockPool block_pool_;

public:
    explicit ObjectPool(size_t capacity)
        : block_pool_(sizeof(T), capacity) {}

    // Custom deleter that returns the block to the pool
    using UniquePtr = std::unique_ptr<T, std::function<void(T*)>>;

    template<typename... Args>
    UniquePtr acquire(Args&&... args) {
        void* mem = block_pool_.allocate();
        if (!mem) throw std::bad_alloc{};

        T* obj = new(mem) T(std::forward<Args>(args)...);

        return UniquePtr(obj, [this](T* p) {
            p->~T();                    // call destructor
            block_pool_.deallocate(p);  // return to pool
        });
    }

    size_t available() const { return block_pool_.available(); }
};

// Usage
struct Connection {
    int fd;
    std::string address;
    Connection(int fd, std::string addr) : fd(fd), address(std::move(addr)) {}
    ~Connection() { std::cout << "Connection " << fd << " closed\n"; }
};

int main() {
    ObjectPool<Connection> pool(100);

    auto conn = pool.acquire(42, "192.168.1.1");
    std::cout << "Connected: " << conn->address << '\n';
    std::cout << "Available: " << pool.available() << '\n';  // 99

    // conn goes out of scope — destructor called, block returned to pool
}

std::unique_ptr with a custom deleter gives pooled objects the same ownership semantics as heap objects: they can be moved, stored in containers and returned from functions, and the block goes back to the pool on every exit path, including exceptions. If the constructor of T throws, the block is not returned in this version; wrapping the placement new in a try that calls deallocate before rethrowing closes that gap.

std::function as the deleter type is convenient but expensive for something meant to be fast: it makes each unique_ptr several pointers wide and adds an indirect call on every release. A small deleter struct holding a pointer to the pool, struct Deleter { ObjectPool* pool; void operator()(T* p) const; };, keeps the unique_ptr at two pointers and lets the compiler inline the release. The lambda also captures this, so the pool must not be moved or destroyed while objects are outstanding; here it cannot be moved at all, because FixedBlockPool deletes its copy operations and therefore has no implicit move either.


Thread-Local Pool

For multi-threaded code, each thread gets its own pool. No locks needed:

class TLSPool {
    static constexpr size_t BLOCK_SIZE = 64;
    static constexpr size_t CAPACITY   = 4096;

    struct ThreadPool {
        FixedBlockPool pool{BLOCK_SIZE, CAPACITY};
    };

    // Each thread has its own pool instance
    static thread_local ThreadPool tl_pool_;

public:
    static void* allocate() {
        return tl_pool_.pool.allocate();
    }

    static void deallocate(void* ptr) {
        tl_pool_.pool.deallocate(ptr);
    }
};

thread_local TLSPool::ThreadPool TLSPool::tl_pool_;

The thread-local pool is lock-free — each thread allocates and frees from its own pool without ever touching another thread’s state.

That statement holds only as long as every block is freed by the thread that allocated it, and that is the part people get wrong. In a producer/consumer design, thread A allocates a message and thread B frees it; with this class, B pushes A’s block onto B’s free list. Nothing crashes immediately, but memory migrates between pools: A eventually runs dry while B accumulates blocks it never uses, and if B’s pool is later destroyed while A still references those addresses, you have corruption. Real thread-caching allocators handle this with a “remote free” path, typically a lock-free list per pool onto which other threads push, drained by the owning thread. The simpler alternative is to design ownership so that objects crossing threads come from a shared, locked pool or from the global allocator.

The second trap is lifetime: a thread_local pool is destroyed when its thread exits, and any block from it that is still in use elsewhere becomes a dangling pointer. For short-lived threads, a thread-local pool also pays the full slab allocation for each thread, which can cost more than it saves. Thread-local pools fit best in thread-per-core servers and worker pools with long-lived threads that allocate and free their own data.


Frame Allocator (Linear / Bump Allocator)

For objects that all share the same lifetime (a single game frame, a request handler), a bump allocator is even faster than a free list — just increment a pointer:

class FrameAllocator {
    char*  buffer_;
    size_t capacity_;
    size_t offset_;  // next free byte

public:
    explicit FrameAllocator(size_t capacity)
        : buffer_(new char[capacity])
        , capacity_(capacity)
        , offset_(0)
    {}

    ~FrameAllocator() { delete[] buffer_; }

    void* allocate(size_t size, size_t align = alignof(std::max_align_t)) {
        // Round up offset to alignment
        size_t aligned = (offset_ + align - 1) & ~(align - 1);
        if (aligned + size > capacity_) return nullptr;

        void* ptr = buffer_ + aligned;
        offset_ = aligned + size;
        return ptr;
    }

    // Reset all allocations at once — O(1)
    void reset() {
        offset_ = 0;
        // All pointers from before reset() are now invalid
    }

    size_t used() const { return offset_; }
    size_t remaining() const { return capacity_ - offset_; }
};

// Game loop usage
int main() {
    FrameAllocator frame(1024 * 1024);  // 1 MB per frame

    for (int frameNum = 0; frameNum < 3; ++frameNum) {
        // All per-frame allocations from the frame allocator
        auto* enemies  = static_cast<int*>(frame.allocate(sizeof(int) * 100));
        auto* messages = static_cast<char*>(frame.allocate(256));

        // Use enemies and messages this frame...
        (void)enemies; (void)messages;

        // End of frame — free everything at once
        frame.reset();
        std::cout << "Frame " << frameNum << " done\n";
    }
}

No individual frees — the whole frame is reset at once. This is extremely cache-friendly since all allocations are contiguous.

The alignment line rounds the current offset up to the next multiple of align, which only works when align is a power of two (true for every alignment C++ types can have). Because reset() only moves the offset back, destructors never run. That is fine for plain data like int arrays and char buffers, and a bug for a std::string or std::vector placed in the frame: their heap buffers leak every frame. Either restrict frame allocations to trivially destructible types (a static_assert(std::is_trivially_destructible_v<T>) in a typed helper enforces it) or record destructors to run at reset.

The main design question is what happens when the frame runs out. Returning nullptr, as here, forces every caller to check; many engines instead assert in debug builds and size the buffer from measured peak usage plus a margin. The standard library’s version of this idea is std::pmr::monotonic_buffer_resource, which also falls back to an upstream allocator when its buffer is exhausted and plugs into std::pmr::vector and friends.


Scenarios for Each Pool Type

ProblemBest Pool
Many objects of the same type (particles, packets)Fixed-block pool
Multi-threaded, high-churn allocationsThread-local pool
Per-frame or per-request temp dataFrame/linear allocator
Heterogeneous small objectsSize-class pools (multiple fixed-block pools)
Already using standard containersstd::pmr with monotonic_buffer_resource

Benchmarking Pool vs new/delete

Always benchmark on your actual hardware and workload. Here’s a simple benchmark pattern:

#include <chrono>
#include <vector>
#include <iostream>

template<typename F>
double measureMs(F&& f, int iterations) {
    auto start = std::chrono::steady_clock::now();
    f(iterations);
    auto end = std::chrono::steady_clock::now();
    return std::chrono::duration<double, std::milli>(end - start).count();
}

struct SmallObj { int data[8]; };

void benchNewDelete(int n) {
    std::vector<SmallObj*> ptrs;
    ptrs.reserve(n);
    for (int i = 0; i < n; ++i) ptrs.push_back(new SmallObj{});
    for (auto* p : ptrs) delete p;
}

void benchPool(int n) {
    FixedBlockPool pool(sizeof(SmallObj), static_cast<size_t>(n));
    std::vector<void*> ptrs;
    ptrs.reserve(n);
    for (int i = 0; i < n; ++i) ptrs.push_back(pool.allocate());
    for (auto* p : ptrs) pool.deallocate(p);
}

int main() {
    const int N = 100'000;
    int reps = 10;

    double nd = 0, pool = 0;
    for (int i = 0; i < reps; ++i) {
        nd   += measureMs(benchNewDelete, N);
        pool += measureMs(benchPool, N);
    }

    std::cout << "new/delete: " << nd/reps   << " ms avg\n";
    std::cout << "pool:       " << pool/reps << " ms avg\n";
    std::cout << "speedup:    " << nd/pool   << "x\n";
}

Results vary significantly by platform, allocator implementation, and contention level, so this article does not quote numbers; measure your specific case. The pool usually comes out ahead in a loop like this, but by how much depends heavily on which malloc you are comparing against, and the gap can be small with a modern thread-caching allocator.

This benchmark is also easy to misread, and it is worth knowing its flaws before trusting it:

  • benchPool includes creating the pool (the aligned_alloc and the loop that links every block) inside the timed region, while the heap has already been set up. Construct the pool outside the timer if you want to measure allocation alone.
  • Since C++14, compilers may elide matching new/delete pairs, and with optimizations on, a loop that allocates and never uses the objects can be simplified. Touching each object (writing to data[0]) and consuming a result keeps the work honest.
  • It is single-threaded and frees in allocation order, which is close to the best case for both sides. Real workloads interleave allocation and deallocation, free out of order, and run on several threads, which is where general allocators do extra work. Benchmark with your real allocation pattern, or replay a trace of it.
  • Wall-clock time for one run is noisy. Repeating the run (as reps does), discarding the first warm-up iteration, and using a framework such as Google Benchmark gives more reliable results.

In a latency-sensitive service, I would look at tail latency (p99 and above) as well as averages: the average speedup of a pool may be modest, while the benefit of never hitting the allocator’s slow path shows up in the tail.


Lifetime, ownership, and size mistakes with pools

Pool outlived by objects: if the pool is destroyed while objects are still in use, any subsequent deallocation writes into freed memory — silent corruption.

// WRONG
ObjectPool<Widget>* pool = new ObjectPool<Widget>(100);
auto w = pool->acquire();
delete pool;  // pool gone
// w's deleter will call pool->deallocate — pool is destroyed!

// CORRECT — ensure pool outlives all acquired objects
ObjectPool<Widget> pool(100);
auto w = pool.acquire();
// w destroyed before pool goes out of scope — correct order

Deallocating to the wrong pool: two pools of the same block size still have separate slabs. Returning a block from pool A to pool B corrupts pool B’s free list.

Object larger than block size: the block pool rounds up to alignment, but it won’t grow. If you allocate 128 bytes from a 64-byte pool, you overflow into adjacent blocks — silent corruption. Add an assertion:

void* allocate(size_t requested) {
    assert(requested <= block_size_ && "Requested size exceeds pool block size");
    return allocate();
}

Frame allocator: using pointers after reset: the frame allocator’s reset() does not zero memory — old data lingers. Treat all pointers from before reset() as invalid.

Double free into a pool: with new/delete, a double free is often caught by the allocator (glibc aborts with free(): double free detected in tcache 2). With this pool, deallocating the same block twice silently puts it on the free list twice, and two later allocate() calls return the same memory to two different owners. The symptom shows up much later as two objects mysteriously overwriting each other. A debug-build check that walks the free list, or a per-block “in use” flag, catches it at the point of the mistake.

Hiding bugs from tools: because the pool reuses memory internally, AddressSanitizer and Valgrind no longer see use-after-free or overflows between blocks (see the FAQ below). I keep a compile-time switch that makes every pool forward to plain new/delete; running the test suite with that switch under ASan has caught more pool-related bugs than any amount of code review.


Choosing a pool, or trying std::pmr first

  • Fixed-block pools eliminate allocator overhead for uniform-sized objects — allocation/deallocation is O(1) with no system calls
  • Thread-local pools are lock-free — each thread owns its pool, no contention
  • Frame allocators are the fastest option when all allocations share the same lifetime — reset is O(1)
  • Always benchmark — pools help when the global allocator is a proven bottleneck, not by assumption
  • Lifetime discipline is critical — the pool must outlive all objects it allocates
  • std::pmr (C++17) provides monotonic_buffer_resource, unsynchronized_pool_resource and synchronized_pool_resource for use with standard containers; try them before writing custom pools

Frequently Asked Questions (FAQ)

Q. Will AddressSanitizer catch use-after-free bugs in objects from my pool?

A. Not by default. ASan tracks memory at the malloc/new level, and the pool carves many blocks out of one large allocation, so a block returned to the free list still looks valid to ASan. Mark free blocks with ASAN_POISON_MEMORY_REGION and unmark them on allocation using ASAN_UNPOISON_MEMORY_REGION from <sanitizer/asan_interface.h>, or add a build option that routes the pool to plain new/delete in sanitizer builds.