What new and delete Actually Do in C++: operator new, delete[] Mismatch, Alignment, Placement new and Arenas

Key takeaways

A new expression is two steps: operator new gets raw bytes, then a constructor runs in them; delete runs the destructor, then operator delete returns the bytes. Almost every manual-memory bug (delete vs delete[], malloc/delete, misaligned placement new, leaks on early return) comes from breaking that pairing, and RAII, reserve and arenas are how you stop paying for it.

Where this article fits

The basic split between stack and heap, and why deep recursion or large locals overflow the stack, is covered in Stack vs Heap in C++. This article is about the heap side: what new and delete do mechanically, why the classic bugs crash where they do, and how to allocate less often.

A new expression is two operations

Widget* w = new Widget(1);

This line does two separate things:

  1. It calls the allocation function operator new(sizeof(Widget)), which returns raw, uninitialized bytes (or throws std::bad_alloc).
  2. It runs Widget’s constructor in those bytes. If the constructor throws, the compiler-generated code calls operator delete on the memory before propagating the exception, so a throwing constructor does not leak.

delete w is the mirror image: run ~Widget(), then call operator delete on the address. You can perform the steps by hand, which is exactly what containers like std::vector do internally:

void* raw = ::operator new(sizeof(Widget));   // 1. bytes only
Widget* w = new (raw) Widget(1);              // 2. construct in place (placement new)
w->~Widget();                                 // 3. destroy
::operator delete(raw);                       // 4. release bytes

Seeing the four steps separately explains the rules that otherwise look arbitrary. malloc does step 1 only; free does step 4 only. Mixing malloc with delete runs a destructor on an object that was never constructed and hands the pointer to a different allocator. Mixing new with free skips the destructor and makes the same allocator mismatch in reverse. It often “works” on one platform because operator new happens to call malloc there, which is exactly why it survives code review and fails after someone replaces the global allocator.

delete vs delete[]: why it corrupts the heap

For an array of a type with a non-trivial destructor, new T[n] must remember n so that delete[] can run n destructors. Common ABIs (Itanium, used by GCC and Clang, and MSVC) store that count in a small header just before the first element, and return a pointer past the header. So:

struct Widget { int id = 0; ~Widget() { std::cout << "~Widget\n"; } };

int main() {
    Widget* arr = new Widget[3];
    delete arr;          // should be delete[]
    std::cout << "survived\n";
}

Built with g++ 10.3 (MinGW) this compiles without a warning, and the process dies with exit code 0xC0000374, STATUS_HEAP_CORRUPTION, before printing “survived”. delete ran one destructor and then passed the address of element 0, not the start of the block with the count header, to the allocator. GCC 11 and later add -Wmismatched-new-delete, which catches the simple cases at compile time, and AddressSanitizer reports alloc-dealloc-mismatch at runtime.

For int and other trivially destructible types there is no header, so delete on new int[10] often appears to work. It is still undefined behavior, and it is the reason this bug can live in a codebase for years: it only explodes when someone changes the element type to one with a destructor.

The modern answer is not “remember the brackets” but “never write either”: std::vector<Widget> or std::make_unique<Widget[]>(n) pick the right deallocation for you.

When new fails

Plain new never returns nullptr. It throws:

try { int* huge = new int[1000000000000]; }                 // about 4 TB
catch (const std::bad_alloc& e) { std::cout << e.what(); }   // std::bad_alloc

long long n = -1;
try { int* p = new int[n]; }
catch (const std::bad_array_new_length& e) { std::cout << e.what(); } // std::bad_array_new_length

Only new (std::nothrow) T returns nullptr on failure, so if (p == nullptr) after a plain new is dead code. Be careful about what bad_alloc means on Linux, though: with the default overcommit policy, a request that is large but not absurd often succeeds, the kernel hands out address space without backing pages, and the process is killed by the OOM killer later when it touches the memory. Catching bad_alloc is not a memory-limit strategy on such systems; bounding your own buffer sizes is.

Alignment and aligned new

Every type has an alignment requirement (alignof(T)), and the allocator must return addresses that satisfy it. Before C++17, operator new only guaranteed alignof(std::max_align_t) (typically 16), so this was silently broken:

struct alignas(64) CacheLine { char data[64]; };
auto* c = new CacheLine;   // C++17: calls operator new(size_t, std::align_val_t{64})

Since C++17 the compiler calls the std::align_val_t overload for over-aligned types, and reinterpret_cast<std::uintptr_t>(c) % 64 == 0 holds (checked with g++ 10.3 in -std=c++17). If you compile as C++14, or write a custom allocator that only implements operator new(size_t), you can get a 16-byte-aligned block for a type that promises 64, and SIMD loads or lock-free code built on that promise misbehave. Memory Alignment and Padding covers the layout side of this.

Placement new: construct into memory you already have

Placement new runs a constructor at an address you supply. The original version of this article used a plain char buffer[...], which is a real bug: a char array has alignment 1, and constructing a type that needs 4 or 8 into it is undefined behavior (on x86 it usually works slowly; on some ARM targets and with SIMD types it faults). The correct form:

alignas(Widget) unsigned char buf[sizeof(Widget) * 2];

Widget* a = new (buf) Widget(2);
Widget* b = new (buf + sizeof(Widget)) Widget(3);

std::destroy_at(b);   // C++17; same as b->~Widget()
std::destroy_at(a);
// no delete: buf was never allocated with new

Three rules come with it: the storage must be large enough and aligned, you must call the destructor yourself (and exactly once), and you use the pointer that placement new returned rather than reinterpreting buf. If you really need a type-erased buffer, std::optional<T> and std::variant already do this correctly and are usually what hand-written placement new code should have been.

Leaks: the early return and the exception

The leak that matters in real code is rarely a forgotten delete at the end of a function; it is an exit path between new and delete:

void process(const Config& cfg) {
    auto* buffer = new char[cfg.size];
    if (!validate(cfg)) return;        // leaks
    parse(buffer, cfg);                // if parse throws, leaks
    delete[] buffer;
}

Every return, throw, and break added later is a new leak. The fix is to tie the lifetime to an object:

void process(const Config& cfg) {
    auto buffer = std::make_unique<char[]>(cfg.size);   // or std::vector<char>
    if (!validate(cfg)) return;                          // freed
    parse(buffer.get(), cfg);                            // freed on exception too
}

That is RAII, and it is covered in depth in C++ RAII Explained. One caveat for code that sets pointers to nullptr after delete as a safety measure: it prevents a double delete through that variable, but any other copy of the pointer still dangles. Ownership that lives in one smart pointer is the actual fix; Heap Corruption: Double Free and Wrong delete shows how those bugs surface far from their cause.

Allocation cost and fragmentation

A general-purpose allocator (malloc, and operator new built on it) has to handle any size, from any thread, in any order. That flexibility costs something on every call: finding a fitting free block, possibly taking a lock or touching a per-thread cache, and bookkeeping for the later free. More important than the per-call cost is what many small, long-lived, interleaved allocations do to the heap over time. Free blocks end up scattered between live ones; the total free space may be large, but no single hole fits a bigger request, and pages that hold even one live block cannot go back to the operating system.

I have chased this in long-running services more than once: memory usage climbing steadily for days, leak checkers reporting nothing, and the cause turning out to be fragmentation plus the allocator keeping freed memory cached, made worse by glibc creating multiple per-thread arenas. It looks exactly like a leak on a dashboard. The fixes that worked were never “find the missing delete”; they were reducing the number of small heap allocations in the hot path, and tuning or replacing the allocator.

Fewer allocations: reserve

The cheapest improvement is often just telling containers how much you need. Counting calls to a replaced global operator new:

std::vector<int> v;
for (int i = 0; i < 1000; ++i) v.push_back(i);   // 11 allocations with libstdc++

std::vector<int> w;
w.reserve(1000);
for (int i = 0; i < 1000; ++i) w.push_back(i);   // 1 allocation

The 11 comes from libstdc++ doubling capacity (1, 2, 4, …, 1024); each growth allocates a new block, moves the elements, and frees the old one, leaving a hole of the old size behind.

Arenas with std::pmr

When many objects share a lifetime (one request, one frame, one parse), an arena allocates by bumping a pointer and frees everything at once. C++17’s polymorphic memory resources make this available to standard containers:

#include <memory_resource>

alignas(std::max_align_t) std::byte arena[64 * 1024];
std::pmr::monotonic_buffer_resource pool(arena, sizeof(arena),
                                         std::pmr::null_memory_resource());

std::pmr::vector<std::pmr::string> names(&pool);
for (int i = 0; i < 100; ++i)
    names.emplace_back(40, char('a' + i % 26));   // strings too long for SSO
// global operator new calls in this block: 0

Both the vector and every string inside it allocate from arena; with the counting operator new in place, the block performs zero global allocations. null_memory_resource() as the upstream makes an overflow throw bad_alloc instead of silently falling back to the heap, which is useful while sizing the arena. Watch out for the easy way to defeat it: my first version built each string with "prefix" + std::to_string(i), and that temporary std::string used the global heap 100 times before being copied into the arena. monotonic_buffer_resource also never reuses freed memory until the resource is destroyed, so it suits “build, use, discard” workloads, not containers that churn forever; unsynchronized_pool_resource is the pmr option for reuse.

For fixed-size objects allocated and freed constantly, a free-list pool is the classic alternative; Custom C++ Memory Pools builds one, and C++ Allocators shows how to plug custom allocation into containers.

Choosing where memory comes from

NeedUse
Small, scoped, size known at compile timeautomatic (stack) variable
Single owned object, size or lifetime dynamicstd::make_unique<T>()
Runtime-sized arraystd::vector<T> (reserve when the size is known)
Many objects sharing one lifetimestd::pmr::monotonic_buffer_resource or a custom arena
Many same-size objects with churnpool / free list, or std::pmr::unsynchronized_pool_resource
Object in storage you manage (optional, variant, containers)placement new + std::destroy_at, aligned storage
Raw new / deleteinside the implementation of one of the above, nowhere else