C++ Buffer Overflows: Causes, Safe APIs, and Security Impact

Key takeaways

Buffer overflows in C and C++: strcpy, memcpy, stack and heap corruption, ASan, strncpy vs string, bounds checks, and secure coding patterns.

What is a buffer overflow?

A buffer overflow happens when a program writes more bytes into a fixed-size block of memory than that block can hold. The extra bytes do not vanish — they land in whatever memory happens to sit immediately after the buffer, silently overwriting it. That neighboring memory might be another local variable, a saved CPU register, a function’s return address, or bookkeeping data the heap allocator uses to track free and allocated chunks. The defining trait of a buffer overflow is that the bug and the crash are often separated in both time and space: the write happens in one function, but the visible failure — a segfault, corrupted output, or a security exploit — can surface much later, in a completely unrelated part of the program.

char buffer[10];
strcpy(buffer, "This is too long");  // overflow

In this example, buffer has room for 10 bytes, but the string literal plus its null terminator needs 18 bytes. strcpy does not know the destination’s capacity — it just keeps copying until it hits the source’s null terminator. The eight extra bytes spill past the end of buffer into whatever memory follows it on the stack. If that happens to be a saved return address, you have moved from “wasted memory” to “attacker-controlled instruction pointer” in a single line of code.

Why C++ never checks bounds for you

New C++ developers coming from Python, Java, or JavaScript are often surprised that arr[15] on a 10-element array does not throw an exception — it just reads or writes garbage memory, or someone else’s data, with no warning. This is not an oversight; it is a deliberate design philosophy carried over from C: “don’t pay for what you don’t use.” Every bounds check costs a branch and a comparison. For code running in a tight loop over millions of elements — a rendering pipeline, a network packet parser, a numerical kernel — that check, multiplied by billions of iterations, is a real and measurable cost. C++ gives you operator[] with zero safety net specifically so you’re not forced to pay for a check you might already know is unnecessary, because you validated the index once, up front, outside the loop.

The trade-off is that this decision is yours to make correctly, every single time, and the compiler will not remind you. Contrast this with std::vector::at(), which does bounds-check and throws std::out_of_range — the standard library gives you both options and lets you choose per call site. Most modern languages made the opposite default choice: safety first, opt-out for performance (unsafe blocks in Rust, unchecked array access in some JIT-compiled languages). C++‘s default is unchecked, opt-in safety. Neither philosophy is “wrong,” but you have to know which one you’re writing in, because the failure mode of getting it wrong in C++ is not an exception — it’s undefined behavior.

Stack buffer overflows: overwriting the return address

A stack frame for a typical function is laid out with local variables, saved registers, and the return address all sitting in the same contiguous block of memory. When a local char array is declared, it lives on that stack frame, and depending on the compiler’s layout decisions, it can sit at a lower address than the saved return address. Writing past the end of that array — often via strcpy, gets, or a hand-rolled loop with an off-by-one — walks forward through memory and can overwrite the return address the CPU will jump to once the current function finishes.

flowchart TB
    subgraph frame["Stack frame (grows downward)"]
        A["local char buffer[64]"] --> B["saved base pointer"]
        B --> C["return address"]
        C --> D["caller's stack frame"]
    end
    E["strcpy writes past end of buffer"] -->|"overflow spills upward"| B
    E -->|"overflow continues"| C
    C -->|"function returns"| F["CPU jumps to return address"]
    F -->|"attacker controls this value"| G["arbitrary code execution"]

This is exactly why stack buffer overflows are a classic remote code execution vector: an attacker who can control the contents written into an undersized buffer can also control what value ends up sitting where the return address used to be. When the vulnerable function returns, the CPU does not jump back to the caller — it jumps to wherever the attacker’s bytes pointed it. This is the mechanism behind decades of “smashing the stack” exploits, from the 1988 Morris worm’s fingerd overflow to modern CTF pwn challenges.

Modern toolchains ship several mitigations, and it’s worth being precise about what they actually do: they are mitigations, not fixes for the underlying bug.

  • Stack canaries place a known sentinel value between local buffers and the return address. Before returning, the function checks whether the canary is still intact; if an overflow overwrote it, the program aborts instead of jumping to corrupted memory. A canary stops the classic “overwrite the return address” exploit, but it does nothing for overflows that only corrupt other local variables, and sufficiently clever attacks can sometimes bypass or leak the canary value.
  • ASLR (Address Space Layout Randomization) randomizes where the stack, heap, and shared libraries are loaded in memory each run, making it harder for an attacker to know the absolute address to jump to. It does not stop the overflow itself, and information leaks elsewhere in the program can defeat it by revealing the randomized base address.
  • NX bit / DEP (No-eXecute / Data Execution Prevention) marks stack and heap memory as non-executable, so even if an attacker gets a pointer to point at injected shellcode sitting on the stack, the CPU refuses to execute it. This pushed attackers toward return-oriented programming (ROP), which chains together existing executable code fragments instead of injecting new code — the mitigation raised the bar without eliminating the underlying vulnerability.

None of these three touch the actual bug: an unchecked write past the end of a buffer. They reduce the odds that the bug becomes a working exploit, and in combination they make exploitation meaningfully harder, but the only real fix is not writing past the buffer in the first place.

Heap buffer overflows: corrupting allocator metadata

Heap overflows are, in some ways, nastier to debug than stack overflows because the corruption is further removed from the crash. Most heap allocators store bookkeeping metadata — chunk size, free-list pointers, allocation flags — immediately before or after the memory block they hand back to you. When you overflow a heap-allocated buffer, you are very often not corrupting “someone else’s variable”; you are corrupting the allocator’s own internal data structures.

void heapOverflow() {
    char* buffer = new char[64];
    strcpy(buffer, longString);
    delete[] buffer;
}

If longString is longer than 64 bytes, the overflow can corrupt the metadata belonging to the next chunk in the heap. The program usually does not crash at the strcpy call — it crashes later, whenever that corrupted chunk is next allocated, freed, or coalesced with a neighbor, because that is when the allocator actually reads the metadata you clobbered. This is why heap corruption bugs are notorious for producing a crash in malloc, free, or an unrelated new call dozens of lines — sometimes seconds of wall-clock time — after the actual overflow. The stack trace at the crash site tells you almost nothing about where the bug actually lives.

I have personally chased this exact failure mode: a service that intermittently died inside glibc’s allocator with a “corrupted size vs. prev_size” abort, always in a code path that had nothing to do with the buffer that was actually at fault. Grepping the crash stack trace for the bug was a dead end because the crash location and the bug location were in different translation units entirely. What actually resolved it was rebuilding with -fsanitize=address and re-running the same workload — AddressSanitizer poisons the “redzone” bytes immediately surrounding every heap allocation, so it caught the one-byte overflow the instant it happened, with a stack trace pointing directly at the offending strcpy, instead of at the innocent allocator call that happened to trip over the damage minutes later. If you only remember one piece of practical advice from this article, it’s this: when a crash looks like “the allocator is broken,” don’t start by reading allocator internals — build a sanitizer instrumented binary first, because it turns a heisenbug into a reproducible, pinpointed one nearly every time.

Off-by-one errors: the single most common root cause

In practice, most real-world buffer overflows are not dramatic security exploits — they’re a boring <= where a < belonged, or a forgotten byte of space for the null terminator. These are worth calling out explicitly because they are so easy to write and so easy to miss in review.

char buf[10];
strcpy(buf, "Long string");
gets(buf);  // never use
int arr[10];
arr[15] = 42;
char* ptr = buffer;
ptr[20] = 'x';
char src[20] = "Hello";
char dst[5];
memcpy(dst, src, 20);

Two patterns show up again and again. First, loop bounds written with <= instead of <: a for (int i = 0; i <= 10; i++) over a 10-element array writes to arr[10], one past the last valid index, on the final iteration. It compiles cleanly, and it will pass any test that doesn’t specifically probe the boundary. Second, forgetting that a C-string of N visible characters needs N + 1 bytes of storage for the terminating '\0'. char buffer[6] looks like room for “Hello” (5 characters) with one byte to spare — and it is, exactly one byte, with zero margin for anything longer. strncpy(buffer, "Hello", 5) is a classic near-miss here too: it copies at most 5 bytes and — this is the part people forget — does not null-terminate the destination if the source is exactly as long as (or longer than) the limit. buffer[5] = '\0'; after the call is not decoration; it is the only thing making the string valid. Skip it, and every subsequent strlen or string-printing call reads past the buffer looking for a terminator that was never written, silently pulling in adjacent stack or heap memory until it happens to find a zero byte somewhere.

gets(buf) deserves a specific callout: it has no length parameter at all — there is no way to tell it the buffer’s capacity, which is precisely why it was removed from the C11 standard library entirely rather than merely deprecated. If you see gets in a codebase, it is not a style nitpick; it is an unconditional bug regardless of how the calling code is structured, because the function cannot be used safely under any circumstances.

I once spent an afternoon on an off-by-one in a fixed-size log-line buffer that had shipped and worked fine for months. Every test fixture, every staging log message, every manually typed debug string fit comfortably under the buffer’s limit, so the missing bounds check never fired. It only overflowed in production when a downstream service started emitting a slightly longer error string than anyone had anticipated — a field got appended to an error payload — and the extra bytes walked straight into an adjacent struct field that happened to hold a status flag. The visible symptom was a status flag flipping to a nonsense value nowhere near the actual copy; the real bug was a char buf[128] sized for “what we’ve seen so far” rather than “the documented maximum,” with no truncation or length check enforcing that assumption anywhere in the code. Fixed-size buffers sized from observed data rather than a real contract are a recurring source of exactly this kind of overflow, and they pass code review constantly because reviewers, like tests, tend to eyeball realistic-looking sample inputs rather than the theoretical maximum.

Security impact

void vulnerable(const char* input) {
    char buffer[64];
    strcpy(buffer, input);
}

If input is attacker-controlled — command-line arguments, network payloads, file contents, environment variables — this single line is a potential remote code execution vulnerability, not just a crash bug. CWE-120 (“Buffer Copy without Checking Size of Input”) and its many CWE-787/CWE-121/CWE-122 relatives sit consistently near the top of MITRE’s list of most dangerous software weaknesses precisely because of how directly an unchecked copy translates into control-flow hijacking, privilege escalation, or information disclosure. This is also why fuzzing tools like AFL and libFuzzer specifically hammer functions that parse untrusted input with malformed and oversized data — that is exactly the class of bug they are built to surface.

Safer String, Array, Input, and Copy Code

String APIs

#include <cstring>
#include <string>
void unsafeCopy(const char* src) {
    char buffer[10];
    strcpy(buffer, src);
}
void safeCopy(const char* src) {
    char buffer[10];
    strncpy(buffer, src, sizeof(buffer) - 1);
    buffer[sizeof(buffer) - 1] = '\0';
}
void useString(const char* src) {
    std::string str = src;
}

safeCopy bounds the copy to sizeof(buffer) - 1 bytes, reserving one byte for the null terminator, and then explicitly writes that terminator itself rather than trusting strncpy to do it. That explicit terminator write is not optional, for the reason explained above. But notice how much simpler useString is: std::string owns its own storage and grows to fit whatever you assign it, so the entire category of “did I reserve enough bytes” bugs disappears. Unless you have a specific, measured reason to manage a fixed-size C buffer — embedded targets with no heap, a hot path where you’ve profiled and confirmed the allocation matters — std::string should be the default, not the exception.

Array access

#include <vector>
#include <stdexcept>
void unsafeAccess(int index) {
    int arr[10];
    arr[index] = 42;
}
void safeAccess(int index) {
    int arr[10];
    if (index >= 0 && index < 10) {
        arr[index] = 42;
    }
}
void vectorAccess(int index) {
    std::vector<int> vec(10);
    try {
        vec.at(index) = 42;
    } catch (const std::out_of_range& e) {
        std::cerr << "Out of range: " << e.what() << std::endl;
    }
}

safeAccess manually validates the index before use — cheap, explicit, and easy to audit, but it relies on the developer remembering to write the check at every single call site. vectorAccess moves the check into the type itself: at() always validates and throws on failure, so it is impossible to forget. The trade-off is the exception-handling overhead and the branch that operator[] skips. A reasonable rule of thumb: use at() while an index comes from untrusted or externally derived input, and reserve unchecked operator[] for indices you’ve already validated or that are provably in range from the surrounding logic (a loop counter bounded by .size(), for instance).

User input

#include <iostream>
#include <string>
void safeInput() {
    char buffer[64];
    if (fgets(buffer, sizeof(buffer), stdin)) {
        buffer[strcspn(buffer, "\n")] = '\0';
    }
}
void cppInput() {
    std::string input;
    std::getline(std::cin, input);
}

fgets is the bounded, safe replacement for gets: it takes the buffer’s capacity as an explicit parameter and will never write past it, though it may leave a trailing newline character in the buffer that gets would have stripped, which is why strcspn is used here to trim it. std::getline sidesteps the whole problem the same way std::string did above — no fixed capacity to overrun, no newline artifact to clean up.

Memory copy

#include <cstring>
#include <algorithm>
void safeCopy(const char* src, size_t srcLen) {
    char dst[64];
    size_t copySize = std::min(srcLen, sizeof(dst) - 1);
    memcpy(dst, src, copySize);
    dst[copySize] = '\0';
}

memcpy is even more dangerous than strcpy in one specific way: it has no concept of a null terminator at all and will happily copy binary data containing embedded zero bytes. It is purely “copy exactly N bytes,” so the caller bears the entire responsibility for computing a safe N. Clamping srcLen against the destination’s capacity with std::min before calling memcpy is the whole fix — the bug in the original, unbounded version is not in memcpy itself, it’s in trusting srcLen without checking it against sizeof(dst).

snprintf, Concatenation, and Null Terminator Pitfalls

sprintf vs snprintf

void safeFormat(int value) {
    char buffer[10];
    snprintf(buffer, sizeof(buffer), "Value: %d", value);
}

sprintf has the exact same defect as strcpy: no capacity parameter, so it writes as many bytes as the format string produces regardless of the destination’s actual size. snprintf takes the buffer size and truncates rather than overflowing — and its return value tells you how many bytes would have been written had truncation not occurred, which is useful for detecting truncation after the fact rather than silently accepting cut-off output.

Concatenation

void cppConcat() {
    std::string str = "Hello";
    str += " World";
}

C-string concatenation with strcat has the same missing-bounds problem as strcpy, compounded by the fact that you also need to know how much of the destination buffer is already used. std::string::operator+= reallocates as needed, which removes the bug class entirely at the cost of occasional reallocation — a cost that is almost always worth paying outside of the tightest hot paths.

Uninitialized buffers

char buffer[64] = {0};
std::array<char, 64> buffer2{};

An uninitialized char array does not start out as zeros — it holds whatever bytes happened to be left on the stack from a previous function call, which is why uninitialized-buffer bugs sometimes leak old data (private data from an earlier function’s stack frame) rather than crashing outright. Zero-initializing with = {0} or std::array’s value-initializing {} closes that leak and also means a missing null terminator degrades to “empty string” instead of “read until we find a stray zero somewhere in memory.”

Off-by-one and null terminator

char buffer[6];
strncpy(buffer, "Hello", 5);
buffer[5] = '\0';

This is the pattern discussed above, shown as a minimal correct example: five bytes of content, a sixth byte reserved and explicitly set for the terminator. Change either number without changing the other and you are back to an overflow or a non-terminated string.

Mitigation quick reference

strncpy(dst, src, sizeof(dst) - 1);
snprintf(buffer, sizeof(buffer), format, args);
fgets(buffer, sizeof(buffer), stdin);

Each of these three functions shares the same shape: the destination’s capacity is passed in explicitly, as data, rather than being implicit in the function’s contract. That single change — turning an implicit assumption into an explicit parameter — is what separates every “banned” C string function from its safer counterpart.

Detection tools

g++ -fsanitize=address -g program.cpp
valgrind --tool=memcheck ./program
clang-tidy program.cpp
cppcheck program.cpp

AddressSanitizer (-fsanitize=address) instruments every memory access at compile time and surrounds heap and stack allocations with “redzones” — small regions of poisoned memory that immediately trigger a fault the instant an overflow touches them. This is why ASan catches overflows at the exact instruction that causes them, with a full stack trace, instead of waiting for the corruption to eventually crash something unrelated later, as described in the heap-corruption story above. Valgrind’s Memcheck does similar work without recompilation but runs your program under a much slower software emulator, which makes it a reasonable choice for a quick one-off check and a poor choice for CI on every commit. Static analyzers like clang-tidy and cppcheck catch a meaningful subset of these bugs — obvious fixed-size copies, suspicious loop bounds — without even running the program, which makes them cheap enough to run on every pull request, but they cannot reason about values that only become unsafe at runtime, so they complement rather than replace ASan.

Input Validation and a Bounds-Checked Array Wrapper

void validateInput(const char* input, size_t maxLen) {
    if (strlen(input) > maxLen) {
        throw std::invalid_argument("input too long");
    }
}
template<typename T, size_t N>
class SafeArray {
    T data[N];
public:
    T& operator[](size_t index) {
        if (index >= N) {
            throw std::out_of_range("index out of bounds");
        }
        return data[index];
    }
};
class Buffer {
    std::unique_ptr<char[]> data;
    size_t size;
public:
    Buffer(size_t s) : data(std::make_unique<char[]>(s)), size(s) {}

    void write(const char* src, size_t len) {
        if (len > size) {
            throw std::length_error("buffer overflow");
        }
        std::memcpy(data.get(), src, len);
    }
};

These three patterns generalize the fixes above into reusable building blocks. validateInput rejects oversized input at the boundary, before it ever reaches a fixed-size buffer downstream — validating early is cheaper than validating at every internal call site, and it produces a clear, actionable error instead of undefined behavior. SafeArray is essentially what std::array already gives you if you swap its operator[] for .at(), but writing it out makes the trade-off explicit: every access now costs a comparison, in exchange for a thrown exception instead of memory corruption on misuse. Buffer wraps a heap allocation with its own size and enforces that size on every write, which is the same idea std::vector and std::string already implement internally — if you find yourself writing a class shaped like Buffer, that is usually a signal to reach for the standard library type it is reimplementing, unless you have a concrete reason (a custom allocator, a fixed memory arena) that the standard container does not accommodate. std::span, added in C++20, is also worth knowing here: it lets a function accept “a pointer plus a length” as one bounds-carrying object instead of two separate, easily-desynchronized parameters, which removes an entire class of “I forgot to pass the right length” bugs at function boundaries.

FAQ

Q1: When do overflows happen?

A: Unsafe C APIs without a length parameter (strcpy, gets, sprintf), array indices that aren’t validated against bounds, and missing or incorrect input-length validation before a fixed-size copy.

Q2: Security risk?

A: Overwriting a stack return address or heap allocator metadata can lead to arbitrary code execution, privilege escalation, or corrupted data — the specific outcome depends on exactly what memory gets overwritten and how it’s later used.

Q3: Prevention?

A: Prefer length-aware, bounds-checked APIs and standard containers (std::string, std::vector, std::span, .at()) over raw C buffers and unbounded C string functions.

Q4: Detection?

A: AddressSanitizer for immediate, precisely-located reports at the moment of the overflow; Valgrind’s Memcheck as a slower recompilation-free alternative; static analyzers (clang-tidy, cppcheck) for catching obvious cases before the program even runs.

Q5: Safer functions?

A: snprintf, fgets, and standard library types (std::string, std::vector) instead of sprintf, gets, strcpy, and raw memcpy on unchecked lengths.

Q6: Resources?

A: Secure Coding in C and C++ (Seacord), the OWASP guidance on buffer overflows, and MITRE’s CWE-120 entry for the formal classification and related weakness IDs.