C++ Segmentation Fault & Core Dump: GDB/LLDB Debugging
Introduction: When you see “Segmentation fault (core dumped)“
For readers who landed here from a search
A segmentation fault means the OS terminated your process after invalid memory access: null dereference, dangling pointers, stack overflow, buffer overrun, etc. Mechanically, the CPU’s memory management unit raised a fault because the address touched was not mapped, or was mapped without the needed permission (writing to a read-only page, executing a non-executable one). The kernel turns that into a SIGSEGV signal, whose default action is to terminate the process and, if allowed, write a core dump: a snapshot of the process’s memory and registers at the moment of the crash.
The important consequence is that a segfault only tells you where the program touched bad memory, not where it went wrong. A null dereference usually crashes at the bug. A use-after-free or buffer overrun often corrupts memory silently and crashes much later in unrelated code, which is why the backtrace from a core dump is the starting point, and AddressSanitizer is often what finds the actual cause. This article walks through enabling core dumps, loading them in GDB/LLDB, reading backtraces, and using AddressSanitizer (ASan) to catch issues earlier. It is a workflow article; the causes themselves, each with a minimal reproduction and the fix, are in C++ Segmentation Fault: Five Causes. See also AddressSanitizer and ThreadSanitizer.
Problem scenarios
Scenario 1: Occasional crashes in production
Logs show only Segmentation fault (core dumped) and no core file.
Cause: ulimit -c is 0, core_pattern discards cores, or systemd limits cores.
Fix: Enable cores → collect with coredumpctl or core_pattern → gdb ./program core and bt.
Scenario 2: Crash inside a third-party library
Cause: Bad pointer arguments, size/type mismatch, use-after-free across the FFI boundary. Fix: bt full, move to caller frames with frame N, inspect arguments; info sharedlibrary for symbols.
Scenario 3: Recursion or huge stack objects
Cause: Default stack size (~8 MB) exceeded. Fix: Repeated frames in bt → recursion; ulimit -s or heap allocation.
Scenario 4: Buffer overrun only on certain inputs
Fix: Re-run under ASan or Valgrind for exact line and size.
Summary
| Scenario | Hint | Approach |
|---|---|---|
| Intermittent prod crash | No core | ulimit, coredumpctl |
Crash in .so | Trace callers | bt full, inspect args |
| Stack overflow | Deep bt | ulimit -s or heap |
| Input-dependent | Bounds | ASan / Valgrind |
Segfault debugging flow
flowchart TB
subgraph Detect[Segfault]
A[Segmentation fault] --> B{Core file?}
end
subgraph NoCore[No core]
B -->|No| C[ulimit -c unlimited]
C --> D[Check kernel.core_pattern]
D --> E[Reproduce]
end
subgraph WithCore[Core available]
B -->|Yes| F[gdb ./program core]
F --> G[bt backtrace]
G --> H[frame 0]
H --> I[print variables]
I --> J{Root cause?}
end
subgraph Prevent[Prevention]
K[ASan build] --> L[Run tests]
L --> M[Fix on first report]
end
E --> F
J -->|Reproducible| K
Enabling core dumps
ulimit
- ulimit -c: if
0, no core file is written. - ulimit -c unlimited: in the shell session that runs the binary.
- For persistence: /etc/security/limits.conf or shell rc files.
ulimit -c
ulimit -c unlimited
echo "ulimit -c unlimited" >> ~/.zshrc
ulimit is a per-process resource limit (RLIMIT_CORE) that children inherit, which is why it must be set in the shell that launches the program. Setting it in one terminal does nothing for a service started by systemd, cron or a container runtime; each of those sets its own limits. An unprivileged process can lower its soft limit or raise it up to the hard limit, so if ulimit -c unlimited fails with “cannot modify limit: Operation not permitted”, the hard limit is the problem and needs to be raised in limits.conf or by whatever starts the process.
Location and naming
- Often core or core.
in the process CWD. - /proc/sys/kernel/core_pattern: e.g. core.%e.%p.%t.
cat /proc/sys/kernel/core_pattern
sysctl kernel.core_pattern
echo "core.%e.%p.%t" | sudo tee /proc/sys/kernel/core_pattern
%e is the executable name, %p the PID and %t the time, so cores from different crashes don’t overwrite each other. A relative pattern like this one writes into the crashing process’s working directory, which silently fails if that directory isn’t writable by the process (a service running as its own user with / as its CWD, for example). An absolute path such as /var/crash/core.%e.%p.%t is more predictable. Writing to /proc/sys lasts until reboot; put kernel.core_pattern=... in /etc/sysctl.d/ to persist it.
If the pattern starts with |, the kernel pipes the core to a program instead of writing a file. That is the most common reason people “can’t find” their core: on systemd distributions it pipes to systemd-coredump (use coredumpctl), and on Ubuntu it often pipes to apport, which may relocate cores or skip binaries that weren’t installed from a package. Check the pattern before assuming no core was written.
systemd
- System-wide default: DefaultLimitCORE=infinity in
/etc/systemd/system.conf. - Per service: LimitCORE=infinity in the unit’s
[Service]section (DefaultLimitCOREis not valid there). - coredumpctl list, coredumpctl debug
or time range.
# /etc/systemd/system/myapp.service
[Service]
LimitCORE=infinity
After editing the unit, run systemctl daemon-reload and restart the service; the limit applies only to processes started after that. With systemd-coredump, cores are stored compressed under /var/lib/systemd/coredump/ and are rotated by size and age according to coredump.conf, so a large core from a crash last week may already be gone.
GDB/LLDB core analysis
Loading a core
Use the same binary build that produced the crash, preferably with -g.
gdb ./your_program core
lldb -c core ./your_program
“Same build” is strict. The core contains memory contents and addresses, not code or symbols; the debugger maps those addresses back to functions using the executable and the shared libraries it loads from your machine. If you rebuild, even from the same commit with different flags, addresses shift and the backtrace becomes nonsense (or GDB warns that the core file may not match the executable). The same applies to shared libraries: analyzing a core from a production server on a laptop with a different glibc version gives wrong frames inside libc. Copy the libraries from the server and point GDB at them with set sysroot or set solib-search-path, or analyze in a container built from the same image.
Backtrace
- bt: call stack; #0 is the faulting frame.
- bt full: locals per frame (GDB); bt -f (LLDB).
Frames and variables
- frame 0, list, print ptr
- *print ptr may fail if the address is invalid in the dump. | Task | GDB | LLDB | |------|-----|------| | Backtrace | bt | bt | | Locals | bt full | bt -f | | Select frame | frame N | frame select N | | Print | print x | frame variable x | | Disassembly | disas | disassemble |
A few habits make backtraces much more useful. In a multi-threaded program, bt shows only the thread that received the signal; thread apply all bt shows every thread, which is essential when the crash is a symptom of a race (another thread freed the object this one is using). In optimized builds, many locals print as <optimized out> because the value only lived in a register that has since been reused; building the reproduction with -O1 -g or RelWithDebInfo keeps most of the program’s behavior while preserving more variables. When a pointer looks valid but crashes on use, info symbol <addr> and info proc mappings (on a live process) tell you whether it points into the heap, a stack, or nowhere.
The first time I opened a core from an optimized release build, frame #0 was in std::vector::operator[], inlined into a function that the source said couldn’t be reached with that index. The backtrace was correct; inlining had simply merged three source functions into one frame. bt marks such frames, and info frame plus disas /s (source interleaved with assembly) is how you work out which inlined call you are really in.
Complete examples
Example 1: Null pointer dereference
// g++ -g -o segfault_demo segfault_demo.cpp && ./segfault_demo
#include <iostream>
int main() {
int* p = nullptr;
std::cout << *p << "\n"; // null deref → segfault
return 0;
}
gdb ./segfault_demo core
(gdb) bt
(gdb) print p
$1 = (int *) 0x0
A null dereference is the easy case: the faulting address is 0 or a small offset from it (for example 0x18 when accessing a member at offset 24 of a null object pointer), and the crash is at the bug. GDB prints the faulting address as si_addr via print $_siginfo._sifields._sigfault.si_addr. Note that with optimization enabled, the compiler may treat dereferencing a known null pointer as undefined behavior and remove or reorder the code, so this tiny demo is only reliable at -O0.
Example 2: Use-after-free
int* p = new int(42);
delete p;
std::cout << *p << "\n"; // UAF
Rebuild with -fsanitize=address and run; ASan prints heap-use-after-free with free and use sites.
Without ASan, this snippet very likely prints 42 and exits normally. delete returns the memory to the allocator but doesn’t unmap it, so the read still hits a mapped page containing stale data. That is what makes use-after-free so dangerous: most of the time nothing happens, and the crash appears only when the allocator has reused the block for another object, often far away in time and code from the actual bug. ASan catches it because it poisons freed memory and delays its reuse (the “quarantine”), so the first illegal access is reported with three stacks: where the memory was allocated, where it was freed, and where it was used.
Example 3: Stack overflow (deep recursion)
bt shows the same function repeated many times → reduce recursion or increase stack / use heap.
The main thread’s stack is typically 8 MB on Linux (ulimit -s), and threads created with std::thread get a default set by the C library, also commonly 8 MB on glibc, but much smaller on some platforms (musl’s default is 128 KB). Besides deep recursion, a large local array (char buf[16 * 1024 * 1024];) overflows the stack on the first call. The core’s backtrace for a deep recursion can be tens of thousands of frames; bt -20 shows the outermost 20 frames, where you’ll see what started the recursion. ASan reports this case as stack-overflow.
Example 4: Buffer overflow
Use ASan: stack-buffer-overflow with file and line. Heap overruns appear as heap-buffer-overflow, and the report says how many bytes past the end of the region the access was, which usually points straight at an off-by-one.
Example 5: Double free
ASan: attempting double-free with allocation and both frees. Without ASan, glibc often detects this itself and aborts with free(): double free detected in tcache 2; that is SIGABRT, not SIGSEGV, but the debugging workflow is the same.
Example 6: LLDB on macOS
lldb -c core ./myapp
(lldb) bt
(lldb) bt -f
(lldb) frame select 0
Enable cores: ulimit -c unlimited; macOS writes cores to /cores/core.<pid> (the directory must be writable). On recent macOS versions, a binary also needs the com.apple.security.get-task-allow entitlement for a core to be written, which debug builds from Xcode have but plain command-line builds may not. ~/Library/Logs/DiagnosticReports/ contains crash reports (.ips files with backtraces of all threads), not cores; for many crashes that report alone is enough.
Common causes
| Cause | GDB / clues | ASan |
|---|---|---|
| Null deref | print ptr → 0x0 | Crash at site |
| UAF | mappings / ASan | heap-use-after-free |
| Stack overflow | Repeated frames | stack-overflow |
| Buffer overrun | Hard to see | stack/heap-buffer-overflow |
| Double free | — | attempting double-free |
Strategies
- Ensure cores or ASan repro.
- bt → frame 0 → print.
- For library crashes, inspect caller frames and arguments.
- CI with ASan tests; Valgrind where ASan is awkward.
flowchart LR
A[Core?] --> B[GDB / print]
A --> C[ASan run]
B --> D[Fix]
C --> D
Defensive patterns
- Check pointers before use; assert in debug.
- delete p; p = nullptr; helps only for that one variable; any copy of the pointer elsewhere still dangles, which is why ownership types are the real fix.
std::unique_ptr/std::shared_ptrmake ownership explicit so that “who frees this” has one answer.- Prefer
snprintf/ std::string / std::span over unchecked C APIs. Avoidstrncpyas a “safe” copy: it doesn’t NUL-terminate when the source is too long, which turns an overflow into a missing terminator and a later overread. - RAII for files, locks, memory.
- Each cause with before-and-after code: C++ Segmentation Fault: Five Causes.
- Use bounds-checked access in debug builds:
_GLIBCXX_ASSERTIONS(libstdc++) or the libc++ hardening modes makeoperator[]on containers check the index and abort with a clear message instead of corrupting memory.
AddressSanitizer
g++ -g -fsanitize=address -fno-omit-frame-pointer -o myapp myapp.cpp
CMake: target_compile_options / target_link_options with -fsanitize=address.
The flag must be passed to both compile and link, because ASan replaces malloc/free with its own runtime; passing it only at compile time fails at link with undefined references to __asan_* symbols. -fno-omit-frame-pointer gives ASan cheap, accurate stack traces. The cost is roughly 2x slowdown and substantially more memory (shadow memory plus the quarantine), which is why it belongs in test and CI builds. ASan can’t be combined with ThreadSanitizer or Valgrind in the same run, and all code that allocates and frees the same objects should be instrumented; mixing an ASan-instrumented binary with an uninstrumented static library can produce false reports at the boundary.
In practice, the first time a team turns on ASan for an existing test suite, it usually reports several bugs that had been “working” for years, typically small overreads of one or two bytes past a buffer. Budget time to fix those before making the ASan job a required CI check.
Production patterns
- systemd: DefaultLimitCORE=infinity, LimitCORE=infinity
- Docker: ulimit / core pattern volume
- Scripts: gdb -batch -ex “bt full” -ex quit on program + core
- coredumpctl: coredumpctl debug, dump -o
- CI: ASan build + ctest
- Split debug info: objcopy —only-keep-debug, debuglink
Two of these need more explanation. In Docker, kernel.core_pattern is a host-wide kernel setting, not a per-container one: a container cannot change it, and if the host pattern is an absolute path, the core is written to that path inside the container’s filesystem, where it vanishes when the container is removed. If the host pattern pipes to systemd-coredump, the core lands on the host instead. Set --ulimit core=-1 on the container and mount a volume at the path the host pattern points to.
Split debug info solves the tension between shipping small, stripped binaries and needing symbols for analysis. Build with -g, then objcopy --only-keep-debug app app.debug extracts the debug sections, strip --strip-debug app removes them from the shipped binary, and objcopy --add-gnu-debuglink=app.debug app records where to find them. Archive the .debug file together with the build ID (readelf -n app shows it); GDB matches a core to the right symbols by build ID, and a debuginfod server can serve them automatically. Without this, a core from a stripped production binary gives a backtrace of raw addresses.
Common tool errors
- Core format mismatch: Same arch, same build machine where possible; file core vs file ./app.
- No symbols: Rebuild with -g, matching binary.
- Cannot access memory: In dumps, some addresses are unmapped—avoid dereferencing in GDB if it fails.
- ulimit cannot change: systemd / limits.conf / Docker —ulimit core=-1
FAQ
Q. No core file?
A. Set ulimit -c unlimited, fix core_pattern, systemd limits, and use coredumpctl on systemd.
Q. No source in GDB?
A. Debug or RelWithDebInfo, -g, same build as the crashing binary.
Q. ASan in production?
A. Typically no—use ASan in test/CI; production uses cores + symbols.
Q. The crash disappears when I run it under GDB or add a print statement. Why?
A. That is typical of memory corruption and uninitialized variables: the debugger or the extra code changes stack layout, timing and heap usage, so the corrupted bytes land somewhere harmless. It is a strong hint to reach for ASan (or MemorySanitizer for uninitialized reads) rather than stepping through code.
Q. Learn more?
A. cppreference, sanitizer docs, GDB/LLDB manuals.
Related Articles
- C++ Segmentation fault: causes and GDB
- C++ debugging basics
- Asio deadlock debugging (#49-3)
- Query optimization (#49-3)
- C++ Runtime Checking
A routine for any segfault
- ulimit -c unlimited (and systemd/Docker as needed).
- gdb ./program core or coredumpctl debug.
- bt → frame 0 → inspect pointers and logic.
- Suspect null, UAF, stack, overrun in order.
- -fsanitize=address when you can reproduce.
- Automate core handling and keep debug symbols.
Next article: CMake link errors LNK2019 / undefined reference Previous article: Custom memory pool (#48-3)
Frequently Asked Questions (FAQ)
Q. The backtrace ends inside free, malloc or memcpy. Is the bug in the standard library?
A. Almost never. A crash inside the allocator usually means the heap was corrupted earlier, for example by a double free or a write past the end of a buffer, and the allocator is only where the damage is noticed. Walk up to the first frame in your own code to see which pointer was passed, then rerun the same input under AddressSanitizer, which reports the original bad write or free instead of the later crash.