Debugging Crashes with Core Dumps: Enabling Them, Reading Them in GDB, Collecting Them in Production

Key takeaways

How to turn on core dumps, load them in GDB to find where and why a process crashed, automate their collection, configure production systems, and use advanced analysis techniques.

Introduction: Why Core Dump is Needed

How do you debug when a program suddenly crashes in production? Core Dump is essential for solving hard-to-reproduce bugs, intermittent Segmentation Faults, and memory corruption issues.

Value of Core Dump:

  • Complete memory snapshot at crash time
  • All variable values and stack traces
  • Post-mortem analysis without reproduction
  • Real data from production environment What This Guide Covers:
  • Core Dump generation configuration
  • Core Dump analysis with GDB
  • Practical debugging scenarios
  • Automation and monitoring

A core dump is most valuable exactly when other tools fail: the crash happens once a week under production load, never under a debugger, and the log line just before it says nothing useful. With the core file and the matching binary, you can open the process as it was at the moment it died and ask it questions, which is often the only way to get from “it segfaults sometimes” to a specific line and a specific bad value.

The catch is that core dumps are disabled or diverted by default on most modern Linux systems, and the most common experience is not a hard analysis but a missing file: Segmentation fault (core dumped) printed, and no core anywhere. Much of this guide is about where the dump actually goes and what stops it, before getting to GDB.

Core Dump Basics

What is Core Dump?

Core Dump is saving memory contents to a file when a program terminates abnormally (crashes).

flowchart TB
    Program[Program Running] --> Crash["Crash Occurs\nSIGSEGV, SIGABRT etc"]
    
    Crash --> Kernel[Kernel Receives Signal]
    
    Kernel --> Check{Core Dump\nEnabled?}
    
    Check -->|Yes| Dump["Generate Memory Dump\ncore.12345"]
    Check -->|No| Exit[Program Terminates]
    
    Dump --> File["Core Dump File\n- Stack\n- Heap\n- Registers\n- Variable Values"]
    
    File --> Debug[Analyze with GDB]

Information in Core Dump

1. Stack Memory
   - Function call stack
   - Local variables
   - Function arguments
2. Heap Memory
   - Dynamically allocated memory
   - Objects created with malloc/new
3. Register State
   - PC (Program Counter)
   - SP (Stack Pointer)
   - General purpose registers
4. Global Variables
   - Static variables
   - Global objects
5. Shared Library Mappings
   - Loaded .so file list
   - Memory address mappings

Signals That Cause Crashes

// Signals that generate Core Dump
SIGQUIT    // Quit (Ctrl+\)
SIGILL     // Illegal Instruction
SIGABRT    // Abort (assert failure)
SIGFPE     // Floating Point Exception (divide by zero)
SIGSEGV    // Segmentation Fault (invalid memory access)
SIGBUS     // Bus Error (alignment error)
SIGTRAP    // Trace/Breakpoint Trap
SIGSYS     // Bad System Call

A core file is an ELF file containing the process’s writable memory segments (stack, heap, data, and anonymous mappings) plus a “notes” section with the registers of every thread, the signal that killed it, and the list of mapped files. Code and read-only data from the executable and shared libraries are normally not copied in, because they can be read back from the files on disk. That is why analysis needs the exact same binary and libraries: GDB combines the memory in the core with the code and symbols from the files. Which segments are written is controlled per process by /proc/<pid>/coredump_filter, which matters when a process has huge shared-memory mappings you do not want in every dump.

Not every crash produces a core. A process killed with SIGKILL (including by the OOM killer) never dumps, and neither does one that exits normally after catching a fatal signal in its own handler. If a crash handler in your code or a framework catches SIGSEGV to log something, make sure it restores the default action and re-raises the signal afterwards, or the core is lost.


Core Dump Generation Setup

ulimit Configuration

# Check current Core Dump size limit
ulimit -c
# 0 → Core Dump disabled
# unlimited → No limit
# Enable Core Dump (current session)
ulimit -c unlimited
# Permanent setting (all users)
echo "* soft core unlimited" | sudo tee -a /etc/security/limits.conf
echo "* hard core unlimited" | sudo tee -a /etc/security/limits.conf
# Specific user only
echo "myuser soft core unlimited" | sudo tee -a /etc/security/limits.conf
# For systemd services
# /etc/systemd/system/myapp.service
[Service]
LimitCORE=infinity

Core Dump Save Location

# Check current setting
cat /proc/sys/kernel/core_pattern
# Default: core (creates 'core' file in current directory)
# Set custom pattern
sudo sysctl -w kernel.core_pattern=/var/crash/core.%e.%p.%t
# Permanent setting
echo "kernel.core_pattern=/var/crash/core.%e.%p.%t" | sudo tee -a /etc/sysctl.conf
# Pattern variables:
# %e - Executable name
# %p - PID
# %t - Timestamp (Unix time)
# %s - Signal number
# %u - UID
# %g - GID
# %h - Hostname
# Example: core.myapp.12345.1234567890

Before changing core_pattern, look at what it already says, because on most distributions it does not start with a file name. If the value begins with |, the kernel pipes the dump to a program instead of writing a file: |/usr/lib/systemd/systemd-coredump ... on systemd-based systems (Fedora, RHEL, Arch, and Ubuntu when systemd-coredump is installed), or |/usr/share/apport/apport ... on a default Ubuntu desktop. Apport is the classic cause of “it said core dumped but there is no file”: it keeps crash reports for packaged programs in /var/crash as .crash files and ignores binaries that are not part of an installed package, such as the one you just compiled. Either install systemd-coredump, which takes over the pattern, or disable apport (sudo systemctl disable --now apport) and set a file pattern.

A plain relative pattern like the default core is written into the crashing process’s current working directory, which must be writable by that process. For daemons that directory is often /, so the dump silently fails. Use an absolute path, such as the /var/crash pattern above, and create the directory with permissions that let the service user write to it. Also note that %e is the process’s command name, which the kernel truncates to 15 characters, so a binary called payment-gateway-worker shows up as payment-gateway; scripts that map core files back to binaries must account for that.

In containers there is one more layer: core_pattern is a host-wide setting, not per container. A pipe pattern runs the handler in the host’s context, so the dump appears in the host’s coredumpctl list, not inside the container. A file pattern is interpreted inside the crashing process’s mount namespace, so the directory must exist in the container. Kubernetes nodes are often configured with neither in mind, and collecting cores from pods usually means setting the pattern on the node image deliberately.

# Install systemd-coredump (Ubuntu)
sudo apt install systemd-coredump
# Configuration
sudo nano /etc/systemd/coredump.conf
[Coredump]
Storage=external
Compress=yes
ProcessSizeMax=2G
ExternalSizeMax=2G
MaxUse=10G
# Core Dump save location
# /var/lib/systemd/coredump/
# List Core Dumps
coredumpctl list
# Specific Core Dump info
coredumpctl info <PID>
# Extract Core Dump
coredumpctl dump <PID> -o core.dump
# Analyze directly with GDB
coredumpctl debug <PID>

systemd-coredump is the easiest setup to live with: dumps are compressed, indexed by PID, executable, signal and time, and old ones are removed automatically once MaxUse is reached. It still honors the crashing process’s core size limit, so a service with LimitCORE=0 (or a shell with ulimit -c 0) shows up in the journal as Resource limits disable core dumping for process 12345 (myapp) and no dump is stored. coredumpctl list shows a COREFILE column: present means the dump is available, missing means it was already removed, and none means only the metadata was recorded (for example because the process exceeded ProcessSizeMax).


GDB Core Dump Analysis

Basic Analysis Flow

# 1. Compile with debug symbols
g++ -g -O0 myapp.cpp -o myapp
# 2. Configure to generate Core Dump
ulimit -c unlimited
# 3. Run program (crash occurs)
./myapp
# Segmentation fault (core dumped)
# 4. Open Core Dump with GDB
gdb ./myapp core.12345
# Or use systemd-coredump
coredumpctl debug 12345

GDB Basic Commands

# After loading Core Dump
# 1. Backtrace (stack trace)
(gdb) bt
#0  0x00005555555551a9 in crash_function () at myapp.cpp:10
#1  0x00005555555551c8 in main () at myapp.cpp:15
# Detailed backtrace
(gdb) bt full
#0  0x00005555555551a9 in crash_function () at myapp.cpp:10
        ptr = 0x0
        value = 42
#1  0x00005555555551c8 in main () at myapp.cpp:15
        argc = 1
        argv = 0x7fffffffe0a8
# 2. Move to specific frame
(gdb) frame 0
(gdb) frame 1
# 3. Check source code
(gdb) list
5       void crash_function() {
6           int* ptr = nullptr;
7           int value = 42;
8           
9           // Segmentation Fault!
10          *ptr = value;
11      }
# 4. Check variable values
(gdb) print ptr
$1 = (int *) 0x0
(gdb) print value
$2 = 42
# 5. All local variables
(gdb) info locals
ptr = 0x0
value = 42
# 6. Register state
(gdb) info registers
rax            0x0                 0
rbx            0x555555557d60      93824992247136
rsp            0x7fffffffddd0      0x7fffffffddd0
rip            0x5555555551a9      0x5555555551a9 <crash_function()+15>
# 7. Check memory contents
(gdb) x/10x $rsp
0x7fffffffddd0: 0xffffddf0      0x00007fff      0x555551c8      0x00005555
# 8. Thread information
(gdb) info threads
  Id   Target Id         Frame
* 1    Thread 0x7ffff7fc1740 (LWP 12345) crash_function () at myapp.cpp:10
  2    Thread 0x7ffff77c0700 (LWP 12346) 0x00007ffff7bc8e2d in poll ()
# 9. Switch to specific thread
(gdb) thread 2
(gdb) bt

The order of those commands is the usual order of an investigation. bt tells you where the crash happened; frame #0 is the innermost function, and it is often inside libc or the standard library (memcpy, std::string internals) even though the bug is yours, so walk up with frame N or up until you reach your own code. Then check the values that fed the faulting line. In this example the answer is immediate (ptr = 0x0), but real crashes usually involve a pointer that is not null, just wrong: a small value like 0x10 suggests a member access through a null object pointer (the 0x10 is the member’s offset), and an address full of a repeated pattern suggests freed or overwritten memory.

On an optimized build, expect <optimized out> for many locals and arguments: the value lived only in a register that was reused before the crash. Inlined functions also appear as separate frames that you cannot inspect as fully. That does not make the dump useless; the backtrace and the registers are still correct, and info registers plus x/i $pc (the faulting instruction) often tell you which pointer was bad. For services where this matters, building with -O2 -g (optimized, but with debug info kept in a separate file, as shown later) is a better trade-off than shipping -O0.

You can also create a core from a process that has not crashed: gcore <pid> (or generate-core-file inside GDB) writes a dump of a live process and lets it continue running. That is useful for a service that is hung or using unexpected memory, where killing it would lose the evidence. It needs ptrace permission for the target, which is where the ptrace: Operation not permitted error in the troubleshooting section comes from.

Practical Example 1: Null Pointer Dereference

// crash_null.cpp
#include <iostream>
void process_data(int* data) {
    std::cout << "Processing: " << *data << std::endl;  // Crash!
}
int main() {
    int* ptr = nullptr;
    process_data(ptr);
    return 0;
}
# Compile (with debug symbols)
g++ -g -O0 crash_null.cpp -o crash_null
# Run
./crash_null
# Segmentation fault (core dumped)
# GDB analysis
gdb ./crash_null core
(gdb) bt
#0  0x00005555555551b2 in process_data (data=0x0) at crash_null.cpp:4
#1  0x00005555555551e1 in main () at crash_null.cpp:9
(gdb) frame 0
(gdb) print data
$1 = (int *) 0x0
(gdb) list
1       #include <iostream>
2
3       void process_data(int* data) {
4           std::cout << "Processing: " << *data << std::endl;  // Crash here!
5       }
# Cause: data is nullptr
# Solution: Add nullptr check

Practical Debugging Scenarios

Scenario 1: Double Free

// crash_double_free.cpp
#include <cstdlib>
#include <iostream>
int main() {
    int* ptr = new int(42);
    
    std::cout << "Value: " << *ptr << std::endl;
    
    delete ptr;
    
    // Double free!
    delete ptr;
    
    return 0;
}
# Compile
g++ -g -O0 crash_double_free.cpp -o crash_double_free
# Run
./crash_double_free
# Value: 42
# free(): double free detected in tcache 2
# Aborted (core dumped)
# GDB analysis
gdb ./crash_double_free core
(gdb) bt
#0  0x00007ffff7a42387 in raise () from /lib/x86_64-linux-gnu/libc.so.6
#1  0x00007ffff7a43a78 in abort () from /lib/x86_64-linux-gnu/libc.so.6
#5  0x0000555555555234 in main () at crash_double_free.cpp:12
(gdb) frame 5
(gdb) print ptr
$1 = (int *) 0x555555559eb0  # Already freed address
# Solution: Set ptr = nullptr after delete

This scenario shows a limitation of core dumps worth understanding. The crash is at the second delete, where glibc’s allocator notices the chunk is already in its free list and calls abort(). The core tells you where the second free happened, not where the first one did. In this toy program they are two lines apart; in a real program the first free may have happened in a different thread, seconds earlier, in unrelated code that believed it owned the object. Heap corruption in general works this way: the crash is where the damage is noticed, not where it was done. Frames #2 to #4, left out above, are glibc’s __libc_message and malloc_printerr, which is how you recognize an allocator abort in a backtrace.

For this class of bug, a core dump is the second tool, not the first. AddressSanitizer (-fsanitize=address) reports the double free with both stacks, the one that freed the memory first and the one freeing it again, and Valgrind does the same without recompiling, at a large speed cost. The core is what you have when the bug only reproduces in production, where those tools are too slow; in that case, the address of the freed object and the type of the enclosing object (found by walking up the frames) are the clues to search the code with.


Automation and Collection

Automatic Core Dump Collection Script

Manually hunting for core files after every crash doesn’t scale — a collection script keeps every crash, its binary metadata, and a first-pass backtrace together for later triage.

#!/bin/bash
# collect_core.sh — run from a crash hook or cron watching /var/crash
CORE_DIR="/var/crash"
ARCHIVE_DIR="/var/crash/archive"
mkdir -p "$ARCHIVE_DIR"
for core in "$CORE_DIR"/core.*; do
    [ -f "$core" ] || continue
    ts=$(date +%Y%m%d_%H%M%S)
    binary=$(echo "$core" | cut -d. -f2)
    dest="$ARCHIVE_DIR/${binary}_${ts}"
    mkdir -p "$dest"
    mv "$core" "$dest/core"
    cp "/usr/local/bin/$binary" "$dest/binary" 2>/dev/null
    gdb -batch -ex "bt full" "$dest/binary" "$dest/core" > "$dest/backtrace.txt" 2>&1
    echo "Collected crash: $dest"
done

The script assumes a file-based core_pattern like /var/crash/core.%e.%p.%t. The binary name taken from the file name is the truncated 15-character command name, and the copy from /usr/local/bin is only correct if that binary has not been replaced by a newer deployment since the crash; archiving the binary at deploy time, keyed by its build ID, is more reliable than copying it at collection time.

Collecting from systemd-coredump

systemd-coredump has no hook that runs your own script per crash; its configuration only controls storage. What you can do is keep its configuration in a drop-in file and let a periodic job (or a .path unit watching /var/lib/systemd/coredump) summarize new entries with coredumpctl:

# /etc/systemd/coredump.conf.d/storage.conf
[Coredump]
Storage=external
Compress=yes
#!/bin/bash
# crash_summary.sh: run from a systemd timer; appends info for recent crashes
coredumpctl list --since=-1h --json=short | jq -r '.[].pid' | while read -r pid; do
    coredumpctl info "$pid" >> /var/log/crash-summary.log
done

Each crash is also written to the journal with the backtrace systemd-coredump produced (journalctl -t systemd-coredump), so if you already ship the journal to a log system, an alert on those entries often replaces a custom collection script entirely.

Grouping Crashes by Signature

Deduplicating crash reports by their top-frame backtrace signature turns thousands of raw core files into a small number of distinct bugs to fix.

#!/bin/bash
# group_crashes.sh — group by top 3 stack frames
for dir in /var/crash/archive/*/; do
    sig=$(grep -A3 "^#0" "$dir/backtrace.txt" | md5sum | cut -d' ' -f1)
    echo "$sig $dir" >> /tmp/crash_signatures.txt
done
sort /tmp/crash_signatures.txt | uniq -c -w32 | sort -rn

Production Environment Setup

Disk Space Management

A crashing service under heavy load can generate core dumps fast enough to fill a disk — cap total usage and compress aggressively.

# /etc/systemd/coredump.conf
[Coredump]
Storage=external
Compress=yes
ProcessSizeMax=2G      # skip dumps larger than this
ExternalSizeMax=2G     # largest single dump kept in external storage
MaxUse=10G              # total space systemd-coredump may use
KeepFree=5G             # always leave this much free

Alerting on New Crashes

Wiring coredumpctl into a monitoring pipeline turns “a crash happened three days ago and nobody noticed” into a page within minutes.

#!/bin/bash
# crash_alert.sh — run periodically via cron/systemd timer
LAST_CHECK_FILE="/var/lib/crash_alert/last_check"
LAST_CHECK=$(cat "$LAST_CHECK_FILE" 2>/dev/null || echo 0)
NOW=$(date +%s)
NEW_CRASHES=$(coredumpctl list --since="@$LAST_CHECK" --no-legend | wc -l)
if [ "$NEW_CRASHES" -gt 0 ]; then
    curl -X POST "$SLACK_WEBHOOK_URL" \
         -d "{\"text\": \"⚠️ $NEW_CRASHES new crash(es) detected on $(hostname)\"}"
fi
echo "$NOW" > "$LAST_CHECK_FILE"

Symbol Management for Stripped Binaries

Production binaries are usually built without -g for size/security reasons — keep a separate debug-symbol package installable alongside the stripped binary so gdb can still resolve a useful backtrace.

# Build: keep symbols in a separate file, strip the shipped binary
g++ -g -O2 myapp.cpp -o myapp
objcopy --only-keep-debug myapp myapp.debug
strip --strip-debug --strip-unneeded myapp
objcopy --add-gnu-debuglink=myapp.debug myapp
# GDB automatically finds myapp.debug next to myapp, or via debug-file-directory
gdb ./myapp core.12345

The link between a core, a binary and a symbol file is the build ID, a hash the linker embeds in the binary (readelf -n myapp shows it as Build ID:). The core records the build IDs of every mapped file, and GDB uses them to find the matching symbols in /usr/lib/debug/.build-id/. That is why the most important production habit is not the strip commands themselves but storing each release’s .debug files somewhere, indexed by build ID, for as long as that release may be running. A core from last month’s build is nearly worthless if the only symbols available are from today’s build: GDB either refuses them or, if they are forced to load, shows plausible-looking but wrong line numbers. Tools like debuginfod serve symbols by build ID over HTTP, and GDB can query them automatically.

Also think about what a core contains before collecting them widely. It holds the process’s memory, which means request bodies, session tokens, decrypted secrets and personal data that happened to be in memory at the time. Treat core files like database backups: restrict who can read them, keep them only as long as needed, and do not attach them to public bug trackers.


Advanced Analysis Techniques

Analyzing Multi-Threaded Crashes

When the crashing thread isn’t obvious, check every thread’s backtrace before assuming the signal-receiving thread caused the actual bug — a race condition often corrupts state from one thread and crashes a different one.

(gdb) thread apply all bt
# Compare stacks across threads; look for the thread holding
# a lock or touching the same memory as the crashing frame
(gdb) info threads
(gdb) thread 3
(gdb) bt full

Inspecting Memory Corruption with x

When a crash is caused by heap corruption rather than a clean null dereference, examining raw memory around the faulting address helps distinguish “use-after-free” (freed-memory poison patterns) from “buffer overflow” (adjacent allocation overwritten).

(gdb) print $rip
(gdb) x/20xg $rsp                 # 20 giant words from the stack pointer
(gdb) x/4xg (void*)ptr - 32       # inspect memory just before a suspect pointer
(gdb) x/s (char*)ptr              # interpret as a string, useful for corrupted buffers

Reverse Debugging with rr

For crashes that are hard to reproduce, Mozilla’s rr records a full execution trace and lets you step backwards through it — instead of guessing where a corruption started, you can run to the crash, then reverse-step to the exact instruction that first wrote bad data.

rr record ./myapp
rr replay
(rr) continue      # run forward to the recorded crash
(rr) reverse-step   # step backwards from the crash point
(rr) watch *ptr     # then reverse-continue to find where *ptr was last written
(rr) reverse-continue

Scripting GDB for Batch Analysis

gdb -batch combined with a command file turns repetitive manual analysis into a one-line report generator, useful when triaging a large archive of collected core dumps.

# analyze.gdb
set pagination off
bt full
info registers
info threads
thread apply all bt
quit
gdb -batch -x analyze.gdb ./myapp core.12345 > report.txt

Troubleshooting

Core Dump Not Generated at All

# Check each layer that can suppress a core dump:
ulimit -c                           # must not be 0
cat /proc/sys/kernel/core_pattern   # must point somewhere writable
df -h /var/crash                    # must have free space
# suid/sgid binaries don't dump by default unless explicitly allowed:
cat /proc/sys/fs/suid_dumpable      # set to 2 ("suidsafe") if needed

”No symbol table is loaded” in GDB

# Confirm the binary was compiled with -g and matches the core exactly
file myapp                    # should mention "with debug_info"
gdb ./myapp core.12345
(gdb) info sharedlibrary      # check if libraries also need debug symbols

Rebuild with -g if the binary is stripped, or install matching -dbg/-debuginfo packages for system libraries (apt install libc6-dbg on Debian/Ubuntu, debuginfo-install glibc on RHEL-family).

Backtrace Shows ?? for Every Frame

Usually means the binary and the core don’t match, or the process was dynamically linked against libraries GDB can’t locate.

(gdb) info sharedlibrary
# Look for "No" in the "Syms Read" column and install the matching debug package

Core File Truncated or Incomplete

Large processes can exceed ProcessSizeMax/ulimit -c, silently producing a partial (unusable) core file.

ulimit -c unlimited
# Also raise the systemd-coredump caps:
sudo sed -i 's/^ProcessSizeMax=.*/ProcessSizeMax=8G/' /etc/systemd/coredump.conf
sudo sed -i 's/^ExternalSizeMax=.*/ExternalSizeMax=8G/' /etc/systemd/coredump.conf

GDB Can’t Attach: “ptrace: Operation not permitted”

Opening a core file never needs ptrace; attaching to a live process or running gcore does. On most distributions the Yama security module limits ptrace to child processes (/proc/sys/kernel/yama/ptrace_scope set to 1), so attaching to an already running process as a normal user fails even when you own it; run as root or temporarily lower the setting. Containerized environments (Docker, Kubernetes) restrict ptrace further by default; grant it explicitly for debugging containers.

docker run --cap-add=SYS_PTRACE --security-opt seccomp=unconfined ...

GDB commands for reading a core file

# Basic
gdb <executable> <core>     # Load Core Dump
bt                          # Backtrace
bt full                     # Detailed backtrace
frame <N>                   # Move to frame N
list                        # Source code
# Variables
print <var>                 # Variable value
print *<ptr>                # Value pointed to by pointer
print <array>[0]@10         # 10 array elements
info locals                 # All local variables
info args                   # Function arguments
# Memory
x/10x <addr>                # Memory (hexadecimal)
x/10s <addr>                # Strings
x/10i <addr>                # Instructions (assembly)
# Threads
info threads                # Thread list
thread <N>                  # Switch to thread N
thread apply all bt         # Backtrace all threads
# Registers
info registers              # All registers
print $rip                  # Specific register
# Other
info sharedlibrary          # Loaded libraries
info proc mappings          # Memory map

When a backtrace is full of ?? frames, the core is usually fine and the problem is the binary you loaded it with. GDB needs the exact executable and shared libraries that produced the core, plus their debug symbols. A rebuilt binary, even from the same commit, can have different addresses, and a core copied from a server to a developer machine will resolve the system libraries against the wrong versions. Keep the stripped debug information for every release you ship (keyed by build ID), and analyze cores on a machine with the same library versions, or point GDB at copies of them with set sysroot.


References


Frequently Asked Questions (FAQ)

Q. Can I analyze a core dump on a different machine than the one where it crashed?

A. Yes, but GDB needs the exact same executable and the same versions of the shared libraries that were loaded at crash time, plus matching debug symbols. If the libraries differ, backtraces through libc or other libraries show wrong frames or ??. Copy the binary and its libraries, or use a container built from the same image, and point GDB at them with set sysroot or set solib-search-path. Keeping build IDs and symbol files for each release makes this much easier.