C++ Compilation Process Explained: Preprocessing, Compiling, Assembling, and Linking

Key takeaways

From C++ source to executable: preprocessing, compilation, assembly, and linking. Where name mangling happens, how symbols resolve, and how Makefiles and include paths fit in.

The compilation pipeline

From source to executable, C++ builds go through preprocessing → compilation → assembly → linking. Name mangling happens during compilation; symbol resolution happens at link time. Makefiles and include paths automate this workflow.

Source (.cpp) → preprocess → compile → assemble → link → executable

Each arrow in this pipeline is a genuinely separate program you can run by hand, which is worth doing once just to see what each stage actually produces — g++ file.cpp -o file hides all four steps behind one command, and that convenience is exactly what makes link errors and preprocessor bugs confusing the first time you hit them, since the error doesn’t tell you which stage it came from unless you know the pipeline.

Preprocessing

// main.cpp
#include <iostream>
#define PI 3.14
int main() {
    std::cout << PI << std::endl;
}
# Inspect preprocessor output
g++ -E main.cpp -o main.i

Opening the resulting main.i file is a genuinely useful debugging habit: #include <iostream> alone expands to thousands of lines of declarations, and PI is replaced with the literal text 3.14 everywhere it appears. If a macro isn’t expanding the way you expect, or you suspect a header is pulling in something unwanted, -E shows you exactly what the compiler sees before any C++ syntax rules are even applied.

Compilation

# C++ to assembly
g++ -S main.i -o main.s

This is the stage where C++ syntax, templates, and overload resolution are all fully resolved and turned into architecture-specific assembly text. It’s also where name mangling happens — int add(int, int) becomes something like _Z3addii in the generated assembly, encoding the parameter types into the symbol name so the linker can later tell overloaded functions apart.

Assembly

# Assembly to machine code
g++ -c main.s -o main.o

The assembler turns human-readable assembly text into actual machine code, packaged into an object file (.o). Critically, this object file is not yet a runnable program — any function or variable this translation unit references but doesn’t define (like std::cout’s implementation) is left as an unresolved symbol, to be filled in at the next stage.

Linking

# Object file to executable
g++ main.o -o main

Linking is where all those unresolved symbols from every object file get matched up against their actual definitions, whether those definitions live in another object file you’re compiling or in a prebuilt library. This is also the stage where the two error classes covered later in this guide — “undefined reference” and “multiple definition” — actually surface, since the compiler itself never sees the full picture across files; only the linker does.

Building Step by Step, Multiple Sources, Libraries, and -O Flags

Example 1: Step-by-step build

# All-in-one
g++ main.cpp -o main
# Step by step
g++ -E main.cpp -o main.i    # preprocess
g++ -S main.i -o main.s      # compile
g++ -c main.s -o main.o      # assemble
g++ main.o -o main           # link

Example 2: Multiple sources

# Compile each file
g++ -c main.cpp -o main.o
g++ -c util.cpp -o util.o
# Link
g++ main.o util.o -o myapp

Compiling each .cpp file into its own .o before linking is what makes incremental builds possible: if only util.cpp changes, a build system only needs to redo the compile step for util.o and then relink, instead of recompiling main.cpp from scratch. This is the entire reason build tools like Make and CMake track per-file object outputs rather than always doing a single all-in-one command.

Example 3: Linking libraries

# Static library
g++ main.o -L./lib -lmylib -o myapp
# Shared library
g++ main.o -L./lib -lmylib -Wl,-rpath,'$ORIGIN/lib' -o myapp

The practical difference shows up after the build finishes. A statically linked library’s code is copied directly into myapp, so the executable runs standalone; a shared library stays a separate .so/.dll file that must be found at runtime, which is what -Wl,-rpath is for — without it, running myapp on a machine where the shared library isn’t on the system’s default search path fails with a runtime “error while loading shared libraries” message, even though the build itself succeeded.

Two details trip people up here. First, a relative rpath such as -Wl,-rpath,./lib is resolved relative to the current working directory when the program starts, not relative to the executable. The program works when launched from the build directory and fails when launched from anywhere else. $ORIGIN (quoted, so the shell does not expand it) means “the directory containing the executable”, which is almost always what you want. Second, the order of arguments matters for static libraries with the GNU linker. It processes files left to right and only pulls from a static library the symbols that are still unresolved at that point:

g++ -L. -lutil main.o -o app   # undefined reference to `add(int, int)'
g++ main.o -L. -lutil -o app   # OK: the library comes after the objects that need it

When the library is listed first, nothing needs it yet, so the linker skips it, and by the time main.o asks for add, the library is behind it. Libraries that depend on each other must also be ordered from most dependent to least dependent, and circular dependencies need -Wl,--start-group ... -Wl,--end-group. The macOS linker and LLD are more forgiving, so a link line that works on a Mac can fail on a Linux CI machine. This is the kind of bug that makes people think the library is broken, when only its position on the command line is.

Example 4: Optimization flags

# Debug symbols
g++ -g main.cpp -o main
# Optimization
g++ -O2 main.cpp -o main
# Warnings
g++ -Wall -Wextra main.cpp -o main

-g and -O2 are often used as if they were opposites, and in practice pushing both together can make debugging confusing: an optimizer is free to reorder instructions, inline functions, and eliminate variables entirely, so a debugger stepping through -O2-compiled code can jump around in ways that don’t match the source line-by-line. Debug builds typically use -g -O0, and only apply real optimization once the logic is verified correct. GCC and Clang also offer -Og, which enables optimizations that do not interfere with debugging, a reasonable default for day-to-day development builds.

Keeping -g in release builds is common and costs nothing at runtime: debug information lives in separate sections that are not loaded when the program runs, and it is what makes crash dumps from production readable. Many teams build with -O2 -g and strip the debug info into a separate file for shipping.

Preprocessor directives

// Include
#include <iostream>   // system header
#include "myheader.h" // project header
// Macros
#define MAX 100
#define SQUARE(x) ((x) * (x))
// Conditional compilation
#ifdef DEBUG
    #define LOG(x) std::cout << x
#else
    #define LOG(x)
#endif
// Include guard
#ifndef MYHEADER_H
#define MYHEADER_H
// contents
#endif

The angle-bracket vs quote distinction in the first two includes isn’t just style: <...> tells the preprocessor to search system include paths first, while "..." searches relative to the current file’s directory first. Getting this backwards is a common source of a header resolving to the wrong file when a project header happens to share a name with a system one.

Issue 1: Missing header

// ❌ No header
int main() {
    std::cout << "Hello";  // error
}
// ✅ Include header
#include <iostream>
int main() {
    std::cout << "Hello";
}
# undefined reference
# → missing implementation or library not linked
g++ main.o -lmissing_lib -o myapp

“Undefined reference” is specifically a linker error, not a compiler error — the compiler was perfectly happy that a function was declared somewhere, and only the linker discovers, once it tries to assemble the final executable, that no object file or library actually provides a matching definition. This distinction matters when debugging: if the error mentions mangled symbol names and comes after “collect2” or mentions ld in GCC’s output, you’re looking at a linking problem, and the fix is almost always either writing the missing function or adding -lname for the library that provides it.

The “almost always” hides four recurring causes worth knowing by sight:

  • Library order (see Example 3): the library is on the command line, but before the objects that need it.
  • Template defined in a .cpp file. A template is compiled only where it is instantiated, and instantiation needs the definition. If twice<T> is declared in a header but defined in twice.cpp, then main.cpp sees only the declaration, emits a call to twice<int>, and nothing ever generates it: undefined reference to 'int twice<int>(int)'. Put template definitions in the header, or add explicit instantiations (template int twice<int>(int);) in the .cpp for the types you support.
  • C code called from C++ without extern "C". The C++ compiler mangles the name it calls (_Z3addii), while the C library exports the plain name add. Wrapping the C header’s declarations in extern "C" { ... } makes the names match.
  • ABI mismatches. With GCC, code compiled with the old and new std::string ABIs (the _GLIBCXX_USE_CXX11_ABI macro) cannot be linked together, and the error mentions std::__cxx11::basic_string in a symbol that you can see is defined. Prebuilt binaries compiled against a different standard library version or compiler produce the same kind of mismatch.

The quickest way to tell which case you have is to look at what the libraries actually export: nm -C libfoo.a | grep add lists the symbols with demangled names. If the symbol is missing entirely, it is a definition or template problem. If it appears with a slightly different signature or namespace (for example with and without __cxx11), it is an ABI or extern "C" problem. That one command has resolved more “impossible” link errors for me than any amount of rereading the build scripts.

Issue 3: Multiple definitions

// header.h
int globalVar = 10;  // ❌ ODR violation if included in multiple TUs
// ✅ Declaration only (header)
extern int globalVar;
// source.cpp
int globalVar = 10;  // single definition

This is the mirror-image problem of Issue 2: instead of zero definitions, the linker finds the same definition duplicated once per translation unit that included the header. The extern keyword tells the compiler “this variable is defined somewhere else — just take my word its type and don’t allocate storage for it here,” deferring the actual allocation to exactly one .cpp file.

Since C++17 there is a simpler option: inline int globalVar = 10; in the header. An inline variable, like an inline function, may be defined in every translation unit that includes it, and the linker merges the copies into one object. Include guards do not help with this error: they prevent a header from being included twice within one .cpp file, but every .cpp file is preprocessed separately and gets its own copy.

Issue 4: Circular dependencies

// a.h
#include "b.h"
// b.h
#include "a.h"  // cycle
// ✅ Forward declaration
class B;  // forward declaration

Include guards (or #pragma once) stop the infinite preprocessing loop this cycle would otherwise cause, but they don’t fix the underlying design problem — each header still ends up seeing an incomplete definition of the other class at the point it needs it. A forward declaration (class B;) is often enough when you only need a pointer or reference to B in a.h; you only need the full #include "b.h" in the .cpp file where B’s members are actually used.

Compiler Options Worth Knowing

# Language standard
g++ -std=c++17 main.cpp
# Warnings
g++ -Wall -Wextra -Werror main.cpp
# Optimization
g++ -O0  # none
g++ -O1  # basic
g++ -O2  # recommended for release
g++ -O3  # aggressive
# Debug
g++ -g main.cpp
# Preprocessor define
g++ -DDEBUG main.cpp

Beyond the classic pipeline: LTO, build speed, and modules

Link-time optimization. Normally each .cpp file is optimized alone, so the compiler cannot inline a function defined in another file. With -flto, the compiler writes an intermediate representation into the object files and the real optimization happens at link time across the whole program. It often gives a few percent of speed and smaller binaries, at the cost of much slower (and more memory-hungry) links. All objects and static libraries that participate should be built with the same compiler and -flto.

Where build time goes. Because every .cpp file re-preprocesses and re-parses every header it includes, large projects spend most of their compile time on headers, not on their own code. The usual remedies are forward declarations instead of includes in headers, precompiled headers, ccache to reuse results when nothing changed, and faster linkers such as lld or mold. Clang’s -ftime-trace produces a per-file timeline that shows exactly which headers and templates are expensive, which is the best starting point before guessing.

C++20 modules change the pipeline itself. A module interface is compiled once into a binary module interface (BMI), and importers read that instead of re-parsing text, so macros from one module no longer leak into another. The trade-off is that build order now matters: a file that imports a module cannot be compiled until that module’s BMI exists, so the build system must scan dependencies before compiling. CMake supports this from version 3.28 with the Ninja generator. C++23 adds import std;, which replaces including standard headers. Compiler and build-tool support has matured quickly, but mixing modules with older libraries is still the area where projects hit the most friction.

-Werror is worth calling out specifically: it promotes every warning to a hard compile error, which sounds aggressive but catches a large class of real bugs (uninitialized variables, signed/unsigned comparison mismatches, unused results) before they ever reach a test suite. Teams that enable it from day one tend to have far fewer of these bugs than teams that try to retrofit it onto an existing codebase full of accumulated warnings.

FAQ

Q1: What are the build stages?

A: Preprocessing → compilation → assembly → linking.

Q2: What is a .o file?

A: An object file. It contains machine code for one translation unit.

A:

  • Missing function definitions
  • Library not linked
  • Undefined symbols

Q4: Which optimization level?

A:

  • -O0: debugging
  • -O2: typical release builds
  • -O3: maximum optimization

Q5: How do I inspect preprocessing?

A: Use g++ -E.

A: Most often library order: the GNU linker only resolves symbols from a static library that are already needed when it reaches it, so list libraries after the objects that use them.