Python Meets C++: High-Performance Engines with pybind11
Why pybind11?
Python dominates data science and AI, but pure Python loops can be orders of magnitude slower than equivalent C++. The typical bottleneck is a preprocessing loop, a custom distance metric, or a simulation kernel that runs millions of times.
pybind11 is the standard way to expose C++ as a native Python module. Unlike ctypes (which calls pre-built C libraries) or Cython (which adds a compiled superset of Python), pybind11 lets you write type-safe C++ binding code that integrates naturally with your existing C++ library. It produces a .so (Linux/macOS) or .pyd (Windows) that Python can import like any other module.
The library is header-only, requires no code generation step, and handles reference counting, exception translation, and STL container conversions automatically.
Before writing bindings, it is worth checking that the boundary is in the right place. Every call from Python into C++ pays a fixed cost for argument conversion and dispatch, on the order of what a Python function call costs. Binding a function that adds two numbers and calling it a million times from a Python loop is therefore not faster than Python; the win comes from moving the loop itself into C++, so one call does a lot of work. The best candidates are functions that take an array or a batch and return a result, not fine-grained per-element helpers. Also check whether NumPy vectorization or Numba already solves the problem, since both avoid a C++ build entirely. pybind11 pays off most when you already have a C++ library you want to expose, or when the kernel needs data structures and control flow that NumPy cannot express.
How pybind11 Works
Python script
└── import example # loads example.cpython-311-x86_64.so
└── PYBIND11_MODULE # your C++ binding code
└── your C++ library functions and classes
When Python calls example.add(1, 2):
- Python’s C API calls the extension module’s function pointer
- pybind11 unpacks Python objects into C++ types (
int,std::string, etc.) - Your C++ function runs
- pybind11 converts the return value back to a Python object
- Python receives the result
The conversion layer handles type checking and throws TypeError if arguments don’t match.
Minimal Example
// example.cpp
#include <pybind11/pybind11.h>
#include <pybind11/stl.h> // needed for std::vector <-> list conversion in mean()
#include <vector>
namespace py = pybind11;
int add(int a, int b) { return a + b; }
double mean(std::vector<double> v) {
double s = 0;
for (double x : v) s += x;
return v.empty() ? 0.0 : s / v.size();
}
PYBIND11_MODULE(example, m) {
m.doc() = "Minimal pybind11 example module";
m.def("add", &add, "Add two integers",
py::arg("a"), py::arg("b")); // named args for Python call sites
m.def("mean", &mean, "Compute arithmetic mean of a list",
py::arg("values"));
}
import example
print(example.add(1, 2)) # 3
print(example.add(a=3, b=4)) # 7 — named arguments work
print(example.mean([1.0, 2.0, 3.0])) # 2.0
Note: the name in PYBIND11_MODULE(example, m) must match the filename (example.so/example.pyd) and the import statement.
The <pybind11/stl.h> include is easy to forget and produces a confusing failure: the module compiles, but calling mean([1.0, 2.0]) raises TypeError: mean(): incompatible function arguments. The following argument types are supported: 1. (values: std::vector<double>) -> float, because pybind11 has no converter for std::vector without that header. Include it in every translation unit that binds functions taking or returning STL containers; including it in some files and not others can even produce inconsistent behavior for the same type.
py::arg names make keyword calls work and improve the signature shown by help(example.add). Arguments are converted by value: int accepts Python integers and raises TypeError for floats rather than truncating them, and a Python integer that does not fit in a C++ int raises an error rather than wrapping silently.
Building with CMake
# CMakeLists.txt
cmake_minimum_required(VERSION 3.20)
project(example)
# Find Python and pybind11
find_package(Python3 COMPONENTS Interpreter Development REQUIRED)
find_package(pybind11 CONFIG REQUIRED) # install: pip install pybind11
# Build the extension module
pybind11_add_module(example example.cpp)
target_compile_features(example PRIVATE cxx_std_17)
# Build
cmake -B build
cmake --build build -j$(nproc)
# Test import
python3 -c "import sys; sys.path.insert(0, 'build'); import example; print(example.add(1,2))"
If find_package(pybind11) fails:
pip install pybind11
# Then set the CMake prefix path
cmake -B build -DCMAKE_PREFIX_PATH=$(python3 -c "import pybind11; print(pybind11.get_cmake_dir())")
Alternative: FetchContent
include(FetchContent)
FetchContent_Declare(pybind11
GIT_REPOSITORY https://github.com/pybind/pybind11.git
GIT_TAG v2.12.0)
FetchContent_MakeAvailable(pybind11)
pybind11_add_module does more than add_library(... MODULE ...): it sets the platform-specific suffix (.cpython-311-x86_64-linux-gnu.so, .pyd), hides symbols by default, enables link-time optimization in release builds, and links against Python correctly for each platform. On Linux and macOS, extension modules are not linked to libpython at all, since the symbols come from the interpreter that loads them; linking it manually is a common cause of crashes when the module is imported by a different Python build.
The most frequent build problem is a Python mismatch: CMake finds one interpreter (say, the system Python 3.10), and you import the module from another (a virtual environment with 3.12). The file name encodes the version, so the import fails with ModuleNotFoundError, or with an ABI error if the name was forced. Pass -DPython3_EXECUTABLE=$(which python) (the hint variable matching the find_package(Python3 ...) call above), run from inside the environment you will import from. For distributable packages, building through pip install . with scikit-build-core, which drives CMake with the right interpreter, avoids the issue altogether.
Exposing Classes
#include <pybind11/pybind11.h>
#include <string>
namespace py = pybind11;
class Matrix {
int rows_, cols_;
std::vector<double> data_;
public:
Matrix(int rows, int cols)
: rows_(rows), cols_(cols), data_(rows * cols, 0.0) {}
double get(int r, int c) const { return data_[r * cols_ + c]; }
void set(int r, int c, double v) { data_[r * cols_ + c] = v; }
int rows() const { return rows_; }
int cols() const { return cols_; }
Matrix operator+(const Matrix& rhs) const {
Matrix result(rows_, cols_);
for (size_t i = 0; i < data_.size(); ++i)
result.data_[i] = data_[i] + rhs.data_[i];
return result;
}
std::string repr() const {
return "Matrix(" + std::to_string(rows_) + "x" + std::to_string(cols_) + ")";
}
};
PYBIND11_MODULE(linalg, m) {
py::class_<Matrix>(m, "Matrix")
.def(py::init<int, int>(), py::arg("rows"), py::arg("cols"))
.def("get", &Matrix::get, py::arg("row"), py::arg("col"))
.def("set", &Matrix::set, py::arg("row"), py::arg("col"), py::arg("value"))
.def_property_readonly("rows", &Matrix::rows)
.def_property_readonly("cols", &Matrix::cols)
.def("__add__", &Matrix::operator+)
.def("__repr__", &Matrix::repr);
}
from linalg import Matrix
m = Matrix(3, 3)
m.set(0, 0, 1.0)
m.set(1, 1, 2.0)
print(m) # Matrix(3x3)
print(m.rows) # 3
py::class_<Matrix> creates a Python type whose instances own a C++ Matrix. When the Python object is garbage-collected, the C++ destructor runs. def_property_readonly turns the rows() getter into an attribute, which is more Pythonic than m.rows(). __add__ works because pybind11 wraps the returned Matrix into a new Python object that owns a moved copy; the alternative spelling .def(py::self + py::self) from <pybind11/operators.h> does the same with less boilerplate.
This class trusts its callers completely. get(5, 5) on a 3x3 matrix reads out of bounds and may crash the Python interpreter, and adding matrices of different sizes reads past the end of rhs.data_. In pure C++ that is a caller bug; exposed to Python, where users expect an IndexError, it turns a typo in a notebook into a segfault that kills the kernel. Validate indices and shapes at the binding boundary and throw std::out_of_range (which pybind11 translates to IndexError) or std::invalid_argument (ValueError).
NumPy Integration
The most common use case: pass a NumPy array from Python, process it in C++, return results without copying data.
#include <pybind11/pybind11.h>
#include <pybind11/numpy.h>
#include <cmath>
namespace py = pybind11;
// Compute dot product of two 1D float64 arrays
double dot(py::array_t<double> a, py::array_t<double> b) {
// Request buffer info: pointer, strides, shape
py::buffer_info buf_a = a.request();
py::buffer_info buf_b = b.request();
if (buf_a.ndim != 1 || buf_b.ndim != 1)
throw std::runtime_error("Expected 1D arrays");
if (buf_a.shape[0] != buf_b.shape[0])
throw std::runtime_error("Array size mismatch");
auto* pa = static_cast<double*>(buf_a.ptr);
auto* pb = static_cast<double*>(buf_b.ptr);
ssize_t n = buf_a.shape[0];
double result = 0.0;
for (ssize_t i = 0; i < n; ++i)
result += pa[i] * pb[i];
return result;
}
// Apply sqrt element-wise, return a new array
py::array_t<double> elementwiseSqrt(py::array_t<double, py::array::c_style> arr) {
py::buffer_info info = arr.request();
auto result = py::array_t<double>(info.shape);
auto* src = static_cast<double*>(info.ptr);
auto* dst = static_cast<double*>(result.request().ptr);
ssize_t n = info.size;
for (ssize_t i = 0; i < n; ++i)
dst[i] = std::sqrt(src[i]);
return result;
}
PYBIND11_MODULE(numops, m) {
m.def("dot", &dot,
"Dot product of two 1D float64 arrays",
py::arg("a"), py::arg("b"));
m.def("sqrt", &elementwiseSqrt,
"Element-wise square root",
py::arg("arr"));
}
import numpy as np
from numops import dot, sqrt
a = np.array([1.0, 2.0, 3.0])
b = np.array([4.0, 5.0, 6.0])
print(dot(a, b)) # 32.0 (1*4 + 2*5 + 3*6)
print(sqrt(np.array([4.0, 9.0, 16.0]))) # [2. 3. 4.]
Always require contiguous arrays when using raw pointer access. Non-contiguous NumPy arrays (slices, transposed views) have non-unit strides:
# This may fail or give wrong results if your C++ code ignores strides
sliced = np.array([1.0, 2.0, 3.0, 4.0])[::2] # non-contiguous
# Safe: force contiguous before passing
dot(np.ascontiguousarray(sliced), ...)
# Or: use py::array::c_style flag in the type — pybind11 will copy if needed
Whether data is copied depends on the parameter type. py::array_t<double> carries the forcecast flag by default, so passing an int64 array or a Python list makes pybind11 create a converted float64 copy; passing a float64 array uses the caller’s memory directly. Adding py::array::c_style also forces a C-contiguous copy when the input is a slice or a transposed view, which makes raw pointer access safe at the cost of a hidden copy. dot above has only forcecast, which is why a strided slice gives wrong results there. The zero-copy promise therefore holds only when dtype and layout already match; for large arrays, it is worth documenting the expected dtype, or passing py::array::c_style | py::array::forcecast explicitly so the behavior is intentional. To reject mismatched input instead of copying, drop forcecast and pybind11 raises TypeError.
For code that needs strides, arr.unchecked<1>() (or mutable_unchecked) returns a proxy that indexes correctly for any layout without bounds checks, and is usually as fast as raw pointers. buf.request() also returns ssize_t shapes, and ssize_t is a POSIX type; pybind11 defines py::ssize_t for portability, which matters when the code must compile with MSVC.
Releasing the GIL
For long-running pure C++ work, release the GIL so other Python threads can run concurrently:
#include <pybind11/pybind11.h>
#include <pybind11/numpy.h>
#include <thread>
#include <vector>
namespace py = pybind11;
// Heavy computation — no Python objects touched
static void heavyWork(double* data, ssize_t n) {
for (ssize_t i = 0; i < n; ++i)
data[i] *= data[i]; // pure C++ — no Python API
}
py::array_t<double> squareAll(py::array_t<double, py::array::c_style> arr) {
py::buffer_info info = arr.request();
auto result = py::array_t<double>(info.shape);
auto* src = static_cast<double*>(info.ptr);
auto* dst = static_cast<double*>(result.request().ptr);
ssize_t n = info.size;
// Copy input to output buffer
std::copy(src, src + n, dst);
{
py::gil_scoped_release release; // release GIL here
// Only raw C++ operations here — never call Python API
heavyWork(dst, n);
// GIL reacquired when 'release' goes out of scope
}
return result;
}
Critical rule: once you have released the GIL, never call any Python C API function until the GIL is reacquired. This includes creating Python objects, calling Python functions, or accessing pybind11 wrapper types. Violations cause crashes and are hard to debug.
The example is careful about this: it extracts raw pointers and sizes from the NumPy arrays before releasing the GIL, and only touches those plain C++ values inside the block. Destructors count too; a py::object going out of scope inside the released region decrements a Python reference count without the GIL, which is exactly the kind of crash that appears only under load. Releasing the GIL has a cost of its own (a lock handoff), so it is worth doing for work measured in milliseconds, not for a function that runs for microseconds.
What releasing buys you is concurrency with other Python threads: a ThreadPoolExecutor calling squareAll on four arrays can now run the four C++ loops truly in parallel, while without the release they would run one at a time. The C++ side must then be thread-safe itself. And the NumPy buffers are still owned by Python objects that another thread could resize or free while the GIL is released; arr is kept alive by the function argument, but code that stores raw pointers for later use must keep a reference to the owning array. With free-threaded CPython builds (3.13+ experimental), the GIL may not exist at all, and extensions have to declare compatibility explicitly.
Exception Handling
C++ exceptions automatically translate to Python exceptions when they propagate through pybind11:
// std::exception → RuntimeError (automatic)
// std::invalid_argument → ValueError, std::out_of_range → IndexError (automatic, built in)
// Custom C++ exception → custom Python exception
class DatabaseError : public std::exception {
std::string msg_;
public:
explicit DatabaseError(std::string msg) : msg_(std::move(msg)) {}
const char* what() const noexcept override { return msg_.c_str(); }
};
PYBIND11_MODULE(mymodule, m) {
// Register the custom exception
py::register_exception<DatabaseError>(m, "DatabaseError");
m.def("connectDB", []() {
throw DatabaseError("Connection refused: host unreachable");
});
}
from mymodule import DatabaseError, connectDB
try:
connectDB()
except DatabaseError as e:
print(f"Caught: {e}") # Caught: Connection refused: host unreachable
Translation happens only when an exception propagates out of a bound function through pybind11. Exceptions thrown on other C++ threads, or inside callbacks from C libraries that are not exception-safe, never reach it; an exception escaping a std::thread calls std::terminate and takes the whole Python process down with it. Catch exceptions at thread boundaries in your C++ code and report them through return values or a stored std::exception_ptr. Also note that register_exception creates a Python exception derived from Exception; pass a base class (PyExc_RuntimeError) as the third argument if callers should be able to catch it as a RuntimeError.
Import errors and undefined symbols
| Error | Cause | Fix |
|---|---|---|
ModuleNotFoundError: No module named 'example' | .so not on Python path | sys.path.insert(0, 'build/') or install with pip |
ImportError: dynamic module does not define module export function (PyInit_example) | Module name in PYBIND11_MODULE differs from filename | Name in macro must match .so filename exactly |
ImportError: ... undefined symbol: _ZN... | A C++ library the module uses was not linked | Add it with target_link_libraries(example PRIVATE mylib) |
find_package(pybind11) failed | pybind11 not installed | pip install pybind11 + set CMAKE_PREFIX_PATH |
| Crash on non-contiguous NumPy array | Ignoring strides, treating data as contiguous | Use py::array::c_style flag or call np.ascontiguousarray() |
Crash after gil_scoped_release | Called Python API without GIL | Only raw C++ operations inside gil_scoped_release block |
TypeError: incompatible types | Python type doesn’t match C++ parameter | Check pybind11 type caster for the type; include <pybind11/stl.h> for STL |
| Crash or double free after returning a pointer/reference | pybind11 guessed the wrong ownership | Set py::return_value_policy (reference_internal for members) |
The last row deserves a word. When a bound function returns a raw pointer, pybind11 must decide who deletes the object. The default for pointers is take_ownership: Python deletes it when the wrapper dies. If the object actually belongs to a C++ container (say, Scene::getNode(i) returns a pointer into a vector), Python deletes memory it does not own, and the program crashes later in an unrelated place. py::return_value_policy::reference_internal tells pybind11 that the result is owned by self and keeps self alive as long as the result exists, which is the right policy for accessors returning references or pointers to members. Returning by value or as std::shared_ptr (with a matching py::class_<T, std::shared_ptr<T>> holder) avoids the question entirely.
STL Container Conversions
Include <pybind11/stl.h> for automatic conversion:
#include <pybind11/pybind11.h>
#include <pybind11/stl.h>
#include <vector>
#include <map>
#include <optional>
namespace py = pybind11;
std::vector<int> makeRange(int n) {
std::vector<int> v(n);
for (int i = 0; i < n; ++i) v[i] = i;
return v; // converted to Python list automatically
}
std::map<std::string, int> wordCount(std::vector<std::string> words) {
std::map<std::string, int> counts;
for (const auto& w : words) ++counts[w];
return counts; // converted to Python dict automatically
}
std::optional<std::string> findUser(int id) {
if (id == 42) return "Alice";
return std::nullopt; // converted to None
}
PYBIND11_MODULE(collections_ext, m) { // not "collections": that would shadow the standard library module
m.def("make_range", &makeRange, py::arg("n"));
m.def("word_count", &wordCount, py::arg("words"));
m.def("find_user", &findUser, py::arg("id"));
}
from collections_ext import make_range, word_count, find_user
print(make_range(5)) # [0, 1, 2, 3, 4]
print(word_count(["a", "b", "a", "c", "b", "a"])) # {'a': 3, 'b': 2, 'c': 1}
print(find_user(42)) # Alice
print(find_user(99)) # None
Note: STL conversions copy the data — modifications to the Python list do not affect the C++ vector.
The copying has consequences beyond performance. If a bound class exposes a std::vector<int> member with def_readwrite, then obj.items.append(5) in Python appends to a temporary list copy and the C++ object is unchanged, which looks like a bug in the binding. Either expose methods that modify the member (add_item), or use PYBIND11_MAKE_OPAQUE(std::vector<int>) together with py::bind_vector from <pybind11/stl_bind.h>, which exposes the C++ vector itself as a Python sequence without copying. For large numeric data, returning a NumPy array instead of a std::vector avoids building a Python list of millions of int objects, which is often slower than the C++ computation that produced it.
Version info and manylinux wheels
Expose version information
PYBIND11_MODULE(mylib, m) {
m.attr("__version__") = "1.2.0";
m.attr("__author__") = "Your Name";
// ...
}
Build manylinux wheels for PyPI
# .github/workflows/wheels.yml
jobs:
build-wheels:
strategy:
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python: ['3.9', '3.10', '3.11', '3.12']
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v4
with: { python-version: '${{ matrix.python }}' }
- run: pip install cibuildwheel
- run: cibuildwheel --output-dir dist/
- uses: actions/upload-artifact@v4
with: { path: dist/*.whl }
cibuildwheel automatically uses manylinux Docker images on Linux to produce wheels compatible with most Linux distributions.
One simplification is possible in this workflow: cibuildwheel builds wheels for every supported CPython version by itself (controlled by CIBW_BUILD, for example cp39-* cp310-* cp311-* cp312-*), so the python matrix dimension only repeats the same work four times per OS. Keep the OS matrix and drop the Python one. The manylinux images are what make Linux wheels portable: they build against an old glibc, so the wheel runs on newer distributions, but not the other way around, which is why building on your own up-to-date machine produces wheels that fail on older servers with GLIBC_2.xx not found. Musl-based distributions such as Alpine need separate musllinux wheels, and macOS wheels for Apple Silicon need CIBW_ARCHS_MACOS to include arm64.
In practice, the build-and-packaging side of a pybind11 project tends to take more effort than the bindings themselves. Settling on scikit-build-core plus cibuildwheel early, and testing the wheels in a clean environment in CI, avoids the “works on my machine” phase that most binding projects go through.
pybind11 rules that avoid crashes and hidden copies
PYBIND11_MODULE(name, m)is the entry point —namemust match the.so/.pydfilename and the Pythonimportstatementpy::array_t<T>withbuf.request()gives direct pointer access to NumPy buffers — zero copy only when dtype and memory layout already match; otherwise pybind11 makes a converted copypy::gil_scoped_releasereleases the GIL for pure C++ work; never call Python API inside the release block<pybind11/stl.h>enables automatic conversion ofstd::vector,std::map,std::optional, etc. — data is copied- Contiguous arrays: check strides or use
py::array::c_styleflag — non-contiguous arrays have stride != sizeof(T) - Exception translation:
std::exception→RuntimeErrorautomatically; usepy::register_exceptionfor custom types - Wheels: build per-Python-version and per-platform; use
cibuildwheelfor CI — manylinux for Linux compatibility
Frequently Asked Questions (FAQ)
Q. Why does import fail with “dynamic module does not define module export function”?
A. Python looks for an init function named after the extension file, so the name in PYBIND11_MODULE(name, m) has to match the built file name (name.so / name.pyd) and the import statement. Renaming the output in CMake or setup.py without changing the macro argument causes exactly this error. If the names match, check that the module was built for the same Python version and architecture as the interpreter importing it.