C++ and Rust Interop: extern "C", cxx vs bindgen, Ownership and Panics at the FFI Boundary

Most C++/Rust projects start the same way: a large C++ codebase that is not going to be rewritten, and one component (a parser, a protocol decoder, anything that eats untrusted input) where Rust’s guarantees are worth the extra toolchain. The hard part is not calling a function across the language boundary, which is a one-liner. The hard part is that every guarantee Rust gives you stops at that boundary: the borrow checker cannot see C++ pointers, C++ cannot see Rust lifetimes, and the two sides use different allocators, different string types and different ideas of what an error is.

This article walks through the boundary itself: the C ABI, the tooling choices, and the four places where hybrid builds actually break — ownership, panics, strings and linking.

The only shared language is the C ABI

C++ and Rust have no common ABI of their own. C++ name mangling, vtable layout and exception tables differ between compilers; Rust’s ABI is deliberately unstable between compiler versions. The one interface both reliably speak is the platform’s C ABI: C calling convention, unmangled symbol names, and #[repr(C)]-compatible layouts.

// cpp_side.h — shared with Rust
#pragma once
#include <cstdint>
#ifdef __cplusplus
extern "C" {
#endif
int32_t checksum(const uint8_t* data, size_t len);
#ifdef __cplusplus
}
#endif
// Rust side: declare what the C++ library exports
unsafe extern "C" {
    fn checksum(data: *const u8, len: usize) -> i32;
}

pub fn checksum_of(bytes: &[u8]) -> i32 {
    // SAFETY: pointer and length come from a valid slice; checksum only reads.
    unsafe { checksum(bytes.as_ptr(), bytes.len()) }
}

And the other direction:

// Rust exports a C symbol. In edition 2024 the attribute must be written
// #[unsafe(no_mangle)]; edition 2021 accepts #[no_mangle].
#[unsafe(no_mangle)]
pub extern "C" fn rs_add(a: i32, b: i32) -> i32 {
    a + b
}

That edition detail trips people up when they copy older examples into a new crate: with edition = "2024", plain #[no_mangle] is a hard error, error: unsafe attribute used without unsafe. The attribute became “unsafe” because two crates exporting the same unmangled symbol is undefined behavior at link time.

What can cross the boundary is limited to what C can express: integers, floats, raw pointers, function pointers, and #[repr(C)] structs and enums of those. &str, String, Vec, Option<Box<T>> aside, most Rust types have no C layout. The compiler tells you when you get this wrong — I compiled an extern "C" fn get_buffer(s: &str) to check, and rustc 1.89 says:

warning: `extern` fn uses type `str`, which is not FFI-safe
  = help: consider using `*const u8` and a length instead
  = note: string slices have no C equivalent

Treat improper_ctypes and improper_ctypes_definitions warnings as errors (#![deny(improper_ctypes_definitions)]). They are the only compile-time check you get on the boundary.

Choosing the tooling: bindgen, cbindgen, or cxx

Hand-writing both sides works for three functions and becomes a drift problem at thirty: the header says int and the Rust declaration says i64, and nothing complains until the stack is corrupted. The tools remove that drift in different ways.

ToolDirectionWhat you getGood fit
bindgenC/C++ header → Rust extern declarationsRaw, unsafe bindings; you write the safe wrapperConsuming an existing C API (or a C facade over C++)
cbindgenRust extern "C" items → C/C++ headerA header that always matches the Rust exportsExposing a Rust library to C++ or other languages
cxxOne #[cxx::bridge] module → both sidesChecked signatures; String, Vec, Box, std::string, std::unique_ptr, CxxVector across the boundaryC++-shaped APIs owned by the same team on both sides

bindgen can parse some C++, but templates, overloads and inline functions either get skipped or generate bindings you cannot use safely. In practice most teams using bindgen on a C++ codebase first write a thin extern "C" facade in C++ and point bindgen at that.

A typical bindgen build script (the bindgen crate as a build dependency, not the CLI):

// build.rs
fn main() {
    cc::Build::new().cpp(true).file("cpp/vector_facade.cpp").compile("vector_facade");
    println!("cargo:rerun-if-changed=cpp/vector_facade.h");
    bindgen::Builder::default()
        .header("cpp/vector_facade.h")
        .allowlist_function("vector_.*")   // don't pull in all of <cstdint>
        .generate()
        .expect("bindgen failed")
        .write_to_file(std::path::PathBuf::from(std::env::var("OUT_DIR").unwrap()).join("bindings.rs"))
        .unwrap();
}
// src/lib.rs
mod sys { include!(concat!(env!("OUT_DIR"), "/bindings.rs")); }

pub struct IntVec(std::ptr::NonNull<sys::Vector>);

impl IntVec {
    pub fn new() -> Self {
        let p = unsafe { sys::vector_new() };
        Self(std::ptr::NonNull::new(p).expect("vector_new returned null"))
    }
    pub fn push(&mut self, v: i32) { unsafe { sys::vector_push(self.0.as_ptr(), v) } }
    pub fn len(&self) -> usize { unsafe { sys::vector_size(self.0.as_ptr()) } }
    pub fn get(&self, i: usize) -> Option<i32> {
        // Bounds check on the Rust side: the C++ facade uses operator[].
        (i < self.len()).then(|| unsafe { sys::vector_get(self.0.as_ptr(), i) })
    }
}

impl Drop for IntVec {
    fn drop(&mut self) { unsafe { sys::vector_delete(self.0.as_ptr()) } }
}

The safe wrapper is where the real work is. The C++ facade (vector_get calling data[index]) would happily read out of bounds; the wrapper is the only place that can stop it. Also note size_t on the C side maps to usize — the original version of this example used int indices, which silently truncates on large inputs.

cxx takes the opposite approach: you describe the interface once and it generates Rust and C++ glue that both sides compile against, so a signature mismatch is a compile error rather than a crash.

// src/lib.rs
#[cxx::bridge(namespace = "geo")]
mod ffi {
    struct Point { x: f64, y: f64 }          // shared struct, same layout on both sides

    unsafe extern "C++" {
        include!("mycrate/include/geometry.h");
        fn distance(a: &Point, b: &Point) -> f64;
    }

    extern "Rust" {
        fn midpoint(a: &Point, b: &Point) -> Point;
    }
}

fn midpoint(a: &ffi::Point, b: &ffi::Point) -> ffi::Point {
    ffi::Point { x: (a.x + b.x) / 2.0, y: (a.y + b.y) / 2.0 }
}
// include/geometry.h
#pragma once
#include "mycrate/src/lib.rs.h"   // generated by cxx; path is <crate>/<path-to-bridge>.rs.h
namespace geo {
double distance(const Point& a, const Point& b);
}
// build.rs
fn main() {
    cxx_build::bridge("src/lib.rs")
        .file("src/geometry.cpp")
        .std("c++17")
        .compile("mycrate-geometry");
    println!("cargo:rerun-if-changed=src/geometry.cpp");
    println!("cargo:rerun-if-changed=include/geometry.h");
}

The trade-off with cxx is scope. It does not support arbitrary templates, and anything outside its supported types has to be opaque (type Foo; used only behind UniquePtr<Foo> or references). If your C++ API is full of std::map<std::string, std::vector<Foo>> you will still write adapter functions — they just live in C++ instead of in extern "C" shims.

My rule of thumb: if the boundary is small and stable, or a Python or C# binding will consume the same API later, go C ABI plus bindgen/cbindgen. If the boundary is wide, C++-shaped, and both sides change every sprint, cxx pays for its learning curve quickly.

Ownership: whoever allocates, frees

This is the bug I see most often in interop examples, including the earlier version of this article: Rust returns Box::into_raw(...) and C++ calls delete on it, or C++ passes a new int and Rust calls Box::from_raw. Both are undefined behavior. Rust’s global allocator and C++ operator new are separate; on Windows they may even be separate CRT heaps. Sometimes it “works” in a small test because both end up in malloc, which makes it worse — it fails later, under a different allocator (jemalloc, mimalloc, a debug CRT), with a heap corruption crash far away from the real bug.

The rule is simple: every pointer that crosses the boundary with ownership comes with a free function from the side that allocated it.

pub struct Parser { /* Rust-only fields */ }

#[unsafe(no_mangle)]
pub extern "C" fn parser_new() -> *mut Parser {
    Box::into_raw(Box::new(Parser { /* ... */ }))
}

/// # Safety
/// `p` must come from `parser_new` and must not be used afterwards.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn parser_free(p: *mut Parser) {
    if !p.is_null() {
        drop(unsafe { Box::from_raw(p) });
    }
}

On the C++ side, don’t let raw handles float around — put them in a unique_ptr with the Rust deleter immediately, and the ownership rule becomes impossible to forget:

extern "C" {
struct Parser;                 // opaque: C++ never sees the layout
Parser* parser_new();
void parser_free(Parser*);
}

struct ParserDeleter { void operator()(Parser* p) const noexcept { parser_free(p); } };
using ParserPtr = std::unique_ptr<Parser, ParserDeleter>;

ParserPtr make_parser() { return ParserPtr{parser_new()}; }

Borrowed data is the other half. A pointer/length pair passed from C++ into Rust is valid only for the duration of the call unless the API says otherwise. If Rust needs to keep the data, copy it (slice.to_vec()). The same applies in reverse: never return a pointer into a Rust String or Vec that lives on the Rust stack — it dangles the moment the function returns.

For buffers that Rust allocates and C++ fills or reads, Box<[u8]> is easier to get right than Vec plus mem::forget, because it has no separate capacity to remember when reconstructing it:

#[unsafe(no_mangle)]
pub extern "C" fn buf_alloc(len: usize) -> *mut u8 {
    Box::into_raw(vec![0u8; len].into_boxed_slice()) as *mut u8
}

#[unsafe(no_mangle)]
pub unsafe extern "C" fn buf_free(ptr: *mut u8, len: usize) {
    if !ptr.is_null() {
        drop(unsafe { Box::from_raw(std::ptr::slice_from_raw_parts_mut(ptr, len)) });
    }
}

Panics and exceptions must not cross

Neither side’s error mechanism can travel across extern "C". A C++ exception thrown through a Rust frame, or a Rust panic unwinding into C++ frames, is undefined behavior historically — and in current Rust, the panic case is defined as an abort. I verified this with rustc 1.89: a panic! inside an extern "C" fn called from Rust, even wrapped in catch_unwind on the caller side, prints

thread 'main' panicked at library\core\src\panicking.rs:225:5:
panic in a function that cannot unwind
...
thread caused non-unwinding panic. aborting.

and the process dies. That behavior has been guaranteed since Rust 1.81. It is safer than the old UB, but it means one unwrap() on bad input takes down the whole C++ host process. So the catch_unwind has to be inside the exported function:

use std::ffi::{c_char, CStr, CString};

#[repr(C)]
pub enum Status { Ok = 0, InvalidArg = 1, Internal = 2 }

#[unsafe(no_mangle)]
pub unsafe extern "C" fn rs_upper(input: *const c_char, out: *mut *mut c_char) -> Status {
    if input.is_null() || out.is_null() {
        return Status::InvalidArg;
    }
    let result = std::panic::catch_unwind(|| {
        let s = unsafe { CStr::from_ptr(input) }.to_str().ok()?;
        CString::new(s.to_uppercase()).ok()
    });
    match result {
        Ok(Some(c)) => { unsafe { *out = c.into_raw() }; Status::Ok }
        Ok(None) => Status::InvalidArg,        // not UTF-8, or contains NUL
        Err(_) => Status::Internal,            // a panic, contained
    }
}

#[unsafe(no_mangle)]
pub unsafe extern "C" fn rs_string_free(s: *mut c_char) {
    if !s.is_null() { drop(unsafe { CString::from_raw(s) }); }
}

catch_unwind does not help if the crate is built with panic = "abort" in Cargo.toml — then every panic aborts regardless. Some teams choose that deliberately for the FFI crate so a bug can never leave the C++ side with half-updated state; just make it a decision, not an accident.

On the C++ side, the mirror rule applies: every extern "C" function that Rust calls must be noexcept in practice.

extern "C" int32_t cpp_decode(const uint8_t* p, size_t n) noexcept {
    try {
        return decode_impl(p, n);            // may throw
    } catch (const std::exception&) {
        return -1;
    } catch (...) {
        return -2;
    }
}

If you genuinely need unwinding across the boundary (rare — usually when Rust callbacks sit between C++ frames that expect exceptions), Rust has the extern "C-unwind" ABI, stable since 1.71. cxx handles this for you differently: a C++ function declared as returning Result<T> in the bridge has its exception caught by the generated glue and turned into a Rust Err.

Strings: three types, three assumptions

C++ std::string, Rust String/&str and C char* disagree on everything that matters:

  • std::string is a byte container with an explicit length; it may contain \0 and need not be valid UTF-8.
  • Rust str is always valid UTF-8, has an explicit length, and may contain \0.
  • A C string is NUL-terminated with no length, so it cannot contain an interior \0.

That gives you the failure modes directly. CString::new returns an error if the Rust string has an interior NUL. CStr::to_str returns an error if C++ hands over bytes that are not UTF-8 — Latin-1 file names on Windows, or binary data someone stuffed into a std::string. If you don’t need UTF-8 semantics, skip both and pass (const uint8_t*, size_t), which round-trips any std::string exactly:

void send_to_rust(const std::string& s) {
    rs_consume(reinterpret_cast<const uint8_t*>(s.data()), s.size());
}

With cxx, &str/String map to rust::Str/rust::String, and &CxxString maps to const std::string&. Converting a CxxString to a Rust &str still goes through to_str(), which can fail on invalid UTF-8 — the check doesn’t disappear, it just moves into a typed API.

Linking: static libs and the order of things

When C++ is the final binary, build the Rust side as a static library:

# Cargo.toml
[lib]
crate-type = ["staticlib"]   # produces librsdemo.a (Unix) / rsdemo.lib (MSVC)

A Rust staticlib bundles the Rust standard library but not the system libraries it depends on. Ask the compiler which ones:

cargo rustc --release -- --print native-static-libs

On my Windows machine (MSVC target) that printed:

note: native-static-libs: kernel32.lib ntdll.lib userenv.lib ws2_32.lib dbghelp.lib /defaultlib:msvcrt

On Linux the list is typically -lgcc_s -lutil -lrt -lpthread -lm -ldl -lc. Leave these off and you get a wall of undefined references to symbols the C++ code never mentions, which is confusing until you know where they come from.

With GNU ld, order matters: a static library only satisfies symbols that were already referenced when the linker reached it. So the Rust library goes after the objects that call it:

g++ -o app main.o -L target/release -lrsdemo -lpthread -ldl -lm   # works
g++ -o app -L target/release -lrsdemo main.o                      # undefined reference to `rs_add'

In CMake, the cleanest approach is an imported target (or the Corrosion project, which drives Cargo from CMake):

add_library(rsdemo STATIC IMPORTED)
set_target_properties(rsdemo PROPERTIES
  IMPORTED_LOCATION "${CMAKE_SOURCE_DIR}/rust/target/release/librsdemo.a")
find_package(Threads REQUIRED)
target_link_libraries(rsdemo INTERFACE Threads::Threads ${CMAKE_DL_LIBS} m)
target_link_libraries(app PRIVATE rsdemo)

Two more traps. First, the toolchains must agree: a Rust x86_64-pc-windows-msvc staticlib links with MSVC, not with MinGW g++, and vice versa for the -gnu target — my own Windows setup has exactly this split (MSVC Rust, MinGW g++), and the two artifacts simply don’t mix. Second, linking two Rust staticlibs into one C++ binary gives you two copies of the Rust standard library and duplicate symbol errors; the fix is a single “umbrella” Rust crate that re-exports both and is built as the one staticlib.

Threads and callbacks

C function pointers are the only callback type that crosses the C ABI. Rust closures don’t; if you need state, pass a void* context alongside the function pointer:

pub type Visit = unsafe extern "C" fn(ctx: *mut std::ffi::c_void, value: i32);

#[unsafe(no_mangle)]
pub unsafe extern "C" fn rs_for_each(data: *const i32, len: usize,
                                     f: Visit, ctx: *mut std::ffi::c_void) {
    if data.is_null() { return; }
    for &v in unsafe { std::slice::from_raw_parts(data, len) } {
        unsafe { f(ctx, v) };
    }
}

Rust’s Send/Sync do not exist on the other side. If a Rust object handed to C++ is not Sync, nothing stops C++ from calling into it from two threads at once. Document the threading contract in the header, and if the object must be shared, make it internally synchronized (a Mutex inside, or atomics) rather than trusting callers.

A checklist that actually catches bugs

When I review an FFI layer, these are the questions that find real problems:

  1. Does every exported function that returns an owned pointer have a matching free function in the same language?
  2. Does every exported Rust function either catch panics or live in a crate built with panic = "abort" on purpose?
  3. Is every C++ function Rust calls noexcept, with a try/catch at the top?
  4. Are all pointer/length pairs checked for null (and zero length handled) before slice::from_raw_parts?
  5. Are improper_ctypes* lints denied, and is the header generated (cbindgen/bindgen/cxx) rather than hand-copied?
  6. Is the native-static-libs list in the build system, and does the Rust library come after its users on the link line?