C++ Alignment and Padding: Why sizeof Surprises You, alignas, #pragma pack and False Sharing

Key takeaways

Every type has an alignment; the compiler pads structs so each member and each array element stays aligned. Order members by alignment to shrink them, use alignas when you need more, and treat #pragma pack as a serialization shortcut with sharp edges.

Sooner or later every C++ programmer writes a struct with three members, calls sizeof, and gets a number larger than the sum of the parts. The extra bytes are padding, and they follow a small set of rules. Once those rules are clear, most layout questions — why reordering shrank a struct, why a packed header crashed on one machine, why two “independent” counters are slow — become predictable.

Sizes and offsets below were printed by g++ 10 (MinGW-w64) on x86-64. Values for int, double and pointers match mainstream 64-bit Linux and Windows; the notable difference is long, covered below.

Every type has an alignment

alignof(T) is the power of two that an object’s address must be a multiple of. On this toolchain:

alignof(char)        // 1
alignof(short)       // 2
alignof(int)         // 4
alignof(long)        // 4 on Windows (LLP64), 8 on 64-bit Linux/macOS (LP64)
alignof(double)      // 8
alignof(void*)       // 8
alignof(std::max_align_t)  // 16

Why the hardware cares: loads and stores are cheapest when a value does not straddle a word or cache-line boundary. x86-64 tolerates misaligned ordinary loads, usually with little penalty unless the access crosses a cache line, but some instructions (aligned SSE/AVX loads, some atomics) and some architectures do fault. More importantly for C++, dereferencing a misaligned int* is undefined behavior regardless of what the CPU tolerates, and optimizers are allowed to assume it never happens.

Where padding comes from

The compiler applies two rules:

  1. Each member is placed at the next offset that is a multiple of its alignment.
  2. The struct’s alignment is the largest member alignment, and its size is rounded up to a multiple of that, so that every element of an array of the struct is aligned too.
struct A { char c; int i; double d; };   // sizeof 16: c@0, pad 1-3, i@4, d@8
struct B { char c; double d; int i; };   // sizeof 24: c@0, pad 1-7, d@8, i@16, tail pad 20-23
struct C { double d; int i; char c; };   // sizeof 16: d@0, i@8, c@12, tail pad 13-15

A is a useful counterexample to a common claim: small-to-large order is not automatically bad. A happens to pack perfectly because the char and int together fill the first 8 bytes. B wastes 11 bytes because the double forces a 7-byte gap and the tail rounding adds 4 more. The reliable rule is to sort members by decreasing alignment, as in C; that minimizes interior padding, although tail padding can remain.

Tail padding is easy to forget when nesting:

struct Tail  { int i; char c; };          // sizeof 8 (3 bytes of tail padding)
struct Outer { Tail t; char extra; };     // sizeof 12, not 9

extra cannot move into Tail’s tail padding, because copying a Tail with memcpy or assignment is allowed to overwrite those bytes.

To see the layout rather than guess, print offsetof for each member, or ask the compiler: clang -Xclang -fdump-record-layouts and MSVC’s /d1reportSingleClassLayout<Name> both print full layouts, and pahole does the same from debug info on Linux. A static_assert(sizeof(Header) == 16) next to structs that must not grow turns an accidental layout change into a compile error.

alignas: asking for more alignment

alignas(N) raises the alignment of a type or variable. N must be a power of two, and it can only increase alignment; asking for less than the natural alignment is an error.

struct alignas(16) Vec2i { int x, y; };   // alignof 16, sizeof 16 (8 bytes of tail padding)

alignas(64) int line[16];                 // address is a multiple of 64

Two consequences to keep in mind. First, raising alignment also raises sizeof to a multiple of it, so a vector<Vec2i> uses twice the memory of a vector of plain pairs. Second, over-aligned types (alignment above alignof(std::max_align_t)) on the heap need C++17: since then new and std::allocator call the aligned operator new, so std::vector<Slot> of an alignas(64) type is correctly aligned. With C++14 and earlier, new only guaranteed alignof(std::max_align_t) and silently ignored the extra alignment. For raw buffers, std::aligned_alloc (C++17; not available in the MSVC runtime, which uses _aligned_malloc) is the C-style option.

The main legitimate uses are SIMD loads that require alignment, lock-free structures that must not share cache lines, and hardware or DMA buffers with alignment requirements. Adding alignas “for performance” elsewhere usually just costs memory.

SIMD: aligned and unaligned loads

_mm256_load_ps requires a 32-byte-aligned address and faults otherwise; _mm256_loadu_ps accepts any address.

#include <immintrin.h>

alignas(32) float data[8] = {1, 2, 3, 4, 5, 6, 7, 8};
__m256 v = _mm256_mul_ps(_mm256_load_ps(data), _mm256_set1_ps(2.0f));
alignas(32) float out[8];
_mm256_store_ps(out, v);                 // 2 4 6 8 10 12 14 16  (compile with -mavx)

On recent x86 cores, an unaligned-load instruction on data that happens to be aligned runs at the same speed as the aligned one, so the usual advice is: use loadu/storeu in library code that receives arbitrary pointers, and align your own buffers so that most loads do not cross cache lines. A loop that processes 8 floats at a time also needs a scalar tail for lengths that are not a multiple of 8; a vector-add loop that advances i += 8 with no tail reads past the end of any array whose length is not a multiple of 8.

#pragma pack: removing padding, and what it costs

#pragma pack(push, 1) (supported by GCC, Clang and MSVC) and GCC/Clang’s __attribute__((packed)) lower member alignment so no padding is inserted:

#pragma pack(push, 1)
struct PacketHeader {
    std::uint8_t  version;     // offset 0
    std::uint16_t length;      // offset 1
    std::uint32_t sequence;    // offset 3
    std::uint64_t timestamp;   // offset 7
};                              // sizeof 15, alignof 1
#pragma pack(pop)

Accessing members by name is fine: the compiler knows the struct is packed and emits byte-safe loads where the target needs them. The trouble starts when a pointer or reference to a member escapes, because an std::uint32_t* carries no “might be misaligned” information:

struct __attribute__((packed)) P { char c; int i; };
int* f(P& p) { return &p.i; }
warning: taking address of packed member of 'P' may result in an unaligned pointer value [-Waddress-of-packed-member]

Passing hdr.sequence to a function taking const std::uint32_t& has the same problem. Whether the warning fires can differ between the __attribute__((packed)) form and the #pragma pack form and between compiler versions, so do not rely on it as your only safety net.

Packing also does nothing about byte order. A header read on a little-endian machine from a big-endian wire format still needs every multi-byte field converted. For that reason I prefer to keep in-memory structs naturally aligned and write explicit (de)serialization: copy bytes out of the buffer with std::memcpy into a properly typed variable, then fix endianness. It is a few more lines, it is well defined on every platform, and compilers turn a fixed-size memcpy into a single load.

unsigned char buf[4] = {0x78, 0x56, 0x34, 0x12};
std::uint32_t v;
std::memcpy(&v, buf, sizeof v);   // 0x12345678 on a little-endian host; no aliasing or alignment UB

The failure mode I have seen most with packed structs is exactly the “works on my machine” kind: code that casts reinterpret_cast<Header*>(buffer + offset) or takes the address of a packed field runs fine on x86, where misaligned loads just work, and then faults or returns garbage when the same code is built for a stricter ARM target, or when the optimizer vectorizes a loop using aligned instructions because it assumed the pointer was aligned.

False sharing: alignment across threads

Caches move data in lines, typically 64 bytes on x86. If two threads repeatedly write to different variables that live in the same line, each write invalidates the other core’s copy, and the line bounces between cores even though the threads share no data logically.

struct Shared {                       // both counters fit in one 64-byte line
    std::atomic<long> a{0};
    std::atomic<long> b{0};
};

struct Split {                        // each counter on its own 64-byte line
    alignas(64) std::atomic<long> a{0};
    alignas(64) std::atomic<long> b{0};
};

If two threads each increment one of the counters in a tight loop, even with memory_order_relaxed, Shared forces every increment to fetch the line in exclusive state from the other core first, while Split lets each core keep its own line in its cache. How large the gap is depends on the CPU and on whether the threads land on cores that share a cache, so time both versions on your own hardware; on Linux, perf c2c is built to find exactly this kind of contended line in a real program.

Per-thread slots in an array follow the same idea: make the element type itself alignas(64).

struct alignas(64) Slot { std::atomic<int> n{0}; };   // sizeof 64
Slot per_thread[4];                                   // each slot on its own line

There is no need to add a manual char padding[60] member; alignas(64) already rounds sizeof up to 64, and a hardcoded padding count breaks as soon as the member types change.

The mistake I have made here is the opposite one: padding everything. A struct of per-thread counters that only one thread ever writes gains nothing from alignas(64), and a vector of 64-byte elements that used to be 4 bytes each is sixteen times larger and far less cache-friendly for readers that scan it. Split lines only where a profiler or a benchmark shows contended writes.

C++17 names the relevant distance std::hardware_destructive_interference_size in <new>. Library support came late — libstdc++ only provides it from GCC 12, so check the feature macro __cpp_lib_hardware_interference_size — and the value is a compile-time guess about the target CPU. A common approach is to use it when available and fall back to 64, or 128 for targets such as Apple Silicon.

Portability traps in “fixed” layouts

  • long is 4 bytes on 64-bit Windows and 8 bytes on 64-bit Linux and macOS, so a struct containing long has different size and offsets on the two. Use <cstdint> types (std::int32_t, std::int64_t) for anything that is written to disk or the network.
  • alignof(long double) and sizeof(long double) vary widely between compilers.
  • Bit-field layout (order, and whether fields straddle storage units) is implementation-defined; do not use bit-fields to describe wire formats.
  • Even with fixed-width types, padding may differ between ABIs. A static_assert on sizeof and on key offsetof values documents the assumption and catches a change at compile time.