C++ Game Engine Basics: ECS, Rendering, Physics, Input, Lua
Introduction: “Characters fall through floors, collisions bounce weirdly”
Problems encountered when building game engines
Building 2D games from scratch without Unity or Unreal, you encounter:
- Entities fall through floors — collision detection order or AABB boundary calculation errors
- Render order shuffles every frame — no Z-index sorting or layer system
- Input feels sluggish tied to frames — polling only, not event-based
- Game logic mixed with engine code, hard to modify — insufficient scripting separation This guide covers 2D game engine basics integrating rendering, physics, input, and scripting based on ECS architecture.
Goals:
- ECS (Entity Component System) architecture
- 2D rendering pipeline (SDL2/SFML)
- Physics simulation (collision detection, rigid body dynamics)
- Input handling and event system
- Lua scripting integration Requirements: C++17+, SDL2 or SFML, Box2D (optional), Lua 5.4
The code here is deliberately simple — a map of components per entity, an O(n²) collision loop, velocity flipping instead of a real solver — so that each system fits on one screen. That makes it good for learning how the pieces talk to each other and a poor starting point for shipping anything with hundreds of moving objects. Where a shortcut causes a known failure (objects sinking into floors, render order flickering, crashes when deleting entities), the text says so and shows what a production engine does instead.
ECS Architecture
Why ECS?
Traditional inheritance-based game objects cause “multiple inheritance hell” and “diamond inheritance” problems. ECS uses composition, attaching only needed components to entities for flexible extension.
The inheritance problem shows up quickly in practice: you start with GameObject → Character → Player, then need a Door that is both Interactable and Animated, and a Turret that is Enemy-like but never moves. Every new combination forces a new class or a fat base class full of unused fields. In ECS an entity is just an ID, components are plain data (position, sprite, velocity), and systems are functions that run over every entity that has a particular set of components. “Can this thing fall?” becomes “does it have a RigidBodyComponent?”, which you can change at runtime by adding or removing the component.
The second motivation — the one large engines care about — is memory layout. If all TransformComponents sit in one contiguous array, the physics system streams through them with good cache behavior. The implementation below does not achieve that: each entity owns a hash map of heap-allocated components, so iteration chases pointers. It is the easiest version to read, but be aware that libraries such as EnTT or flecs store components in packed arrays per type, and that is where ECS gets its performance reputation.
flowchart TB
subgraph ECS[ECS Architecture]
E[Entity]
C1[TransformComponent]
C2[SpriteComponent]
C3[RigidBodyComponent]
C4[ColliderComponent]
E --> C1
E --> C2
E --> C3
E --> C4
end
subgraph Systems[Systems]
S1[RenderSystem]
S2[PhysicsSystem]
S3[InputSystem]
end
C1 --> S1
C2 --> S1
C3 --> S2
C4 --> S2
Core Component Definitions
#include <glm/glm.hpp>
#include <SDL2/SDL.h>
#include <memory>
#include <string>
#include <typeindex>
#include <unordered_map>
using EntityID = uint32_t;
class Component {
public:
virtual ~Component() = default;
};
// Position, rotation, scale — needed for all render/physics entities
struct TransformComponent : Component {
glm::vec2 position{0, 0};
float rotation = 0.0f;
glm::vec2 scale{1, 1};
};
// Sprite rendering
struct SpriteComponent : Component {
std::string texture_id;
SDL_Rect src_rect;
int z_index = 0; // Render order (lower = drawn behind)
};
// Physics velocity, mass
struct RigidBodyComponent : Component {
glm::vec2 velocity{0, 0};
float mass = 1.0f;
bool is_static = false; // Fixed objects like floors, walls
};
// Collision area (AABB) — required by physics system
struct ColliderComponent : Component {
float width = 32.0f;
float height = 32.0f;
bool is_trigger = false; // Passable area (trigger)
};
class Entity {
EntityID id_;
std::unordered_map<std::type_index, std::unique_ptr<Component>> components_;
public:
explicit Entity(EntityID id) : id_(id) {}
template <typename T, typename... Args>
T& add_component(Args&&... args) {
auto component = std::make_unique<T>(std::forward<Args>(args)...);
auto* ptr = component.get();
components_[typeid(T)] = std::move(component);
return *ptr;
}
template <typename T>
T* get_component() {
auto it = components_.find(typeid(T));
return it != components_.end() ? static_cast<T*>(it->second.get()) : nullptr;
}
template <typename T>
bool has_component() const {
return components_.find(typeid(T)) != components_.end();
}
EntityID get_id() const { return id_; }
};
class EntityManager {
std::unordered_map<EntityID, std::unique_ptr<Entity>> entities_;
EntityID next_id_ = 1;
public:
Entity& create_entity() {
auto id = next_id_++;
auto entity = std::make_unique<Entity>(id);
auto* ptr = entity.get();
entities_[id] = std::move(entity);
return *ptr;
}
void destroy_entity(EntityID id) {
entities_.erase(id);
}
template <typename... Components>
std::vector<Entity*> get_entities_with() {
std::vector<Entity*> result;
for (auto& [id, entity] : entities_) {
if ((entity->has_component<Components>() && ...)) {
result.push_back(entity.get());
}
}
return result;
}
};
Caution: get_entities_with allocates new vector every frame. For high performance, optimize with per-component indices.
A few details in this code are worth understanding before building on it. components_[typeid(T)] keys the map by the exact static type, so get_component<Component>() will not find a TransformComponent — lookups must use the concrete type. Calling add_component<T> twice silently replaces the first component and destroys it, which invalidates any T* you were holding. And get_entities_with iterates an unordered_map, whose order is unspecified and can change when the map rehashes after new entities are added; anything that depends on iteration order (like drawing order for equal z-values) needs its own sort key, as Issue 2 below shows. The fold expression (entity->has_component<Components>() && ...) is C++17 and short-circuits, so the check stops at the first missing component.
Rendering System
Rendering Pipeline Flow
sequenceDiagram
participant GameLoop
participant RenderSystem
participant SDL
GameLoop->>RenderSystem: update(entities)
RenderSystem->>RenderSystem: Z-index sort
RenderSystem->>SDL: RenderClear
loop Each entity
RenderSystem->>SDL: RenderCopyEx
end
RenderSystem->>SDL: RenderPresent
Implementation
SDL’s 2D renderer draws in painter’s order: whatever is drawn last ends up on top, and there is no depth buffer to fix mistakes. That is why the system sorts by z_index before issuing any draw calls. The whole frame is built into a back buffer between SDL_RenderClear and SDL_RenderPresent, and only Present makes it visible, so the player never sees a half-drawn frame. Textures are cached by ID because IMG_Load decodes a PNG from disk and SDL_CreateTextureFromSurface uploads it to the GPU — both far too slow to do per frame.
#include <algorithm>
#include <SDL2/SDL_image.h>
class RenderSystem {
SDL_Renderer* renderer_;
std::unordered_map<std::string, SDL_Texture*> textures_;
public:
explicit RenderSystem(SDL_Renderer* renderer) : renderer_(renderer) {}
void update(EntityManager& entities) {
auto entities_to_render =
entities.get_entities_with<TransformComponent, SpriteComponent>();
// Sort by z_index ascending (lower value = drawn behind)
std::sort(entities_to_render.begin(), entities_to_render.end(),
[](Entity* a, Entity* b) {
return a->get_component<SpriteComponent>()->z_index <
b->get_component<SpriteComponent>()->z_index;
});
SDL_RenderClear(renderer_);
for (auto* entity : entities_to_render) {
auto* transform = entity->get_component<TransformComponent>();
auto* sprite = entity->get_component<SpriteComponent>();
auto it = textures_.find(sprite->texture_id);
if (it == textures_.end()) continue; // Skip if no texture
SDL_Rect dest_rect = {
static_cast<int>(transform->position.x),
static_cast<int>(transform->position.y),
static_cast<int>(sprite->src_rect.w * transform->scale.x),
static_cast<int>(sprite->src_rect.h * transform->scale.y)};
SDL_RenderCopyEx(renderer_, it->second, &sprite->src_rect,
&dest_rect, transform->rotation, nullptr,
SDL_FLIP_NONE);
}
SDL_RenderPresent(renderer_);
}
void load_texture(const std::string& id, const std::string& path) {
SDL_Surface* surface = IMG_Load(path.c_str());
if (!surface) return;
SDL_Texture* texture = SDL_CreateTextureFromSurface(renderer_, surface);
SDL_FreeSurface(surface);
if (texture) textures_[id] = texture;
}
~RenderSystem() {
for (auto& [id, tex] : textures_) SDL_DestroyTexture(tex);
}
};
Two things to watch. SDL_RenderCopyEx takes the rotation in degrees, clockwise, and rotates around the center of dest_rect when the center argument is nullptr; if your physics code stores radians (as Box2D does), convert before drawing or sprites will spin wildly. Second, this class owns raw SDL_Texture* handles and destroys them in its destructor, but it has the compiler-generated copy operations. Copying a RenderSystem gives two objects that both destroy the same textures, which crashes on shutdown. Either delete the copy constructor and copy assignment, or wrap each texture in a std::unique_ptr with SDL_DestroyTexture as the deleter. Textures must also be destroyed before the renderer that created them, which is why the engine below destroys the render system before calling SDL_DestroyRenderer.
Physics Simulation
Simple Physics Engine (AABB Collision)
The update follows the usual order for a semi-implicit (symplectic) Euler integrator: apply forces to velocity first, then move positions using the new velocity, then fix up collisions. Updating velocity before position is noticeably more stable than the reverse for the same timestep, which is why most simple game physics uses this ordering. Note that gravity_{0, 9.8f} is in whatever units your positions use — here pixels per second squared — so a falling object gains under 10 pixels/s per second and looks like it is floating. Games typically use a few hundred to a couple of thousand pixels/s², or keep physics in meters and convert at render time as the Box2D section does.
class PhysicsSystem {
glm::vec2 gravity_{0, 9.8f};
public:
void update(EntityManager& entities, float dt) {
auto physics_entities = entities.get_entities_with<
TransformComponent, RigidBodyComponent, ColliderComponent>();
// 1. Apply gravity
for (auto* entity : physics_entities) {
auto* rb = entity->get_component<RigidBodyComponent>();
if (!rb->is_static) {
rb->velocity += gravity_ * dt;
}
}
// 2. Update positions
for (auto* entity : physics_entities) {
auto* transform = entity->get_component<TransformComponent>();
auto* rb = entity->get_component<RigidBodyComponent>();
if (!rb->is_static) {
transform->position += rb->velocity * dt;
}
}
// 3. Detect and resolve collisions
check_collisions(physics_entities);
}
void check_collisions(const std::vector<Entity*>& entities) {
for (size_t i = 0; i < entities.size(); ++i) {
for (size_t j = i + 1; j < entities.size(); ++j) {
if (check_collision(entities[i], entities[j])) {
resolve_collision(entities[i], entities[j]);
}
}
}
}
bool check_collision(Entity* a, Entity* b) {
auto* ta = a->get_component<TransformComponent>();
auto* tb = b->get_component<TransformComponent>();
auto* ca = a->get_component<ColliderComponent>();
auto* cb = b->get_component<ColliderComponent>();
if (!ca || !cb) return false;
// AABB collision: check if two rectangles overlap
return ta->position.x < tb->position.x + cb->width &&
ta->position.x + ca->width > tb->position.x &&
ta->position.y < tb->position.y + cb->height &&
ta->position.y + ca->height > tb->position.y;
}
void resolve_collision(Entity* a, Entity* b) {
auto* rba = a->get_component<RigidBodyComponent>();
auto* rbb = b->get_component<RigidBodyComponent>();
if (!rba || !rbb) return;
// Triggers have no physical response
auto* ca = a->get_component<ColliderComponent>();
auto* cb = b->get_component<ColliderComponent>();
if (ca->is_trigger || cb->is_trigger) return;
// Reverse velocity on collision with static object
if (rbb->is_static) {
rba->velocity.y = -rba->velocity.y * 0.8f; // Elasticity
} else if (rba->is_static) {
rbb->velocity.y = -rbb->velocity.y * 0.8f;
} else {
// Both dynamic: swap velocities (simple elastic collision)
auto temp = rba->velocity;
rba->velocity = rbb->velocity;
rbb->velocity = temp;
}
}
};
The AABB test is four comparisons: two boxes overlap only if they overlap on the x axis and on the y axis. Using strict < and > means boxes that merely touch edges do not count as colliding, which matters for a player standing exactly on a floor.
The weak point is resolve_collision. It only changes velocity; it never moves the overlapping objects apart. By the time the collision is detected, the player has already moved partway into the floor, and flipping the velocity does not undo that. With gravity pulling down every frame, the player sinks a little, bounces with 80% speed, sinks again, and you see jitter or a slow slide through the floor — often the real cause of “entities fall through floors” rather than a missing collider. A proper resolution computes the penetration depth on each axis, pushes the dynamic object out along the axis of smallest overlap, and zeroes (or reflects) only the velocity component along that axis. It also reflects only velocity.y regardless of whether the hit came from the side, so running into a wall makes the object bounce vertically. The “swap velocities” branch is the correct result only for equal masses in a head-on elastic collision; it ignores mass entirely. These shortcuts are fine for a first prototype, and they are exactly where a library like Box2D earns its place.
When I write a hand-rolled platformer, the first thing that goes wrong is almost always this resolution step, not detection: the player catches on the seam between two adjacent floor tiles, because at the seam the smallest-overlap axis briefly flips to horizontal and the player gets pushed sideways. The usual fix is to resolve the two axes separately — move on x, resolve x collisions, then move on y, resolve y collisions — which is also how many tile-based games handle it.
Input and Events
Input System (Polling + Events)
Games need two kinds of input queries. State questions — “is the right arrow held down right now?” — drive continuous movement and are answered from key_states_. Edge events — “was jump pressed this frame?” — should fire once and are delivered by callbacks. Using state for edge actions is a common mistake: checking is_key_pressed(SDLK_SPACE) to jump makes the character jump again every frame while the key is held.
#include <functional>
#include <vector>
class InputSystem {
std::unordered_map<SDL_Keycode, bool> key_states_;
glm::vec2 mouse_position_{0, 0};
bool quit_requested_ = false;
using KeyCallback = std::function<void(SDL_Keycode)>;
std::vector<KeyCallback> key_down_callbacks_;
std::vector<KeyCallback> key_up_callbacks_;
public:
void update() {
SDL_Event event;
while (SDL_PollEvent(&event)) {
switch (event.type) {
case SDL_QUIT:
quit_requested_ = true;
break;
case SDL_KEYDOWN:
key_states_[event.key.keysym.sym] = true;
for (auto& cb : key_down_callbacks_) cb(event.key.keysym.sym);
break;
case SDL_KEYUP:
key_states_[event.key.keysym.sym] = false;
for (auto& cb : key_up_callbacks_) cb(event.key.keysym.sym);
break;
case SDL_MOUSEMOTION:
mouse_position_ = {static_cast<float>(event.motion.x),
static_cast<float>(event.motion.y)};
break;
}
}
}
bool is_key_pressed(SDL_Keycode key) const {
auto it = key_states_.find(key);
return it != key_states_.end() && it->second;
}
glm::vec2 get_mouse_position() const { return mouse_position_; }
bool is_quit_requested() const { return quit_requested_; }
void on_key_down(KeyCallback cb) { key_down_callbacks_.push_back(std::move(cb)); }
void on_key_up(KeyCallback cb) { key_up_callbacks_.push_back(std::move(cb)); }
};
SDL_PollEvent must be drained every frame: SDL only receives OS messages when events are pumped, so a loop that skips it makes the window go “Not Responding” on Windows even though the game is running. The loop above handles that, but it has one trap in the callbacks: when a key is held, the OS sends auto-repeat SDL_KEYDOWN events, and SDL marks them with event.key.repeat != 0. Without checking that field, a key-down callback fires many times per second while the key is held. Also note the difference between keysym.sym (the keycode, which follows the keyboard layout) and keysym.scancode (the physical key position). WASD movement should use scancodes, otherwise players on AZERTY keyboards get Z/Q/S/D mapped to the wrong physical keys. For pure state queries, SDL_GetKeyboardState returns SDL’s own up-to-date key array and makes a hand-maintained map unnecessary.
Scripting Integration
Lua API Registration Example
Scripting splits the code into two speeds of change: the engine (rendering, physics, memory) is compiled C++ that changes rarely, while game rules — what a coin does, how an enemy behaves — live in Lua files that designers can edit and reload without a rebuild. The C++/Lua boundary is a virtual stack: C++ pushes arguments, Lua reads them by index, and results come back the same way. Each C function exposed to Lua has the signature int f(lua_State*) and returns how many values it pushed.
The code passes the EntityManager* as an upvalue — a value bound to the C closure — rather than a global variable. A captureless lambda converts to the plain function pointer Lua needs, but it cannot capture this, so the upvalue is how the function gets its context. Light userdata is just a raw pointer with no lifetime tracking: if the EntityManager is destroyed while the Lua state is still alive, the next script call dereferences a dangling pointer.
extern "C" {
#include <lua.h>
#include <lualib.h>
#include <lauxlib.h>
}
class ScriptingSystem {
lua_State* L_;
EntityManager* entities_ = nullptr;
public:
ScriptingSystem() {
L_ = luaL_newstate();
luaL_openlibs(L_);
}
void set_entity_manager(EntityManager* em) { entities_ = em; }
void register_api() {
if (!entities_) return;
// create_entity() -> returns entity_id
lua_pushlightuserdata(L_, entities_);
lua_pushcclosure(L_, [](lua_State* L) -> int {
auto* ud = static_cast<EntityManager*>(lua_touserdata(L, lua_upvalueindex(1)));
if (!ud) return 0;
auto& e = ud->create_entity();
lua_pushinteger(L, static_cast<lua_Integer>(e.get_id()));
return 1;
}, 1);
lua_setglobal(L_, "create_entity");
// set_position(entity_id, x, y)
// Implementation omitted: lua_tointeger, get_entity, get_component<Transform>, etc.
}
bool run_script(const std::string& script) {
if (luaL_dostring(L_, script.c_str()) != LUA_OK) {
fprintf(stderr, "Lua error: %s\n", lua_tostring(L_, -1));
lua_pop(L_, 1);
return false;
}
return true;
}
~ScriptingSystem() { lua_close(L_); }
};
Lua Game Logic Example
-- game_init.lua: run at game start
local player_id = create_entity()
add_transform(player_id, 100, 200)
add_sprite(player_id, "player", 0)
add_rigidbody(player_id, 0, 0, 1, false)
add_collider(player_id, 32, 32)
Calling C++ Callbacks from Lua
To handle game logic in Lua and call C++ functions on specific events, use lua_pcall and table-based callback registration.
-- Lua side: register on_collision
function on_collision(a_id, b_id)
if get_entity_tag(a_id) == "player" and get_entity_tag(b_id) == "coin" then
add_score(10)
destroy_entity(b_id)
end
end
// C++ side: call Lua callback
void PhysicsSystem::on_collision_detected(Entity* a, Entity* b) {
lua_getglobal(L_, "on_collision");
if (lua_isfunction(L_, -1)) {
lua_pushinteger(L_, a->get_id());
lua_pushinteger(L_, b->get_id());
if (lua_pcall(L_, 2, 0, 0) != LUA_OK) {
fprintf(stderr, "Lua callback error: %s\n", lua_tostring(L_, -1));
lua_pop(L_, 1); // remove the error message
}
} else {
lua_pop(L_, 1); // remove the non-function value
}
}
Stack hygiene is the part people get wrong here. lua_getglobal always pushes something — nil if the global does not exist — and a failed lua_pcall leaves the error message on the stack. Without the two lua_pop calls, every frame with a collision and no handler would leak one stack slot; eventually Lua fails with a “stack overflow” error or memory keeps growing for no visible reason. Always use lua_pcall rather than lua_call from engine code: an error in an unprotected call longjmps out of C++ frames, skipping destructors. Also be careful what a callback does: here destroy_entity(b_id) is called from inside the physics loop, while check_collisions is still iterating a vector of raw Entity*. Destroying immediately would leave a dangling pointer in that vector — this is exactly the case the deferred destruction in Issue 7 exists for.
Putting It Together: The Game Loop
Game Loop Flow
flowchart LR
subgraph Frame[One Frame]
A[Input processing] --> B[Physics update]
B --> C[Script update]
C --> D[Rendering]
D --> E[Frame limiting]
end
E --> A
Game Loop Integration
The loop measures real elapsed time with SDL_GetPerformanceCounter (a high-resolution counter, unlike the millisecond SDL_GetTicks) and passes it to physics as dt, so movement speed is the same whether the game renders at 30 or 144 frames per second. dt is clamped to two target frames: after a stall — dragging the window on Windows blocks the loop, as does hitting a breakpoint — the first dt can be several seconds, and feeding that into physics teleports objects through walls.
RenderSystem is held through a unique_ptr because it cannot be constructed until the SDL renderer exists, which happens in init(), after the GameEngine constructor has already run. Holding it by value would require a default constructor and a later assignment, and that assignment would copy the texture handles (see the ownership note in the rendering section).
class GameEngine {
SDL_Window* window_ = nullptr;
SDL_Renderer* renderer_ = nullptr;
EntityManager entities_;
std::unique_ptr<RenderSystem> render_system_; // created after the renderer exists
PhysicsSystem physics_system_;
InputSystem input_system_;
ScriptingSystem script_system_;
bool running_ = true;
const float target_dt_ = 1.0f / 60.0f;
public:
bool init() {
if (SDL_Init(SDL_INIT_VIDEO) != 0) return false;
window_ = SDL_CreateWindow("2D Engine", SDL_WINDOWPOS_CENTERED,
SDL_WINDOWPOS_CENTERED, 800, 600, 0);
if (!window_) return false;
renderer_ = SDL_CreateRenderer(window_, -1, SDL_RENDERER_ACCELERATED);
if (!renderer_) return false;
render_system_ = std::make_unique<RenderSystem>(renderer_);
script_system_.set_entity_manager(&entities_);
script_system_.register_api();
create_sample_scene();
return true;
}
void create_sample_scene() {
// Player
auto& player = entities_.create_entity();
player.add_component<TransformComponent>().position = {100, 200};
player.add_component<SpriteComponent>();
auto& spr = *player.get_component<SpriteComponent>();
spr.texture_id = "player";
spr.src_rect = {0, 0, 32, 32};
spr.z_index = 1;
player.add_component<RigidBodyComponent>();
player.add_component<ColliderComponent>();
// Floor
auto& floor = entities_.create_entity();
floor.add_component<TransformComponent>().position = {0, 500};
auto& floor_spr = floor.add_component<SpriteComponent>();
floor_spr.texture_id = "floor";
floor_spr.src_rect = {0, 0, 800, 100};
floor_spr.z_index = 0;
auto& floor_rb = floor.add_component<RigidBodyComponent>();
floor_rb.is_static = true;
floor.add_component<ColliderComponent>().width = 800;
floor.get_component<ColliderComponent>()->height = 100;
}
void run() {
Uint64 last = SDL_GetPerformanceCounter();
while (running_) {
Uint64 now = SDL_GetPerformanceCounter();
float dt = static_cast<float>(now - last) / SDL_GetPerformanceFrequency();
last = now;
input_system_.update();
if (input_system_.is_quit_requested()) break;
physics_system_.update(entities_, std::min(dt, target_dt_ * 2));
render_system_->update(entities_);
// Frame limiting
float elapsed = static_cast<float>(SDL_GetPerformanceCounter() - now) /
SDL_GetPerformanceFrequency();
if (elapsed < target_dt_) {
SDL_Delay(static_cast<Uint32>((target_dt_ - elapsed) * 1000));
}
}
}
void shutdown() {
render_system_.reset(); // destroy textures before their renderer
if (renderer_) SDL_DestroyRenderer(renderer_);
if (window_) SDL_DestroyWindow(window_);
SDL_Quit();
}
};
SDL_Delay is a coarse way to cap the frame rate. It sleeps at least the requested time, and on Windows the default timer resolution can make a 3 ms request take closer to 15 ms, so the frame rate wobbles. Passing SDL_RENDERER_PRESENTVSYNC to SDL_CreateRenderer lets SDL_RenderPresent block until the display refresh instead, which gives smoother pacing and no tearing — at the cost of tying the frame rate to the monitor (60, 120, 144 Hz), which is fine because the loop already uses real dt.
Common Issues and Solutions
Issue 1: Entities fall through floor
Cause: Missing ColliderComponent, or check_collision requires ColliderComponent but not added to floor. If the floor does have a collider and the player still sinks slowly, the cause is the velocity-only collision response described in the physics section — the overlap is never corrected.
Solution:
// ❌ Wrong: floor missing ColliderComponent
auto& floor = entities_.create_entity();
floor.add_component<TransformComponent>();
floor.add_component<RigidBodyComponent>().is_static = true;
// ColliderComponent missing!
// ✅ Correct
floor.add_component<ColliderComponent>().width = 800;
floor.get_component<ColliderComponent>()->height = 100;
Issue 2: Render order changes every frame
Cause: std::sort is not stable, and its input comes from an unordered_map whose iteration order can change as entities are added. Sprites with the same z_index therefore swap places between frames, which shows up as flickering where two sprites overlap. Adding a tie-breaker that is fixed per entity (its ID) makes the order deterministic; with a total order like this, plain std::sort would work too.
Solution:
// ✅ Stable sort + secondary sort by entity ID
std::stable_sort(entities_to_render.begin(), entities_to_render.end(),
[](Entity* a, Entity* b) {
int za = a->get_component<SpriteComponent>()->z_index;
int zb = b->get_component<SpriteComponent>()->z_index;
if (za != zb) return za < zb;
return a->get_id() < b->get_id(); // Fixed order on same z_index
});
Issue 3: Lua script “attempt to call a nil value”
Cause: The script calls a global that was never registered. In this code that happens when register_api() runs before set_entity_manager() (it returns early without registering anything), or when a script calls functions such as add_transform that the C++ side has not exposed yet. The full message, e.g. attempt to call a nil value (global 'create_entity'), names the missing function. A related bug is registering with lua_register and a function that reads lua_upvalueindex(1): lua_register creates a closure with no upvalues, so the pointer comes back null and the call does nothing or crashes.
Solution:
// ✅ Pass context via upvalue
void register_api() {
lua_pushlightuserdata(L_, entities_);
lua_pushcclosure(L_, [](lua_State* L) -> int {
auto* em = static_cast<EntityManager*>(lua_touserdata(L, lua_upvalueindex(1)));
// Use em
return 1;
}, 1);
lua_setglobal(L_, "create_entity");
}
Issue 4: Physics “tunneling” at high speeds (fast objects pass through walls)
Cause: Single-frame movement distance exceeds collider size, skipping collision detection. Solution: Use continuous collision detection (CCD) or substeps.
A quick sanity check is speed × dt versus the thinnest collider: a bullet at 2,000 px/s with a 16 ms frame moves 32 px per step and will skip a 20 px wall. Substeps reduce the per-step distance but multiply the cost of the collision loop, so they suit a handful of fast objects better than a whole scene. CCD instead sweeps the box along its path and finds the time of first contact; Box2D does this for bodies marked as bullets.
// ✅ Substeps: divide dt for multiple physics updates
const int substeps = 4;
float sub_dt = dt / substeps;
for (int i = 0; i < substeps; ++i) {
physics_system_.update(entities_, sub_dt);
}
Issue 5: Black screen on texture load failure
Cause: Not checking IMG_Load failure returning nullptr before calling SDL_CreateTextureFromSurface.
Solution:
// ✅ Error checking
void load_texture(const std::string& id, const std::string& path) {
SDL_Surface* surface = IMG_Load(path.c_str());
if (!surface) {
SDL_Log("Failed to load %s: %s", path.c_str(), IMG_GetError());
return;
}
SDL_Texture* texture = SDL_CreateTextureFromSurface(renderer_, surface);
SDL_FreeSurface(surface);
if (!texture) {
SDL_Log("Failed to create texture: %s", SDL_GetError());
return;
}
textures_[id] = texture;
}
Issue 6: Physics “looks slow” when game runs below 60 FPS
Cause: Large dt increases single-frame movement distance, destabilizing physics. On low-spec PCs, dt can exceed 1/30 second. The clamp in the main loop keeps physics stable, but it also means that when frames take longer than the clamp, the game world runs slower than real time.
Solution: Run physics at a fixed timestep and consume real time from an accumulator.
A fixed step makes physics deterministic — the same inputs give the same results regardless of frame rate — which is also a prerequisite for replays and lockstep networking. The accumulator must be a member variable that survives between frames; the leftover fraction carries over to the next frame. Rendering then happens once per frame, and advanced engines interpolate positions by accumulated / target_dt_ to hide the stutter of physics and display running at different rates. The cap on how much time is added per frame prevents the “spiral of death”, where a slow frame schedules more physics steps, which makes the next frame slower still.
// ✅ dt clamping + fixed timestep
const float max_dt = 1.0f / 30.0f; // Physics based on max 30 FPS
float accumulated = 0; // make this a member: it must persist across frames
accumulated += std::min(dt, max_dt);
while (accumulated >= target_dt_) {
physics_system_.update(entities_, target_dt_);
accumulated -= target_dt_;
}
Issue 7: Crash on entity deletion (use-after-free)
Cause: After destroy_entity call, other systems continue referencing that entity pointer.
Solution: Use deferred destruction pattern.
// ✅ Delete at next frame start
std::vector<EntityID> to_destroy_;
void mark_for_destruction(EntityID id) {
to_destroy_.push_back(id);
}
void process_destruction() {
for (auto id : to_destroy_) {
entities_.destroy_entity(id);
}
to_destroy_.clear();
}
// Call process_destruction() at run() loop start
The reason immediate deletion crashes is that systems hold raw Entity* lists for the duration of their update, and callbacks (collision handlers, Lua scripts) can request deletion from inside that update. Deferring keeps every pointer valid until the frame ends. The remaining hole is stale IDs held across frames — for example an enemy storing its target’s EntityID. Since next_id_ never reuses numbers here, a lookup of a destroyed ID simply fails, which is safe; engines that recycle IDs add a generation counter to each ID so an old handle cannot silently refer to a new entity.
Performance Optimization Tips
Optimize queries with component indices
get_entities_with traverses all entities every frame. Maintaining entity ID lists per component type enables near-O(1) lookup.
// Update index on component add/remove
std::unordered_map<std::type_index, std::vector<EntityID>> component_index_;
Render batching
Grouping sprites using same texture and drawing together reduces draw calls. Be aware of two limits. Since SDL 2.0.10 the SDL renderer already batches consecutive draws internally, so the gain from manual grouping is smaller than with raw OpenGL. And grouping by texture conflicts with z-ordering: if sprites from two textures interleave in depth, drawing all of texture A and then all of texture B draws some sprites in the wrong order. Batching must happen within each z-layer, or you pack sprites into a texture atlas so that one texture serves most of the scene.
// Group by texture_id then batch render
std::map<std::string, std::vector<Entity*>> by_texture;
for (auto* e : entities_to_render) {
by_texture[e->get_component<SpriteComponent>()->texture_id].push_back(e);
}
for (auto& [tex_id, list] : by_texture) {
SDL_Texture* tex = textures_[tex_id];
for (auto* e : list) {
// Repeat SDL_RenderCopyEx (same texture)
}
}
Optimize collision detection with spatial partitioning
Instead of O(n²) collision checks, use spatial hash or Quad Tree to compare only objects in same area.
// Simple grid-based spatial partitioning
// std::pair has no std::hash specialization, so pack the cell coordinates into one key
std::unordered_map<std::uint64_t, std::vector<Entity*>> spatial_grid_;
// key = (uint64_t(uint32_t(cx)) << 32) | uint32_t(cy)
// Divide into cells like 64x64, check collisions only in same/adjacent cells
The O(n²) pair loop is fine for a few dozen objects: 50 bodies are 1,225 checks. At 1,000 bodies it is about half a million checks per physics step, and substeps multiply that. A uniform grid works well when objects are roughly the same size; choose a cell size around the size of a typical object, insert each object into every cell its box touches, and test only pairs that share a cell (deduplicating pairs that share several cells). Quadtrees adapt better when object sizes vary a lot, at the cost of rebuilding or updating the tree as things move.
Object pooling
Reuse entities/components from pool instead of new/delete every time.
std::vector<std::unique_ptr<Entity>> entity_pool_;
// On destroy, return to pool instead of actual deletion; on create, fetch from pool
Box2D Integration (Production-Grade Physics)
Custom AABB physics suits simple games. For complex collisions (circles, polygons, joints) or stable rigid body simulation, use Box2D.
Box2D and ECS Integration
#include <box2d/box2d.h>
class Box2DPhysicsSystem {
b2World world_{{0, 9.8f}};
std::unordered_map<EntityID, b2Body*> entity_to_body_;
public:
void sync_to_physics(EntityManager& entities) {
for (auto* entity : entities.get_entities_with<TransformComponent, RigidBodyComponent, ColliderComponent>()) {
if (entity_to_body_.count(entity->get_id())) continue; // body already created
auto* t = entity->get_component<TransformComponent>();
auto* rb = entity->get_component<RigidBodyComponent>();
auto* col = entity->get_component<ColliderComponent>();
b2BodyDef def;
def.position.Set(t->position.x / 100.0f, t->position.y / 100.0f); // Pixels→meters
def.type = rb->is_static ? b2_staticBody : b2_dynamicBody;
b2Body* body = world_.CreateBody(&def);
b2PolygonShape box;
box.SetAsBox(col->width / 200.0f, col->height / 200.0f);
b2FixtureDef fix;
fix.shape = &box;
fix.density = 1.0f;
body->CreateFixture(&fix);
entity_to_body_[entity->get_id()] = body;
}
}
void step(float dt) {
world_.Step(dt, 6, 2); // velocityIterations, positionIterations
}
void sync_from_physics(EntityManager& entities) {
auto physics_entities = entities.get_entities_with<
TransformComponent, RigidBodyComponent, ColliderComponent>();
for (auto* entity : physics_entities) {
auto it = entity_to_body_.find(entity->get_id());
if (it == entity_to_body_.end()) continue;
b2Body* body = it->second;
auto* t = entity->get_component<TransformComponent>();
auto* rb = entity->get_component<RigidBodyComponent>();
auto pos = body->GetPosition();
t->position = {pos.x * 100.0f, pos.y * 100.0f};
auto vel = body->GetLinearVelocity();
rb->velocity = {vel.x * 100.0f, vel.y * 100.0f};
}
}
};
Caution: Box2D uses meter units. Define pixel-to-scale ratio (e.g., 100 pixels = 1 meter) for conversion.
The unit conversion is not cosmetic. Box2D’s tolerances (linear slop, sleep thresholds) are tuned for moving objects roughly 0.1 to 10 meters in size; feeding pixel values directly makes a 32-pixel sprite a 32-meter object, and everything falls in slow motion and behaves like a skyscraper. Other details in this integration: SetAsBox takes half-widths, hence the division by 200 rather than 100; Box2D positions a body by its center, while this engine’s TransformComponent uses the top-left corner, so a real integration also offsets by half the size in both directions; and the body angle comes back in radians. sync_to_physics must create each body only once (hence the early continue), and destroying an entity must call world_.DestroyBody and erase the map entry, or Box2D keeps simulating a ghost. Finally, this code uses the Box2D 2.4 C++ API; Box2D 3.0 switched to a C API with handles (b2CreateWorld, b2CreateBody), so the calls differ if you install the latest version.
Build and Dependencies
CMakeLists.txt Example
cmake_minimum_required(VERSION 3.16)
project(game_engine LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
find_package(SDL2 REQUIRED)
find_package(SDL2_image REQUIRED)
find_package(glm CONFIG REQUIRED)
find_package(Lua REQUIRED)
add_executable(game_engine
main.cpp
engine/entity.cpp
engine/render_system.cpp
engine/physics_system.cpp
)
target_include_directories(game_engine PRIVATE
${SDL2_INCLUDE_DIRS}
${LUA_INCLUDE_DIR}
)
target_link_libraries(game_engine
SDL2::SDL2
SDL2_image::SDL2_image
glm::glm
Lua::Lua
)
Installing Dependencies with vcpkg
# After installing vcpkg
vcpkg install sdl2 sdl2-image glm lua
cmake -B build -DCMAKE_TOOLCHAIN_FILE=[vcpkg root]/scripts/buildsystems/vcpkg.cmake
cmake --build build
Platform-Specific Notes
| Platform | Notes |
|---|---|
| Windows | SDL2 supports DLL dynamic linking or static linking. Static linking simplifies distribution |
| macOS | brew install sdl2 sdl2_image lua or vcpkg |
| Linux | apt install libsdl2-dev libsdl2-image-dev liblua5.4-dev |
Production Patterns
Config File Loading (JSON/YAML)
// config.json: {"gravity": [0, 9.8], "target_fps": 60}
// Load with nlohmann/json, etc., then inject into PhysicsSystem, GameEngine
Scene Transition System
class SceneManager {
std::string current_scene_;
std::function<void(EntityManager&)> load_scene_;
public:
void load(const std::string& name) {
entities_.clear(); // Or destroy_all
load_scene_ = scene_registry_[name];
load_scene_(entities_);
current_scene_ = name;
}
};
Save/Load (Serialization)
// Serialize Transform, RigidBody, etc. components to JSON/binary
void save_game(const std::string& path) {
nlohmann::json j;
for (auto& [id, entity] : entities_) {
j["entities"].push_back(serialize_entity(*entity));
}
std::ofstream f(path);
f << j.dump();
}
Debug Overlay
// Display FPS, entity count, physics computation time with ImGui or SDL
void render_debug_overlay() {
ImGui::Text("FPS: %.1f", 1.0f / dt_);
ImGui::Text("Entities: %zu", entities_.size());
}
Implementation Checklist
- Install and link SDL2/SDL_image (vcpkg:
vcpkg install sdl2 sdl2-image) - Install Lua 5.4 (
vcpkg install lua) - Handle texture path errors (relative path vs execution path)
- Adjust viewport/scale on window resize
- Check memory leaks (Valgrind/ASan)
- Verify NDEBUG and optimization flags in release build
Fix the timestep before tuning collisions
Many of the physics symptoms in the introduction, objects falling through floors or bouncing differently on fast and slow machines, come from running the simulation with the frame’s variable delta time. A long frame moves a fast object far enough to skip past a thin collider entirely, and behavior changes with frame rate. The standard fix is a fixed physics timestep: accumulate real elapsed time, run the physics update in constant steps (for example 1/60 s) while the accumulator holds at least one step, and render with whatever remains, optionally interpolating between the last two physics states.
Cap the accumulated time as well. If a frame takes longer than the physics can catch up on, the next frame has even more steps to run, and the game spirals into a freeze. Dropping the excess time after a few steps keeps the game responsive at the cost of briefly running in slow motion. Only after the timestep is fixed does it make sense to tune substeps, collider thickness or continuous collision detection for very fast objects.
References
Frequently Asked Questions (FAQ)
Previous: [C++ Practical Guide #50-2] Building REST API Server
Q. Why do fast objects pass straight through walls?
A. This is tunneling: in a single frame the object moves farther than the collider it should hit, so the discrete AABB check never sees an overlap. Split the physics step into substeps (divide dt and run the physics update several times per frame) or use continuous collision detection (CCD).
Q. When should I replace the custom AABB physics with Box2D?
A. The simple AABB system in this post is fine for axis-aligned boxes. Once you need circles, polygons, joints, or stable rigid-body stacking, integrate Box2D (or a similar physics library): map each entity to a Box2D body, convert pixels to meters when creating bodies, and sync positions between the bodies and your Transform components every frame.