Tutorial 78: Multi-threading Considerations
What you’ll learn
- Which parts of CNA are safe to touch from a worker thread.
- A job queue built on C++23
std::jthread. - Splitting logic onto workers while drawing stays on the main thread.
- Double-buffering game state to avoid races.
Before you start — Tutorial 77: Asset Streaming and Async Loading (the same main-thread constraint, in a narrower case) and Tutorial 71: Memory Management in C++ (ownership across threads). Audio is the case the page treats separately.
What is thread-safe in CNA?
Almost nothing in CNA is thread-safe by default, exactly like the XNA API it mirrors (XNA’s own reference documents static members as thread-safe and instance members as not guaranteed). These are the rules:
| API / Object | Thread-safe? | Notes |
|---|---|---|
GraphicsDevice | No | Owner-thread only: use and dispose it on the thread that owns it, the game thread |
SpriteBatch, Texture2D, buffers, render targets | No | Not concurrently thread-safe, and drawing stays on the game thread. On the GL renderers (EasyGL) constructing, SetData, GetData and Dispose may run on a loading thread: each takes the GL context and waits for the gap between frames. On the other renderers keep them on the game thread |
Effect, BasicEffect | No | Issues GPU state commands |
Texture2D::GetData | No | Reads GPU memory — the game thread, or a loading thread on EasyGL (it takes the GL context and runs between frames) |
ContentManager::Load | No | Creates GPU resources, and the asset cache has no lock: never load through one manager from two threads. On EasyGL one loading thread may own a manager (the content reader holds the GL context for a whole asset); elsewhere load on the game thread |
Keyboard, Mouse, GamePad, TouchPanel | No | Input is a single-threaded (game-loop thread) API, as in XNA: the state is refreshed once per tick on the loop thread. Read it there and pass plain copies to workers |
Game component collection (Add/Insert) | Yes | The game’s component lists are guarded by a mutex, so a loading thread may add a component; use the std::shared_ptr overload so the collection keeps it alive |
Accelerometer, Gyroscope (property getters, Start/Stop) | Yes | A CNA guarantee beyond the XNA baseline, documented in CNA’s devices thread-safety contract |
SoundEffectInstance::Play/Stop | Not promised | CNA gives no cross-thread contract for audio instances (its mixer has its own internal lock against the audio callback thread, which is not the same thing); see the audio section |
Vector3, Matrix, etc. | Yes | Value types, no shared state (separate instances) |
CNA::Diagnostics counters and profiler zones | Yes | Designed for use from any thread: counters are atomic, zones use a per-thread event ring |
std::vector of game objects | No (by default) | Wrap in a mutex or use separate read/write buffers |
The fundamental constraint is that drawing belongs to the thread that owns the GraphicsDevice, the game thread: never draw from a worker. On the GL renderers (EasyGL) resource creation, upload and readback from a loading thread are supported, because each of those calls takes the GL context lease and waits for the gap between frames. That makes one pattern a deadlock: the game thread waiting, inside LoadContent, Update or Draw, for a worker that is creating graphics resources, since the game thread holds the context for the whole frame (FNA and MonoGame deadlock the same way). On the other renderers, keep every graphics call on the game thread.
Threads on the web. A default Emscripten build is single-threaded. Worker threads need -DCNA_ENABLE_EMSCRIPTEN_THREADS=ON (Emscripten only, default OFF), which changes the WebAssembly ABI for the whole application, and the hosting page must be cross-origin isolated for the browser to provide shared memory. Code in this tutorial that starts std::jthreads therefore needs a fallback (run the jobs inline) for the default web build. A threaded build also uses WasmFS by default, which cannot host CNA’s IndexedDB save storage; -DCNA_EMSCRIPTEN_USE_WASMFS=OFF brings persistence back at the cost of the legacy file system, which can wait on the browser’s main thread during background loads.
Worker threads for logic
Safe candidates for worker thread offloading:
- Physics simulation (position integration, broad-phase collision)
- AI pathfinding (A* graph search, behaviour tree evaluation)
- Procedural mesh generation (marching cubes, noise terrain)
- Asset decompression and decoding (PNG pixel extraction, OGG decode)
- Frustum culling / visibility determination (read-only scene data)
- Animation blending (matrix palette computation)
All of these produce plain C++ data (positions, matrices, vertex arrays) that the main thread can consume safely if you use double-buffering or a mutex-protected queue.
Job queue with std::jthread (C++23)
std::jthread (C++23) is a thread that automatically joins on destruction, eliminating the need to call join() manually and preventing join-on-destroyed-thread undefined behaviour.
// JobQueue.hpp
#pragma once
#include <thread>
#include <queue>
#include <mutex>
#include <condition_variable>
#include <functional>
#include <vector>
#include <atomic>
#include <algorithm>
class JobQueue {
public:
using Job = std::function<void()>;
explicit JobQueue(int threadCount = 0) {
if (threadCount <= 0)
threadCount = static_cast<int>(
std::thread::hardware_concurrency()) - 1;
threadCount = std::max(1, threadCount);
for (int i = 0; i < threadCount; ++i) {
workers_.emplace_back([this](std::stop_token st) {
for (;;) {
Job job;
{
std::unique_lock<std::mutex> lock(mutex_);
// This wait overload also returns when a stop is requested;
// it returns false only if stopped with the queue empty.
if (!cv_.wait(lock, st, [&] { return !queue_.empty(); }))
return;
job = std::move(queue_.front());
queue_.pop();
}
job();
--pending_;
}
});
}
}
// JobQueue destructor automatically requests stop and joins all jthreads
~JobQueue() = default;
// Submit a job. Thread-safe.
void Submit(Job job) {
++pending_;
{
std::lock_guard<std::mutex> lock(mutex_);
queue_.push(std::move(job));
}
cv_.notify_one();
}
// Wait until all submitted jobs have completed.
void WaitAll() {
while (pending_.load() > 0)
std::this_thread::yield();
}
int Pending() const { return pending_.load(); }
private:
std::queue<Job> queue_;
std::mutex mutex_;
std::condition_variable_any cv_;
std::atomic<int> pending_{0};
// Declared LAST on purpose: members are destroyed in reverse order, so the
// threads are joined while the queue, mutex and condition variable still exist.
std::vector<std::jthread> workers_;
};
Game logic on workers, draw on main
The standard pattern is to run expensive Update work on the job queue, then collect results on the main thread and draw:
// In Game::Update — dispatch expensive logic to workers
void Update(GameTime& gt) override {
Game::Update(gt); // components, FrameworkDispatcher
float dt = static_cast<float>(gt.getElapsedGameTimeProperty().getTotalSecondsProperty());
// Submit independent AI jobs (each entity is independent)
for (auto& entity : entities_) {
jobQueue_->Submit([&entity, dt]() {
entity.UpdateAI(dt); // read-only world, writes to entity only
entity.IntegratePhysics(dt);
});
}
// Meanwhile, do main-thread-only work:
ProcessInput();
UpdateCamera(dt);
// Wait for all worker jobs to complete before drawing
jobQueue_->WaitAll();
// Now safe to read all entity positions for culling/rendering
BuildRenderList();
}
void Draw(const GameTime&) override {
auto& gd = getGraphicsDeviceProperty();
gd.Clear(Color::CornflowerBlue);
for (auto& entry : renderList_) {
DrawObject(gd, entry);
}
// no Present(): Game presents after Draw() returns
}
Lockless ring buffer for frame data
For high-frequency data exchange between a worker thread and the main thread (e.g. streaming audio sample positions, or particle positions), a single-producer single-consumer lock-free ring buffer avoids mutex overhead entirely:
// LocklessRingBuffer.hpp — SPSC, power-of-two capacity
#include <array>
#include <atomic>
#include <cstddef>
template <typename T, size_t N>
class LocklessRingBuffer {
static_assert((N & (N - 1)) == 0, "N must be a power of two");
public:
// Called from producer thread only
bool TryPush(const T& value) {
size_t head = head_.load(std::memory_order_relaxed);
size_t next = (head + 1) & (N - 1);
if (next == tail_.load(std::memory_order_acquire)) return false; // full
buffer_[head] = value;
head_.store(next, std::memory_order_release);
return true;
}
// Called from consumer thread only
bool TryPop(T& value) {
size_t tail = tail_.load(std::memory_order_relaxed);
if (tail == head_.load(std::memory_order_acquire)) return false; // empty
value = buffer_[tail];
tail_.store((tail + 1) & (N - 1), std::memory_order_release);
return true;
}
private:
std::array<T, N> buffer_{};
std::atomic<size_t> head_{0};
std::atomic<size_t> tail_{0};
};
Audio: trigger sounds from the main thread
CNA’s SoundEffectInstance plays through a mixer (SDL3_mixer with the default CNA_AUDIO_PLATFORM=SDL3, or CNA’s own mixer with ALSA; the NULL audio platform has no mixer at all), and the mixer runs its own audio callback thread. The audio module protects the state it shares with that thread with an internal lock, but CNA makes no promise that calling Play(), Stop() or the other instance members from your worker threads is safe, and neither does the XNA API it mirrors. Treat audio like the rest of CNA: call it from the game thread.
Workers can still cause sounds. Have them post small plain-data requests into the lockless ring buffer above and let the main thread play them from Update. The same goes for constructing a SoundEffect (it decodes the file and registers it with the mixer): create it on the main thread in LoadContent or via the BackgroundLoader approach in Tutorial 77 (decode on a worker, construct on the main thread).
struct SoundRequest { int soundId; float volume; };
// Worker side (physics thread): only posts a request
class PhysicsSystem {
public:
explicit PhysicsSystem(LocklessRingBuffer<SoundRequest, 64>& out) : out_(out) {}
void OnCollision(const CollisionEvent& ev) {
if (ev.impactSpeed > 5.0f)
out_.TryPush({kCollisionSound, 1.0f}); // full queue: the sound is simply dropped
}
private:
static constexpr int kCollisionSound = 0;
LocklessRingBuffer<SoundRequest, 64>& out_;
};
// Main thread (in Game::Update): drain the queue and touch the audio API
SoundRequest req;
while (soundRequests_.TryPop(req)) {
SoundEffectInstance& inst = *instances_[req.soundId]; // created on the main thread
inst.setVolumeProperty(req.volume);
inst.Play();
}
Avoiding data races: double-buffering game state
If workers write entity positions while the main thread reads them for rendering, you have a data race. The cleanest solution is double-buffering: workers write to a "back" state buffer while the renderer reads the "front" buffer. At the end of each frame, swap the pointers:
struct EntityState { Vector3 position; float rotation; };
// Two copies: front (renderer reads) and back (workers write)
std::vector<EntityState> stateA_, stateB_;
std::vector<EntityState>* frontState_ = &stateA_;
std::vector<EntityState>* backState_ = &stateB_;
void Update(GameTime& gt) override {
float dt = static_cast<float>(gt.getElapsedGameTimeProperty().getTotalSecondsProperty());
// Workers write to backState_ — frontState_ is untouched
for (int i = 0; i < static_cast<int>(entities_.size()); ++i) {
jobQueue_->Submit([this, i, dt]() {
(*backState_)[i] = SimulateEntity(entities_[i], dt);
});
}
jobQueue_->WaitAll();
// Swap: back becomes new front
std::swap(frontState_, backState_);
}
void Draw(const GameTime&) override {
// Read from frontState_ — safe, no workers are writing to it now
for (int i = 0; i < static_cast<int>(entities_.size()); ++i) {
DrawEntityAt(entities_[i], (*frontState_)[i].position);
}
}