Tutorial 77: Asset Streaming and Async Loading

CNA Tutorials  ·  Architecture

ℹ

What you’ll learn

  • Why a synchronous load stalls the loop, and moving it onto std::async.
  • A thread-safe queue between the loader and the game.
  • Keeping the GPU upload on the main thread even when decode is off it, and where the GL renderers let a loading thread upload too.
  • The loading-screen pattern, and what ContentManager does not guarantee about threads.

Before you start — Tutorial 45: ContentManager and Asset Pipeline (the loader being moved off the main thread) and Tutorial 05: The Game Loop (the loop it must not block).

The problem with loading stalls

Synchronous asset loading in LoadContent blocks the main thread. For a small game with a few textures this is fine — loading takes under a second. For a large game with gigabytes of textures, meshes, and audio, synchronous loading produces freezes that break immersion: the game hangs for several seconds between areas, or shows a frozen loading screen.

The solution is to load assets on a background thread and upload them to the GPU on the main thread when they are ready. The two-phase approach is the portable one. Drawing always belongs on the thread that owns the GraphicsDevice, which is the game thread, and outside the GL renderers everything created from the device (textures, buffers, render targets) should be created, used and disposed there too. The GL renderers (EasyGL: OPENGLES3, OPENGL33, WEBGL2) relax that: constructing a texture, buffer or render target, SetData, GetData and Dispose take the GL context on whatever thread they run on and wait for the gap between frames, so there a loading thread may do the upload as well (see the ContentManager note below).

std::async and std::future for background loading

C++23's standard std::async / std::future provide a simple way to run a function on another thread and retrieve its result later:

// Provided by you (SDL_image, stb_image, ...): decodes a PNG file into RGBA pixels.
// Pure CPU/disk work: no CNA GPU objects and no OpenGL/Vulkan calls in here.
std::vector<Color> LoadPNGToMemory(const std::string& path,
                                   int* width = nullptr, int* height = nullptr);

// Start loading a PNG from disk on a background thread.
// Returns raw pixel data — no GPU operations yet.
std::future<std::vector<Color>> LoadRawPixelDataAsync(
    const std::string& path) {
    return std::async(std::launch::async, [path]() -> std::vector<Color> {
        return LoadPNGToMemory(path);  // pure CPU/disk operation
    });
}

// In LoadContent or wherever you start the load:
auto future = LoadRawPixelDataAsync("assets/level2/terrain.png");

// In Update — check if the load is done without blocking:
if (future.valid() &&
    future.wait_for(std::chrono::seconds(0)) == std::future_status::ready) {
    auto pixels = future.get();  // retrieve result, non-blocking now
    // Upload to GPU on main thread (width/height known from the decoder):
    terrainTex_ = std::make_unique<Texture2D>(
        gd, texWidth_, texHeight_,
        false, SurfaceFormat::Color);
    terrainTex_->SetData(pixels.data(),        // const Color*
                         static_cast<int>(pixels.size()));  // element count, not bytes
}

Thread-safe resource queue

For loading many assets simultaneously, use a producer-consumer queue. Background threads push completed CPU-side data into the queue; the main thread drains the queue each frame to do the GPU upload:

// BackgroundLoader.hpp
#pragma once
#include <future>
#include <queue>
#include <mutex>
#include <functional>
#include <string>
#include <vector>
#include <atomic>
#include <memory>
#include <unordered_map>

struct LoadedAsset {
    std::string              name;
    int                      width = 0;
    int                      height = 0;
    std::vector<Color>       pixels;      // decoded RGBA, CPU side only
    bool                     failed = false;
    std::function<void()>    onComplete;  // optional; invoked later, on the main thread
};

class BackgroundLoader {
public:
    // Enqueue a texture path for background loading. Call from the main thread.
    // onComplete is invoked from PumpUploads(), i.e. on the main thread.
    void EnqueueTexture(const std::string& name,
                        const std::string& path,
                        std::function<void()> onComplete = {}) {
        ++totalQueued_;
        ++pending_;
        futures_.push_back(
            std::async(std::launch::async,
                       [this, name, path, onComplete]() {
                LoadedAsset asset;
                asset.name       = name;
                asset.onComplete = onComplete;
                try {
                    asset.pixels = LoadPNGToMemory(path,
                                       &asset.width, &asset.height);
                } catch (...) {
                    asset.failed = true;  // still counts as finished, or the loading screen never ends
                }
                std::lock_guard<std::mutex> lock(mutex_);
                ready_.push(std::move(asset));
                --pending_;
            }));
    }

    // Call from the main thread each frame.
    // Returns true if all previously enqueued assets have been processed.
    bool PumpUploads(GraphicsDevice& gd,
                     std::unordered_map<std::string,
                                        std::unique_ptr<Texture2D>>& out,
                     int maxPerFrame = 2) {
        for (int uploaded = 0; uploaded < maxPerFrame; ++uploaded) {
            LoadedAsset asset;
            {   // hold the lock only to take the item, never during the GPU upload
                std::lock_guard<std::mutex> lock(mutex_);
                if (ready_.empty()) break;
                asset = std::move(ready_.front());
                ready_.pop();
            }
            if (!asset.failed) {
                auto tex = std::make_unique<Texture2D>(
                    gd, asset.width, asset.height,
                    false, SurfaceFormat::Color);
                tex->SetData(asset.pixels.data(),
                             static_cast<int>(asset.pixels.size()));
                out[asset.name] = std::move(tex);
            }
            if (asset.onComplete) asset.onComplete();
        }
        std::lock_guard<std::mutex> lock(mutex_);
        return pending_.load() == 0 && ready_.empty();
    }

    float Progress() const {
        std::lock_guard<std::mutex> lock(mutex_);
        const int notDone = static_cast<int>(pending_.load() + ready_.size());
        return totalQueued_ > 0
             ? static_cast<float>(totalQueued_ - notDone) / static_cast<float>(totalQueued_)
             : 1.0f;
    }

private:
    int                            totalQueued_ = 0;   // main thread only
    std::atomic<int>               pending_{0};
    mutable std::mutex             mutex_;
    std::queue<LoadedAsset>        ready_;
    std::vector<std::future<void>> futures_;
};

GPU upload on the main thread

The PumpUploads method above is called from the game's Update or Draw. It dequeues up to maxPerFrame assets to avoid uploading so much data in one frame that the game stutters, and it holds the queue lock only while taking an item, never during the upload. Adjust maxPerFrame based on how large each asset is — for small textures set it to 4 or 8; for large ones (4K textures) set it to 1.

Loading screen pattern

Implement a state machine in your game that switches between STATE_LOADING and STATE_PLAYING. During loading, render a progress bar using SpriteBatch (no 3D assets needed — just a solid colour quad) and call PumpUploads each frame:

void Update(GameTime& gt) override {
    Game::Update(gt);
    if (state_ == State::Loading) {
        bool done = loader_->PumpUploads(
            getGraphicsDeviceProperty(), textures_, 4);
        loadProgress_ = loader_->Progress();
        if (done) state_ = State::Playing;
        return;
    }
    // Normal game update...
}

void Draw(const GameTime& gt) override {
    auto& gd = getGraphicsDeviceProperty();
    gd.Clear(Color::Black);

    if (state_ == State::Loading) {
        // Draw a loading bar
        int barW = static_cast<int>(640.0f * loadProgress_);
        spriteBatch_->Begin();
        spriteBatch_->Draw(*whitePixel_,
            Rectangle(80, 350, barW, 20), Color::CornflowerBlue);
        spriteBatch_->Draw(*whitePixel_,
            Rectangle(80, 350, 640, 20), Color::White * 0.3f);
        spriteBatch_->DrawString(*font_,
            "Loading... " +
            std::to_string(static_cast<int>(loadProgress_ * 100)) + "%",
            Vector2(80, 310), Color::White);
        spriteBatch_->End();
        return;   // no Present(): Game presents after Draw() returns
    }
    // Normal game draw...
}

ContentManager thread safety note

CNA's ContentManager::Load<T>() is not safe for concurrent use: one manager keeps its asset cache in plain containers with no lock, so never load through one manager from two threads at once, and loading a Texture2D, Model or Effect creates GPU resources. The BackgroundLoader pattern above bypasses ContentManager intentionally — it reads and decodes raw file data on background threads and only touches CNA GPU types on the main thread, which works on every renderer. (The process-wide reader registries that ContentManager consults are guarded separately, but that does not make a manager safe to share.)

On the GL renderers a loading thread may own a manager. With EasyGL, one loading thread can create its own ContentManager and call Load itself — XNA’s loading-screen pattern: the content reader holds the GL context for a whole asset, between frames, while the game thread keeps drawing the loading screen. Use one manager per loading thread. The one thing that deadlocks is the game thread waiting, inside LoadContent, Update or Draw, for that thread to finish: the game thread holds the GL context for the whole frame, so the two wait for each other (FNA and MonoGame behave the same way). Start the work, keep running frames, and pick up the results once the thread signals that it is done. This applies to the EasyGL identities only; on the other renderers keep Load on the game thread.

If you want to use ContentManager for level streaming on the other renderers, create a separate ContentManager per level on the main thread, but trigger the loading from a background thread using a flag / condition variable to signal when to start. The actual Load calls must happen on the main thread there. ContentManager has no decode/upload split of its own, so loading through it always stalls the frame that calls it; spread Load calls over several frames if a level has many assets.

⚠

Web builds have no threads by default. A default Emscripten build of CNA is single-threaded, so std::async(std::launch::async, ...) cannot start a worker there. Shared-memory pthread support is an application-wide WebAssembly ABI choice you opt into with -DCNA_ENABLE_EMSCRIPTEN_THREADS=ON (Emscripten only, default OFF), and a browser only offers shared memory to cross-origin-isolated pages. For a web target, either enable that option and host accordingly, or load in slices on the main thread. A threaded build uses WasmFS by default, which cannot host CNA’s IndexedDB save storage; -DCNA_EMSCRIPTEN_USE_WASMFS=OFF restores persistent saves, but CNA’s documentation keeps WasmFS recommended for games that load content on worker threads, because the legacy file system can wait on the browser’s main thread during background loads.