Tutorial 73: Performance Profiling and Optimization
What you’ll learn
- Measuring frame time from
GameTime, and counting draw calls — by hand and with CNA’s built-inCNA_DIAGNOSTICScounters and profiler zones. - Cutting texture binds with an atlas, and the rules that keep
SpriteBatchbatching. - Buffer orphaning for dynamic geometry.
- Telling a CPU bottleneck from a GPU one, with RenderDoc and Vulkan validation layers.
Before you start — Tutorial 05: The Game Loop (frame timing) and Tutorial 21: SpriteBatch Deep Dive (the batching rules being measured).
GameTime for frame timing
Every Update and Draw call receives a GameTime reference. Read it through its property accessors:
gameTime.getElapsedGameTimeProperty().getTotalSecondsProperty()— the time step of this call. With a variable step (IsFixedTimeStep = false) it is the measured time since the previous frame. With the default fixed step it is not a measurement: everyUpdatesees exactlyTargetElapsedTime(1/60 s by default;0on the game’s very firstUpdate), andDrawsees the target multiplied by the number of steps that ran in that tick.gameTime.getTotalGameTimeProperty().getTotalSecondsProperty()— accumulated game time: the sum of the completed updates’ elapsed times, not wall-clock time sinceGame::Run().gameTime.getIsRunningSlowlyProperty()— set totrueby the fixed-step loop once it has fallen several steps behind and is running catch-upUpdatecalls; it clears when the lag is gone.
Use getIsRunningSlowlyProperty() as an early warning: if it is true on most frames, an Update-plus-Draw tick no longer fits inside the step and the game cannot hold its target rate. See Tutorial 48 for the details of the clock.
Because the fixed step hands you the step rather than a stopwatch reading, a frame-rate counter built only from ElapsedGameTime reports the simulation rate. To measure real frame time, either set IsFixedTimeStep = false while benchmarking, or use the frame duration from the built-in diagnostics described below, which is timed with std::chrono::steady_clock.
FPS counter and draw call counter
A minimal performance overlay can be rendered with SpriteBatch over the game scene. Track frames per second using a rolling average and maintain a per-frame draw call counter that you reset at the start of each Draw:
// PerformanceOverlay.hpp
#pragma once
#include "Microsoft/Xna/Framework/GameTime.hpp"
#include "Microsoft/Xna/Framework/Color.hpp"
#include "Microsoft/Xna/Framework/Vector2.hpp"
#include "Microsoft/Xna/Framework/Graphics/SpriteBatch.hpp"
#include "Microsoft/Xna/Framework/Graphics/SpriteFont.hpp"
#include <string>
#include <deque>
using namespace Microsoft::Xna::Framework;
using namespace Microsoft::Xna::Framework::Graphics;
class PerformanceOverlay {
public:
explicit PerformanceOverlay(SpriteFont& font) : font_(font) {}
void BeginFrame() {
drawCallCount_ = 0;
}
void CountDraw(int primitives = 1) {
++drawCallCount_;
totalPrimitives_ += primitives;
}
void EndFrame(const GameTime& gt) {
float dt = static_cast<float>(gt.getElapsedGameTimeProperty().getTotalSecondsProperty());
frameTimes_.push_back(dt);
if (frameTimes_.size() > SAMPLE_COUNT) frameTimes_.pop_front();
float avg = 0.0f;
for (float t : frameTimes_) avg += t;
avg /= static_cast<float>(frameTimes_.size());
fps_ = (avg > 0.0f) ? 1.0f / avg : 0.0f;
}
void Draw(SpriteBatch& sb, const Vector2& pos) {
std::string text =
"FPS: " + std::to_string(static_cast<int>(fps_)) +
" Draws: " + std::to_string(drawCallCount_) +
" Prims: " + std::to_string(totalPrimitives_);
sb.DrawString(font_, text, pos, Color::Yellow);
totalPrimitives_ = 0;
}
private:
static constexpr size_t SAMPLE_COUNT = 60;
SpriteFont& font_;
std::deque<float> frameTimes_;
float fps_ = 0.0f;
int drawCallCount_ = 0;
int totalPrimitives_ = 0;
};
Wrap every DrawPrimitives / DrawIndexedPrimitives call with overlay.CountDraw(primitiveCount) to track the GPU workload across frames. (This overlay counts only the calls you wrap. SpriteBatch’s own draw calls and the engine’s effect and render-target changes are counted for you by the built-in diagnostics below.)
Built-in diagnostics: CNA_DIAGNOSTICS
CNA ships a renderer-independent, in-process observation layer (CNA::Diagnostics) that already counts the things this tutorial tells you to measure. It is off by default and costs nothing when off; you choose the level when you configure the build:
-DCNA_DIAGNOSTICS= | What you get |
|---|---|
OFF (default) | Every instrumentation macro compiles away. Its arguments are type-checked but never evaluated, so leaving instrumentation in shipping code is free. |
STATS | Counters, gauges and per-frame counters, a bounded history of recent frames (duration and frames per second), and resource metadata for textures, render targets and vertex and index buffers. |
FULL | Everything in STATS plus CPU profiler zones and markers, a bounded event history, and recording that can be written out as a Chrome/Perfetto trace. |
cmake -S . -B build -DCNA_DIAGNOSTICS=FULL # or STATS; the level is shared with your game
cmake --build build
The level is exported to your game as the compile definition CNA_DIAGNOSTICS_LEVEL (0, 1 or 2), so engine and application always agree. With the level at STATS or higher, the engine itself publishes, among others (the complete list is in the Diagnostics reference): the frame scope and the Game/Update and Game/Draw zones (FULL), Runtime/UpdateCount and Runtime/DrawCount, Graphics/DrawCalls (with indexed, non-indexed and indirect variants), Graphics/SubmittedPrimitives, Graphics/EffectChanges, Graphics/RenderTargetChanges, Graphics/TextureBindingChanges, Graphics/SpriteSubmissions, and the audio counters Audio/AllocatedVoices and Audio/VoiceCreations. GPU timings, input metrics and process memory are not published.
Add your own measurements with the macros from CNA/Diagnostics/Instrumentation.hpp, and read the results from the process-wide provider:
#include <fstream>
#include "CNA/Diagnostics/Diagnostics.hpp"
#include "CNA/Diagnostics/Instrumentation.hpp"
void Simulate() {
CNA_PROFILE_SCOPE_CATEGORY("Physics/Simulate", CNA::Diagnostics::Category::Update); // FULL only
CNA_DIAGNOSTICS_FRAME_COUNTER_ADD("Physics/Contacts", 12); // STATS and FULL
CNA_DIAGNOSTICS_GAUGE_SET("Physics/ActiveBodies", 340);
}
// Somewhere you can afford a copy, e.g. once per second:
auto snap = CNA::Diagnostics::GetProvider().CaptureSnapshot();
if (!snap.recentFrames.empty()) {
const auto& frame = snap.recentFrames.back();
double frameMs = static_cast<double>(frame.durationNs) / 1.0e6; // real, monotonic-clock time
for (const auto& m : frame.metrics) // this frame's counters
if (m.name == "Graphics/DrawCalls") { /* m.value draw calls */ }
}
// FULL only: record a few thousand events and open the result in Chrome or Perfetto
auto rec = CNA::Diagnostics::StartRecording(4096);
/* ... run some frames ... */
auto trace = CNA::Diagnostics::StopRecording(rec);
std::ofstream out("trace.json");
(void) trace.WriteChromeTrace(out);
Two limits worth knowing: the event history is bounded (a full producer ring drops new events and counts them, the process ring overwrites the oldest), and CNA’s own CI builds only the default OFF configuration, so the STATS and FULL levels are exercised by unit tests rather than by a CI lane of their own. The Diagnostics reference documents the modes, the complete macro and API list, and the trace format.
Texture atlas to reduce binds
Every time the next sprite (in submission order) uses a different Texture2D than the previous one, the run of sprites that could share a draw call ends. A scene with 500 sprites spread across 100 textures can therefore cost up to one draw call per texture change — roughly 100 when the sprites are grouped by texture, up to 500 when they alternate. Packing the sprites into a single texture atlas removes the texture changes, so the whole scene can be submitted as one run regardless of sprite count. Confirm it with the Graphics/DrawCalls and Graphics/TextureBindingChanges counters rather than assuming.
Build a texture atlas offline with a tool such as TexturePacker or at startup with a custom packer. Store a Rectangle for each sprite's region within the atlas. Pass that rectangle as the sourceRectangle argument to SpriteBatch::Draw:
// Atlas sprites drawn in a single batch (one draw call)
spriteBatch_->Begin();
for (auto& sprite : sprites_) {
spriteBatch_->Draw(
*atlas_, // single shared texture
sprite.position,
sprite.atlasRegion, // Rectangle within the atlas
Color::White,
sprite.rotation,
sprite.origin,
sprite.scale,
SpriteEffects::None,
sprite.depth);
}
spriteBatch_->End(); // one texture, so one run of sprites (measure the actual draw-call count)
SpriteBatch batching rules
Between Begin() and End() SpriteBatch queues your Draw calls (except in Immediate mode), orders them according to the sort mode, and hands them to the active renderer in that order at End(). What ends a run of sprites that can share one GPU draw call:
- A different texture than the previous sprite in the submitted order (in
SpriteSortMode::Texturemode SpriteBatch sorts by texture first to minimise this). SpriteBatch::End(). Blend, sampler, depth-stencil and rasterizer state and the custom effect are fixed perBegin()/End()pair, so every change of state means another pair, and each pair ends a batch. Group sprites by state, not only by texture.- Any internal capacity limit of the active renderer’s sprite path. That limit differs between renderers and is not part of the API, so do not rely on a particular batch size; count draw calls with the diagnostics counters instead.
Sort mode summary:
| SpriteSortMode | Batching | Draw order |
|---|---|---|
| Deferred (default) | Batches by order of Draw calls | Submission order |
| Texture | Sorts by texture — fewest draw calls | Texture-grouped |
| BackToFront | Sorts by depth descending — correct for transparency | Back to front |
| FrontToBack | Sorts by depth ascending — early-Z efficiency | Front to back |
| Immediate | No queue and no sorting — each Draw is submitted at once | Submission order |
Buffer orphaning for dynamic geometry
When you rewrite the vertices of a buffer the GPU may still be reading for the previous frame, naively overwriting it can stall the CPU until the GPU finishes. XNA’s answer is DynamicVertexBuffer with SetDataOptions::Discard (“orphaning”: you promise not to need the old contents, so the driver can hand back fresh memory and let the GPU finish with the old region in parallel). CNA follows XNA’s rule here: a plain VertexBuffer that is currently bound to the device cannot be rewritten (SetData throws System::InvalidOperationException) unless you use a DynamicVertexBuffer and pass Discard or NoOverwrite.
// Created once, sized for the worst case
dynamicVB_ = std::make_unique<DynamicVertexBuffer>(
gd, VertexPositionColor::getVertexDeclarationStatic(),
kMaxVerts, BufferUsage::WriteOnly);
// Every frame: Discard = "I no longer need the previous contents"
dynamicVB_->SetData(
newVertices.data(),
0, // startIndex into newVertices
static_cast<int>(newVertices.size()),
SetDataOptions::Discard); // key: avoids a GPU stall
Use SetDataOptions::NoOverwrite if you are appending to a region of the buffer that the GPU is not currently reading (for example a particle ring buffer). Be aware that the windowed overload, SetData(offsetInBytes, data, startIndex, elementCount, vertexStride, options), accepts the options for conformance but CNA composes such a write in a CPU-side copy and uploads the whole buffer, so the result is always correct but the cost is not XNA’s. How much orphaning helps in practice depends on the renderer; profile before adopting it everywhere.
CPU vs GPU bottleneck identification
Before optimising, determine whether your frame time is CPU-bound or GPU-bound:
- CPU-bound: Removing all
DrawPrimitivescalls (comment out the Draw method body) still leaves a high frame time. The bottleneck is in Update logic, physics, or AI. - GPU-bound: Frame time is proportional to viewport size. Halving the resolution halves the frame time. The bottleneck is pixel throughput, texture bandwidth, or geometry throughput.
- Driver-bound: Frame time is proportional to draw call count but not geometry count. Reducing draw calls (batching) helps more than reducing polygon count.
RenderDoc integration
RenderDoc is an open-source GPU frame debugger that works with the graphics APIs it supports — for example CNA’s OPENGLES3 and VULKAN renderers on Linux and Windows. To capture a frame:
- Launch your CNA binary through RenderDoc (File > Launch Application).
- Press F12 (or the configured hotkey) in-game to capture a frame.
- Open the captured frame and inspect each draw call, shader inputs, and render targets.
RenderDoc shows the contents of each RenderTarget2D after each draw call, which is invaluable for debugging post-processing pipelines (bloom, deferred rendering). No source changes are needed — RenderDoc injects into the process via the driver.
Vulkan validation layers
The VULKAN renderer requests VK_LAYER_KHRONOS_validation automatically in a debug build and disables validation when NDEBUG is present. At instance creation it checks whether the layer is installed; if not, it reports that fact and continues without it. There is no CNA_ENABLE_VULKAN_VALIDATION CMake option or public CnaVulkanConfig API in this snapshot.
# Install your distribution's Khronos validation-layer package, then build Debug.
cmake -S . -B build-vulkan-debug -G Ninja \
-DCMAKE_BUILD_TYPE=Debug \
-DCNA_GRAPHICS_RENDERER=VULKAN
cmake --build build-vulkan-debug
Read the renderer's stderr/debug-messenger output and confirm that the layer is active before treating a zero-message run as evidence. Release builds intentionally omit this diagnostic cost.