Tutorial 60: Instanced Rendering
What you’ll learn
- What instanced rendering saves, and calling
DrawInstancedPrimitives. - Supplying per-instance data through a second
VertexBufferand aVertexBufferBindingarray. - Packing a world matrix and colour per instance.
- Choosing between static and dynamic instance buffers.
Before you start — Tutorial 38: Vertex Buffers and Index Buffers (bindings and buffer usage) and Tutorial 52: Writing Custom Shaders (ShaderEffect) (the instanced vertex shader reads the extra stream). Requires a 3D-capable renderer that supports instancing, such as OPENGLES3 or OPENGL33; the 2D-only SDL_RENDERER throws on 3D calls by default, and STUB draws nothing. The table below covers all 14 identities.
Instancing is a per-renderer capability, not a given. DrawInstancedPrimitives needs hardware instanced draw support (OpenGL 3.3+ / ES 3.0+, or an equivalent pipeline), and it does not consult the Instancing capability on your behalf: on a renderer without it the call fails. STUB has no instancing at all, and the 2D-only SDL_RENDERER throws on 3D calls entirely.
Query the capability, and know its limits. Ask gd.SupportsCapability(GraphicsCapability::Instancing) before choosing this path. Only DIRECTX9 inherits the shared default (true) without an override of its own — it does have a real instanced path — while DIRECTX11, SDL_GPU, VULKAN, METAL and the others answer explicitly, and FNA3D asks its driver at run time. Whether the instanced draw also supports your custom shader is a separate question, answered per renderer in the table below.
Requirements: profile limits. CNA’s default GraphicsProfile is Reach, and this snapshot enforces its ceilings on every renderer. DrawInstancedPrimitives checks both primitiveCount and instanceCount against the per-draw ceiling: 65,535 on Reach, 1,048,575 on HiDef. The 10,000 blades in the example fit inside Reach. A field with more than 65,535 instances — or the “one million billboards” mentioned under Performance Benefit — needs HiDef, which you request in the Game constructor before Initialize() applies your preferences. 32-bit index buffers are HiDef-only too, so keep the blade mesh on 16-bit indices. The full list of profile ceilings, and the errors you will see when one is exceeded, is in Tutorial 152: Reach vs HiDef.
GrassGame() : graphics_(this) {
graphics_.setGraphicsProfileProperty(GraphicsProfile::HiDef);
// Or for the whole project, before the Game is constructed:
// CNA::SetProjectGraphicsProfileEXT(GraphicsProfile::HiDef); // "CNA/ProjectGraphicsProfile.hpp"
}
| Renderer identities | Instancing capability |
Custom ShaderEffect with a per-instance stream (this tutorial’s example) |
|---|---|---|
OPENGLES3, OPENGL33, WEBGL2 | Yes | Yes. GLSL; instance attributes are placed right after the last per-vertex attribute |
VULKAN | Yes | Yes. The program must be SPIR-V |
WEBGPU | Yes | Yes. WGSL; per-instance layouts are built for custom effects |
SDL_GPU | Yes | No. CNA declares “ShaderEffect instancing is not implemented”; stock effects only |
DIRECTX11 | Yes | Not verified for custom HLSL; the stock instanced effect is verified |
DIRECTX9 | Yes (inherited default, real instanced path) | Not verified |
FNA3D | Driver-dependent, queried at run time | No custom ShaderEffect (CustomEffects is false); stock effects refuse instanced draws |
SOFTWARE | Yes (CPU) | No: CustomEffects is false, stock path only |
HEADLESS | Yes (validated and traced, no pixels) | Recorded only |
METAL | Yes (stock effects, multi-stream and base-instance draws; the per-instance world matrix is applied before the effect’s World) | No: custom effects run only in SpriteBatch |
STUB | No | No |
SDL_RENDERER | No (2D only) | No (3D calls throw) |
The shader in this tutorial is GLSL ES 3.00 (#version 300 es), which runs as written on OPENGLES3 and WEBGL2; OPENGL33 needs #version 330 core, VULKAN needs SPIR-V and WEBGPU needs WGSL. See Tutorial 52 for the per-renderer shader dialects.
What is Instanced Rendering?
Instanced rendering is a GPU technique that draws the same mesh geometry many times in a single draw call, with each copy reading its own per-instance data (world transform, colour tint, animation frame, etc.) from a separate GPU buffer. Without instancing you must loop over every object in C++ and issue one draw call per object — each draw call carries significant CPU and driver overhead, and hundreds of draw calls per frame can easily become the rendering bottleneck even on fast hardware.
With instancing the GPU itself handles the repetition. The driver submits one command to the GPU and the hardware replicates the mesh, fetching per-vertex data from the mesh vertex buffer and per-instance data from a second vertex buffer. The result is that drawing 10,000 identical grass blades costs only marginally more GPU time than drawing a single blade, and the CPU cost of the draw submission becomes almost negligible.
Typical use cases include:
- Vegetation — grass, flowers, shrubs, trees with the same mesh but different positions and scales.
- Particle systems — thousands of small quads or sprites with per-particle transforms.
- Crowd rendering — many characters using the same skeleton and mesh but different positions and animation states.
- Rocks, debris, coins — any repeated world decoration object.
- Voxel chunks — large numbers of cube-shaped blocks with per-block transforms and colours.
Instancing is one of the most impactful GPU optimisations available in a real-time 3D engine. It is strongly recommended any time you need to draw more than a few dozen copies of the same mesh per frame.
DrawInstancedPrimitives
CNA exposes instanced rendering through GraphicsDevice::DrawInstancedPrimitives, which directly maps to glDrawElementsInstanced on OpenGL ES 3.0 / OpenGL 3.3, to the equivalent Vulkan instanced draw command, and to the corresponding instanced draw of the Direct3D, WebGPU and CPU renderers.
The full signature is:
void GraphicsDevice::DrawInstancedPrimitives(
PrimitiveType primitiveType,
int baseVertex,
int minVertexIndex,
int numVertices,
int startIndex,
int primitiveCount,
int instanceCount);
| Parameter | Meaning |
|---|---|
primitiveType | Topology: TriangleList, TriangleStrip, LineList, etc. |
baseVertex | Offset added to each index value when fetching vertices. |
minVertexIndex | Lowest index in the index buffer range (for driver range hints). |
numVertices | Number of vertices in the range. |
startIndex | First index to read from the index buffer. |
primitiveCount | Number of primitives per instance (e.g. triangle count). |
instanceCount | How many instances to draw. |
This call must be issued after binding both the per-vertex and per-instance vertex buffers via SetVertexBuffers, as described below, and after an effect has been applied and an index buffer bound; without them it throws InvalidOperationException. instanceCount and primitiveCount must be positive and inside the profile ceiling described under Requirements.
Instance Data via a Second VertexBuffer
The per-instance data lives in a second VertexBuffer — one element per instance. You define a C++ struct for the per-instance data and register a VertexDeclaration that describes its layout to the GPU driver.
A typical per-instance struct contains:
- The world transform for this instance, stored as four
Vector4rows of a 4x4 matrix (a fullMatrixis 64 bytes; storing it as fourVector4s means four vertex attribute slots of 16 bytes each, matching GLSLlayout(location = N) in vec4). - An optional per-instance colour or tint as a
Vector4. - Any other per-instance scalars: animation blend factor, roughness scale, wind phase offset, etc.
Create the instance buffer with BufferUsage::WriteOnly when the CPU only ever writes it, which is almost always; BufferUsage::None additionally allows GetData. Neither value means “static” or “dynamic”: for data you rewrite every frame use a DynamicVertexBuffer (see the last section).
VertexBufferBinding Array
To tell CNA that two vertex buffers are active simultaneously — one per-vertex and one per-instance — use SetVertexBuffers that takes an array of VertexBufferBinding structs:
void GraphicsDevice::SetVertexBuffers(
const std::vector<VertexBufferBinding>& vertexBuffers);
Each VertexBufferBinding has three fields, passed to its constructor as VertexBufferBinding(VertexBuffer*, int vertexOffset, int instanceFrequency):
- VertexBuffer pointer — the buffer to bind.
- vertexOffset — an element (vertex) offset within the buffer, not a byte offset (usually 0).
- instanceFrequency —
0for per-vertex data,1for per-instance data. A value of2would advance one element every 2 instances, but per-instance (1) is by far the most common.
Under OpenGL ES 3.0 this translates to glVertexAttribDivisor(attrib, instanceFrequency) calls for each attribute from that binding. One per-vertex stream plus one per-instance stream, as here, does not need the MultiStreamVertexInput capability; two or more per-vertex streams do.
Per-Instance Data: World Matrix and Colour
The GLSL vertex shader must declare one input attribute for each field in the per-instance struct. A 4x4 world matrix occupies four vec4 attribute slots. For a custom ShaderEffect on OPENGLES3, OPENGL33 and WEBGL2, attribute locations follow declaration element order: the per-vertex elements come first, and the per-instance elements start immediately after the last per-vertex one. The mesh in this example declares three elements (position, normal, texture coordinate at locations 0–2), so the four matrix rows are locations 3 through 6 and the tint colour is location 7. The usage indices you write into a VertexElement do not change where an attribute lands. (CNA’s own stock instanced shaders use fixed locations 12–15 instead.)
Inside main() the four rows are assembled back into a mat4. Do not transpose. CNA’s Matrix is row-major with the XNA row-vector convention, and Matrix::ToColumnMajor simply copies that memory; GLSL reads it as columns, so every uniform matrix in these tutorials reaches the shader as the transpose of the C++ matrix, which is exactly the form M * v needs. Four rows placed straight into mat4(row0, row1, row2, row3) arrive the same way. An extra transpose() would undo that and push the translation into the wrong part of the matrix.
#version 300 es
precision highp float;
// Per-vertex attributes (binding 0)
layout(location = 0) in vec3 a_position;
layout(location = 1) in vec3 a_normal;
layout(location = 2) in vec2 a_texcoord;
// Per-instance attributes (binding 1, instanceFrequency = 1). They follow the
// three per-vertex elements, so they are locations 3-7.
layout(location = 3) in vec4 i_row0;
layout(location = 4) in vec4 i_row1;
layout(location = 5) in vec4 i_row2;
layout(location = 6) in vec4 i_row3;
layout(location = 7) in vec4 i_color;
uniform mat4 u_view;
uniform mat4 u_projection;
out vec2 v_texcoord;
out vec4 v_color;
out vec3 v_normal;
void main() {
// Reconstruct the per-instance world matrix from the four row vectors of
// CNA's row-major Matrix. Placed as GLSL columns they are already the form
// "world * position" needs (the same as ToColumnMajor uniforms): no transpose().
mat4 world = mat4(i_row0, i_row1, i_row2, i_row3);
vec4 worldPos = world * vec4(a_position, 1.0);
gl_Position = u_projection * u_view * worldPos;
// Transform normal by the inverse-transpose of the upper 3x3.
// For uniform scale, the world matrix upper 3x3 works directly.
mat3 normalMat = mat3(transpose(inverse(world)));
v_normal = normalize(normalMat * a_normal);
v_texcoord = a_texcoord;
v_color = i_color;
}
Complete Example: 10,000 Grass Blades
The following example places 10,000 grass blade instances randomly across a 100 x 100 metre field. Each instance gets a random world translation and a slightly varied green tint. The mesh vertex buffer contains a single grass blade geometry (a few triangles), and the instance buffer holds the per-blade data. The entire field is rendered with one DrawInstancedPrimitives call.
The shader is a ShaderEffect built from renderer-native GLSL source text. This snapshot also has a renderer-qualified path for compiled XNA/FNA Effect Framework binaries (see Tutorial 128), but CNA does not compile HLSL .fx source at run time and that is not the path used by this example. ShaderEffect has no Parameters collection, so uniforms are set with SetUniformXxx(), calling Apply() first. See Tutorial 52.
#include "Microsoft/Xna/Framework/Graphics/ShaderEffect.hpp"
#include "System/IO/File.hpp"
struct GrassInstance {
// 4 rows of the world matrix (row-major, matches CNA Matrix layout)
Vector4 row0, row1, row2, row3;
// Per-blade colour tint
Vector4 color;
// Describes the memory layout to CNA. Note the name: calling the member
// `VertexDeclaration` would shadow the type inside this class scope, so
// the stock vertex types use getVertexDeclarationStatic() and so do we.
static const VertexDeclaration& getVertexDeclarationStatic();
};
// 5 x Vector4 = 5 x 4 floats = 80 bytes per instance. GrassInstance is a plain,
// trivially copyable struct (no virtual functions), so SetData<T> can upload it.
// On EasyGL the shader locations follow element order (3..7 here);
// the usage indices are labels. TextureCoordinate 1..4 matches the layout that
// CNA's stock instanced effects read on the Direct3D renderers.
const VertexDeclaration& GrassInstance::getVertexDeclarationStatic()
{
static const VertexDeclaration decl(80, {
VertexElement(0, VertexElementFormat::Vector4, VertexElementUsage::TextureCoordinate, 1),
VertexElement(16, VertexElementFormat::Vector4, VertexElementUsage::TextureCoordinate, 2),
VertexElement(32, VertexElementFormat::Vector4, VertexElementUsage::TextureCoordinate, 3),
VertexElement(48, VertexElementFormat::Vector4, VertexElementUsage::TextureCoordinate, 4),
VertexElement(64, VertexElementFormat::Vector4, VertexElementUsage::Color, 0),
});
return decl;
}
// A minimal camera helper used by the examples in this series. It is application
// code, not a CNA type; Tutorial 34 builds a fuller FpsCamera.
struct Camera {
Matrix view = Matrix::CreateLookAt(Vector3(0.0f, 12.0f, 60.0f), Vector3::Zero, Vector3::Up);
Matrix projection = Matrix::CreatePerspectiveFieldOfView(
MathHelper::ToRadians(60.0f), 16.0f / 9.0f, 0.1f, 300.0f);
const Matrix& View() const { return view; }
const Matrix& Projection() const { return projection; }
void Update(const GameTime&) {}
};
class GrassGame final : public Game {
GraphicsDeviceManager graphics_;
std::unique_ptr<VertexBuffer> meshVB_; // grass blade geometry
std::unique_ptr<VertexBuffer> instanceVB_; // per-instance transforms
std::unique_ptr<IndexBuffer> meshIB_;
std::unique_ptr<ShaderEffect> grassEffect_;
Camera camera_;
int instanceCount_ = 10000; // inside Reach's 65,535 ceiling (see Requirements)
int triCount_ = 4; // triangles per grass blade
public:
GrassGame() : graphics_(this) {}
protected:
void LoadContent() override {
auto& gd = getGraphicsDeviceProperty();
// Build the grass blade mesh (positions + normals + UVs).
// A simple cross-quad: two intersecting quads in world space.
buildGrassMesh(gd, meshVB_, meshIB_, triCount_);
// Build the per-instance data.
std::vector<GrassInstance> instances(instanceCount_);
std::mt19937 rng(42);
std::uniform_real_distribution<float> posDist(-50.0f, 50.0f);
std::uniform_real_distribution<float> colDist(-0.05f, 0.05f);
for (auto& inst : instances) {
float x = posDist(rng);
float z = posDist(rng);
float angle = posDist(rng) * 0.062f; // random Y rotation
Matrix world = Matrix::CreateRotationY(angle)
* Matrix::CreateTranslation(x, 0.0f, z);
// Store rows of the CNA row-major Matrix.
inst.row0 = Vector4(world.M11, world.M12, world.M13, world.M14);
inst.row1 = Vector4(world.M21, world.M22, world.M23, world.M24);
inst.row2 = Vector4(world.M31, world.M32, world.M33, world.M34);
inst.row3 = Vector4(world.M41, world.M42, world.M43, world.M44);
inst.color = Vector4(0.20f + colDist(rng),
0.60f + colDist(rng),
0.10f + colDist(rng),
1.0f);
}
instanceVB_ = std::make_unique<VertexBuffer>(
gd,
GrassInstance::getVertexDeclarationStatic(),
instanceCount_,
BufferUsage::WriteOnly);
// GrassInstance is a plain trivially copyable struct, so the generic
// SetData<T> uploads it; no stock vertex type or SetDataRaw is needed.
instanceVB_->SetData(instances.data(), instanceCount_);
// ShaderEffect takes shader SOURCE TEXT, not a content name or a path.
// This example uses renderer-native ShaderEffect source. Compiled XNA/FNA
// Effect Framework bytecode is a separate, renderer-qualified path.
// See Tutorial 52.
grassEffect_ = std::make_unique<ShaderEffect>(
gd,
System::IO::File::ReadAllText("Content/effects/grass_instanced.vert.glsl"),
System::IO::File::ReadAllText("Content/effects/grass_instanced.frag.glsl"));
// The constructor does not throw on a compile failure.
// GetCompileErrorEXT() returns the compiler log (also written to stderr).
if (!grassEffect_->IsEffectValid()) {
// The shader did not compile. Do not draw with it.
}
}
void Draw(const GameTime&) override {
auto& gd = getGraphicsDeviceProperty();
gd.Clear(Color::SkyBlue);
// Bind mesh buffer (per-vertex, frequency 0) and
// instance buffer (per-instance, frequency 1) together.
// GraphicsDevice::SetVertexBuffers takes a std::vector, not a raw array.
std::vector<VertexBufferBinding> bindings = {
VertexBufferBinding(meshVB_.get(), 0, 0), // per-vertex
VertexBufferBinding(instanceVB_.get(), 0, 1), // per-instance
};
gd.SetVertexBuffers(bindings);
gd.SetIndexBuffer(meshIB_.get());
// Set shared uniforms (view and projection; world comes from instance data).
// Apply() first is the portable order for a ShaderEffect.
float viewCM[16], projCM[16];
camera_.View().ToColumnMajor(viewCM);
camera_.Projection().ToColumnMajor(projCM);
grassEffect_->Apply();
grassEffect_->SetUniformMat4("u_view", viewCM);
grassEffect_->SetUniformMat4("u_projection", projCM);
grassEffect_->SetUniformFloat("u_time", totalTime_);
gd.DrawInstancedPrimitives(
PrimitiveType::TriangleList,
/*baseVertex*/ 0,
/*minVertexIdx*/ 0,
/*numVertices*/ meshVB_->getVertexCountProperty(),
/*startIndex*/ 0,
/*primitiveCount*/triCount_,
/*instanceCount*/ instanceCount_);
// No gd.Present(): Game presents after Draw() returns, so a manual call
// would present the frame twice.
}
float totalTime_ = 0.0f;
void Update(GameTime& gt) override {
totalTime_ += (float)gt.getElapsedGameTimeProperty().getTotalSecondsProperty();
camera_.Update(gt);
}
};
GLSL Instanced Vertex Shader
The following is the full vertex shader for the instanced grass. It reconstructs the per-blade world matrix from the four row-vector attributes and applies a simple wind animation by displacing the top vertices along the X axis using a sine wave keyed to u_time and the blade's world position (so adjacent blades are not in phase). With the matrix assembled as above, world[3][0] and world[3][2] are the blade’s translation X and Z.
#version 300 es
precision highp float;
// Per-vertex (binding 0)
layout(location = 0) in vec3 a_position;
layout(location = 1) in vec3 a_normal;
layout(location = 2) in vec2 a_texcoord;
// Per-instance (binding 1): locations follow the three per-vertex elements
layout(location = 3) in vec4 i_row0;
layout(location = 4) in vec4 i_row1;
layout(location = 5) in vec4 i_row2;
layout(location = 6) in vec4 i_row3;
layout(location = 7) in vec4 i_color;
uniform mat4 u_view;
uniform mat4 u_projection;
uniform float u_time;
out vec2 v_texcoord;
out vec4 v_color;
out vec3 v_normal;
void main() {
// The rows of CNA's row-major Matrix, placed as columns: no transpose().
mat4 world = mat4(i_row0, i_row1, i_row2, i_row3);
// Wind: displace upper part of blade (a_texcoord.y = 0 at root, 1 at tip)
vec3 pos = a_position;
float windStrength = a_texcoord.y * a_texcoord.y; // stronger at tip
float phase = world[3][0] * 0.3 + world[3][2] * 0.2; // unique per blade
pos.x += sin(u_time * 2.0 + phase) * windStrength * 0.15;
vec4 worldPos = world * vec4(pos, 1.0);
gl_Position = u_projection * u_view * worldPos;
mat3 normalMat = mat3(transpose(inverse(world)));
v_normal = normalize(normalMat * a_normal);
v_texcoord = a_texcoord;
v_color = i_color;
}
Performance Benefit
The performance advantage of instancing grows dramatically with instance count. The figures below are an illustrative order-of-magnitude sketch, not a CNA measurement; real numbers depend on the renderer, driver and GPU, so profile on your own target. Roughly, for 10,000 grass blades on a mid-range integrated GPU:
| Approach | Draw calls / frame | Approx. CPU time / frame | Approx. GPU time / frame |
|---|---|---|---|
| One draw call per blade | 10,000 | ~15 ms | ~3 ms |
| Instanced (this tutorial) | 1 | ~0.1 ms | ~3 ms |
The GPU time is roughly the same in both cases — the GPU still processes the same number of vertices and fragments. The enormous saving is entirely on the CPU side: no loop over 10,000 objects, no 10,000 uniform uploads, no 10,000 driver state validations. On integrated GPUs and mobile targets the driver overhead per draw call can be even higher, making instancing even more valuable.
On modern dedicated desktop GPUs, a single draw call for one million simple instances (e.g., a quad billboard) running at 60 fps is entirely feasible on hardware paths, and needs HiDef to get past the 65,535-instance Reach ceiling (up to 1,048,575 per draw). Practical limits are usually set by vertex throughput and fill rate rather than draw call overhead. The SOFTWARE renderer instances on the CPU, so it saves API overhead but not per-instance work.
Instancing stock effects
Instancing is not limited to custom shaders. On OPENGLES3, OPENGL33, WEBGL2, DIRECTX11 and METAL the stock effects such as BasicEffect accept an instanced draw, reading a per-instance world matrix from the second stream: on EasyGL through fixed attribute locations 12–15 (deliberately the last four slots of GLES 3’s guaranteed 16), on Direct3D from TEXCOORD1–TEXCOORD4, on Metal from the first four elements of the per-instance streams (a missing column reads (0, 0, 0, 1)), and the Direct3D renderers throw NotSupportedException if one of the four columns is missing from the declaration. That is what the TextureCoordinate usage indices 1–4 in the GrassInstance declaration line up with. Other renderers were not verified for this route.
Dynamic vs. Static Instance Data
If instance transforms change every frame (e.g., moving particles), do not rewrite the ordinary VertexBuffer that the previous frame’s draw still has bound: SetData on a currently bound buffer throws InvalidOperationException. Use a DynamicVertexBuffer and pass SetDataOptions::Discard (or NoOverwrite when you only append), which is what allows a bound buffer to be rewritten:
// LoadContent: a dynamic instance buffer instead of a plain VertexBuffer
instanceVB_ = std::make_unique<DynamicVertexBuffer>(
gd, GrassInstance::getVertexDeclarationStatic(),
instanceCount_, BufferUsage::WriteOnly);
// Update instance buffer each frame for dynamic instances
void UpdateInstanceBuffer(const std::vector<GrassInstance>& instances) {
instanceVB_->SetData(instances.data(),
0, // first element to read
static_cast<int>(instances.size()),
SetDataOptions::Discard);
}
(The member is declared std::unique_ptr<DynamicVertexBuffer> instanceVB_; in that variant; DynamicVertexBuffer derives from VertexBuffer, so it goes into VertexBufferBinding unchanged.) For completely static environments (a forest that never changes) keep a plain VertexBuffer with BufferUsage::WriteOnly and upload once in LoadContent. For partially dynamic scenes (most grass static, some animated) consider splitting into a static instance buffer and a smaller dynamic one, and issuing two instanced draw calls.
Deep dives on this topic
Long-form pages that explain the exact semantics, invariants and evidence behind this subject.
- Vertex declarations, bindings and stream composition — From C++ vertex values to the renderer boundary: stream layouts, VertexDeclaration rules and profile limits, index widths, dynamic updates, VertexBufferBinding, semantic composition, the minimum-offset fold and draw validation order.