Tutorial 62: Multiple Render Targets (MRT)

CNA Tutorials  ·  Advanced Rendering

ℹ

What you’ll learn

  • Binding several render targets at once with SetRenderTargets().
  • Laying out a G-buffer as a RenderTarget2D array.
  • Writing to multiple outputs from one fragment shader.
  • The HiDef profile MRT requires, and which renderers support it.

Before you start — Tutorial 23: Render Targets for Off-Screen Rendering (a single render target first) and Tutorial 52: Writing Custom Shaders (ShaderEffect) (MRT is only useful with your own shader). Requires a 3D-capable renderer with multiple-render-target support, such as OPENGLES3 or OPENGL33 (the table below lists all 14 identities); the 2D-only SDL_RENDERER throws on 3D calls by default, and STUB draws nothing.

⚠

MRT is a per-renderer capability. On the GL family it uses core glDrawBuffers (ES 3.0 / GL 3.0 and later, no ARB extension), elsewhere multiple framebuffer attachments. STUB has no MRT, and the 2D-only SDL_RENDERER throws on 3D calls, including SetRenderTargets with more than one binding. SOFTWARE supports four ordered CPU targets, and METAL binds up to eight.

Query the capability, and know its limits. gd.SupportsCapability(GraphicsCapability::MultipleRenderTargets) is the renderer’s answer and the profile ceiling: it is false on every renderer under Reach. Only DIRECTX9 has no SupportsCapability override (its answer is the inherited default plus the profile); DIRECTX11, SDL_GPU, VULKAN, WEBGPU, METAL and the rest answer explicitly. The public numeric limit is gd.GetRendererLimitEXT(CNA::RendererLimit::MaxColorAttachments), which is populated on VULKAN, SDL_GPU (4) and WEBGPU and reports “unknown” elsewhere (Tutorial 133 reads the whole capability profile).

⚠

Requirements: HiDef profile. CNA’s default GraphicsProfile is Reach, which allows one simultaneous render target, and this snapshot enforces that on every renderer. SetRenderTargets with the three G-buffer bindings below throws NotSupportedException (“SetRenderTargets: 3 render targets exceeds GraphicsProfile.Reach’s own maximum of 1”) until you request HiDef (maximum four) in the Game constructor, before Initialize() applies your preferences:

DeferredGame() : graphics_(this) {
    graphics_.setGraphicsProfileProperty(GraphicsProfile::HiDef);
    // Or for the whole project, before the Game is constructed:
    //   CNA::SetProjectGraphicsProfileEXT(GraphicsProfile::HiDef);   // "CNA/ProjectGraphicsProfile.hpp"
}

The full list of profile ceilings, and the errors you will see when one is exceeded, is in Tutorial 152: Reach vs HiDef.

More than four bindings throw on every renderer, HiDef or not.

Renderer identities MRT (HiDef) A RenderTargetCube face inside a multi-target set
OPENGLES3, OPENGL33, WEBGL2Yes, up to 4 (bounded by the GL limits)Implemented
VULKANYes when maxColorAttachments is above 1, up to 4Implemented (tested)
SOFTWAREYes, 4 ordered CPU targetsImplemented
FNA3DYes (shared default), up to 4Implemented
WEBGPUYes, 2 to 4 RenderTarget2D targetsRefused
SDL_GPUYes, 4 independently writable, mixed-format targetsRefused
DIRECTX11YesImplemented
DIRECTX9Yes (inherited default; the profile limit still applies)Refused
HEADLESSAccepted (validated and traced, no pixels)Refused
METALYes, up to 8 (the stock effects write the first target only)Implemented
STUBNo—
SDL_RENDERERNo (2D only, 3D calls throw)—
⚠

Cube faces inside an MRT set are refused on several otherwise-capable 3D renderers. Plain 2D MRT — several RenderTarget2D objects bound at once, which is what this tutorial builds — works on every MRT renderer in the table above. What does not work everywhere is binding a RenderTargetCube face as one slot of a multi-target set. It is refused with NotSupportedException on DIRECTX9, SDL_GPU and WEBGPU (and with a runtime_error on HEADLESS). It is implemented on the EasyGL family (OPENGLES3, OPENGL33, WEBGL2), VULKAN, DIRECTX11, SOFTWARE, FNA3D and METAL; CNA’s tree has a dedicated cube-MRT test for VULKAN, CNA’s Windows records show the Direct3D 11 implementation passing on one physical Intel machine, METAL’s shared MRT contract passes on a Mac mini M4, and the EasyGL path exists in code without a dedicated test. If your design calls for rendering cube faces and 2D targets in one pass — a point-light shadow cube alongside a G-buffer, say — and must run everywhere, split it into separate passes.

What is MRT?

Multiple Render Targets (MRT) is a GPU feature that lets a single fragment shader invocation write to several textures simultaneously. Normally, the output of the fragment shader goes to one colour attachment — the active render target. With MRT, the shader can declare multiple output variables (one per attachment) and the GPU writes each to a separate texture in a single draw call.

MRT is the foundation of deferred rendering, where the scene is rendered in two major passes:

  1. Geometry pass (G-buffer fill) — draw all opaque geometry once, writing albedo, normals, metalness, roughness, and depth into separate textures. This is where MRT is used.
  2. Lighting pass — draw a fullscreen triangle or quad, reading the G-buffer textures and computing all lighting in screen space. One fullscreen pass instead of one pass per light.

The advantage over forward rendering becomes apparent with many dynamic lights: a forward renderer must re-draw each object once per light it is affected by, while a deferred renderer separates geometry from lighting so each light only costs one fullscreen rectangle pass regardless of how many objects are in the scene.

Other uses of MRT include:

  • Screen-Space Ambient Occlusion (SSAO) — write depth and normals to G-buffer targets, then sample them in the SSAO compute pass.
  • Velocity buffer — write per-pixel motion vectors alongside colour for motion blur or temporal anti-aliasing.
  • OIT accumulation — order-independent transparency techniques that accumulate weighted colour and alpha into two separate targets.

GraphicsDevice::SetRenderTargets()

CNA exposes MRT through a single method on GraphicsDevice, which takes a vector, not a pointer and count:

void GraphicsDevice::SetRenderTargets(
    const std::vector<RenderTargetBinding>& renderTargets);

std::vector<RenderTargetBinding> GraphicsDevice::GetRenderTargets() const;

RenderTargetBinding associates one render target texture with one output slot. The XNA-shaped constructors take a pointer to the concrete render-target type; the 2D one is deliberately implicit, so a RenderTarget2D* converts to a binding on its own:

RenderTargetBinding(RenderTarget2D* renderTarget);                        // implicit, as in XNA
RenderTargetBinding(RenderTargetCube* renderTarget, CubeMapFace face);    // one face of a cube target

CNA also has two CNAEXT forms that take a generic Texture* — explicit RenderTargetBinding(Texture*, int arraySlice = 0) and RenderTargetBinding(Texture*, CubeMapFace) — and getArraySliceProperty(); a RenderTarget2D array slice other than 0 is refused. Because the 2D constructor is implicit you can write std::vector<RenderTargetBinding> targets = { rtA_.get(), rtB_.get() };; the examples below still name the type at each element for readability.

ⓘ

CNA exposes both capability queries and a numeric limit: gd.GetRendererLimitEXT(CNA::RendererLimit::MaxColorAttachments) returns a value with a known flag, known on VULKAN, SDL_GPU and WEBGPU and unknown elsewhere. XNA’s general ceiling is four, GraphicsProfile::Reach lowers it to one on every renderer, and a renderer can impose a smaller native cap. Treat four as a ceiling, not a guaranteed floor.

To restore the default back buffer as the only render target, call the single-target overload with nullptr:

gd.SetRenderTarget(nullptr); // restore back buffer

The same single-target overload is the convenient form when only one target is needed:

gd.SetRenderTarget(rtAlbedo_.get()); // single target

RenderTarget2D Array for the G-Buffer

⚠

This portable MRT layout uses SurfaceFormat::Color. Non-Color render-target formats (float, half-float, Rgba1010102 and the like) are HiDef-only and classified per renderer: OPENGLES3, OPENGL33, WEBGL2, VULKAN, DIRECTX11, SDL_GPU, WEBGPU and SOFTWARE probe them and METAL accepts every float, half and wide render-target format, while FNA3D, DIRECTX9 and HEADLESS are Color-only. A RenderTarget2D that asks for a format the profile or renderer refuses silently becomes Color, so wider-range quantities are encoded into RGBA8 here. Where you need better precision, ask gd.SupportsSurfaceFormatAsRenderTargetEXT(...) first and read getFormatProperty() on the target you actually got. Mixed formats across the slots of one set work on VULKAN and SDL_GPU.

Create one RenderTarget2D per G-buffer slot. This example deliberately fixes the portable format and focuses on how to encode each quantity into RGBA8:

G-Buffer SlotContentsFormatEncoding
RT0 (Albedo)RGB albedo colour + roughness in alphaSurfaceFormat::ColorDirect; albedo is already in [0,1].
RT1 (Normal)World-space XYZ normal + metalness in alphaSurfaceFormat::Colorn * 0.5 + 0.5 per channel. 8 bits per axis is visibly coarse on smooth surfaces; octahedral encoding into RG with metalness in B is a better use of the same four bytes.
RT2 (Depth)Linear depth (or clip-space Z)SurfaceFormat::ColorPack a normalised float across all four channels — see the packDepth/unpackDepth pair in Tutorial 59.
RT3 (Emission)Emissive colour for glowSurfaceFormat::ColorDirect, but LDR only. Store an exposure scale in alpha if you need values above 1.0.

All render targets in a single MRT bind must have matching width, height and applied sample count, and the same subresource cannot occupy two slots. Slot 0 owns the depth buffer for the pass. MRT with multisampling is supported (with matching applied sample counts) on OPENGLES3/OPENGL33/WEBGL2 and VULKAN. Format availability follows the active renderer; Color is the portable common denominator. One more rule worth knowing: stock effects (BasicEffect and friends) write attachment 0 only, so writing several outputs needs a custom ShaderEffect like the ones below.

G-Buffer Pattern

The standard G-buffer for a PBR (Physically Based Rendering) deferred pipeline stores:

  • Albedo (RT0) — the base colour of the surface, RGB. The alpha channel can store roughness or AO factor to save a render target slot.
  • Normal (RT1) — the world-space surface normal, XYZ. Normals are signed values in [−1, +1], so they must be remapped to [0, 1] on write and back on read.
  • Depth (RT2) — the raw clip-space Z value from gl_FragCoord.z, packed across all four RGBA8 channels. Used by the lighting pass to reconstruct world-space position from the depth and the inverse view-projection matrix.

The geometry pass renders every opaque mesh once. Alpha-tested surfaces can also be included in the geometry pass. Transparent surfaces must be rendered in a separate forward pass after the deferred lighting pass, since deferred renderers do not handle multiple transparent layers at a given pixel.

Reading Multiple Outputs in the Fragment Shader

In GLSL, declare one out vec4 variable per render target. The layout(location = N) qualifier maps each variable to the corresponding MRT attachment index (0-based):

#version 300 es
precision highp float;

in vec3 v_worldPos;
in vec3 v_normal;
in vec2 v_texcoord;

uniform sampler2D u_albedoTex;
uniform float     u_roughness;
uniform float     u_metalness;

// MRT outputs — one per render target
layout(location = 0) out vec4 gAlbedo;       // RT0
layout(location = 1) out vec4 gNormal;       // RT1
layout(location = 2) out vec4 gLinearDepth;  // RT2

// Every target is RGBA8, so depth is spread across all four channels.
vec4 packDepth(float depth) {
    const vec4 bitShift = vec4(1.0, 255.0, 65025.0, 16581375.0);
    const vec4 bitMask  = vec4(1.0 / 255.0, 1.0 / 255.0, 1.0 / 255.0, 0.0);
    vec4 res = fract(depth * bitShift);
    res -= res.gbaa * bitMask;
    return res;
}

void main() {
    // Sample surface albedo from texture
    vec4 albedo = texture(u_albedoTex, v_texcoord);

    // RT0: albedo in RGB, roughness in A
    gAlbedo = vec4(albedo.rgb, u_roughness);

    // RT1: encode normal from [-1,+1] to [0,1] range; metalness in A
    vec3 n = normalize(v_normal);
    gNormal = vec4(n * 0.5 + 0.5, u_metalness);

    // RT2: clip-space depth, packed into RGBA8 for position reconstruction
    gLinearDepth = packDepth(gl_FragCoord.z);
}

Writing vec4(gl_FragCoord.z, 0.0, 0.0, 1.0) instead would quantise depth to the 256 levels an 8-bit red channel can hold, which is far too coarse to reconstruct world position from. The packing above is what makes an RGBA8 depth slot usable.

Complete Example: Deferred G-Buffer Setup

⚠

Both shaders in this example are ShaderEffects. This snapshot’s separate XNA/FNA Effect Framework bytecode path is renderer-qualified: it works on FNA3D always and on ten more identities only when the matching default-OFF CNA_*_COMPILED_EFFECTS option was enabled at configure time (11 of the 14 identities in all). Here ShaderEffect takes renderer-native source and has no Parameters collection; uniforms go through SetUniformXxx() and textures through SetTexture(unit, tex), with Apply() called first. The GLSL ES 3.00 shown runs as written on OPENGLES3 and WEBGL2; OPENGL33 needs #version 330 core, VULKAN needs SPIR-V, WEBGPU needs WGSL and DIRECTX11 needs HLSL (with SV_Target0..2 outputs); METAL’s custom effects run only in SpriteBatch, so on Metal a G-buffer pass needs a compiled XNA effect (CNA_METAL_COMPILED_EFFECTS=ON). See Tutorial 52 and Tutorial 128.

#include "Microsoft/Xna/Framework/Graphics/ShaderEffect.hpp"
#include "System/IO/File.hpp"

// A minimal camera helper used by the examples in this series. It is application
// code, not a CNA type; Tutorial 34 builds a fuller FpsCamera.
struct Camera {
    Vector3 position = Vector3(0.0f, 4.0f, 10.0f);
    Matrix view = Matrix::CreateLookAt(position, Vector3::Zero, Vector3::Up);
    Matrix projection = Matrix::CreatePerspectiveFieldOfView(
        MathHelper::ToRadians(60.0f), 16.0f / 9.0f, 0.1f, 200.0f);
    const Vector3& Position() const   { return position; }
    const Matrix&  View() const       { return view; }
    const Matrix&  Projection() const { return projection; }
};

class DeferredGame final : public Game {
    GraphicsDeviceManager           graphics_;
    Camera                          camera_;
    std::unique_ptr<RenderTarget2D> rtAlbedo_;   // RT0: albedo + roughness
    std::unique_ptr<RenderTarget2D> rtNormal_;   // RT1: normal + metalness
    std::unique_ptr<RenderTarget2D> rtDepth_;    // RT2: linear depth
    std::unique_ptr<ShaderEffect>   geoEffect_;  // G-buffer fill shader
    std::unique_ptr<ShaderEffect>   lightEffect_;// Deferred lighting shader
    std::unique_ptr<VertexBuffer>   fullscreenVB_;

    // ShaderEffect's matrix setter takes raw column-major floats.
    static void SetMat4(ShaderEffect& fx, const char* name, const Matrix& m) {
        float cm[16];
        m.ToColumnMajor(cm);
        fx.SetUniformMat4(name, cm);
    }

public:
    DeferredGame() : graphics_(this) {
        // Three simultaneous render targets need HiDef (Reach allows one).
        graphics_.setGraphicsProfileProperty(GraphicsProfile::HiDef);
    }

protected:
    void LoadContent() override {
        auto& gd = getGraphicsDeviceProperty();
        // HiDef is necessary but not sufficient: STUB and the 2D-only
        // SDL_RENDERER have no MRT at all.
        if (!gd.SupportsCapability(CNA::GraphicsCapability::MultipleRenderTargets)) {
            // Fall back to a forward-rendered path here.
        }
        int w = gd.getPresentationParametersProperty().getBackBufferWidthProperty();
        int h = gd.getPresentationParametersProperty().getBackBufferHeightProperty();

        // Create G-buffer render targets.
        // Only RT0 needs a depth buffer (shared across the geometry pass).
        rtAlbedo_ = std::make_unique<RenderTarget2D>(
            gd, w, h, false,
            SurfaceFormat::Color,
            DepthFormat::Depth24Stencil8);

        // Every target is SurfaceFormat::Color for portability (HiDef also allows
        // float targets on renderers that classify them; see the note above).
        // Normals are encoded to [0,1]; depth is packed to RGBA8.
        rtNormal_ = std::make_unique<RenderTarget2D>(
            gd, w, h, false,
            SurfaceFormat::Color,
            DepthFormat::None);

        rtDepth_ = std::make_unique<RenderTarget2D>(
            gd, w, h, false,
            SurfaceFormat::Color,
            DepthFormat::None);

        // Three arguments: device, vertex source, fragment source. Not file paths.
        geoEffect_ = std::make_unique<ShaderEffect>(
            gd,
            System::IO::File::ReadAllText("Content/effects/gbuffer_fill.vert.glsl"),
            System::IO::File::ReadAllText("Content/effects/gbuffer_fill.frag.glsl"));
        lightEffect_ = std::make_unique<ShaderEffect>(
            gd,
            System::IO::File::ReadAllText("Content/effects/deferred_light.vert.glsl"),
            System::IO::File::ReadAllText("Content/effects/deferred_light.frag.glsl"));

        // Neither constructor throws on a compile failure -- check both.
        // GetCompileErrorEXT() returns the compiler log (also written to stderr).
        if (!geoEffect_->IsEffectValid() || !lightEffect_->IsEffectValid()) {
            // The shader did not compile. Do not draw with it.
        }

        // Sampler unit assignments for the lighting pass, set once.
        lightEffect_->Apply();
        lightEffect_->SetUniformInt("u_albedo", 0);
        lightEffect_->SetUniformInt("u_normal", 1);
        lightEffect_->SetUniformInt("u_depth",  2);

        buildFullscreenTriangle(gd, fullscreenVB_);
    }

    void Draw(const GameTime&) override {
        auto& gd = getGraphicsDeviceProperty();

        // ---- Geometry pass: fill the G-buffer -------------------------
        // RenderTargetBinding(RenderTarget2D*) is implicit; naming the type at each
        // element just reads clearly. SetRenderTargets takes a vector, and throws
        // on Reach (one target) or above four.
        const std::vector<RenderTargetBinding> targets = {
            RenderTargetBinding(rtAlbedo_.get()),
            RenderTargetBinding(rtNormal_.get()),
            RenderTargetBinding(rtDepth_.get()),
        };
        gd.SetRenderTargets(targets);
        gd.Clear(Color::Black);

        // Apply() first is the portable order for a ShaderEffect.
        geoEffect_->Apply();
        SetMat4(*geoEffect_, "u_view",       camera_.View());
        SetMat4(*geoEffect_, "u_projection", camera_.Projection());
        drawSceneGeometry(gd, *geoEffect_);

        // ---- Lighting pass: deferred shading on fullscreen triangle ---
        gd.SetRenderTarget(nullptr);     // restore back buffer
        gd.Clear(Color::Black);

        const Vector3 eye      = camera_.Position();
        const Vector3 lightDir = Vector3::Normalize(Vector3(-1.0f, -2.0f, -1.0f));

        lightEffect_->Apply();

        // Bind G-buffer textures on the units assigned in LoadContent().
        // SetTexture takes a reference, and RenderTarget2D derives from Texture2D.
        lightEffect_->SetTexture(0, *rtAlbedo_);
        lightEffect_->SetTexture(1, *rtNormal_);
        lightEffect_->SetTexture(2, *rtDepth_);

        // Camera parameters needed to reconstruct world positions from depth
        SetMat4(*lightEffect_, "u_invViewProj",
                Matrix::Invert(camera_.View() * camera_.Projection()));
        lightEffect_->SetUniformVec3("u_cameraPos", eye.X, eye.Y, eye.Z);

        // Submit light data (example: one directional light)
        lightEffect_->SetUniformVec3("u_lightDir",
                                      lightDir.X, lightDir.Y, lightDir.Z);
        lightEffect_->SetUniformVec3("u_lightColor", 1.0f, 0.95f, 0.85f);

        gd.SetVertexBuffer(fullscreenVB_.get());
        gd.DrawPrimitives(PrimitiveType::TriangleList, 0, 1);
        // No gd.Present(): Game presents after Draw() returns, so a manual call
        // would present the frame twice.
    }
};

Deferred Lighting Pass Shader

The lighting pass samples all three G-buffer textures and applies a simple Lambertian + specular lighting model. In a production deferred renderer you would loop over all lights (with each light contributing one additive pass, or using a clustered approach).

#version 300 es
precision highp float;

in vec2 v_texcoord;

uniform sampler2D u_albedo;
uniform sampler2D u_normal;
uniform sampler2D u_depth;

uniform mat4  u_invViewProj;
uniform vec3  u_cameraPos;
uniform vec3  u_lightDir;   // normalised, pointing toward light
uniform vec3  u_lightColor;

out vec4 fragColor;

// Inverse of the geometry pass's packDepth().
float unpackDepth(vec4 rgba) {
    const vec4 bitShift = vec4(1.0,
                               1.0 / 255.0,
                               1.0 / 65025.0,
                               1.0 / 16581375.0);
    return dot(rgba, bitShift);
}

// Reconstruct world-space position from depth texture and NDC coords.
vec3 reconstructWorldPos(vec2 uv, float depth) {
    vec4 ndc = vec4(uv * 2.0 - 1.0, depth * 2.0 - 1.0, 1.0);
    vec4 world = u_invViewProj * ndc;
    return world.xyz / world.w;
}

void main() {
    // Sample G-buffer
    vec4  albedoRough = texture(u_albedo, v_texcoord);
    vec4  normalMetal = texture(u_normal, v_texcoord);
    float rawDepth    = unpackDepth(texture(u_depth, v_texcoord));

    vec3 albedo    = albedoRough.rgb;
    float roughness = albedoRough.a;
    // Decode normal from [0,1] back to [-1,+1]
    vec3 N = normalize(normalMetal.xyz * 2.0 - 1.0);
    float metalness = normalMetal.a;

    // Reconstruct world position
    vec3 worldPos = reconstructWorldPos(v_texcoord, rawDepth);
    vec3 V = normalize(u_cameraPos - worldPos);
    vec3 L = normalize(-u_lightDir);
    vec3 H = normalize(V + L);

    // Simple Lambertian diffuse
    float NdotL = max(dot(N, L), 0.0);
    vec3 diffuse = albedo * u_lightColor * NdotL;

    // Blinn-Phong specular (approximation; replace with GGX for PBR)
    float shininess = mix(4.0, 128.0, 1.0 - roughness);
    float NdotH = max(dot(N, H), 0.0);
    vec3 specular = u_lightColor * pow(NdotH, shininess) * (1.0 - roughness);

    // Ambient
    vec3 ambient = albedo * 0.05;

    fragColor = vec4(ambient + diffuse + specular, 1.0);
}

Combining Additional Passes

A full deferred pipeline stacks additional passes after the geometry and lighting passes:

  • SSAO pass — reads the depth and normal G-buffer targets and writes an ambient occlusion factor to another render target.
  • Multiple light accumulation — one additive fullscreen pass per dynamic point or spot light using scissor rectangles to limit cost.
  • Forward transparency pass — transparent objects rendered with standard forward blending on top of the deferred result.
  • Post-processing chain — bloom, colour grading, FXAA, motion blur applied to the final composited image.

Each of these passes can read the G-buffer textures as needed. Because they are ordinary RenderTarget2D objects on the CPU side, they can be passed as Effect parameters like any other texture.