Audio System

Microsoft::Xna::Framework::Audio — SoundEffect, MediaPlayer, XACT classes, microphone input, and the three CNA_AUDIO_PLATFORM implementations

ⓘ

Implementation status: SoundEffect, SoundEffectInstance, DynamicSoundEffectInstance, MediaPlayer, and Song have real playback paths whenever the audio implementation provides a mixer — that is SDL3 (the default) and, in this snapshot, the new native Linux ALSA implementation. XACT is a real parser and player, not a facade: CNA parses XGS/XWB/XSB data, applies FACT's volume formula and RPC curves, and implements the named boundaries below. 3D audio uses FAudio's Doppler calculation and F3DAudio attenuation. This snapshot has 38 audio test source files with 765 static GoogleTest-family definitions (the site's static count method); the executed subset depends on the chosen build. Honest partials are listed in Tier 2.

Overview

CNA's audio system is split into two tiers. Tier 1 covers playback of sound effects and music. Its device selection is independent through CNA_AUDIO_PLATFORM, but the four values are not feature-equivalent: SDL3 (the default) and ALSA define SOUND_ENABLED and provide a mixer — SDL3_mixer for the first, CNA's own mixer for the second. NULL provides a low-level IAudioDevice implementation and selection/conformance coverage, while the mixer-dependent XNA playback facade is compiled without its real engine. Tier 2 covers XACT — cue-based audio driven by .xgs/.xsb/.xwb content — and real playback likewise requires a mixer-backed choice (SDL3 or ALSA).

It is worth being precise about what "real XACT" means here, because most XNA reimplementations stop at the API shape. CNA ships a real parser for the three XACT binary formats, and the playback side reproduces FACT's actual volume formula and evaluates real RPC (runtime parameter control) curves rather than approximating them. The 3D positional path uses FAudio's exact Doppler calculation and F3DAudio's attenuation model. The honest partials below are the places where that fidelity currently stops.

Both mixers use the same track-based model behind one internal facade (MixerEngine.hpp): each SoundEffectInstance and each Song played through MediaPlayer occupies its own mixer track, rather than writing to a global shared output. SDL3_mixer is vendored as a Git submodule and built alongside CNA; it uses the MIX_Track model natively, which avoids the global-state pitfalls of older SDL_mixer versions and maps cleanly onto the per-instance XNA API. The ALSA implementation carries its own mixer (CnaMixer) that follows the same semantics: per-track resampling, gain, a per-track mix callback for filter and pan, then a master mix and a post-mix callback.

All audio classes live in the Microsoft::Xna::Framework::Audio namespace, with MediaPlayer and Song in Microsoft::Xna::Framework::Media, matching XNA 4.0 exactly. The examples on this page use the real property-style call names (setVolumeProperty, getStateProperty and so on) and were checked against this snapshot's headers.

Audio implementations (CNA_AUDIO_PLATFORM)

The audio device layer is a separate CMake axis from the platform (CNA_PLATFORM) and the renderer. Three values are implemented, SDL3 is the default on every target, and OPENAL and WASAPI are reserved names that fail configuration instead of falling back. Values are case-sensitive.

AspectSDL3 (default)NULLALSA
StatusImplemented, defaultImplemented; deterministic silent transportImplemented, native, no SDL
TargetsEvery target CNA supports (Linux, Windows, macOS, iOS, Android, Emscripten)AllLinux only (configure fails on any other target)
Playback deviceSdl3AudioDeviceNullAudioDevice (paced thread that discards samples)AlsaAudioDevice: libasound.so.2 loaded at run time, roughly 10 ms periods, four buffered
Defines SOUND_ENABLEDYesNoYes
Mixer engineSDL3_mixer (MIX_Track model)noneCNA's own CnaMixer (linear-interpolation resampler, same track and callback semantics)
SoundEffect construction and DurationYesYes for effects built from a raw PCM buffer (including .xnb and .cnb content); an effect built from a path or FromStream carries no audio and reports a zero DurationYes
SoundEffect.Play() / SoundEffectInstance (volume, pitch, pan, loop, pause)YesPlay() returns false; no mixingYes
DynamicSoundEffectInstanceYesSameYes
MediaPlayer / SongYesSameYes (a Song is decoded as it plays)
XACT AudioEngine, SoundBank, WaveBank, CueParse and playSameParse and play
3D audio (Apply3D)Applied to the mixer trackSameApplied to the CnaMixer track
Capture (Microphone)Yes (SDL3 recording devices)NoYes (ALSA capture)
Decoders on the play pathSDL3_mixer's decoders; the optional external Opus, WavPack, mpg123, libvorbisfile, libFLAC, FluidSynth, GME and xmp codec libraries are switched off at build; no AACnoneWAV (PCM 8/16/24/32-bit, IEEE float, MS-ADPCM, IMA-ADPCM), Ogg Vorbis (stb_vorbis), MP3 (dr_mp3), FLAC (dr_flac); no Opus, WMA/xWMA, XMA or AAC
Extra link inputsSDL3, SDL3_mixernoneALSA headers at build time (libasound2-dev), dl; vendored stb and dr_libs
Environment knobsSDL_AUDIODRIVER (SDL's own)noneCNA_AUDIO_DEVICE, CNA_AUDIO_RECORDING_DEVICE
Automatic CI evidencePlatform-matrix cells (with SDL_AUDIODRIVER=dummy) and the Apple lanesThe Headless and Terminal cells of platform-ci.yml configure with itThe SDL-free Headless + ALSA job of platform-ci.yml: WavDecoderTest, CnaMixer, CnaMixerXna, AlsaAudioDevice, AlsaAudioRecordingDevice, Microphone, AudioCategory, AudioEngine and the device-conformance suites on ALSA's null device

A WAV decoder that does not depend on any of the four (DecodeWavToPcm16) is compiled into every build, so the content tools and content tests link without SDL. It is not a playback path under NULL.

Choosing and configuring

# Default: SDL3 device + SDL3_mixer (every desktop, mobile and web target)
cmake -S . -B build -DCNA_AUDIO_PLATFORM=SDL3

# Native Linux audio with CNA's own mixer (needs libasound2-dev); the window still comes from SDL3
cmake -S . -B build -DCNA_AUDIO_PLATFORM=ALSA -DCNA_GRAPHICS_RENDERER=OPENGLES3

# No SDL anywhere: a windowless HEADLESS (or TERMINAL) build with ALSA audio
cmake -S . -B build-nosdl \
  -DCNA_ENABLE_SDL=OFF -DCNA_PLATFORM=HEADLESS \
  -DCNA_AUDIO_PLATFORM=ALSA -DCNA_GRAPHICS_RENDERER=HEADLESS

# Deterministic silent build with no mixer: device-contract and logic work only
cmake -S . -B build-headless \
  -DCNA_PLATFORM=HEADLESS -DCNA_AUDIO_PLATFORM=NULL -DCNA_GRAPHICS_RENDERER=HEADLESS
  • Normal games: keep SDL3. It is the only choice that provides a mixer on macOS, iOS, Android and Emscripten.
  • Linux without SDL audio: ALSA. It gives you the full XNA audio facade — effects, songs, dynamic instances, XACT and capture — without linking SDL or libasound (the library is opened at run time; a machine without it still starts the program, and playback then reports NoAudioHardwareException).
  • Tests and CI that must be silent: either the default SDL3 with SDL_AUDIODRIVER=dummy, or ALSA with CNA_AUDIO_DEVICE=null. Both keep a functioning mixer. NULL is the choice only when you do not need any mixer behaviour at all.
  • Windows and macOS without SDL: there is no SDL-free audio device there. CNA_ENABLE_SDL=OFF refuses SDL3 audio, so an SDL-free Windows build (a windowless HEADLESS build) has NULL audio only. OPENAL and WASAPI are reserved and rejected.
⚠

Audio selection is ahead of audio parity. In modules/CMakeLists.txt, SOUND_ENABLED is defined only for CNA_AUDIO_PLATFORM=SDL3 and CNA_AUDIO_PLATFORM=ALSA. NULL builds omit the mixer engine, its decoders and the mixer-dependent tests; the XNA audio facade remains present (SoundEffect, SoundEffectInstance, DynamicSoundEffectInstance, MediaPlayer and Microphone compile) but does not gain production playback from those device classes. NULL is therefore useful for deterministic configuration and device-contract work, not proof that a silent mixer consumes normal game audio.

Environment variables (ALSA)

VariableMeaning
CNA_AUDIO_DEVICEThe ALSA PCM to play on. Unset (or empty) means default. Any ALSA PCM name works: hw:0,0, null for a silent device that still paces itself in real time, or file:FILE=out.raw,FORMAT=raw to record exactly what was played.
CNA_AUDIO_RECORDING_DEVICEThe ALSA PCM to capture from. When set, exactly that PCM is offered as the default microphone (null keeps tests from recording a room). Otherwise CNA lists ALSA's default capture PCM first, then each sound card's capture devices opened through plughw so the requested format is converted.

There is no CNA_AUDIO_PLATFORM environment variable: the implementation is fixed at configure time. On a PipeWire desktop ALSA's default device is PipeWire (through pipewire-alsa), on a PulseAudio desktop it is PulseAudio (through the ALSA plugins), and on a bare system it is the sound card through dmix, so one backend reaches all three provided the distribution installs the corresponding ALSA plugin. The step-by-step guide is Tutorial 137: Native Linux Audio with ALSA.

Tier 1 — Implemented playback (SDL3_mixer or CNA's own mixer)

SoundEffect

SoundEffect represents a loaded audio asset — typically a short sound effect stored as WAV. It is an immutable, move-only handle to decoded audio data. To play a sound once, call Play() directly. To control playback parameters, call CreateInstance() and manipulate the resulting SoundEffectInstance.

MemberDescription
Play()Fire-and-forget one-shot playback at default volume, pitch, and pan; returns false if the sound could not start (the effect is disposed, there is no native audio as under NULL, or the mixer track could not be created, bound or played); Play() itself enforces no instance limit
Play(volume, pitch, pan)One-shot playback with explicit volume (0.0–1.0), pitch (−1.0–1.0), and pan (−1.0–1.0)
CreateInstance()Returns a SoundEffectInstance by value for controlled playback
static FromStream(stream)Loads a SoundEffect from an open stream (WAV data); returns a heap pointer the caller owns
getDurationProperty()Total length of the audio clip as a TimeSpan (reported on every audio implementation for effects built from a raw PCM buffer, including .xnb and .cnb content; under NULL an effect built from a path or FromStream carries no audio and reports zero)
getNameProperty()Asset name set by ContentManager at load time

The preferred way to load a SoundEffect is through ContentManager, but FromStream() is provided for loading from arbitrary sources such as network streams or embedded resources.

// One-shot play — simplest usage (inside a Game subclass)
auto boom = getContentProperty().Load<SoundEffect>("audio/explosion");
boom.Play();

SoundEffectInstance

SoundEffectInstance is a controllable playback handle obtained from SoundEffect::CreateInstance(). Unlike a fire-and-forget Play() call, an instance lets you adjust volume, pitch, and pan at any time, pause and resume playback, and query the current playback state. Each instance holds its own mixer track.

MemberType / RangeDescription
Volumefloat — 0.0–1.0Playback amplitude; 1.0 is full volume
Pitchfloat — −1.0–1.0Pitch shift in octaves; 0.0 is unmodified
Panfloat — −1.0–1.0Stereo pan; −1.0 is full left, 1.0 is full right
IsLoopedboolWhen true the clip loops until Stop() is called; cannot be changed once playback has started
StateSoundStateCurrent state: Playing, Paused, or Stopped
Play()—Starts or resumes playback
Pause()—Pauses playback, retaining position
Resume()—Resumes from the paused position
Stop()—Stops playback and rewinds to the beginning
Apply3D(...)—Positions the instance in 3D; see 3D positional audio

Every property is a getXProperty()/setXProperty() pair: for example setVolumeProperty(0.7f) and getStateProperty().

// SoundEffectInstance with volume, pitch, and pan control
auto sfx = getContentProperty().Load<SoundEffect>("audio/laser");
auto instance = sfx.CreateInstance();     // returned by value

instance.setVolumeProperty(0.7f);
instance.setPitchProperty(-0.3f);   // slightly lower pitch
instance.setPanProperty(0.5f);      // panned right
instance.setIsLoopedProperty(false);
instance.Play();

// Later, in response to a game event:
if (instance.getStateProperty() == SoundState::Playing) {
    instance.Pause();
}

DynamicSoundEffectInstance

DynamicSoundEffectInstance allows streaming audio data from application-managed buffers rather than from a preloaded file. The engine raises the BufferNeeded event whenever its internal buffer queue runs low, signalling the application to call SubmitBuffer() with the next chunk of PCM data. This is suitable for procedurally generated audio, network audio streams, or decoded-on-the-fly music. Tutorial 118 walks through it end to end.

MemberDescription
DynamicSoundEffectInstance(sampleRate, channels)Constructs an instance with the given sample rate (Hz) and AudioChannels (Mono or Stereo); both are fixed for the instance's life
SubmitBuffer(bytes)Enqueues a block of headerless, little-endian, signed 16-bit PCM samples (a std::vector<SharpRuntime::bytecs>, channels interleaved)
BufferNeededEvent raised when fewer than three buffers are pending; subscribe with +=
Play(), Pause(), Stop()Playback control, same semantics as SoundEffectInstance
getPendingBufferCountProperty()Number of buffers currently queued but not yet consumed
// DynamicSoundEffectInstance with a BufferNeeded callback (inside a Game subclass)
dynSfx_ = std::make_unique<DynamicSoundEffectInstance>(44100, AudioChannels::Stereo);

dynSfx_->BufferNeeded += [this](System::Object* sender, const System::EventArgs&) {
    // Generate or decode the next chunk of PCM data
    auto* dyn = static_cast<DynamicSoundEffectInstance*>(sender);
    dyn->SubmitBuffer(GenerateNextAudioChunk());   // std::vector<SharpRuntime::bytecs>
};

// Submit an initial buffer before calling Play()
dynSfx_->SubmitBuffer(GenerateNextAudioChunk());
dynSfx_->Play();

Song

Song represents a music track loaded through the ContentManager. Unlike SoundEffect, a Song is played exclusively through the MediaPlayer static class; only one song plays at a time. Neither mixer pre-decodes a song: SDL3_mixer decodes the audio file progressively, and CNA's own mixer keeps the compressed bytes in memory and decodes them as they play — either way a Song is suitable for large music files that would be impractical to decode entirely into memory.

MemberDescription
getDurationProperty()Total length of the track as a TimeSpan
getNameProperty()Asset name set by ContentManager at load time

MediaPlayer

MediaPlayer is a static class that manages playback of a single Song at a time. It mirrors the XNA 4.0 Microsoft.Xna.Framework.Media.MediaPlayer API. Volume, muting, and looping are controlled through static properties, and the current playback state is available via getStateProperty(). Play() takes a Song* (or a SongCollection).

MemberType / RangeDescription
Play(song)static voidStarts playing the given Song*, stopping any currently playing track
Pause()static voidPauses the currently playing song
Resume()static voidResumes from the paused position
Stop()static voidStops playback and rewinds
Volumefloat — 0.0–1.0Music volume; independent of SoundEffect volume
IsMutedboolSilences output without altering Volume
IsRepeatingboolWhen true the song loops automatically when it ends
StateMediaStateCurrent state: Playing, Paused, or Stopped
// MediaPlayer playing a Song (theme_ is a std::optional<Song> member)
theme_ = getContentProperty().Load<Song>("music/main_theme");

MediaPlayer::setIsRepeatingProperty(true);
MediaPlayer::setVolumeProperty(0.8f);
MediaPlayer::Play(&*theme_);

// Mute on focus loss, restore on focus gain
void OnFocusLost()  { MediaPlayer::setIsMutedProperty(true);  }
void OnFocusGained(){ MediaPlayer::setIsMutedProperty(false); }

Loading audio via ContentManager

Both SoundEffect and Song are loaded through the standard ContentManager pipeline. Place audio files in your content directory; the path passed to Load<T>() is relative to the manager's root directory (getRootDirectoryProperty()) and may omit the file extension. Load<T> returns the asset by value. When the manager tries extensions itself, it looks for a compiled .xnb and then a .cnb first; a SoundEffect then falls back to .wav, and a Song to .mp3, .ogg, .wav, .flac, .opus, .aac and .wma in that order — but a name resolving does not mean the file decodes: the audio implementation decides what it can play (see Decoders and file formats).

// Loading audio assets via ContentManager (inside a Game subclass)
auto& content = getContentProperty();          // root directory "Content" by default

// Load a short sound effect (WAV)
auto jumpSfx = content.Load<SoundEffect>("audio/jump");

// Load a music track (OGG recommended for large files)
auto bgMusic = content.Load<Song>("music/level1");

// Alternatively, load a SoundEffect from a raw stream
std::ifstream file("assets/custom.wav", std::ios::binary);
std::unique_ptr<SoundEffect> custom(SoundEffect::FromStream(file));

A standalone ContentManager (outside a Game) is constructed with a service provider and a root directory: ContentManager(System::IServiceProvider*, const std::string& rootDirectory). A SoundEffect loaded through the manager is never cached — each Load call returns its own independently owned instance, because the type is move-only.

Decoders and file formats

What can be played is decided by the audio implementation, not by the file extension. CNA's own mixer identifies a file by its content (RIFF/WAVE, OggS, fLaC, an ID3 tag or an MPEG sync word) and refuses anything else with an error naming the formats it does play.

FormatSDL3 (SDL3_mixer)ALSA (CnaMixer)NULL
WAV: PCM, floatYesYes (8/16/24/32-bit PCM and IEEE float)No playback
WAV: MS-ADPCM, IMA-ADPCMYes, through the mixer facade (XACT wave banks take this route; the XNB SoundEffect reader does not: it decodes to 16-bit PCM in the content library with CNA's own WAV decoder, in every build)Yes, decoded by CNA (the build-time content pipeline's AudioContent::ConvertFormat can also encode MS-ADPCM)No playback
Ogg VorbisYesYes (stb_vorbis 1.22)No playback
MP3YesYes (dr_mp3 0.7.3)No playback
FLACYesYes (dr_flac 0.13.3, native or in Ogg)No playback
Opus, WavPack, MIDI, tracker modulesNo for Opus, WavPack and tracker modules (codec libraries switched off at build). MIDI is not switched off: CNA disables only FluidSynth and leaves the pinned SDL_mixer’s built-in Timidity decoder at its default (on), so it is probably compiled in (read from CMake, not built), but it needs a Timidity patch set on the host, which CNA does not ship, and no CNA test plays MIDINo; Ogg Opus is refused by nameNo playback
AAC / .m4a, WMA/xWMA, XMANoNoNo playback

XACT wave banks add their own codecs: 8-bit and 16-bit PCM and MS-ADPCM decode, XMA and WMA do not (see Honest partials). An .xnb SoundEffect is decoded to 16-bit PCM by the content library itself (16-bit PCM directly, 8-bit PCM widened, 32-bit float, MS-ADPCM and IMA-ADPCM through CNA's own WAV decoder, which is compiled into every build) and built through the raw-PCM constructor, so it does not depend on the selected mixer; XMA2 is rejected with a ContentLoadException that names the format. MediaLibrary track durations come from an FFmpeg-based probe: in a build without FFmpeg they report 0 (unknown). See Video Playback for the optional FFmpeg layer.

3D positional audio

3D audio uses the same three pieces as XNA: an AudioListener (the ear), an AudioEmitter (the source) and SoundEffectInstance::Apply3D. Doppler is computed exactly (FAudio's formula); distance and stereo placement are approximated by one attenuation value and one pan value because the mixers have a single stereo gain pair rather than per-listener output matrices. Tutorial 119 covers the attenuation curve, Doppler and the globals in depth.

OverloadBehaviour in this snapshot
Apply3D(const AudioListener&, const AudioEmitter&)One listener.
Apply3D(const AudioListener*, int listenerCount, const AudioEmitter&)Accepts any positive count. Every listener is evaluated and the dominant one — the nearest to the emitter, i.e. the one that hears it loudest — decides attenuation, pan and Doppler. A null array throws ArgumentNullException; a zero or negative count throws ArgumentOutOfRangeException.
Apply3D(const std::vector<AudioListener>&, const AudioEmitter&)New in this snapshot. The same dominant-listener rule; an empty vector throws ArgumentOutOfRangeException.
Cue::Apply3D(const AudioListener&, const AudioEmitter&)Single listener (XNA's own signature); forwards to every wave the cue is playing.

The old “exactly one listener” limit is gone: multi-listener calls no longer throw NotSupportedException. What CNA cannot reproduce is XACT's own multi-listener DSP, so the dominant-listener rule is an approximation — it is deliberately not “use listeners[0]”, and moving a second, closer listener changes the result. Two more behaviours follow XNA: the effective Doppler scale is the emitter's DopplerScale multiplied by the global SoundEffect one (a stationary sound is not pitch-shifted), and an instance's mode is fixed while it plays — before its first Play() you can still switch between Apply3D and setPanProperty, but on a playing instance calling Apply3D on a pan-mode instance, or setPanProperty on a 3D one, throws InvalidOperationException. Stop() makes the instance re-aimable.

Tier 2 — XACT (AudioEngine/SoundBank/WaveBank/Cue)

ⓘ

XACT is a real parser and player, not a facade. AudioEngine, SoundBank, and WaveBank genuinely parse the XGS/XWB/XSB binary formats, and Cue plays the result back through the selected mixer (SDL3_mixer or CNA's own CnaMixer) using FACT's actual volume formula and real RPC curves. 3D cues are positioned with FAudio's exact Doppler calculation and F3DAudio attenuation. XACT files are read through the platform's title-content reader, so packaged assets and ordinary desktop paths share one loader.

Class Header Status Notes
AudioEngine Present Implemented Real XGS parsing; drives RPC curve evaluation
SoundBank Present Implemented Real XSB parsing, PlayCue, GetCue
WaveBank Present Implemented Real XWB parsing. PCM and MS-ADPCM entries decode; XMA and WMA entries log to stderr and produce no sound.
Cue Present Implemented Real playback through the selected mixer with FACT's volume formula. Only the first PlayWave per track is honored.

XACT-driven audio works directly in CNA — you do not need to migrate AudioEngine/SoundBank/Cue calls away. It needs a mixer-backed audio implementation: with NULL the banks parse but nothing is audible. For new CNA-only projects without existing XACT content, SoundEffect and SoundEffectInstance remain the simpler Tier 1 API and integrate directly with ContentManager. Tutorial 120 is the end-to-end guide.

Honest partials

Three limitations are worth knowing before you point a real XACT project at CNA:

  • Only the first PlayWave per track is honored. Cues authored with multiple waves on a single track will not play them all.
  • XMA and WMA wave-bank entries are silent, not fatal. A wave-bank entry in either codec does not throw: CNA writes a diagnostic naming the bank, wave index and format to stderr and returns no sound, so the cue simply makes no sound. Re-encode to PCM or MS-ADPCM. (A .xnb SoundEffect in XMA2 is different: the XNB reader rejects it with a ContentLoadException.)
  • The reverb send is a no-op. The API accepts it, but no reverb is applied to the signal.

Microphone

Microphone is implemented on the two mixer-backed implementations: real SDL3 capture devices under SDL3 and ALSA capture under ALSA. Device enumeration (Microphone::getAllProperty() lists the machine's devices, getDefaultProperty() returns nullptr when there are none), Start()/Stop(), GetData() and BufferReady delivery all work against genuine hardware input. Capture is mono, signed 16-bit at a requested 44,100 Hz, and BufferDuration accepts 100–1000 ms in steps of 10 ms. Under NULL there is no recording provider, so the microphone list is empty. With ALSA, CNA_AUDIO_RECORDING_DEVICE chooses the PCM and the capture thread queues up to eight seconds so a game polling once a frame loses nothing between polls.

Tests and CI evidence

The audio module's own test executable is CnaAudioTests. The ALSA path is exercised by the SDL-free Headless + ALSA job of platform-ci.yml, which builds a CNA_ENABLE_SDL=OFF configuration with CNA_AUDIO_PLATFORM=ALSA and runs the WavDecoderTest, CnaMixer, CnaMixerXna, AlsaAudioDevice, AlsaAudioRecordingDevice, Microphone, AudioCategory and AudioEngine suites on ALSA's null device; it also asserts that the built executables do not link SDL and do not link libasound. That is evidence about the mixer and the device layer on a machine with no sound card, not a listening test on real hardware. Tests that read SDL3_mixer track handles remain SDL3-only, and mixer-dependent test sources are excluded from NULL builds. See What CI actually covers and Verification & Known Issues.