
Soundscapes for Web Games: The Layer Beneath Your SFX, Part 1
· Last updated
For years I shipped web games that went dead the moment you stood still. The samples were fine and the music was fine. Something was still missing, and it took me an embarrassingly long time to hear what it was.
You’ve heard it on a video call. The other person is talking and the line sounds normal. Then they mute, and their whole room disappears with them: no air conditioning, no chair creak, no car going past the window. The line isn’t quiet. It’s empty.
Most web games mute the room on purpose. They play a footstep, a click and a music loop, and they leave the space between those events silent. The fix is to leave the room on, and I call that layer a soundscape. Flip between the two versions below and you’ll hear the difference straight away.
Layered comparison
Music / Bed / Emitters / SFX
Music
Composed loop, market square cafe
Bed
Restaurant room tone, looping under everything
Emitters
Cash register, espresso, cups, phone, on randomised timers
SFX
Click the strip below to fire a UI sound
What is a soundscape?
It’s the third kind of game audio, and it isn’t music or SFX. If you mix all three on one bus, the game ends up sounding like a YouTube video with the ads left in.
| Layer | What it is | Triggered by | Wants |
|---|---|---|---|
| Music | A composed, looping track | Scene or beat | Its own bus, fade in/out, ducks rarely |
| SFX | One-shot reactions to gameplay | Player or game | Tight latency, no fades |
| Soundscape | Ambient bed plus randomized non-musical emitters | A timer, not you | A bus that ducks under dialog |
A soundscape has two parts. The bed is a looping ambience: room tone, wind, distant traffic. The emitters are short, non-musical samples that fire on a timer with a little randomness in their pan, volume and pitch. A floorboard creaks, a bird calls once, a door closes somewhere down the hall.
Together they fill the gaps between your SFX, so the room already sounds alive before the player has done anything.
Why is this worse on the web?
None of this is new. Native engines have shipped it for years. Unity gives you AudioMixerGroup, Unreal gives you Submixes and Sound Cues, and you wire up bus routing and priorities in an editor once and then forget about it.
The web gives you an AudioContext 🔗 and a GainNode 🔗, and that’s your entire mixer.
The libraries don’t fill the gap either. Phaser ships WebAudioSound 🔗, three.js ships PositionalAudio 🔗, and drei wraps that as <PositionalAudio> 🔗 for r3f. None of them give you scheduled ambience, or a bus that can duck a whole category of sounds while a voice-over plays.
So most web games never get past music plus SFX. The work isn’t hard. Nothing in the tooling nudges you toward it, so it quietly never happens.
Isn’t ambience just a quiet music loop?
I used to think so. Drop a 30-second room-tone MP3 into the music slot, set the volume to 0.2, and call it ambience.
The trouble is that the ear clocks a static loop fast. Within about three repeats your brain has found the seam, and once it has, it can’t stop hearing it. A single unchanging loop ends up with the mood of a fridge hum: the room is on, but nobody’s home.
But isn’t a quiet loop better than nothing?
Yes, it is. A bed on its own beats silence. It’s the baseline, and emitters are what you put on top of it.
Here’s what thirty seconds of the café demo actually does. The bed runs the whole time. Every few seconds an emitter picks a sample, a volume and a pan, and fires once.
A short crossfade at the loop point hides the restart, and the emitters give the ear something more interesting to follow than the seam. The demo below has both. The red flash marks the exact moment the loop wraps, so try to hear it with the emitters off and then on.
Seam audition
Hear the loop seam
Stand still in a real room for a minute and just listen. What you hear isn’t one texture. A fridge compressor kicks in, a neighbour’s door closes, the floor settles, a gust hits the window. None of those are looping. They’re events on top of a near-silent bed, and a soundscape system is exactly that, on a timer.
The knobs
There are three, and skipping any one of them turns the layer back into a tape loop.
A sample pool per emitter, with last-picked exclusion. Each emitter holds three to six short samples. Every fire picks one at random, but never the one it just played. The ear catches an immediate repeat faster than anything else, which makes this the biggest win for the least work.
Random pan, volume and pitch, within tight ranges. Pan ±0.4, volume 0.7 to 1.0, playback rate 0.95 to 1.05. That’s tight enough that nothing sounds broken and wide enough that no two fires are identical. Pitch is the easy one to overdo: past about ±5%, samples start sounding chipmunky or sluggish. Ask me how I know.
A concurrency cap. maxConcurrentEmitterInstances stops pile-ups. Every so often the dice will fire three emitters within 200ms, and without a cap that moment craters the mix and clips the master. Cap it at three or four. It’ll kick in maybe once a minute, and you won’t hear it when it does.
Everything else is plumbing. Play with all three in the playground below.
Parameter playground
The three knobs, live
Pool
- cups-clanging
- cash-register
- phone-ringing
- espresso-machine
Same engine, different emphasis
The same three knobs show up whatever you’re building. What changes is which one you lean on.
2D Phaser games. Tie the bed to the scene and schedule emitters from the scene’s update loop or from timers. The listener rarely moves, so pan is decoration more than space. Three or four samples per emitter with last-picked exclusion is plenty.
3D three.js or r3f games. The bed stays 2D, because room tone doesn’t come from anywhere; it’s everywhere. Emitters can stay 2D and randomized, or become real positional sources with PositionalAudio 🔗 when they should pan as the camera turns. Ducking matters more here, because dialog gets buried quickly when music, footsteps and UI are all competing for the same space.
Narrative scenes. This is where ducking pays off most. Without it, the voice-over fights the ambience and you end up cranking the VO bus until the room is a whisper. Instead, duck the ambience bus over about 120ms when a line starts and release it over about 450ms when the line ends. The voice sits on top, and the room comes back when it’s done.
Ducking A/B
Voice-over over the ambience bed
“You finally made it to the tavern.”
What’s the catch?
It’s more authoring. Three to six samples per emitter is real work compared to one, and you’ll feel that budget the first time a designer asks for “one more bird.”
The concurrency cap means an emitter you wanted to hear will occasionally get dropped. That’s almost always fine, but once in a while the dropped fire was the one sound that sold the room.
Ducking adds buses, and bugs like to hide in bus routing. If you have one bus today, you’ll have four tomorrow.
And the whole layer is invisible when it works. Nobody on the team will notice it, which is the point, but it makes for a tough sprint review when your demo is “the room, except now it sounds like a room.”
What should you not do?
I’ve done every one of these at least once.
Start the bed at full volume with no fade-in. On first load the player hears a hard cut into ambience. A 500ms fade-in costs nothing and removes it.
Use hardcoded <audio> tags. They give you no bus routing, no ducking and no per-instance volume, and you can’t duck what you can’t address. Use Web Audio, or a wrapper like Howler, for anything that has to share the stage with other sounds.
Create a new Howl for every emitter fire. That leaks memory and re-decodes the sample each time. Construct each Howl once at load, then call play() per fire.
Duck by pausing the bed. Pausing loses the loop’s position and you get a click on resume. Lower the bus gain instead.
Forget to clear timers on scene change. A setTimeout armed in scene A will happily fire an emitter in scene B. Track every timer id, clear them all on dispose, set a disposed flag, and check that flag before re-arming inside the recursion.
If you only take one of these home, take the timers one. It’s the one that finds you at 11pm.
Part 2
Part 2 builds this for real: one engine-agnostic class, a six-function bus adapter, and bindings for React, Phaser and plain three.js. The engine itself is under a hundred lines. What changes from game to game is the bus underneath it and the lifecycle wiring above it.
If you only have an afternoon, start with last-picked exclusion and a concurrency cap. Those two knobs are what stop a bed from sounding like a tape.
Stay in touch
Don't miss out on new posts or project updates. Hit me up on X for updates, queries, or some good ol' tech talk.
Follow @zkmake