A Synthesizer in a Garbage-Collected Language
Synthr is a modular software synthesizer I wrote in C#. Seventeen modules, ASIO output, MIDI input, a Dear ImGui interface, and both a standalone application and a VST plugin. Around six and a half thousand lines.
It works. It also does, inside its audio callback, three things that every piece of real-time audio guidance says never to do.
Both halves of that are the post.
The Deadline Is Not Negotiable
Audio runs at 48 kHz here. The driver hands over a buffer, and the callback has until that buffer finishes playing to fill the next one. At a typical ASIO buffer size that's somewhere between three and ten milliseconds, every time, forever.
Missing it isn't a dropped frame. There's no equivalent of a stutter that most people won't notice. The hardware plays whatever is in the buffer, and if the buffer isn't ready it plays silence or the previous contents, which is an audible click. One late callback in a thousand is a click every few seconds.
That deadline is what makes audio the strictest environment in ordinary software, and it's why the rules around it are so absolute.
What the Callback Actually Does
The chain is short. The driver calls a sample provider, which calls into polyphony:
public int Read(float[] buffer, int offset, int sampleCount)
{
lock (instrument.polyphony.noteLock)
{
instrument.polyphony.FillBufferASIO(buffer, offset, sampleCount);
}That's the first one. A mutex, taken on the audio thread. If the interface thread or the MIDI handler holds that lock when the callback fires, the audio thread waits for them, and the audio thread cannot afford to wait for anything.
Inside, the second:
if (parallelBuffer.Length != voices.Count * sampleCount)
{
parallelBuffer = new float[voices.Count, sampleCount];
}A heap allocation on the audio thread. It's guarded, so it only happens when the voice count changes, which is exactly when somebody plays a chord. An allocation can trigger a collection, and a collection can pause the thread that allocated.
And the third:
ParallelLoopResult result = Parallel.For(0, AggressiveMax(threads, voices.Count), i => { ... });Dispatching to the thread pool from inside a real-time callback. The pool is a general-purpose scheduler with no knowledge of the deadline, shared with everything else in the process, and Parallel.For blocks until all its work completes.
There's no GC configuration anywhere in the project. No server garbage collector, no low-latency mode, nothing. It runs on the defaults.
Every One Came From a Correct Instinct
That list reads like a catalogue of mistakes, and each item is a right idea reaching for the nearest tool.
The lock exists because the problem is real. Note state is written by the MIDI thread and read by the audio thread. That is genuine shared mutable state across threads and it does need protecting. Noticing that, rather than discovering it later as an intermittent crash, is the harder half.
The answer real-time audio uses instead is a single-producer single-consumer queue of note events, written by MIDI and drained by audio, where neither side ever blocks. The audio thread reads what's there and moves on. Same problem, same correctness, no waiting.
The allocation exists because the buffer genuinely has to resize. A buffer sized to voices times samples is the right shape, and the guard shows the author knew allocating every callback would be worse. The missing step is preallocating for the maximum voice count at startup, so the buffer is large enough for any chord anyone can play and the resize never happens.
The parallel dispatch exists because voices are genuinely independent. Sixteen voices are sixteen separate computations with no shared state, which is exactly what parallelizes well, and noticing that is correct. The answer is a small pool of threads created once, pinned, and signalled, rather than a scheduler that decides when to run and shares its workers with the file system.
Three right diagnoses, three wrong instruments. That's a pattern I keep finding in my own old code, and here it's concentrated in the fifty lines where it matters most.
The Rules I Later Wrote Down
Years after this, I spent a while reading a large engine codebase and wrote up what transferred. The organizing idea was that code should be sorted into tiers with different rules, and the strictest tier had this list: zero managed allocation, no LINQ, no closures, no async, no locks.
The case study I used for that tier was real-time audio, with its hard deadline at 48 kHz.
Synthr is why I knew to write that. Those rules aren't received wisdom I copied from an article, they're the list of things I'd already done wrong in the one place they can't be got away with. That's a more useful way to hold a rule than having read it, and a considerably slower way to acquire one.
What Is Actually Good Here
The audio path is the weakest part of a design that's otherwise nicely put together.
A module is one class holding its DSP, its own interface panel, and its own serialization:
public class Module
{
[JsonIgnore] public int height = 1;
[JsonIgnore] public int width = 1;
public bool active = true;
protected string name = "Module";Width and height in grid units, laid out automatically into a rack. Fields without [JsonIgnore] are saved, so a patch file is just the modules' state and adding a parameter makes it persistent for free. Adding a module means writing one class, and it arrives with a panel, a saved state, and a place on the grid.
Routing is by name rather than by wire. A Bus provides sixteen named channels, and modules write to and read from them:
public const int numChannels = 16;
float[] channels = new float[numChannels];That's a patch bay rather than a node graph, and for a rack-style synth it's the better choice. No module needs a reference to any other module, connections are two dropdowns rather than a cable-dragging interface, and a saved patch is a list of names instead of a serialized graph. The cost is that arbitrary topologies can't be expressed and feedback paths need care, which for this instrument is a cost worth paying.
The module list is a real instrument: oscillators, three envelope types, an LFO, filters, delay, reverb, distortion, bit crusher, compressor, limiter, a sequencer, a drum machine, and a spectrum analyser. That's not a demo, it's something a person can make music with.
What I'd Change
Replace the lock with a lock-free event queue. MIDI writes note events, the audio thread drains them at the top of the callback, neither blocks. This is the single change that matters most, and everything else is smaller.
Preallocate the voice buffers for the maximum polyphony at startup and never resize them again.
Replace Parallel.For with a fixed set of worker threads created once and woken by a signal, or process voices serially and measure whether the parallelism was ever needed. Sixteen voices of subtractive synthesis is well within what one modern core can do in three milliseconds, and the parallel version brings a scheduler into the deadline to buy something that may not have been necessary.
And set the garbage collector's latency mode for the duration of playback, which is the cheapest possible mitigation and costs one line.
None of that changes the architecture. The module model, the bus, and the serialization are all sound. It's fifty lines at the boundary between the interface and the audio hardware, which is the usual place for this kind of problem, because it's the one place where the rules are different from everywhere else in the program.