Time-Sliced Asset Streaming Without Threads

This is the first of a series taking one system at a time out of the custom engines I've written and looking at the idea inside it, separately from the engine it happened to live in. The chronological series covered how the engines evolved. This one is about the parts worth stealing.

Streaming is the right place to start, because it's the system where the constraint produced a design I'd still use.

The Problem

A game ships its assets in a compressed pack. Loading one means seeking to an offset, decompressing a header, then decompressing pixel or vertex data. Doing that when the asset is first needed produces a visible hitch. Doing all of it at startup produces a long load and a memory footprint sized by the whole game rather than by what's on screen.

The usual answer is a loader thread. That's a reasonable answer, and it brings a queue, a mutex, a lifetime problem where the main thread can free something the loader is still reading, and a graphics API that mostly wants to be talked to from one thread anyway.

The answer I built instead has no threads at all.

Bounded Work Per Frame

The loader is a function called once per frame from the game loop, handed the actual frame duration:

updateStreamingTextures((float)duration);

It advances whatever work is outstanding by a bounded amount and returns. Not "finishes the current item," just "makes some progress." A texture might take twenty frames to arrive, and that's fine, because nothing is waiting on it.

This is cooperative scheduling, and its properties are the reason to prefer it here. There's no shared state, so there are no locks. Everything happens on the thread that owns the graphics context, so uploads are legal. The lifetime problem doesn't exist, because nothing is ever half-owned by two threads. And the whole system can be stepped in a debugger, which a threaded loader effectively cannot.

The cost is that a single asset arrives later than a dedicated thread would deliver it. For a game that hides arrival behind a fade, that cost is invisible.

Tuning the Budget Against a Clock

The hard question in bounded-work-per-frame is how much work a frame gets. A constant is wrong on most hardware: too small and loading crawls on a fast disk, too large and it stutters on a slow one.

So the budget measures itself. Each unit of work is timed, and the amount attempted next frame moves in response:

float timeout = 10.0f;
if (duration < timeout && attempted == budget) budget *= 2;
if (duration > timeout && attempted == budget && budget > 1) budget /= 2;

Doubling on success and halving on overshoot converges quickly and needs no knowledge of the machine. It's the same control shape as TCP congestion control, arrived at for the same reason: the right rate is unknowable in advance and discoverable continuously.

The attempted == budget guard matters more than it looks. It only adjusts when the full budget was actually used, so a short final chunk near the end of a file doesn't get read as evidence that the disk got faster.

Decompress Into the Destination

The upload path is where a streaming loader usually leaks memory bandwidth. The straightforward version decompresses into a heap buffer, hands that buffer to the graphics API, and the driver copies it again.

Mapping the destination first removes a copy:

glBindBuffer(GL_PIXEL_UNPACK_BUFFER, bufferId);
glBufferData(GL_PIXEL_UNPACK_BUFFER, width * height * channels, NULL, GL_STREAM_DRAW);
bufferMemory = (GLbyte *)glMapBuffer(GL_PIXEL_UNPACK_BUFFER, GL_WRITE_ONLY);

Decompression then writes directly into memory the driver owns. The upload is the decompression rather than a step after it.

The generalizable version of this has nothing to do with graphics. When a pipeline has a producer, a transform, and a consumer that owns memory, ask whether the transform can write into the consumer's memory rather than into its own. It usually can, and the copy it removes is proportional to the data.

Seeking Through a Compressed Stream

Streaming out of a gzip pack has one genuinely awkward property: the stream can't be randomly accessed. Reaching an asset's offset means decompressing forward to it.

That would be a blocking operation of unbounded length, which defeats the whole design. So seeking is itself time-sliced. The loader moves forward a bounded number of bytes per frame, keeps its file position across frames, and resumes next time. The seek is just another kind of work the budget governs.

There's a lesson in that beyond compression. Any step that can't be bounded needs to be made resumable before it can live in a frame loop, and resumability is usually a matter of storing the position and re-entering rather than of a fundamentally different algorithm.

Forward seeking is the tolerable case. Seeking backward isn't a seek at all, it's a rewind to byte zero and a re-decompress of everything up to the target, which is why the pack is deliberately built in the order assets are first requested. That ordering is its own system and gets its own post later in this series.

Residency, Not Just Loading

A loader that only loads is half a system. This one tracks two timers per asset: how long since it arrived, and how long since it was used.

#define TEXTURE_POP    250    // fade in over a quarter second
#define TEXTURE_UNLOAD 5000   // release after five seconds unused

The fade-in hides the moment of arrival, which is what makes the latency of cooperative loading acceptable in the first place. The unload timeout bounds the working set by recency rather than by a fixed budget.

Those two constants are doing something subtler together. The gap between them is hysteresis. An asset that flickers in and out of view isn't reloaded on every flicker, because five seconds of grace covers the gap, and the fade means even a genuine reload doesn't pop. Choosing an unload timeout much longer than the fade-in is what stops the system thrashing at the boundary.

With one caveat that only the commit history shows. The general unload path shipped switched off. It arrived on 2011-06-08 behind a FULL_TEX_MGMT define, and the define was commented out later the same day in favour of a narrower policy that evicted backgrounds only. That narrower one then picked up an unconditional return at the top of its body seven weeks later. The constants are real and the code is real, and neither ran in the build people bought.

That doesn't damage the design, but it changes what the design is evidence of. Two eviction policies were written and both were turned off, which says the pressure that justified the loader never arrived for the evictor. A game that preloads 131 textures at startup and holds a few hundred more has a working set that fits in memory, so eviction by recency was answering a question the game never asked. Building it, measuring it, and then switching it off is the right sequence.

The Failure Mode Is the Point

What makes this design worth reusing is what it does when it can't keep up.

A blocking loader that falls behind stalls the frame. A threaded loader that falls behind grows a queue and eventually the memory holding it. This one just delivers assets later. The frame rate is unaffected, because the budget is enforced against the clock, and the game keeps running with lower-detail or missing content until the loader catches up.

Degrading rather than stalling is the property to design for in anything on a frame budget. It's also what makes the adaptive budget safe: the worst case of a bad estimate is slower loading, not a dropped frame.

What I'd Change

The first version kept its state in twenty globals, which meant exactly one streaming operation could exist. The second version made it a class, and the moment it was a class the states became named methods and a second concurrent stream became possible. Anything resumable wants an object, because the object is where "where was I" lives.

I'd also separate the budget from the work. In both versions the adaptive timing is entangled with the file reading, so the same scheduling logic can't govern anything else. A small Budget type that hands out permission and takes back a duration would let decompression, uploading, and unrelated background work share one policy.

And I'd measure the whole thing rather than each step. The current design times individual operations against a fixed ten milliseconds, which doesn't know how much of the frame is already gone. A loader that asked how much time was left in this frame, rather than whether this operation took under ten milliseconds, would fill the gaps more precisely and never overshoot a frame it was already late in.