Part 1 of 5
Four Engines, Fifteen Years
603 Globals and a Texture Streamer
This is the first of five posts working through four custom C++ engines I built between roughly 2008 and now, in the order I built them. I've kept all the source, so this isn't a memory exercise. The numbers come from counting the actual repositories.
The first one worth writing about is Gumball Engine, which shipped a puzzle game on Steam around 2010. Its globals.h contains 603 declarations. Its source files average 475 lines. It began life as somebody else's tutorial code.
It also streams textures out of a compressed asset pack on a per-frame time budget, throttles its own I/O against a wall-clock measurement, and decompresses straight into memory mapped from the GPU. It has two texture eviction policies as well, and ships with both of them switched off.
Somebody Else's Basecode
The file that starts everything is NeHeGL.cpp, and its header comment is intact:
Jeff Molofee's Revised OpenGL Basecode
Huge Thanks To Maxwell Sayles & Peter Puck
http://nehe.gamedev.net
2001
A generation of self-taught graphics programmers learned from those tutorials. The window creation, the pixel format selection, the message loop, the multisampling setup, all of it arrived as working code that I modified until it was mine. The Mac support came the same way, through a Cocoa GameShell conversion carrying its own attribution to two more people.
Starting from working code means something is on screen on day one instead of week three, and that difference decides whether a project survives. What a basecode can't teach is how to organize the ten thousand lines that come after, and the interesting question is what a person builds in that vacuum.
The Streaming System
The answer, in this case, is a texture streamer.
The game ships its assets in a single gzip-compressed pack. Loading a texture means finding its offset in that pack, seeking there, decompressing a TGA header, then decompressing pixel data. Doing that synchronously when a texture is first needed means a visible hitch, and doing all of it at startup means a long load and a memory footprint sized by the whole game rather than by what's on screen.
So updateStreamingTextures(float milliseconds) runs every frame, from both the Windows loop and the Mac one, and is handed the actual frame duration. It advances whatever work is outstanding and returns. There are no threads anywhere in it.
It seeks incrementally. A gzip stream can't be randomly accessed cheaply, so reaching a texture's offset means decompressing forward to it. Rather than doing that in one blocking call, the streamer moves forward by a bounded number of bytes per frame, holding its position in the file across frames and resuming next time.
It tunes its own budget. This is the part I'd still defend. Each seek is timed with a performance counter, and the byte budget adapts:
float timeout = 10.0f;
if (duration < timeout && obytesToSeek == bytesToSeek) streamingTextureByteFactor *= 2;
if (duration > timeout && obytesToSeek == bytesToSeek && streamingTextureByteFactor > 1) streamingTextureByteFactor /= 2;Under ten milliseconds, do twice as much next frame. Over it, halve. The engine doesn't need to know whether it's running on a fast disk or a slow one, because it discovers that continuously and converges on whatever the machine can sustain. In 2010, when the same game might be on an SSD or a laptop drive or an optical disc, that's the difference between one tuned constant that's wrong on most hardware and no constant at all.
It decompresses into GPU memory. Once the header is known, the streamer allocates a pixel buffer object and maps it:
glBindBuffer(GL_PIXEL_UNPACK_BUFFER, streamingTextureBufferId);
glBufferData(GL_PIXEL_UNPACK_BUFFER, width * height * channels, NULL, GL_STREAM_DRAW);
streamingTextureBufferMemory = (GLbyte *)glMapBuffer(GL_PIXEL_UNPACK_BUFFER, GL_WRITE_ONLY);Decompressed pixels are written directly into that mapping, so there's no staging copy in system memory that later gets handed to the driver. The upload is the decompression. That's the correct way to do it, and it isn't the obvious way.
It manages residency, and then it stops. Every streamed texture carries a fade-in timer
and a last-used timer, against two constants: TEXTURE_POP at 250 milliseconds and
TEXTURE_UNLOAD at 5000. A texture fades in over a quarter second when it arrives, which
hides the moment it appears, and is released after five seconds unused. That's a working set
policy with a hysteresis window, in a game from a solo developer.
It is also switched off in the build that shipped, and the commit history is unusually clear
about how that happened. Three commits land on 2011-06-08. The first adds
unloadUnusedBackgrounds, which evicts background textures and, in its commit message,
"stays at 100 meg". The second adds the general policy, unloadUnusedTextures, behind a
FULL_TEX_MGMT define. The third is one line:
#ifndef FULL_TEX_MGMT
//#define FULL_TEX_MGMT 1
#endifThe general policy was written, tried, and traded back for the narrow one inside a single day.
Turning it off also turns preloadTextures back on, and that function warms 131 textures at
startup so they never stream at all. Seven weeks later, on the same day the atlas work landed,
unloadUnusedBackgrounds gained an unconditional return at the top of its body.
So the shipped game streams textures in and never releases any of them. Both eviction policies are still in the source, one behind a commented-out define and one behind an early return, and both were switched off deliberately, within a day of the alternative being measured.
There's also a low-resolution asset path, selected by a smalltextures flag that switches the pack directory. Two quality tiers, chosen at runtime.
Right Instinct, Wrong Structure
The streamer isn't the only advanced thing in here. The others are worth more, because of how they fail.
A call-tree profiler. There's a hand-built hierarchical profiler with a profilefunc node carrying a parent pointer, a vector of children, and an exclusiveTime field. Understanding that a profiler should build a call tree, and that self time and total time are different questions, is not beginner knowledge. Sixty-one call sites across the renderer feed it.
The implementation undoes it. The entry point is void profileStart(string func), taking a std::string by value, called as profileStart(__FUNCTION__) at the top of hot functions. Every profiled call heap-allocates a string, copies it into a node, and pushes that node into a vector. Then it takes the address of the element it just pushed and stores it as the current node, so the next sibling push can reallocate the vector and leave every parent pointer in the tree dangling.
The commented-out lines are the tell. Inside profileStart there are dead filters that would have restricted profiling to three or four named functions. Somebody noticed the profiler was too expensive and reached for measuring fewer things, rather than for measuring the same things without allocating.
A redundant state filter. glOptionEnable exists because redundant glEnable calls cost something, which is a real and non-obvious thing to know. It checks whether the state is already set before calling into the driver.
It checks by walking a std::vector<GLenum> with an iterator, comparing one element at a time, on every state change. To avoid a driver call that is often little more than setting a flag, it does a linear scan. The optimization is correct and the data structure makes it a pessimization, and the distance between those two is the whole story of this engine.
A best-fit atlas allocator. findEmptyChunk walks the atlas looking for the free chunk that wastes the least area, tracking the smallest overhead seen and returning immediately on an exact fit. That's best-fit allocation, arrived at independently for texture packing.
It identifies a free chunk by testing whether its filename string is empty, and it scans every chunk on every allocation. Right algorithm, and the free list is a linear walk over strings.
All three go the same way. The idea is one a good engine programmer would have. The structure underneath it is whatever was reachable without knowing that the structure is where the win actually lives.
The Globals Are the State Machine
Twenty of those 603 globals are the streaming system's state. The open file handle, the temporary TGA header, the width, height, channels, stride, bit depth, the target byte, the adaptive byte factor, the bytes read so far, the buffer id, the mapped pointer, and three flags tracking whether the header is loaded, whether a seek is in progress, and whether decompression has started.
Read that list again as a description rather than a complaint. Those are the member variables of a coroutine. They're exactly the state a resumable operation has to keep between suspensions: where it is, what it's doing, and what it has already produced.
C++ would not get coroutines for another decade. What this code needed was an object that owned an in-progress streaming operation, and what it had was a header file. So the coroutine frame became globals, one per field, because globals were the only storage with the right lifetime.
That reframes the whole codebase for me. The 603 globals aren't the absence of design. A meaningful fraction of them are designs that had nowhere to live. The engine wasn't naive about streaming, it was naive about ownership, and those are different failures with different fixes.
Immediate Mode, Everywhere
The tension is real, though. The same codebase that maps a pixel buffer and streams into it also draws with 98 calls to glBegin and glVertex in engine_graphics.cpp and another 76 in visuals.cpp.
Those two facts sit about a thousand lines apart. One subsystem got years of attention because it was the thing standing between the game and shipping. The other mostly stayed at tutorial level because a puzzle game with a few hundred cubes never made it hurt.
Mostly, but not entirely, and the exception is the more interesting case. On 2011-05-11 the
text renderer was converted to vertex buffer objects, then converted again to use
glBufferSubData, then reverted to immediate mode, all within one day. The commit that ends
the sequence is called "working version after ATI crap". A fourth commit that day rolled
engine_graphics.cpp back to an earlier revision, and it finished the day with four more
glVertex calls in it than it started with.
That reframes some of the immediate-mode drawing. It isn't all untouched tutorial code. Part of it is code that moved off immediate mode and was moved back, because the buffer path broke on a driver a real customer was running. The word for that is a retreat, not an oversight, and in 2011, with a game selling on Steam and Impulse against whatever hardware buyers owned, it was the right retreat. ATI is a recurring character in this log. It has its own fix in 2010-10-21, its own workaround for static texture upload in 2011-06-09, and a texture bind cache that was "implemented but disabled" on 2011-06-16.
That's what a solo engine looks like. Sophistication is not evenly distributed, it's concentrated wherever the pain was, and the distribution is a record of what the project actually struggled with.
What the Globals Actually Cost
Not crashes and not performance. The game shipped and worked.
The cost was that changing anything required knowing everything. With 603 globals, any function might depend on any state, so "is it safe to change this" has no local answer. Fine for one person holding the whole model. A wall for anyone else, and for the same person two years later.
The second cost was that nothing could be constructed in isolation, so nothing could be tested in isolation. There's no test project and there was never going to be one.
The third cost was reuse, and the reuse wasn't zero. Both of the designs worth keeping did survive into the next engine. The best-fit atlas allocator reappears there almost line for line, and the streaming state machine reappears as a class with one method per state. What didn't survive was this code. Every idea had to be rewritten from scratch to escape the header it lived in, and rewriting from scratch is what a design costs when it has no owner.
What I'd Keep
I wouldn't tell anyone starting out to avoid this path. Beginning from working basecode and hacking at it until it belongs to the person writing it is how a lot of us got here, and reading about architecture before shipping anything produces people who can discuss engines without having built one.
The change I'd make is small and specific, and the streaming system is the argument for it. The moment a piece of work needs to remember where it is between frames, that state wants an owner. Not a manager, not an architecture, just a struct that something passes around. Had those twenty globals been one StreamingTexture struct, the system would have been portable to the next engine, testable without a window, and capable of running more than one stream at a time.
That's the whole gap. The design was already there. It just had nowhere to live.
The next engine is where I overcorrected on exactly that, and split one program into twenty-four projects.