One Schema, Six Implementations

Adding a new kind of content to a game engine usually means writing the same six things. A class to hold it. Code to read it from disk. Code to write it back. A way to copy it. A way to compare two of them. And an editor panel so somebody other than a programmer can change it.

None of those six are interesting. All of them are derivable from the same information: what fields does this thing have, and what type is each one.

The current engine derives them. Thirty-five JSON files describing schemas and data produce 9,315 lines of C++ across 96 files, and adding a content type means editing JSON.

This is the fifth post taking one system at a time out of my custom engines.

The Declaration

A schema names a class and lists typed members. There's nothing clever in the format, and that's deliberate:

{ "Classes": [ { "Name": "AdventuringGear",
    "Members": [ { "Name": "name",   "Typename": "string" },
                 { "Name": "cost",   "Typename": "int" },
                 { "Name": "weight", "Typename": "float" } ] } ] }

A description field sits alongside each member, unused by the compiler and read by whoever opens the file next.

The whole system rests on that being boring. A schema format with conditionals, inheritance, or expressions becomes a language, and a language needs a parser, error messages, and documentation. This one is a list of names and types, which is the smallest thing that can carry the information the six implementations need.

What Comes Out

The generated header is where the payoff is visible:

class AdventuringGear
{
public:
    AdventuringGear* DeepCopy() const;
    bool DeepEquals(AdventuringGear* other) const;
    void Deserialize(Json::Value& jsonObj);
    void Serialize(Json::Value& jsonObj) const;
    static AdventuringGear* Make();
    void Destroy();
    void ImGui();

Six operations from one declaration, and the last one is the one that changes how the project feels.

ImGui() draws an editor for the type. Not a generic property grid reflecting over something at runtime, an actual generated function that emits the right widget for each field, in declaration order, with the field's name as its label.

Which means a new content type arrives with its editor already built. There's no step where someone decides whether this type is worth building a UI for, and no backlog of types that are only editable by hand-writing JSON because nobody got around to the tool. The tool is a consequence of the declaration.

That's the thing I'd carry into any data-heavy project. Generating the class saves typing. Generating the editor changes who can work on the content, and that's a different order of benefit.

Why Deep Equality Is There

DeepEquals looks like the least useful of the six until it's connected to the editor.

An editor needs to know whether the thing being edited has changed, to enable a save button, to prompt before discarding, to mark a document dirty, or to skip rebuilding a cache. The usual implementations are a dirty flag set by every mutation, which is one forgotten call away from silently wrong, or a hash, which is a dirty flag with extra steps and a collision story.

Generated deep equality gives an exact answer by comparing against a copy of the original. No mutation site has to remember anything, because nothing is tracking mutations. The comparison is the check.

This is the same mechanism I'd used years earlier in a particle system, where a transition kept backup copies of its inputs and rebuilt its lookup table when they stopped matching. I didn't recognize it as a pattern at the time. Keeping a known-good copy and comparing against it is a general answer to "has this changed," and it trades memory for the entire class of bugs where something changed and nothing noticed.

The Ratio Is the Argument

Thirty-five declarations, 9,315 lines of implementation.

The number that matters isn't the volume, it's what can no longer happen. Those nine thousand lines can't disagree with the declaration, because they don't survive the next generation run. The serializer can't drift from the class. The editor can't show a field the class doesn't have. Deep copy can't miss a member added last week.

Every one of those is a real bug I'd written by hand before. Adding a field and forgetting to add it to the copy constructor produces an object that's subtly incomplete after a copy, and the symptom appears far from the cause. Generation removes the category rather than the instance.

The generated directories carry a header saying not to edit them, which is the one place this approach reliably fails. Someone edits generated code to fix something urgent, the fix works, and the next generation run silently deletes it. The header is a convention where a read-only file attribute or a build-time check would be enforcement, and that gap is the same one I've written about elsewhere in these engines.

Where Generation Stops

The boundary needs to be precise, because "generate everything" fails in its own way.

What's generated is the mechanical consequence of a data shape: storage, transport, copying, comparison, and presentation. What isn't generated is behavior. No generated code decides what an item does when equipped, because that isn't derivable from a field list, and a schema format extended far enough to express it would have become the language I said the format shouldn't be.

Behavior lives in scripts and in hand-written C++. The split is that the schema owns what the data is and the code owns what happens to it, and keeping that line clean is what stops the generator from growing into a bad programming language.

The other limit is that generation is a build step, which means a code generator is now part of the build and a broken generator is a broken build. The configuration for it is a single JSON file naming the input and output directories and the projects to regenerate, which is about as much ceremony as the idea can carry before it costs more than it saves.

What Transfers

Find the six things being written by hand for every new type. In most codebases it's some subset of construct, copy, compare, serialize, deserialize, and edit.

Derive them from one declaration, and keep the declaration format boring enough that it never needs a parser worth the name.

Generate the editor, not just the class. That's the part that changes who can contribute.

And enforce the do-not-edit rule with something stronger than a comment, because the one time someone edits generated code will be the time it matters.