The Source Control Heartbeat: Caching Branch State for Unreal Editors
An Unreal editor with source control enabled asks the same question on a fixed heartbeat: which assets have changed on the branches I care about, so I can warn someone before they lock one. Answering it locally costs a git fetch, then a git log and a git diff per monitored branch. That happens on every workstation, roughly every thirty seconds, and it produces an identical answer on all of them.
I built ktsu.GitBranchStateCache to do that work once, next to the forge, and serve the result. The service is small. The interesting part is how many of its design decisions are forced by things that have nothing to do with caching.
Repeated Work, Once Per Editor
This isn't the usual "our CI is slow" complaint.
The work is duplicated across people rather than across time. Twenty artists with the editor open are twenty clients running the same fetch against the same remote, computing the same diff, on the same heartbeat. Caching it locally on each workstation saves nothing, because the answer is already local. The redundancy only exists when the workstations are viewed together.
The cost also lands unevenly. A fetch is round trips, and round trips are what a residential connection is worst at. The engineers who notice the problem least are the ones in an office on a fast symmetric link. The people who feel it are the artists working from home, who are also the people the locking workflow exists to protect. That asymmetry is the same one the install-experience problem has, and it goes unreported for the same reason.
Once the work is understood as duplicated across people, the fix follows. Move it to one place where the duplication collapses into a cache hit.
Blob Ids Instead of a Verdict
The service returns blob ids per path per branch. It does not return "this file changed."
That distinction matters. A verdict computed from a log and a diff approximates the answer, and it gets the changed-then-reverted case wrong: a file modified in one commit and restored in the next appears in the log, so it reads as changed, when the content at the branch tip is identical to what the client already has. Comparing blob ids catches that exactly, because an identical blob has an identical id, however many commits it took to get back there.
The client compares the returned ids against its own working tree. That comparison is local, exact, and cheap, and it means the service never has to know anything about the client's state beyond one commit id.
There's a second benefit I didn't design for and noticed afterward. For an asset tracked by Git LFS, the blob in the tree is the pointer file, and the pointer file contains the LFS object id. Comparing pointer blobs is comparing LFS object ids. The service needs no LFS awareness at all, which for a studio whose entire content directory is LFS-tracked is the difference between a small service and a large one.
A Blobless Mirror
The service keeps a bare mirror of each repository, cloned with --filter=blob:none. That fetches commits and trees but no file content.
This is the property that makes the disk cost predictable. The volume is sized by the allow-list rather than by traffic, so it can be provisioned in advance instead of watched. A mirror of a game repository with years of binary history would be enormous. A blobless mirror of the same repository is a few percent of that, because the parts that are large are exactly the parts it doesn't fetch.
It works because the question never needs blob content. git diff-tree in raw mode reports blob ids, not blob contents, so nothing in the request path ever asks for a blob and nothing triggers a lazy fetch. To make sure a mistake surfaces loudly rather than quietly downloading a hundred gigabytes, GIT_NO_LAZY_FETCH=1 is set on every invocation. A bug there becomes an error instead of a bandwidth incident.
The diff itself uses git diff-tree -r -z --no-renames rather than git diff --raw. Plumbing rather than porcelain, so the output format can't change with configuration. NUL-delimited, so no path has to be unquoted, which matters because asset paths contain spaces and occasionally worse. And renames reported as a delete plus an add, so the answer never depends on a similarity heuristic that might score two assets differently on two different days.
Authorization Without a Service Account
The obvious way to build this is a background poller with a service credential. I rejected that, and here is why.
A service credential is a credential that can read the whole repository, held by a process whose job is to answer questions about the repository for anyone who asks. The moment it exists, the service's own access control is the only thing standing between a caller and the source. Get that wrong and the service becomes a way to read code without permission to read code.
Instead, every upstream operation runs under the requesting client's own credential. Authorization is a git ls-remote against the upstream, using the caller's credential, before anything is served. That's forge-agnostic, it's one cheap round trip, and it proves exactly the read access being requested rather than approximating it.
Two consequences follow, and both are deliberate.
It can't fail open. If the forge is unreachable, nothing is served, because there's no path through admission that admits without a successful upstream call. A cache that keeps answering when it can no longer check permissions is a cache that leaks.
And a background poller becomes impossible, because there's no credential to poll with. Everything is driven by requests, leader-pays and coalesced, so the first client to ask pays for the fetch and the rest join the result.
The caller's credential also never reaches a git command line. It's passed through the environment instead, because on Linux a command line is world readable through /proc and an environment block isn't. This process handles many different people's forge credentials at once, so that distinction stops being academic.
The Allow-List Before Admission
Repository allow-listing is required, and there's deliberately no pattern meaning "everything." Every pattern has to name at least one literal path segment.
The reason is that a request for an unlisted repository wouldn't fail cheaply. It would clone a permanent mirror of that repository onto a shared volume, sized by the repository rather than by the request, that nothing ever evicts. One curious request becomes a permanent disk cost.
The ordering matters more than the list. The allow-list is checked before admission, not after. Reversed, an unlisted repository would still be probed against the forge with the caller's credential before being refused, and the difference between "refused" and "refused after a successful probe" is observable in the timing. That turns the service into an oracle for which repositories a given credential can read. Checking the list first means an unlisted repository produces no upstream call at all.
My opinion is that this ordering question is the most common way a well-intentioned access check goes wrong. The check happens, it returns the right answer, and it leaks anyway because of what it had to do to get there.
Deployment Next to the Forge
The service deploys adjacent to the forge, not on-premises at the studio. That's the opposite of where an object cache goes, and the reasoning needs stating.
An object cache moves bytes, so it belongs near the people receiving the bytes. This service's cost is round trips against the forge, and its clients are worst served on residential links. Putting it near the forge means the expensive leg happens over a fast link, and the client's leg is a single request and response.
Each replica keeps its own mirrors and its own diff cache, so replicas multiply fetch traffic and disk without improving the hit rate. Start at one. Scaling this horizontally makes it worse, which is unusual enough to write down next to the deployment manifests rather than leaving for someone to discover.
The diff cache is keyed on the merge base rather than on the client's commit. Every artist sits on a different commit, but a team shares an integration point, so one computed diff per branch serves all of them. When that assumption breaks the metric shows it: a low hit ratio means the team is spread across many integration points, which is a fact about the team's branching, not about the cache.
What It Deliberately Isn't
It's not a version control system and not a git server. It never serves object content, never accepts writes, and is never a source of truth for anything. Every answer it gives is verifiable against the forge, and a client that doesn't trust it can compute the same answer the old way.
That framing is what the security posture rests on. A service holding read-only copies of source is a more attractive target than a blob cache, and the mitigation isn't "it's secured," it's "there's nothing here that isn't already in the forge, and nothing that can be written through it."
Limits
The service is new, and I'm not claiming it's proven. It has run against two forges, GitHub and Azure DevOps, and the reason it works identically against both is that it uses no forge REST API at all. ls-remote, fetch, merge-base, and diff-tree behave the same everywhere and have no file-count cap. GitHub's compare endpoint caps at 300 changed files and truncates silently, which for this question is indistinguishable from a complete answer, and that alone ruled out the API approach.
The unknown_base counter is the one I'd alert on. Every increment is a client falling back to the local computation the service exists to replace, usually because it declared a base the mirror has never seen. That's the failure that looks like success from outside, because the editor keeps working and only the bandwidth bill notices.
The largest open question is whether the thirty-second heartbeat is the right thing to optimize at all, or whether the editor should be asking less often. Caching the answer to a question makes the question cheap, and cheap questions get asked more. I've made the current behavior affordable rather than making it correct, and those aren't the same achievement.