Caching Git LFS: Why the Bytes and the Locks Want Opposite Homes

A studio with its repositories in the cloud pays for the same bytes repeatedly. Continuous integration pods clone the same repository over and over. Artists on home connections pull the same assets their colleagues pulled an hour ago. Upstream bandwidth and clone latency dominate everything else about the experience, and none of it is work that needed doing twice.

ktsu.GitLfsCache sits between those clients and the forge. It relays every Batch API call upstream with the client's own credentials, rewrites the transfer URLs in the response to point back at itself, and serves object bytes from a local content-addressed store. A miss is fetched once and written on the way through.

That much is an ordinary caching proxy. What I didn't anticipate is that the Git LFS API has a second job hiding in it, and that job wants to be deployed somewhere else entirely.

The Two Costs

Object transfer is bandwidth-bound. The cost is bytes, the fix is to put the bytes closer to whoever is receiving them, and the right placement is near the client.

Locking is round-trip-bound. The LFS locking API is part of the same HTTP API, addressed as a suffix of whatever lfs.url points at, so a client already sends its lock traffic to the same place. But the cost there isn't bytes at all. It's the number of times a client has to talk to the forge and wait. The right placement for that is near the forge.

These pull in opposite directions, and with repositories in the cloud and clients spread across an office and a lot of home connections, one deployment can't be right for both. That's the whole reason the service has a metadata-only mode, which I'll come back to.

Relaying Every Batch Call

The proxy never short-circuits a Batch call, even when it already holds every object being requested.

Nobody should optimize this away, because the optimization is obvious and the failure it causes is silent. Batch is where the forge decides whether this client may read these objects. If the proxy answers from cache because it recognizes the object ids, then holding a cached object becomes a way to skip an access check, and the cache has quietly become an access control bypass.

So upstream stays the sole authority on who may read what. Every Batch call reaches it. The proxy's contribution is what happens after the answer comes back: the transfer URLs get rewritten, and the bytes come from local disk instead of the internet.

The same principle appears in the store. Content is digested while it streams and published only if it hashes to the object id it claims, so a truncated or tampered transfer can never become a cache hit. A cache that stores whatever arrives is a cache that serves whatever arrived.

The Token in the URL

Each rewritten transfer URL carries an opaque token holding the upstream URL, its headers, the object id, the size, the upstream key, and an expiry. The payload is encrypted, authenticated with a message authentication code over the whole envelope, and Base64url encoded into a query parameter.

That one decision buys three separate things.

Replicas need no shared state. Everything required to serve a URL travels inside the URL, so any replica can serve any request with no sticky sessions and no shared cache.

Holding a valid token is itself proof that upstream approved that object for that client. The token can only have come from a Batch response the forge authorized, so cached bytes never reach a client the forge didn't clear. The access check and the cache hit stay connected without the cache having to remember anything about who asked.

And because the token always carries the upstream action, an object evicted between the Batch call and the transfer still resolves. The proxy falls back to fetching it, rather than failing a request that was valid when it was issued.

The Locking API Nobody Optimizes

GET .../locks is answered from an in-memory snapshot per repository. The first client to ask after the snapshot goes stale walks every upstream cursor page and publishes the result, and everyone else reads it. A client polling on a timer costs upstream one walk per ListTtl, however many clients are running.

The proxy also assembles the cursor pages itself, so a client gets the whole listing in one round trip instead of walking pages over a wide-area link. For a studio where every editor polls lock state on a timer, that pair of changes is the largest single reduction in upstream traffic the service makes, and it has nothing to do with object bytes.

The cache never decides who may read locks. It only remembers, briefly, what upstream already decided. A caller with no current admission is checked with a single-page request under its own credential before it sees anything, and an admission only ever exists because a real upstream call succeeded. A credential revoked upstream can still read listings for up to AdmissionTtl.

Creation and release are never cached, only relayed, because upstream is the only thing that may grant or release a lock. A successful one drops the snapshot immediately, so a client that just took a lock sees it.

Staleness here costs a retry, not a conflict. Two clients can both see a file unlocked within ListTtl and both try to lock it. Upstream grants one and refuses the other, exactly as it would have without the cache. What the proxy can't see is a lock taken outside it, through a forge's web interface or a client not configured through the proxy, and ListTtl is the only thing bounding how long that stays invisible. That's why it's short, and why it isn't a knob to turn up casually.

Batched Locking, and What It Can't Promise

git-lfs issues one lock request per path, one after another. Locking several hundred assets over a wide-area link is several hundred sequential round trips, which is the kind of number that turns a routine operation into a coffee break.

POST .../locks/batch is a proxy extension that takes them at once and fans them out in parallel under a rate limiter. Each item is still one upstream call carrying the caller's own credential, so upstream decides every lock individually under the caller's identity. The proxy chooses only the order and the concurrency.

It always returns 200 when the request was well formed, with one result per item in the order they were sent. Partial success is the normal outcome here, not an error case, and failing the whole request would throw away the half that worked.

This is not atomic and it cannot be, over an API with no transaction. A client that disconnects part way through leaves the locks already granted still granted. The result array is authoritative, and a client has to reconcile against it rather than assuming the request either happened or didn't.

Forge rate limits are the expected failure here rather than an edge case. A refusal carrying Retry-After pauses every call to that upstream for the stated duration. Anything but zero on the throttled-items counter means the concurrency is above what that forge tolerates.

Two Planes, Opposite Placement

Store:Enabled: false runs the same binary with the object plane switched off. No store, no volume, no eviction, and batch responses relayed unrewritten, so clients receive the forge's own hrefs and transfer straight from it. Object bytes never cross the process.

That mode exists entirely because of the placement conflict. A full deployment near the client caches the bytes. A metadata-only deployment near the forge handles the lock listings and the batched locking. Neither one is a degraded version of the other, they're two different jobs that happened to arrive through the same API.

The two are interchangeable from a client's point of view when they share their token keys and public base URL, because a rewritten transfer URL carries everything needed to serve it. One hostname can resolve to a different instance for different clients without anyone needing two credential entries. A client that batches against one and transfers against the other misses, and the object is fetched from upstream, which is correct rather than a failure.

A metadata-only deployment needs no persistent volume and can be a plain Deployment rather than a StatefulSet, which makes it cheap enough to put wherever it's useful.

The Tension With Not Locking at All

I've argued in a couple of places that studios should be reducing how much their workflow depends on file locking, and that lock-free approaches to binary assets are worth the investment. Building infrastructure that makes locking faster looks like it cuts against that.

My opinion is that both are true at once, and treating them as opposites is the mistake. Reducing lock dependence is a multi-year structural change that touches asset organization, tooling, and habits. The locking a studio does today is happening regardless of what its roadmap says, and making it cost one round trip instead of several hundred is worth doing this week. One is the destination and the other is the current condition, and refusing to improve the current condition because it isn't the destination is how teams end up with neither.

Limits

The hit and miss counters are the pair to watch on the object plane. A low hit ratio means the proxy is costing latency without saving bandwidth, which usually means the volume is too small for the working set rather than that the idea is wrong.

On the locking plane, lock list hits against refreshes is directly the multiple by which upstream lock traffic has been reduced. If that number is near one, the snapshot isn't being shared and the deployment is in the wrong place.

Two things I'd flag before anyone adopts this. Batched locking is an extension, so a stock git-lfs client won't use it and something has to be taught to. And access times are stamped explicitly on every hit rather than read from the filesystem, because noatime and relatime mounts make filesystem access times unreliable and eviction depends on them. That last one is the sort of detail that works perfectly in testing and then quietly evicts the wrong objects in production.