Deciding Between Failing Fast and Carrying On
Code meets data that should not exist. A null where the schema said there would be a value, a negative duration, a version field naming a shape that no longer exists. There are two familiar answers. Stop immediately, or clamp it, default it, skip the record and keep going.
The argument between those two is usually conducted as taste, and it's the wrong argument, because crash versus continue is the wrong axis.
The axis that matters is loud versus quiet. A crash is loud by construction. Continuing can be loud too, if the anomaly is counted and somebody watches the count. What is never defensible is continuing silently, and that is the option most code picks without anybody choosing it, because it's what writing nothing produces.
Two Frameworks That Encode This
Unreal ships three assertion macros and the distinction between them is the entire argument in miniature.
check halts execution when its expression is false. ensure doesn't halt: it informs the crash reporter and the engine keeps running, and to avoid flooding that reporter it only reports once per session. checkSlow behaves like check but operates only in Debug builds, so it's gone from Development, Test and Shipping.
Three levels, and they are not three degrees of severity. They are three answers to who finds out. check tells the user by stopping. ensure tells the developer while letting the user continue. checkSlow tells nobody outside a debug build, which is correct for a condition that costs too much to evaluate in production and wrong for anything else.
.NET has the same split, but by accident rather than design. Debug.Assert is marked [Conditional("DEBUG")], so it is removed from a release build entirely. A team that asserts its invariants everywhere, and does nothing else, ships a binary that continues silently on every single one of those conditions. Nobody decided that. It follows from a language feature interacting with a build configuration, and the release binary's behavior was never selected by a person.
That's the problem. Silence is rarely chosen. It's a default.
Continuing on Purpose, and the Silence Underneath
ktsu.AppDataStorage is a library of mine for application settings, and it continues on purpose. When the settings file cannot be deserialized:
catch (JsonException)
{
// file was corrupt or could not be deserialized
// delete and try load a backup
AppData.FileSystem.File.Delete(newAppData.FilePath);
return LoadOrCreate(subdirectory, fileName);
}Delete the bad file, recurse, and the read path will restore from the .bk backup. For a truncated write after a power cut, that is the right call: the user gets their settings back and never learns anything went wrong, which is what they want.
Now trace a different input through the same code. Not corruption, but a settings type that changed shape between releases.
The deserialize fails, so the main file is deleted. The recursion restores from .bk and renames the backup to a timestamped file. The deserialize fails again, because the schema is what changed and the backup has the old schema too. The main file is deleted again. This time there's no .bk left to restore, so the read returns an empty string, and an empty string means "no settings yet", so a fresh default instance is created and saved.
The user's settings are gone. They are recoverable, sitting in timestamped files nobody mentioned. No error, no log line, no counter. The application opens with defaults and the person using it concludes that it forgot.
I wrote that, and I'd argue the deserialize-and-fall-back design is right. What's missing is not the fallback. It's that the fallback happens without a sound.
Three Failures That Have Been Silent for Years
Because it would be easy to make this abstract, three from my own code, all of them continuing quietly.
A removed public method ships as a minor version. My build tool decides semantic versions by running eight regular expressions over a git diff. Two of those patterns match a removed public or protected member. And the decision they feed is:
return hasApiChanges
? (VersionType.Minor, "Public API changes detected (additions, removals, or modifications)")
: (VersionType.Patch, "Found changes warranting at least a patch version");Minor or patch. Major is not reachable from that path at all. It only happens when a human types [major] in a commit message. So deleting a public method, which is the definition of a breaking change, ships as a minor version bump and every consumer takes it automatically. The reason string even says "removals" out loud.
The output is a version number. Nobody re-reads a version number.
An impression cap compares against the wrong bound. An advertising platform I wrote capped how often one viewer sees one ad, and the check was if (count($results) > 1) where the intent needed > 0. Every viewer saw each ad twice per window instead of once, and the reporting counted what was served, so the number the advertiser was billed against and the number the code intended were different and nothing in the system could tell.
A conflict path abandons ten locked tables. A data editor of mine takes LOCK TABLES on ten tables at line 35 of its merge handler. On the conflict path it redirects and calls exit() at line 128. UNLOCK TABLES is at line 192, and the conflict path never reaches it.
That one is the purest example in the set, because nothing observable ever goes wrong. MySQL releases table locks when the connection closes at the end of the request, so the bug self-heals every single time, which is precisely why it survived. It is a real defect with no symptom.
Every one of these is best-effort continuation with no counter behind it. Every one was silent for years.
The Third Option
Most code has two branches available at the point of decision, throw or don't, and the interesting one is missing.
Continue and increment something. Clamp the value, default the field, skip the record, and add one to a counter that has a name and a rate. That converts an invisible defect into a number on a dashboard, and a number on a dashboard is a thing somebody can notice going from zero to nonzero.
This matters more than it sounds because of how the arithmetic of rare failures works. A defect that fires once per thousand sessions is effectively invisible to any test plan a team can afford, and reaches a thousand people a day at a million users. Testing cannot find it. A counter can, on the first day, because production has the volume that testing never will. Instrumentation is not a consolation prize for failing to test properly, it's the only instrument with enough samples.
Where to Draw the Line
Neither question that decides it is about how the failure feels.
Can the bad value reach persistent state or money? A wrong pixel is gone next frame. A wrong row is there forever, and every later read compounds it. A wrong payment is a phone call. Continuing is cheap when the damage lasts one frame and expensive when it outlives the process, so the decision follows the durability of the damage rather than the severity of the input.
Who is holding the build? In a build a debugger can attach to, stopping is nearly free and it's the fastest possible feedback. In a build in somebody else's hands, stopping is the most expensive thing that can be done to them, and the goal shifts to preserving their work and telling the author. That's exactly the check and ensure split, and it's why a codebase wants both rather than a house style.
Domains genuinely disagree here, and it is not a maturity difference. A renderer that drops a frame has lost nothing. A ledger that drops a transaction has lost money. A game and a payment system should not share an error policy, and the mistake is having one policy at all rather than having the wrong one.
What Transfers
Ask who finds out, not whether to stop. Every one of these decisions is really about routing a signal to a person. Once the question is "who learns about this, and how", the crash-versus-continue argument mostly dissolves, because both options can be made to answer it.
Silence is a default, not a decision. Nobody in any of my four examples chose to be quiet. They wrote a fallback, or an assertion that compiles out, or a redirect, and quiet is what was left over. Assume that anywhere code handles a bad case without emitting anything, no one has decided anything.
A counter is cheaper than an argument. Teams spend more time debating fail-fast versus fail-safe than it takes to add an incrementing metric that makes the question empirical. If nobody knows how often the bad branch is taken, the debate is about intuitions, and after a week of a counter it's about a number.
Grep the codebase for the quiet branches. catch blocks with no logging, ?? default, clamps, and early returns on data that does not match the schema. I found three of the four above by looking rather than by being told, and every one had been running for years with nothing reported against it.