Thirty Seven Percent of the Time, You See Nothing
There is a class of defect that is effectively invisible during development and enormous in production, and the gap between those two facts is not a failure of diligence. It's arithmetic, and the numbers are more uncomfortable than the intuition.
Take a bug that fails one session in a hundred. One percent.
A Full Day of Testing
Five testers, twenty sessions each. A hundred sessions, which is a solid day of deliberate manual testing on a feature.
The chance that nobody hits it is 0.99 to the power of 100, which is 36.6%.
The chance that exactly one person hits it, once, is 37.0%.
Those two outcomes together are three quarters of all possible days. And they are the same outcome in practice, because a single occurrence of an intermittent failure, with no reproduction steps and no second sighting, becomes a ticket that says cannot reproduce and gets closed. That isn't negligence. It's the correct triage decision given one data point.
So the most likely result of a hard day's testing against a one-percent bug is that the bug survives, and the second most likely result is that the bug survives with a closed ticket attached to it.
The Same Bug in Production
A million daily active users, one percent, is ten thousand people a day.
Not ten thousand total. Ten thousand a day, every day, until it's fixed. If a tenth of them care enough to contact support, that's a thousand tickets a day about something the team has already investigated and closed.
The ratio between those two worlds is the whole problem. A hundred-session test day expects one occurrence. Production produces ten thousand. The same code, the same defect, the same rate, and a four-order-of-magnitude difference in how often it happens, purely because of how many times the dice get rolled.
Making It Rarer Doesn't Help
The instinct at this point is that one percent is quite high for a bug, and a serious defect would be rarer than that. Push it down and see what happens.
One in a thousand, and a test plan ten times bigger, a thousand sessions:
0.999 to the power of 1000 is 36.8%.
Exactly the same answer. Ten times rarer, ten times more testing, and the probability of seeing nothing has not moved.
That is not a coincidence, it's a limit. As the rate gets smaller, (1 − p) to the power of 1/p converges on 1/e, which is 36.8%. Whenever test volume is the reciprocal of the failure rate, the bug goes unseen about 37% of the time, no matter what the rate is. The arithmetic is scale-invariant, so no amount of shrinking the failure rate buys visibility, as long as testing scales the same way.
Meanwhile the production side has not become safe. One in a thousand at a million users is still a thousand people a day.
The Structure Underneath
Rarity and volume are independent variables, and the two halves of the process sample different ones.
Testing samples rarity. A test plan has a fixed session budget, and what it can detect is bounded by that budget. Below a certain rate, a defect is not rare from the test plan's point of view, it is invisible, and no amount of care changes that. Testers cannot find a bug they statistically will not encounter.
Production samples volume. Users do not care about the rate. They experience occurrences, and the occurrence count is the rate multiplied by a number that is many orders of magnitude larger than any test budget.
So there is a band of failure rates that is simultaneously too rare to find and too common to accept, and every product of any size has defects living in it right now. The band is not a gap in the process. It's a consequence of the process having a session budget at all.
What the Numbers Are Actually Good For
The useful thing about doing this as arithmetic is that it converts vague dread into decisions.
It sizes a test plan realistically. A 95% chance of seeing a one-percent bug at least once takes 299 sessions. That's three full days of the five-tester team above. For a one-in-a-thousand bug it's 2,995 sessions, which is thirty days, which nobody is going to authorise. Knowing that is better than hoping, because it names which defects the test plan is structurally capable of finding, and for most teams that means rates above about one in a hundred.
It changes what not reproducible means. A single unreproducible report of an intermittent failure is not an absence of evidence. Against a hundred sessions, one occurrence is what a one-percent bug most often produces, at 37.0%. Treating that report as noise is throwing away the only signal the test budget was ever going to produce. The correct response is not to reproduce it, it's to record it and count.
It argues for instrumentation over reproduction. If a bug cannot be found by testing more, the remaining option is to make each occurrence carry more information. A crash report with state attached, an error counter with a rate on it, a structured log with a correlation ID: these turn production into the test plan that testing could never afford, and they work precisely because production has the volume.
It reframes the staged rollout. A one-percent rollout to a million users is ten thousand sessions, which detects a one-percent bug with certainty and a one-in-a-thousand bug at 99.995%. That's not a risk reduction measure, it's the only test plan with enough sessions in it, and the reason to stage a rollout is that it's the first point in the process where the arithmetic works.
What Transfers
A rate below the session budget's reciprocal is invisible, not rare. Once a failure rate drops below roughly one over the number of sessions a team can afford, testing stops being a detection mechanism for it. Finding where that line sits is a one-line calculation, and it names which class of defect the team is choosing not to look for.
Count occurrences, don't chase reproductions. The instinct with an intermittent bug is to reproduce it, and for the rates that matter that instinct spends the budget in the one way guaranteed to fail. Counting is cheap and rate estimates converge fast, and a rate is what says how bad it is in production.
Never close an intermittent report as noise on n = 1. For any rate where this whole problem exists, a single sighting is the expected outcome rather than an anomaly. The ticket that says cannot reproduce, closing is the process working exactly as designed and reaching the wrong conclusion.
Multiply before arguing. Most disagreements about whether a rare defect is worth fixing are conducted in adjectives. It takes one multiplication to turn "this is pretty rare" into "ten thousand people a day", and that number ends the argument in whichever direction it deserves to be ended.