The Silent Failure
Every dashboard is green. Every alert is quiet. And something has been rotting underneath for six weeks.
The most expensive failure I ever chased didn't break anything at all. There was a recommendation system, and one March it started serving suggestions that were slightly worse - a little less relevant, slightly less likely to actually sell something - and it just carried on like that, every single day, for three months. What stopped it was a finance analyst who was reconciling the quarter's revenue and noticed the numbers were softer than they ought to have been, and she went digging for why. Nothing had fired. No log line had gone red, and there was nothing for one to go red about, because nothing had crashed. The thing had just got worse, quietly, while everything anyone was actually looking at stayed a nice healthy green.
Our instinct on this is backwards. A crash is more or less a gift. It wakes somebody up, it tells you what it is, it leaves a stack trace pointing at where it happened, and then it has the good manners to stop. We've built up a whole profession's worth of reflexes around the failures that scream and basically nothing around the ones that whisper, and the whispering ones are the ones that bleed you out over months. There are roughly three quiet ways a system goes wrong. It can degrade - get slower or worse in small steps, none of them big enough to notice on the day. It can lie, handing back an answer that sounds confident and is just wrong. Or it can rot, quietly mangling something underneath while everything on top stays calm. None of those three trips a wire. The clock on the damage starts the day the thing begins and only stops the day some human actually looks, and those two dates can be a quarter apart.
The commonest hiding place for this stuff is inside an average. Your median response time is forty milliseconds and has honestly never looked better, and that is precisely why you can't see that one request in every hundred is taking two whole seconds. An average is a device for wiping out the outliers, and the outliers are your customers walking away. The pain is out in the tail of the distribution, the worst one percent, and the tail is the first thing any summary chucks overboard. A dashboard with a healthy mean on it is not telling you everything is fine. It's telling you most things are fine, which is a weaker thing to say, and the gap between those two is where the quiet failures get to breed.
And the tail doesn't politely stay at one in a hundred as a system grows. It takes over. Two engineers at Google, Jeffrey Dean and Luiz Barroso, set this out very clearly in a 2013 paper called "The Tail at Scale". Say a server answers in ten milliseconds most of the time, but one time in a hundred it takes a full second. One slow request in a hundred, you'd think, is easy enough to live with. Except a real request these days doesn't hit just one server. It fans out to dozens, or hundreds, and it has to wait for the slowest one of them to come back before it's done. Send it to a hundred servers like that and the odds that at least one of them is in the middle of its one-in-a-hundred bad moment aren't one percent. They're sixty-three percent. So the system as a whole is slow on nearly two requests in every three, and yet not one of its parts looks ill when you go and check it on its own. The failure is real, you can measure it, and it's invisible to every single per-component gauge you've got.
Go underneath the software and it gets odder, because once you're at enough scale even the hardware starts lying to you. We build the whole edifice on the assumption that the processor does its sums right - that if the chip tells you two and two are four, then they are. At fleet scale that just isn't true. In 2021 both Google and Facebook reported finding processors in their data centres that were quietly doing their arithmetic wrong. Not all the time. Now and again, on certain inputs, with no error and no flag raised, corrupting data that then went on downstream looking exactly like perfectly good data. Google called them mercurial cores and reckoned on a few in every several thousand machines. Facebook described hundreds of bad chips spread across hundreds of thousands of its own. This is the real reason serious systems checksum their data all the way from end to end and check it against some independent witness - a layer that marks its own homework is always going to tell you it passed.
All of which crashes into one hard fact: you can't set an alarm for a failure you never thought of. Monitoring answers the questions you already knew to ask - is the disk full, is the queue backing up, is the error rate going up - a whole wall of gauges for the failures you saw coming. The silent failures are the ones you didn't see coming, more or less by definition, the questions it never occurred to you to write down in the first place. The harder thing, and the more useful one, is being able to take a brand new question to a system that's already running, something you never anticipated at all, and pull an answer out of the data it's been keeping anyway. There's a trap on the other side of this too. Heap on enough alerts that most of them turn out to be noise and you've trained every human in earshot to tune them out, so when the one that actually matters goes off it gets ignored right along with all the rest. Hospitals have been studying this for years and they call it alarm fatigue, after patients came to harm lying next to machines that had cried wolf so many times the staff had simply stopped hearing them. With too few signals you're blind to what's happening. Pile on too many and the noise leaves you deaf to the one that counts.
So the discipline is to measure before you do anything, because under pressure what you really want to do is optimise whatever you happen to be able to see, and the thing that's actually slowing you down is hardly ever the thing you assumed it was - a guess dressed up as a fix just shoves the pain off somewhere darker. But there's one last turn here, and it's the one that stops all this collapsing into some cult of dashboards. The popular line is that you can't manage what you can't measure. W. Edwards Deming, who often gets the credit for saying it, spent his whole career arguing the opposite - that the line is an expensive myth, that running an organisation off its visible figures alone is a disease, and that the most important numbers of the lot are unknown and unknowable. The job isn't to worship the metrics. It's to make the system tell you the truth, and never once let yourself mistake the part you can measure for the whole that actually matters.
In the manifesto, this is tenets (XVII) and (XVIII).
Sources
- [Dean & Barroso 2013] Jeffrey Dean & Luiz André Barroso, "The Tail at Scale". CACM 56(2), 2013. https://www.barroso.org/publications/TheTailAtScale.pdf. Tail latency and a stranger's worst day, invisible to every per-component average; tenets V, XVII.
- [Deming 1986] W. Edwards Deming, "Out of the Crisis". MIT Center for Advanced Engineering Study, 1986. https://deming.org/explore/seven-deadly-diseases/. Running on visible figures alone is a deadly disease; the most important figures are unknown and unknowable; tenet XVII.
- [Dixit et al. 2021] Harish Dattatraya Dixit et al. (Facebook), "Silent Data Corruptions at Scale". arXiv:2102.11245, 2021. https://arxiv.org/abs/2102.11245. CPUs in the fleet corrupting data with no error reported; the failure that announces nothing; tenet XVIII.
- [Hochschild et al. 2021] Peter H. Hochschild et al. (Google), "Cores that don't count". HotOS '21, 2021. https://research.google/pubs/cores-that-dont-count/. Mercurial cores: rare hardware that computes the wrong answer silently, so output must be checked, not trusted; tenet XVIII.
- [Joint Commission SEA 50 2013] The Joint Commission, Sentinel Event Alert Issue 50: "Medical device alarm safety in hospitals". 2013. https://www.jointcommission.org/en-us/knowledge-library/newsletters/sentinel-event-alert/issue-50. Alarm fatigue: when most alarms need no action, people stop hearing them; tenet XVIII.
- [Kalman 1960] Rudolf E. Kalman, "On the general theory of control systems". Proc. First IFAC Congress (Moscow), 1960. https://en.wikipedia.org/wiki/Observability. Origin of 'observability': internal state known only from external output; tenet XVIII.
One of a series of field notes on building software for the way minds actually work: tired, distractible, ordinary, and now partly machine. They all lead back to the manifesto behind them, The Shape of the System.