Blast Radius
Everything you build will fail one day. The only thing you get to decide in advance is how far the wreckage spreads.
A fuse and a house fire start out the same: too much current going through a wire that's too thin to carry it. What sets them apart is that with a fuse, someone decided ahead of time, long before any spark, which bit of metal ought to be the one that gives way, and they pointed it at nothing much so that when it went it went harmlessly. So a fuse is really just a fire that got planned for and sized and aimed at the floor. And a fire is a fuse that nobody bothered to fit. The physics is the same and the spark is the same. The two mornings are not the same at all.
Software has no shortage of sparks. Somebody types a command and gets it slightly wrong. An eleven-line helper. One server out of a fleet of eight that didn't get the message everyone else got. You can't stop the sparks coming, there's always going to be another one. What you do get to decide is what happens in the half second after one jumps, and the strange thing is you decide that well in advance, in a hundred little choices about what's allowed to take down what.
Every postmortem goes off hunting for the root cause, as if naming the spark told you anything about the fire. It doesn't. A dropped match in a wet field is nothing, nobody writes it up. The same match in a dry barn burns the barn down, and the thing that made the difference wasn't the match. It was the barn. We spend nearly all our attention trying to stop sparks, which can't be done, and almost none on the one question we could actually sit down and answer on a quiet afternoon, which is: when this thing catches, and it's going to catch, how far does it get to run before something stops it.
You don't get to choose whether your components fail. What you choose is what shares fate with what. A network call you wrote without a deadline on it is a promise to wait forever for an answer that might never come. A credential that's broader than the job it's doing is a key that opens more doors than it needs to. Every customer sharing one unguarded queue with every other customer is a wall you decided not to build. Your system has a blast radius right now, this minute, and you didn't sit down and set it deliberately. It got set for you, by a hundred defaults nobody ever stopped to argue about. The work isn't stopping the spark. It's drawing the lines that the fire isn't allowed to cross.
And those lines are specific things you can point at and name, not some general air of being a good engineer. A timeout is a wall. It stops one slow dependency from holding all your threads hostage while it sits there thinking. A limit on what a single caller is allowed to ask for is a wall, so that no one request can demand a gigabyte of memory and bring the whole box down with it. Least privilege is a wall too. It's the difference between somebody nicking a hotel room key and somebody nicking the master key the cleaner walks around with. Both of them get stolen sooner or later. Only one of them ends up as a press release. Keeping one customer's work off another customer's queue is a wall: when a job runs away with itself it stalls the person who started it and leaves everybody else alone.
Leave the walls out and ordinary helpfulness turns into something lethal. Think about a nice friendly shared pool of fifty threads that every outbound call borrows from. One day a downstream service slows right down to a crawl. The calls to it don't fail, they just wait, and one after another all fifty threads end up parked against that one slow service, and every other feature in the application, none of which had anything to do with it, goes dark at the same time. The slow service didn't take you down. Your decision to let it borrow all your threads did that.
The outages that cost the most are hardly ever the big dramatic fault you'd have braced for and defended against. They're some small thing with nothing built around it. Knight Capital, a trading firm, pushed new code out one August morning in 2012, to seven of its eight servers. The eighth one still had a routine on it that had been sitting dormant since 2003, and a reused configuration flag woke the thing up. The market opened and that one machine started firing orders into it by the million, and there was no limit sitting anywhere that could say "this is obviously wrong, stop now". In something like forty-five minutes the firm lost about four hundred and forty million dollars. It wasn't finished off there and then, the way people usually tell it. It got rescued a few days later and then swallowed by a competitor before the year was out. But the lesson is the eighth server. One box with no wall round it and nobody's hand near the switch, and that was the whole disaster.
Tiny things travel just as far as big ones. In March 2016 a developer unpublished a package of eleven lines that padded out a string, and builds broke right across half the open-source world inside a few minutes, because thousands of projects were leaning on it without ever knowing they were. Eleven lines turned out to have an enormous blast radius, because there was nothing at all sitting between those lines and everybody who depended on them.
Worst of all is the kind of failure that eats its own cure. In February 2017 an engineer at Amazon was debugging a billing system and ran a command to pull a few servers out of service, and got one of the inputs wrong, so a far bigger set of servers went down than was meant to. As it happened that was the storage a wide stretch of the web had quietly been built on top of, and getting it back took something like four hours. Here's the detail worth holding onto. Amazon's own status dashboard, the page the entire world refreshes to work out whether the problem is them or the cloud, was itself being served out of the very storage that had gone down. So the one page whose whole job was to announce the outage couldn't be updated to admit there was one. Recovery has a blast radius of its own. When Facebook withdrew the wrong routes during routine maintenance back in October 2021 and rubbed itself off the internet for around six hours, the engineers who could have put it right reportedly couldn't badge in through the doors, because the locks ran on the same network that had just vanished. If your alarms and your dashboards and your deploy pipeline and your actual building all hang off the thing that's breaking, then you haven't built a system with a fuse in it. You've built one where the fire gets to the fire brigade before anyone else.
Containment isn't only walls. It's also deciding in advance how you're going to bend. A system that gets asked for more than it can give can drop an arm and keep its heart beating. Switch off the recommendations, say, but still take the payment. Or hand back data that's a bit stale instead of no data at all. Falling back to read-only is better than falling over completely. The other option, the one where you don't bend, is the stampede. One service hiccups, and a thousand clients all decide to have another go in the same instant, and that second wave, all of it arriving together, is what actually kills the thing. The first fault was survivable. The rescue attempt was the murder. Getting everyone to stagger their panic instead of all retrying at once is the whole of what backoff and jitter were ever about. And whatever else you do, make the failure loud. A wall that gives way silently saves nobody, because nobody comes running until the damage is already done.
You won't be stood over the system, rested and sharp, on the morning it all goes wrong. It'll pick the worst possible moment and the corner nobody documented and the most junior person on the team, and that's just how the base rate works out. So build for that morning now, while it's still a calm Tuesday and you've got time. Sit with what might be the cheapest question in the whole of engineering, the one almost nobody asks until it's already too late: when this one piece fails, how much does it drag down with it. The fuse costs you next to nothing today. The fire costs the whole company.
In the manifesto, this is tenets (IV), (VI), (VII), (XVI) and (XIX).
Sources
- [AWS 2019] Amazon Web Services, "Reducing blast radius with cell-based architectures" (re:Invent 2019, ARC411-R). 2019. https://d1.awsstatic.com/events/reinvent/2019/REPEAT_1_Reducing_blast_radius_with_cell-based_architectures_ARC411-R1.pdf. 'blast radius' and cell isolation as fault-isolation; tenets XIX, PREAMBLE.
- [AWS S3 2017] "Summary of the Amazon S3 Service Disruption in the Northern Virginia (US-EAST-1) Region". Amazon Web Services, 2017. https://aws.amazon.com/message/41926/. An input entered incorrectly in a foundational service cascades across the internet, even taking down AWS's own status dashboard; tenets XIX, XIII.
- [Brooker 2015] Marc Brooker, "Exponential Backoff And Jitter" (AWS Architecture Blog). 2015. https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/. Jittered backoff so synchronised retries don't herd; tenets VII, XIX.
- [Knight Capital 2012] SEC, "Order Instituting Administrative and Cease-and-Desist Proceedings: Knight Capital Americas LLC". SEC Release No. 70694, 2013. https://www.sec.gov/litigation/admin/2013/34-70694.pdf. Dormant Power Peg code revived by a reused flag; about 45 minutes of unstoppable erroneous orders (the SEC order puts the loss near $460M; Knight's reported pre-tax loss was about $440M); tenets XIX, VII.
- [left-pad 2016] npm left-pad incident. 2016. https://en.wikipedia.org/wiki/Npm_left-pad_incident. Unpublishing eleven lines breaks Babel, React, and thousands of builds; a trivial dependency with an outsized blast radius; tenets XIX, VII.
- [Meta BGP 2021] "More details about the October 4 outage". Engineering at Meta, 2021. https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/. A backbone maintenance command withdraws every BGP route, disconnecting Facebook's data centres and DNS from the internet; tenet XIX.
- [Nygard 2007] Michael T. Nygard, "Release It!: Design and Deploy Production-Ready Software". Pragmatic Bookshelf, 2007. https://pragprog.com/titles/mnee2/release-it-second-edition/. Timeouts, Circuit Breaker, Bulkheads, Fail Fast, and the blast-radius/cascading-failure framing; tenets VI, VII, XIX, V, XIII.
- [Saltzer & Schroeder 1975] Jerome H. Saltzer & Michael D. Schroeder, "The Protection of Information in Computer Systems". Proc. IEEE 63(9), 1975. https://web.mit.edu/Saltzer/www/publications/protection/. Least privilege, fail-safe defaults, complete mediation, separation of privilege; tenets XVI, IV, XXI, XI, XXV, PREAMBLE.
One of a series of field notes on building software for the way minds actually work: tired, distractible, ordinary, and now partly machine. They all lead back to the manifesto behind them, The Shape of the System.