Let It Crash
The most reliable systems in the world stopped trying not to fall over, and got very good at standing back up.
A few minutes into the first Moon landing the computer that was flying two men down to the surface flashed an alarm and rebooted. Then it did it again. Five times in that final descent the machine gave up on whatever it had been doing, dumped it, and started again from scratch, and each time mission control told them to go, and they landed. The thing is, the computer wasn't broken when it rebooted. The reboot was the point. It had been given more work than it could get through, and instead of carrying on in some muddled half-finished state it threw away the jobs that didn't matter, held on to the few that were keeping the spacecraft pointed the right way, and came back from somewhere it knew was clean. So the most dangerous twelve minutes anyone had ever flown were got through by a computer that had been taught how to give up well. And what we tend to build is the other thing entirely. Software that will try anything rather than fall over, and then comes apart in some way nobody practised on the day it finally does.
Every error path is a small invitation to fix it right there where you're standing. Catch the exception, patch up the broken value, carry on. But each of those little rescues leaves the program a bit less certain about where it really is. Joe Armstrong spent a career on this and he made the point that the error-handling code can grow until it's as big as the actual logic it was supposed to be protecting, and the system you've built with all that work to never fall over turns out to be one that's quietly piling up states it was never meant to reach. A clean crash tells you one thing. It's gone, start it again. A caught-and-patched error tells you a thousand things, and you only find out which one you've got three steps later, somewhere downstream, in a state nobody can piece back together.
The heresy in Armstrong's idea was never really the word crash. It was about where the recovery lives. You don't fix the error inside the process that failed, because that process has just shown you it can't be trusted to reason about anything any more, and the thing it can reason about least of all is its own broken state. You let it die. The fixing gets done somewhere else, by a separate process that's still healthy and that was keeping an eye on things. Detection and recovery come out of the part that broke and move up to something standing over it. People miss that bit when they repeat the slogan. It doesn't mean be reckless. It means stop asking the patient to do their own surgery.
And there's a sharper point under all of this. If the only way you've got to stop your software is to crash it, then the crash is the best-tested path in the whole thing, and recovery is just your normal startup happening a bit more often than usual. Most systems are the wrong way round. They've got a graceful shutdown that someone has lovingly looked after and a crash path that's never once been rehearsed, which means the one road they're certain to go down when something really fails is the road nobody trusts. George Candea and Armando Fox gave the alternative a name, crash-only software. Make killing the thing the only way to stop it, so that coming back from a kill is a road you drive every day, in daylight, with nothing wrong. A recovery you do all the time is worth more than an uptime you cross your fingers over.
Two things have to be in place or none of it works, and the first is isolation. The part that crashes has to be walled off so that when it dies it doesn't slop its failure all over the things sitting next to it. One part goes, the whole doesn't. The second thing is a supervisor. Something outside the broken process that spots it has died, starts it again from a state known to be good, and works out whether to sound the alarm or just quietly get on with it. Arrange a load of these in a tree and you've got the shape of the most resilient software that's ever shipped. But the slogan leaves out something you have to be honest about, which is that a restart only rescues you from the random and the temporary. There was another machine, Mars Pathfinder, sitting in the dust back in 1997 rebooting itself again and again, and that was because the fault underneath it was a design flaw and not some passing glitch. Restarting a design flaw just gives you the same crash again, round and round, for ever, on another planet. Crash-and-restart buys you resilience against bad luck. It does nothing for you when the trouble is your own logic.
The goal, then, was never zero failures. Those famous reliability numbers people quote about these systems, the long runs of nines, didn't come from machines that never went down. They came from machines built to go down well. To fail in small pieces you could survive, and to come back so fast and so dependably that from outside it looked like nothing had gone wrong at all. Reliability isn't some heroic never going down. It's how quickly and how surely you get back up. Your own body lives by the same deal. Something like fifty billion of your cells quietly kill themselves every single day, each one folding itself up and getting carried off without bothering the cells around it, and the whole keeps running. Death, when it's done properly and cleanly and on a schedule, isn't the system failing. It's part of how the system keeps itself alive.
In the manifesto, this is tenets (XIII) and (XIX).
Sources
- [Alberts et al. 2002] Bruce Alberts et al., "Molecular Biology of the Cell" (4th ed.), 'Programmed Cell Death (Apoptosis)'. Garland Science, 2002. https://www.ncbi.nlm.nih.gov/books/NBK26873/. Healthy tissue destroys huge numbers of its own cells on a schedule to keep the whole alive; tenet XIX.
- [Armstrong 2003] Joe Armstrong, "Making Reliable Distributed Systems in the Presence of Software Errors" (PhD thesis). KTH, 2003. https://erlang.org/download/armstrong_thesis_2003.pdf. Let-it-crash: detect and recover in a separate supervising process, isolate failures, arrange supervisors in a tree; tenets XIII, XIX.
- [Candea & Fox 2003] George Candea & Armando Fox, "Crash-Only Software". HotOS IX, 2003. https://www.usenix.org/conference/hotos-ix/crash-only-software. Make crashing the only way to stop, so recovery is the routine, well-tested path; tenets XIII, XII.
- [Hall 1996] Eldon C. Hall, "Journey to the Moon: The History of the Apollo Guidance Computer". AIAA, 1996. https://en.wikipedia.org/wiki/Apollo_Guidance_Computer. The restartable, priority-scheduled executive shed low-priority jobs under the radar overload and kept guidance running through the 1201/1202 alarms; tenets XIII, XIX.
- [Reeves 1997] Glenn E. Reeves, "What Really Happened on Mars?". JPL, 1997. https://www.cs.unc.edu/~anderson/teach/comp790/papers/mars_pathfinder_long_version.html. The Pathfinder resets came from a priority-inversion design fault, so restarting alone could not fix it; tenet XIII.
One of a series of field notes on building software for the way minds actually work: tired, distractible, ordinary, and now partly machine. They all lead back to the manifesto behind them, The Shape of the System.