When the Code Runs Out
Types and tests guard the ground inside the compiler; past its edge, the only thing between you and disaster is a process someone bothered to build.
Take the best pilot in the world. Ten thousand hours in the seat, the sort of person who could fly the thing with his eyes shut, and he still stops before takeoff and reads a card out loud. It isn't that he's forgotten how to fly. The people who built modern aviation worked out something the hard way, and it gets relearned every so often and the relearning costs lives: the steps a system leans on are the very steps a competent person will skip now and then, on a perfectly ordinary day, and the better they are at the job the more certain they'll be that they didn't skip anything. Software has a border like this too. There's a line where skill stops being enough, and most of us step over it most weeks without even noticing we've left the safe ground behind.
Types and tests put up a fence. Inside it a whole class of mistakes just can't happen, because the compiler won't let them, and a decent test suite keeps hold of every decision you ever bothered to write down. That's real safety and a lot of careful work sits quite happily behind that fence. But it has an edge. The minute a step lives in someone's head, or in the order two people happen to do things, or in a deploy you can't take back, or a change made straight against the live system, the compiler hasn't got an opinion. You've walked off the end of what code can hold for you. Out there, being clever and being careful are about all you've got left, and the trouble is they both give out at the exact moment you need them most.
The things that go wrong out there are hardly ever strange. It's the step everyone knows by heart. The migration that ran before the backup had actually finished. The flag flipped on the wrong environment because two terminal windows looked the same. The delete you can't undo, fired at the live database and not the copy. Nobody had forgotten their job. They did the thing they'd done a thousand times before, except on this one run the attention slipped for a second or two, and there was nothing underneath to catch it. Skill doesn't help you here. Skill is the voice telling you the check is beneath someone like you.
On the thirtieth of October 1935, at Wright Field in Ohio, the United States Army watched the future of air power kill its own test pilot. The Boeing Model 299 was the prototype that became the B-17, the most capable bomber anyone had put together, and it took off, climbed a few hundred feet, stalled, came back down and burned. Two of the five men aboard died. It wasn't the aircraft. Someone had left a lock engaged on the control surfaces - a guard meant to hold them still while the thing sat on the ground - and in the rush of a first flight nobody had taken it off. The machine was so far past anything before it that the papers asked out loud whether it was simply too much aeroplane for one man to fly. The people who built it landed on a better answer. The aeroplane wasn't too complex to fly. It was too complex to leave to a pilot's memory. So they wrote the pre-flight checklist, a plain card listing the things you do every single time, and the aircraft went on to fly something close to two million miles without a serious accident. A checklist isn't an insult to an expert. It just takes the view that expertise and memory aren't the same thing, and that the second one fails on a schedule the first can't do anything about.
Now, an objection is probably forming already, and it deserves a straight answer, because most of the process a working engineer has run into has properly earned the contempt. The approval that's only there so somebody has cover. The form that protects the org chart. The sign-off where nobody is actually asked to look at anything. That's bureaucracy and it's worthless, and the risk is that we let it sour the whole idea of process, badly enough that we talk ourselves clean out of the one sort that would have saved us. There's a simple test that sorts the two. Does the step stop a real failure that the code can't, or does it just pass the blame around once one's already happened? Good process gets built with the same care you'd give a function. It has an input, a named failure it's there to stop, and a cost you can point to, and a bad one gets debugged and deleted the same way bad code does. A twenty-line checklist nobody reads is no more use than a twenty-line function nobody calls.
You could write off that 1935 crash as a charming one-off, except the lesson keeps turning up again in the most expert work people do. In 2009 a team published the results of a checklist. Nineteen plain items, read aloud at three pause points, tried across eight hospitals from Toronto and London out to New Delhi and rural Tanzania. Major complications after surgery dropped from eleven in a hundred to seven. Deaths fell from one and a half in a hundred to under one. Nineteen items, read out loud, in operating theatres run by some of the most trained people alive. There's a quieter version of all this in software, and it's about the best process we've actually got: a second pair of eyes on a dangerous change before it ships. Not review as a rubber stamp. A fresh person who hasn't been staring at the same diff for six hours, and who catches the dropped "not", or the live hostname sitting where the test one should be, or the missing way back, and catches it precisely because they aren't tired in the same spot you are. The runbook is that same idea once more, the checklist for the job at three in the morning when no amount of fresh memory is going to help. Where the stakes are high and the one doing the work is a human being, you put some structure between the person and the thing that can't be undone.
What you want to do after a near miss is promise yourself you'll be more careful next time, and that promise is worth nothing, and somewhere in you you already know it, because the next failure is going to land on a day you were already trying as hard as you possibly could. The answer isn't more effort. It's to build the thing that doesn't need you sharp. The step written down so your memory stops carrying the weight. The second reader you have to have, so one tired pair of eyes is never the last line standing. The irreversible action put behind a confirmation you have to mean, or a way back that's been tested. The code stops at a fixed border. What you put on the far side of it isn't heroism, it's just a card you read out loud on the days you're sure you don't need to.
In the manifesto, this is tenet (XXV).
Sources
- [Beyer et al. 2016] Betsy Beyer, Chris Jones, Jennifer Petoff & Niall Richard Murphy (eds.), "Site Reliability Engineering: How Google Runs Production Systems". O'Reilly, 2016. https://sre.google/sre-book/table-of-contents/. Runbooks, on-call ownership, and the human process around the system; tenets VI, XVIII, XIX, XXV.
- [Boeing 299 1935] Boeing Model 299 (XB-17 prototype) crash, Wright Field, 30 October 1935 (gust locks left engaged). https://en.wikipedia.org/wiki/Boeing_B-17_Flying_Fortress. Too complicated to leave to memory; the origin of the pre-flight checklist; tenet XXV.
- [Gawande 2009] Atul Gawande, "The Checklist Manifesto: How to Get Things Right". Metropolitan Books, 2009. https://en.wikipedia.org/wiki/The_Checklist_Manifesto. The discipline of writing the steps down so they survive the worst day; tenet XXV.
- [Haynes et al. 2009] Alex B. Haynes et al., "A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population". New England Journal of Medicine 360(5), 2009. https://doi.org/10.1056/NEJMsa0810119. The WHO checklist cuts complications and deaths across eight hospitals; tenet XXV.
One of a series of field notes on building software for the way minds actually work: tired, distractible, ordinary, and now partly machine. They all lead back to the manifesto behind them, The Shape of the System.