A four-box review that separates a bad decision from bad luck

By Meet Patel · 2026-10-03 · 6 min read

Summary

Resulting is judging a decision by its outcome. To separate decision quality from outcome, rate the decision from the pre-decision record, place it in a four-box review (good or bad decision, good or bad outcome), and check calibration across at least 20 decisions.

Key Metrics & Takeaways

26 seconds, 1-yard line
the situation when Pete Carroll called the pass in Super Bowl XLIX, with Seattle trailing 28 to 24 and holding one timeout (Super Bowl XLIX summary, Wikipedia)
692 participants
in a replication of Baron and Hershey's 1988 outcome-bias experiment, which reproduced the effect
About 35%
chance of 4 or more failures in 10 independent bets at 70 percent success each (binomial arithmetic, worked illustration)

With 26 seconds left in Super Bowl XLIX, the Seattle Seahawks trailed the New England Patriots 28 to 24. They had the ball at the 1-yard line on second down with one timeout and Marshawn Lynch in the backfield. Coach Pete Carroll called a pass, Malcolm Butler intercepted it at the goal line, and the game was effectively over (Super Bowl XLIX summary). Annie Duke recalls the headlines the next day as “Worst play in Super Bowl history.”

Duke, a former professional poker player, uses the call as the lead example in Thinking in Bets, and her argument is worth taking out of football. The result of a single play is a poor measure of the quality of the call, and most companies review their decisions as if it were a good one.

What resulting is

Duke’s term comes from poker. In her words: “Poker players have a word for this: ‘resulting’.” Experienced players warned her against changing strategy because a few hands went badly in the short run (quoted from Thinking in Bets). Her point is that over a short run luck dominates what you see, so the outcome carries little information about how well you played.

Duke sets Carroll’s call beside a different play. The Philadelphia Eagles’ successful trick play in Super Bowl LII drew praise as a stroke of genius, and she argues that the two calls were similarly unexpected, creative and well reasoned, with one result for each. The reaction to each followed the result, and the quality of the thinking had little to do with it (Annie Duke, on judging decisions only by outcomes).

The bias has a research record. In 1988 Jonathan Baron and John Hershey gave people descriptions of medical and monetary decisions made under uncertainty, with the outcomes attached. Participants knew they had all the information the decision maker had. They still rated the thinking as better, and the decision maker as more competent, when the outcome was favorable (Baron and Hershey, “Outcome bias in decision evaluation,” Journal of Personality and Social Psychology, 1988). A replication with 692 online participants reproduced the effect, including among people who said outcomes should not affect their judgment (replication of Baron and Hershey Experiment 1).

How resulting damages a company

A review is training data for the next decision. Take a hypothetical product team that enters a new customer segment after estimating a 60 percent chance of hitting its first-year target. The team reaches 40 percent of the target and the quarterly review treats it as a failure of judgment. The lesson the team absorbs is to avoid 60 percent bets. A few quarters later it proposes only bets that look like 90 percent, or it states 90 percent when it believes 60.

Both adaptations are costly. The first removes the portfolio of uncertain, high-payoff choices that a company needs, which is the territory covered in the asymmetric bet framework. The second corrupts the probabilities, and probabilities are the raw material for every later review.

The reverse error is quieter. A launch that succeeds because a competitor stumbled gets celebrated, the team gets promoted, and the process that produced it gets copied. Nobody asks whether the same process would work again.

The four-box review

I would use a two-by-two to separate the decision from the result. One axis rates the decision, judged on what was knowable when it was made. The other records the outcome. Each box has a different action.

Two hypothetical placements show the method at work. A team raises prices 8 percent for new customers after a test across 400 accounts showed no drop in conversion. Churn rises that quarter because a competitor ran a heavy discount. The test was sound, the downside was sized, and the competitor’s move was unforecastable, so this is a good decision with a bad outcome, and the pricing process stays. A second team skips the test because the founder feels sure, and revenue rises because of a seasonal spike. The outcome is good and the decision was a guess, so it goes in the dangerous box, and the next price change gets a test.

The weight of the method sits in how you rate the decision. I would use five questions, answered only from the record made before the outcome:

  1. What was known at the time, and what was not?
  2. Which alternatives were considered, and why were they rejected?
  3. What probability of success was stated, and does it fit the evidence listed?
  4. Was the downside sized so that the company survives the bad case?
  5. Was the decision made by the named owner, by the date?

This needs a record written at the time, which is why the sibling post on keeping a decision log matters. Memory rewrites the past once the result is known, so a review that depends on recollection will always drift toward the outcome.

A test and a number

Two checks keep the four-box honest, since “bad luck” is an easy label to apply to anything that hurt.

The swap test. Take the decision record and imagine the opposite outcome. If you would score the decision the same way, you are rating the decision. If your score moves with the imagined result, you are rating the outcome. Have a second person rate the record without being told how it ended, and compare scores.

The base-rate check. Suppose a team makes 10 independent bets in a quarter and honestly estimates each at a 70 percent chance of success. A perfectly calibrated team still fails on 3 of them on average. The binomial distribution says that 4 or more failures will occur about 35 percent of the time, and 3 or more about 62 percent of the time. So in roughly one quarter out of three, a team with perfect judgment looks like it had a bad quarter. These figures illustrate the arithmetic and describe no real team.

The practical consequence is that a single outcome says very little, and a batch says a lot more. Review calibration across at least 20 decisions: group them by stated probability, then compare the stated numbers with the share that worked. If the 70 percent bets succeed 70 percent of the time, the team’s judgment is sound, whatever happened last quarter. If they succeed 45 percent of the time, the judgment needs attention, even if last quarter went well.

Running the review

A monthly review of this kind takes about an hour with three ground rules.

First, score the decision before the outcome is discussed. The person presenting reads out the pre-decision record and the group scores it on the five questions, then the outcome is revealed. Second, put each reviewed decision into one of the four boxes and write a single action beside it, which for the good decision, bad outcome box is usually “nothing”. Third, keep a running table of stated probabilities and results, so that the calibration check becomes a habit.

Unmade decisions cannot be reviewed at all, which is the cost described in the decisions your company is not making. A company that scores decisions also has to make them on a date, with a written estimate, or there is nothing to score.

None of this excuses bad calls. Luck is a hypothesis, and the swap test and the calibration table are how you test it. What the method removes is the reflex to learn from the result alone, which in a short run teaches the wrong lesson about as often as the right one. A review that cannot tell luck from skill will punish good bets and reward lucky ones.

Perspectives

“Poker players have a word for this: 'resulting'.”

— Annie Duke, Author, Thinking in Bets; former professional poker player

Frequently asked questions

What is resulting in decision making?

Resulting is Annie Duke's term, borrowed from poker, for judging the quality of a decision by how it turned out. A good decision can produce a bad result through luck, and a poor one can get lucky. Resulting leads teams to abandon sound processes after one bad outcome and to copy careless ones after one good outcome.

How do you separate decision quality from outcome in a review?

Rate the decision from the record written before the outcome was known, using questions about what was known, which options were weighed, the stated probability and the size of the downside. Then record the outcome separately. Placing each decision in a four-box grid shows whether a bad result was bad luck or a bad process.

How many decisions do you need to judge whether a team has good judgment?

A single outcome says very little. Even a perfectly calibrated team that makes 10 bets at 70 percent odds will see four or more failures about 35 percent of the time. Compare stated probabilities with actual results across at least 20 decisions before drawing a conclusion about the quality of the team's judgment.

Sources

Written by Meet Patel — startup operator and growth strategist in Dubai.

Read on themeetpatel.com