Prioritization without RICE theatre: score evidence and reversibility
By Meet Patel · 2026-10-03 · 6 min read
Summary
RICE (Reach x Impact x Confidence / Effort) is useful, but its inputs are menu-picked estimates, so one notch of Impact can flip a ranking. Test by nudging Impact, then prioritize by evidence strength and reversibility: build strong and reversible, test weak and reversible.
A RICE spreadsheet ranks twenty ideas to one decimal place. The first row scores 360, the second 333, the third 300, and the roadmap meeting ends with a list that looks like a measurement. Underneath, every input was picked from a short menu. In Sean McBride's Intercom post, Impact has five allowed values and Confidence has three. A score with that much rounding behind it cannot support a ranking that fine, and the team often does not notice, because the output has a decimal point.
RICE is a sensible tool, and its author is more candid about its limits than most of the people who use it. This post explains where the scores turn into theatre, gives a five-minute test that shows whether your ranking carries information, and offers a lighter method based on how strong the evidence is and how reversible the decision is.
What RICE actually says
McBride's formula is Reach times Impact times Confidence, divided by Effort. Reach is how many people the idea affects in a defined period. Impact is chosen from a multiple-choice scale: massive is 3, high is 2, medium is 1, low is 0.5 and minimal is 0.25. Confidence is 100 percent for high, 80 percent for medium and 50 percent for low. Effort is in person-months, with a minimum of half a month.
His worked example has 500 customers a month reaching a step in the signup flow, 30 percent of whom choose a given option, which he turns into a reach of 450 customers per quarter. He gives that project an effort of 2 person-months and a confidence of 100 percent because it has quantitative metrics for reach, user research for impact and an engineering estimate for effort.
He also writes that “Impact is difficult to measure precisely” and, at the end, that “RICE scores shouldn't be used as a hard and fast rule.” His stated purpose for Confidence is “to curb enthusiasm for exciting but ill-defined ideas.” Those are fair statements. The trouble starts when a team keeps the arithmetic and drops the caveats.
Three ways the scores become theatre
One notch flips the order. Most steps on the Impact scale double or halve the value, so a single notch up or down can double or halve a score. Take two hypothetical ideas. Idea A reaches 450 customers a quarter, has impact 2, confidence 80 percent and effort 2 person-months, which scores 450 times 2 times 0.8 divided by 2, or 360. Idea B reaches 2,000 customers, has impact 1, confidence 50 percent and effort 3, which scores 2,000 times 1 times 0.5 divided by 3, or 333. A leads by 8 percent. If the group rates A's impact as 1 instead of 2, A falls to 180 and B leads by almost double. Impact is the input with the least evidence behind it, and it controls the result.
Small effort numbers dominate. Effort sits in the denominator, and the scale allows half a month. An idea scored at 0.5 person-months earns four times the score of an identical idea scored at 2. Cheap ideas rise to the top because they are cheap, and the list drifts toward polish while the larger bets sink.
Confidence is set by the person with the idea. A multiplier of 0.8 against 0.5 is a swing of 1.6 times, about the size of a rounding error on Impact. When the proposer picks both numbers, the multiplier mostly records how much they like the idea. Confidence only does its job if someone other than the proposer sets it from a defined evidence standard.
A five-minute test for your own ranking
Take the top five items on your list. For each one, move Impact up one notch and then down one notch, and recompute. If the order of the top five changes under that small nudge, the order is not information, and the useful output of the exercise was the conversation about why the numbers differ. If the order survives, the ranking is robust to the input you trust least, and you can use it with more confidence.
A lighter method: evidence strength and reversibility
Two questions carry most of what the formula was reaching for. The first is how strong the evidence is. The second is how hard the decision is to undo.
For evidence I would use a four-step ladder. Level 1 is an anecdote, such as one customer or one salesperson. Level 2 is a repeated pattern in qualitative research, such as five of eight interviews describing the same problem. Level 3 is behavioral data, such as funnel drop-off or usage logs. Level 4 is the result of an experiment on the idea itself. This scale is my own and has no published source, so adjust the levels to what your team can actually observe.
For reversibility I borrow Jeff Bezos's distinction in his 2015 letter to Amazon shareholders between one-way and two-way doors. He writes that some decisions are “consequential and irreversible or nearly irreversible” and that most decisions are “changeable, reversible.” The letter warns that applying heavyweight process to reversible choices produces “slowness, unthoughtful risk aversion, failure to experiment sufficiently.” I explored the related idea of capping downside in the asymmetric bet framework.
Together the two questions give four cases.
- Strong evidence, easy to reverse: build it now. A score would only slow you down.
- Weak evidence, easy to reverse: run the cheapest test you can finish within a week, then re-rank with the result.
- Strong evidence, hard to reverse: build it, with a staged rollout and one named owner for the final call.
- Weak evidence, hard to reverse: stop. The next piece of work is getting evidence, and the idea does not compete for build capacity until then.
Reach and effort stay on the page as plain numbers beside each idea and are not multiplied into a score. A reader can see that one idea touches 2,000 customers and costs 3 person-months without a formula doing the combining.
The same two ideas, re-read (hypothetical)
Return to ideas A and B. Suppose A rests on funnel data showing where 450 customers a quarter drop out, so it sits at evidence level 3, and it ships behind a flag, so it is easy to reverse. B rests on one enterprise prospect's request, which is level 1, and it requires a change to the data model that would be painful to undo. The rule says to build A now. B gets a cheap test, such as a prototype shown to five more customers, and does not join the build queue before that.
The ranking is the same as the RICE sheet gave on its first pass, but it no longer flips when someone changes a number from 2 to 1. Its reasons are written in words that the group can challenge.
Using it in the meeting
Three habits keep the method honest. Write the evidence level next to each idea along with the source, such as the interview count or the dashboard. Have someone other than the proposer assign it. And keep a short list of ideas that were stopped for weak evidence, so that the test results have somewhere to land. This fits with what I wrote in kill your roadmap, where bets are revisited as the market moves and are not defended because they were scored in January.
McBride closes his post by noting that a scoring system lets you see when you are making a trade-off, for instance when you work on a lower-scoring project first. That is the best use of any framework, including this one. The score exists to make a trade-off visible, and a decision that rests on written evidence and a stated cost of being wrong is one the team can revisit when the evidence changes.
Perspectives
“RICE scores shouldn't be used as a hard and fast rule.”
— Sean McBride, Product manager, Intercom (author of the RICE post)
“To curb enthusiasm for exciting but ill-defined ideas, factor in your level of confidence about your estimates.”
— Sean McBride, Product manager, Intercom (author of the RICE post)
Frequently asked questions
How is a RICE score calculated?
Multiply Reach, Impact and Confidence, then divide by Effort. In Sean McBride's Intercom version, Reach is people affected in a period, Impact is chosen from 3, 2, 1, 0.5 or 0.25, Confidence is 100, 80 or 50 percent, and Effort is in person-months with a minimum of half a month. McBride notes the scores should not be used as a hard and fast rule.
What is wrong with RICE scoring?
Its output looks more precise than its inputs. Impact and Confidence are picked from short menus, one Impact notch can double or halve a score, small effort values inflate cheap ideas, and the proposer often sets Confidence. A quick test is to move Impact up and down one notch for your top five items and see whether the order changes.
What is a simpler alternative to RICE?
Rank ideas by how strong the evidence is and how reversible the decision is. Build now when evidence is strong and the decision is easy to undo. Run a cheap test when evidence is weak but reversible. Stage the rollout when evidence is strong but hard to reverse. Gather evidence first when it is weak and hard to reverse.
Sources
Written by Meet Patel — startup operator and growth strategist in Dubai.