Prioritization with evidence, not with AI's guesswork
AI builds the RICE spreadsheet and calculates the score in seconds, but it also fills in reach and impact with a number that 'seems reasonable' when you don't supply the real data. The premise is yours, and every number behind the prioritization is suspect until audited.
You open AI to prioritize next quarter's roadmap, and in minutes you have a RICE ranking all ready for tomorrow's meeting. The scores add up, the spreadsheet looks pretty, everything seems ready to become a decision. The problem shows up when you look closely: the forty percent user reach it assigned to a feature didn't come from your Amplitude, it came from "seems reasonable for this kind of feature". AI builds the entire spreadsheet in one request; the premise behind each number is still yours, and it's what decides whether the meeting starts from real data or from a well-formatted guess.
A manager asked AI for a twelve-month cash flow projection, with an optimistic, base, and pessimistic scenario. In the pessimistic scenario, it assumed a default rate that "seemed reasonable", but the company was already running at double that this month.
You're deciding whether to enter a new market now or wait two years, and you ask AI for three market-sizing scenarios for the board deck, with a conversion rate that "seems reasonable" for some generic market.
You ask AI to prioritize which campaigns to relaunch next quarter based on expected ROI, and it assumes a paid-media conversion rate that's a market average, not your real funnel's.
You're about to open two positions and ask AI to project cost and time to hire, assuming in the base scenario an offer-acceptance rate that doesn't match your real recruiting history.
You open AI to prioritize next quarter's roadmap, and in minutes you have a RICE ranking all ready for tomorrow's meeting: reach, impact, confidence, and effort, calculated and ranked. The problem shows up when you look closely: the forty percent of the base it attributed as reach to feature X didn't come from your Amplitude, it came from "seems reasonable for this kind of feature in similar products". The eighty percent confidence it gave feature Y also has no backing at all, it was filled in because the PRD's text "sounded convincing". The whole spreadsheet closes nicely and lies in at least two columns at once.
You're closing the quarter's forecast and ask AI to project the pipeline, applying a proposal-to-close conversion rate that's a market average, not your CRM's real one.
You're about to decide whether to bring a shipping process in-house, and you ask AI for three cost and SLA scenarios, assuming a damage rate that "seems reasonable" for a healthy operation, not for yours, which is running worse.
You need to size the risk of a regulatory exposure, and AI assumes a moderate enforcement probability for a generic case, without knowing your real history with the regulator.
You're about to decide between migrating to new infrastructure now or in six months, and AI assumes a cost drop that has no relation to your contracts or your real traffic volume.
You're about to decide whether to redo the checkout flow, and AI assumes the new flow reduces abandonment by a proportion that "seems reasonable", without having run any test with your real users.
You're deciding whether to enter a new market now or wait two years, and you ask AI for three market-sizing and penetration scenarios for the board deck. Where did the conversion rate it assumed come from? It doesn't have your real data.
Man, think about it for a second. Prioritizing the roadmap was always the part that generated the most argument: open the spreadsheet, debate reach, fight over impact, and in the end you had a ranking nobody was sure reflected reality. AI really changes this game. It builds the RICE spreadsheet, calculates each item's score, and even writes the justification while you have your coffee. The problem is the speed is deceptive. It hands you the whole quarter's ranking in the time it takes for a single item, but if the reach or impact premise is made up, you didn't speed anything up. You just made the mistake faster, with a pretty spreadsheet on top.
The core idea of this lesson. AI is the best prioritization-matrix builder you've ever had, but it doesn't own the data. In roadmap prioritization, the rule is double and non-negotiable: the reach, impact, and confidence premises are YOURS, anchored in real data from your product, not guesswork; and every number AI spits out is suspect until you audit it against the source. The machine calculates; you decide what each variable means and sign off.
01The choreography: who does what
Prioritization isn't guesswork, it's a dance with defined steps. AI is good at three things, and only those three. It structures the model: it builds the spreadsheet with the right RICE columns (reach, impact, confidence, effort). It calculates the score: from the premises you give it, it runs the formula and orders the ranking. And it writes the explanation: it turns the dry spreadsheet into a paragraph the team understands. That's the grunt work of modeling, and the machine does in minutes what used to take you the afternoon.
What's still yours, and gets more expensive precisely because the easy part disappeared, is the premise. How many people does this feature really reach, measured in your analytics? What's the expected impact, based on what evidence? How much confidence do you have in that number, and why? That's not arithmetic, it's a reading of your product. AI doesn't have your Amplitude open, and when it doesn't, it makes up a number that "seems reasonable".
02Sensitivity: not every RICE column weighs the same
Here's the question that separates people playing with a spreadsheet from people prioritizing with method. RICE's four variables (reach, impact, confidence, effort) don't weigh the same in reliability. Reach usually has real data available: your Amplitude knows how many users pass through that flow. Effort usually has a reasonable estimate from the technical team. Impact and confidence are the two variables easiest for AI to fill in with a number that "seems reasonable", because they're the most subjective and the ones with the least auditable source directly in the dashboard.
The choreography is simple: pin down what has real data (reach, effort) and question more rigorously what's an estimate (impact, confidence). Ask "where did this ten percent impact come from, is it in some previous test or is it a guess?". If the answer is a guess, mark that number as the ranking's weakest, and that's where you apply the most scrutiny before deciding.
03The central danger: AI makes up reach and impact with the look of certainty
Now the point that names the risk, and that you need to carry with you, engraved. AI doesn't fail like a broken calculator, which freezes and warns you. It fails with confidence. It spits out a forty percent reach, well formatted, inside a confident sentence, and that number simply didn't come from your analytics. It assumed a "reasonable" impact because the PRD "sounded convincing", while the real data, had you checked Amplitude, would show something else.
This is where product's frame becomes non-negotiable. A roadmap built on top of a made-up number isn't just a wasted meeting, it's months of engineering effort aimed at the wrong problem. That's why the rule is harsh and simple: every number behind a RICE score is suspect until you audit the math. If a feature's reach is listed as forty percent, open Amplitude and check how many users actually pass through that flow. If it doesn't match, the whole ranking might be prioritizing the wrong thing.
Learn more: why "confidence" is the most dangerous variable to outsource
In RICE, the confidence column exists precisely so you can admit when the data is weak. It's tempting to let AI fill in this column too, because it always sounds confident when it writes. But high confidence assigned by AI to an untested assumption is the worst kind of mistake: it pushes up the priority list an item that should actually be marked as uncertain. The practical rule is: if you can't point to where the confidence number came from, the correct value is low, not whatever AI suggested.
Do it now
Take a real prioritization that's on your desk (your real task works well): two or three roadmap items competing for the next cycle. Run the choreography in three steps:
- YOUR PREMISE: for each item, write by hand the real reach (check the analytics), the expected impact and where that expectation comes from, and the effort estimated by the technical team.
- RANKING: ask AI to build the RICE spreadsheet and calculate the score using EXACTLY your numbers. Say it explicitly: "use these values, don't estimate any of them yourself".
- AUDIT: pick the item at the top of the ranking and check, at the source, whether the reach and impact match what you supplied. Only after that does the ranking get permission to become a roadmap decision.
If you got stuck on step 1 because you didn't have the reach data on hand, great: that was exactly the expensive part of the work, and now you know it.
Practice
1. In RICE prioritization, which variable usually has the most reliable data, straight from your analytics?
2. AI returned a RICE ranking where a feature has 'confidence: high', but you can't point to any test or data backing that value. What's the correct rule?
For the board
On the columnsnot every variable carries the same weight. Reach usually has a real source; impact and confidence are the most fragile.
On confidenceif you cannot point to where the number came from, the correct value is low.
On the effectinflated confidence pushes to the top of the queue an item that should be flagged as uncertain.
Thanks for the feedback. It helps sharpen the next lesson.