How to Prioritize A/B Tests (PIE and ICE Frameworks)
You can only run so many tests a year, so picking the right ones is everything. How to score a test backlog with PIE and ICE, ground it in research, and stop testing things that never mattered.


If your store can only run a handful of meaningful tests a year, the question that decides your results isn't "how do we test?" - it's "what do we test first?" Most teams answer it with whoever argued loudest in the meeting. There's a better way.
Why prioritization is the whole game
Every test costs the same precious thing: a slot in your limited annual bandwidth. Spend it on a homepage hero tweak that moves nothing and that's a month you won't get back. Spend it on the checkout friction quietly killing 20% of orders and you fund the next quarter. The teams that win at CRO aren't the ones with the cleverest variants - they're the ones that consistently test the right things.
PIE: Potential, Importance, Ease
PIE scores each idea on three dimensions, usually 1-101:
- Potential - how much room to improve is there? A page already converting well has less headroom than a leaky one.
- Importance - how valuable and trafficked is the page? A 10% lift on your top collection beats 10% on a page nobody visits.
- Ease - how hard is it to build and ship? Political and technical friction count.
Average the three for a score, sort the backlog, test top-down.
ICE: Impact, Confidence, Ease
ICE is the faster cousin, popular with growth teams2:
- Impact - how much could this move the metric?
- Confidence - how sure are you it'll work? This is where evidence earns its keep.
- Ease - how cheap is it to run?
Both frameworks force the same trade-off: upside versus effort. The clearest way to see it is a simple impact-vs-effort matrix - the top-left quadrant is where your limited slots belong.
The trap: scoring on opinion
Both frameworks share one weakness. Confidence and Potential are only as good as the evidence behind them. Score them on gut feel and you've just built a more official-looking version of the loudest-voice meeting.
That's why prioritization sits downstream of research, not instead of it. Before you score anything, gather:
- Analytics - where do users actually drop off in the funnel?
- Session replay and heatmaps - where do they hesitate, misclick, rage-click?
- Voice of customer - surveys, support tickets and reviews telling you what confuses or blocks people.
Now "Confidence" means something. A test backed by three independent signals pointing at the same friction deserves a high score. "I read it on a CRO blog" does not.
A worked scoring example
Three ideas, scored with ICE (1-10):
- Add trust badges to checkout - Impact 6, Confidence 8 (support tickets mention payment doubt), Ease 9. Score ≈ 7.7.
- Redesign the homepage hero - Impact 7, Confidence 3 (no evidence, just taste), Ease 3. Score ≈ 4.3.
- Simplify the 3-step form to 1 step - Impact 8, Confidence 7 (replays show form abandonment), Ease 5. Score ≈ 6.7.
The trust-badge and form tests rise to the top - not because they're flashy, but because evidence backs the confidence and they're cheap to run. The hero redesign, the idea most likely to dominate a meeting, sinks.
A simple workflow
- Research first. Mine analytics, replays and voice-of-customer for real friction.
- Turn findings into hypotheses. "Because [evidence], we believe [change] will cause [outcome], measured by [metric]."
- Score the hypotheses with PIE or ICE - whichever your team will actually keep using.
- Test top-down, respecting your bandwidth.
- Feed results back in. Every test - win or lose - teaches you something that re-ranks the backlog.
The takeaway
Prioritization isn't bureaucracy; it's how you make a limited number of tests count. Use PIE or ICE to force the trade-off between upside and effort, but ground the scores in research so "confidence" is earned, not assumed. Do that and your win rate climbs - not because you got luckier, but because you stopped testing things that were never going to matter.
This is exactly the process we run for ecommerce stores. See how our CRO program works.
References
- 1.Goward, C. (2012). PIE Prioritization Framework. Conversion (WiderFunnel). https://conversion.com/framework/pie-framework/
- 2.Ellis, S. (2017). ICE Scoring Model. GrowthMentor / GrowthHackers. https://www.growthmentor.com/glossary/ice-scoring-model/
Related research
Want a team to run this for you?
See how we help