Playbook

Kill it or scale it: a decision framework

Every creative review ends in the same argument, held with total confidence on both sides, and nobody has the number that would settle it.

7 min read · Updated September 2026

Every creative review ends in the same argument. Someone wants to kill an ad that's three days old and £200 in. Someone else wants to give it room. Both positions are defensible, both are held with total confidence, and neither side has the number that would settle it.

There is a number. It's uncomfortable.

Your test is almost certainly underpowered

The definitive work here is Lewis and Rao's The Unfavorable Economics of Measuring the Returns to Advertising, published in the Quarterly Journal of Economics in 2015. They examined 25 digital advertising experiments covering $2.8 million of spend. The median campaign reached over a million people.

3 of 25
campaigns had enough statistical power to reliably distinguish a 50% ROI from zero. None could reliably tell a 10% ROI from zero. Lewis & Rao (2015), QJE 130(4), 1941–1973.

The median experiment would need to be 9× larger to separate a 50% ROI from breakeven, and 62× larger to detect a 10-point difference in ROI.

The reason is variance, not sample size alone. In their worked example, a profitable campaign needed to detect a sales impact of about $0.35 per person against a standard deviation in individual sales of roughly $75. Individual spending is enormously noisy compared to the effect any single ad has on it — the standard deviation of individual-level sales typically runs about ten times the mean over a normal campaign window.

If campaigns reaching a million people can't resolve a 10% ROI difference, an ad with 40 conversions cannot tell you anything reliable about a 15% difference in CVR. That is not pessimism. It's arithmetic.

What this does and doesn't say

This is about measuring incremental return — the hard question. Relative comparisons between creatives running in the same account at the same time are easier, because much of the noise is shared and cancels out. The finding still holds directionally: small differences on small spend are not signal, and treating them as signal is how teams end up rotating creative on noise.

A framework that respects the noise

The practical move is to stop asking "is this ad good?" and start asking "do I have enough evidence to act?" — which has three possible answers, not two.

Not enough evidence yet

Below a meaningful spend and volume floor, the honest verdict is no verdict. The ad hasn't earned a decision. The failure mode here is asymmetric and expensive: killing a good ad on noise is invisible, because you never see what it would have done, while scaling a bad ad on noise shows up in the CAC a fortnight later.

A workable floor: the ad should have spent enough to produce at least 25–30 of whatever the campaign's primary event is, or roughly 2% of the account's spend in the window — whichever is stricter for your volumes.

Enough evidence, and it's decisive

Scale or kill. The evidence for scaling is a metric meaningfully ahead of the account average on real volume; the evidence for killing is a metric meaningfully behind it on spend large enough to matter. "Meaningfully" is doing work in both sentences — a 5% gap is not a finding.

Enough evidence, and it's mixed

This is the interesting bucket and the one most frameworks collapse. An ad with a strong hook and a weak hold is not a bad ad; it's a good opening attached to a weak middle. An ad with good CTR and poor CVR is usually a landing-page or offer problem wearing a creative costume. These get iterated, and the iteration should be specific — which is only possible if you know which component is failing.

Three failure modes this prevents

How Knack handles this

Knack never issues a verdict an ad hasn't earned. Creatives below the spend and volume floor are flagged too early to call rather than ranked, so a cheap winner on forty impressions is held back instead of promoted. Verdicts are graded against the metric each campaign's objective implies — ROAS and CAC for lower funnel, CTR and CPC for mid, hook and hold for upper — and every flag is a specific, checkable condition rather than a composite score, so you can see exactly why a creative landed where it did and disagree with it if you want to.

The uncomfortable conclusion

Most creative decisions are made on less evidence than the people making them believe. The fix isn't more dashboards. It's being explicit about which decisions the data can support, and being willing to say "we don't know yet" out loud in a meeting where everyone else is certain.

Sources

  1. Lewis, R. A., & Rao, J. M. (2015). The Unfavorable Economics of Measuring the Returns to Advertising. The Quarterly Journal of Economics, 130(4), 1941–1973. 25 experiments, $2.8m of spend. Full text (PDF).
  2. Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press — on statistical power and the cost of underpowered experiments generally.
Next up