Judge any creative on eight criteria in a fixed order: the first two seconds, the claim, the audience fit, the proof, the product presence, the single action, the format fit, and the durability. The first six decide whether an ad can work at all, and the last two decide whether it survives being scaled. A scorecard cannot tell you which surviving ad will win, and that is not its job: its job is to stop taste deciding which ads get made, and to say which specific component failed when one does not.
- A winning ad is profitable volume over time, not a strong result in a three-day test.
- A scorecard rules ads out reliably and cannot predict which survivor wins.
- The first six criteria are close to necessary; the last two decide whether it scales.
- In a competitor library, score the longest-running ads, because duration is the only spend signal you can see.
- Rebuild the failing component rather than commissioning a new ad, or you discard what worked.
A winning ad is one that produces profitable volume at scale over time. That definition rules out most of what gets called a winner, because the usual evidence is a few days of spend on a small budget, a window narrow enough that noise regularly outperforms signal. Last updated: August 2026.
The reason a scorecard is worth having is that creative review without one is a taste conversation, and taste conversations are won by whoever is most senior. Omniconvert has measured what separates stores that compound creative performance from stores that restart every quarter, across 70,000+ experiments over 13 years in eCommerce, and the difference is rarely budget. It is that one group can say which component of a failed ad failed, and the other group can only say it did not work.
The Ecommerce Benchmark scores stores against real competitors in their category and country across six dimensions, and Creative and Ads is the one where the gap between teams is widest. This article is the instrument behind that dimension: eight criteria, applied in order, to any ad you are about to make or any competitor ad you are about to learn from. For the wider competitive method, see the DTC competitor research stack.
Why judging a winning ad by results alone fails
Consider what a purely results-based process teaches. An ad runs for four days, returns a cost per acquisition below target, and is declared a winner. The next three ads built in the same style return nothing. The team concludes that creative is unpredictable, which is the wrong lesson: the first result was probably within the range of variation for that sample, and no component of it was ever identified as the cause.
Now add a scorecard. The same ad is scored before it runs, so its components are on record. When it succeeds, you know which criteria it satisfied. When the next three fail, you can see that two of them lost the checkable claim and one lost the single action. The failures become informative, and the file starts to be worth something.
This is also the only defensible way to use a competitor library. Copying an ad that has run for six months gives you the surface. Scoring twenty long-running ads in your category tells you which criteria your buyers respond to, which transfers to creative you have not made yet.
The eight criteria
- The first two seconds. Can a stranger, with no sound and no context, name what this ad is about? Not what it is selling, what it is about. An opening that requires patience to decode has already lost most of the people it needed, and no strength later recovers them.
- The claim. Is one specific, checkable thing asserted? Lasts a season, fits a standard case, replaces three products, works without a subscription. A general claim of quality is not a claim, because nothing about it could be wrong, and an assertion that cannot be wrong cannot persuade.
- The audience fit. Does it address one identifiable buyer situation? An ad for everyone who might buy is written to a composite person who does not exist. Naming the situation, first apartment, third child, cold climate, small kitchen, is what makes a stranger recognise themselves.
- The proof. Is there something a sceptic would accept? A demonstration in one shot, a number with a source attached, a customer artefact, a visible test. Claims without proof shift the burden to the landing page, where most of the audience will never arrive.
- The product presence. Is the product visible doing its job, rather than appearing at the end as a logo card? Ads where the product performs its function outlast ads where it is revealed, because the demonstration is the reason to watch and the reveal is a payment demanded at the end.
- The single action. Is exactly one next step requested, in words matching what the landing page delivers? Two actions halve the response to both. A mismatch between the ad promise and the page headline is the most common and most expensive break in the chain.
- The format fit. Was this built for the placement, or resized into it? Vertical video cropped from horizontal, text positioned where an interface element will cover it, a first frame designed for a feed and shown in a story. Cheap to check, and it silently taxes every impression.
- The durability. Does the idea survive thirty viewings? An ad built on a surprise works once per person, which is fine at low frequency and fails exactly when you scale. Ads that survive are usually built on a demonstration or a specific situation rather than on a twist.
What each criterion predicts, and what it does not
| Criterion | Symptom when it fails | Necessary or scaling | Cost to fix |
|---|---|---|---|
| First two seconds | High impressions, almost no watch time | Necessary | Low, reshoot the opening |
| The claim | Watched through, nobody clicks | Necessary | Low, rewrite the line |
| Audience fit | Broad reach, weak response everywhere | Necessary | Medium, new angle |
| Proof | Clicks arrive, landing page does not convert | Necessary | Medium, needs a demonstration |
| Product presence | Engagement without purchase intent | Necessary | Medium, restructure the edit |
| Single action | Clicks that bounce immediately | Necessary | Low, match ad and page wording |
| Format fit | One placement underperforms the rest | Scaling | Low, rebuild per placement |
| Durability | Strong start, collapse as frequency rises | Scaling | High, the idea itself is the limit |
Read the cost column and the working order becomes obvious. Four of the six necessary criteria are cheap to fix, which means most failing creative is repairable rather than replaceable, and replacing it is what teams do by default.
How to score a competitor ad library
The method takes an afternoon. Pull the ads that have run longest for the brands you compete with, take the top twenty, and score each one present or absent on all eight criteria. Do not evaluate whether you like them. You are collecting which criteria the survivors share.
The pattern is what you are after. If eighteen of the twenty lead with a demonstration in the first two seconds, that is a category fact about your buyers and it applies to creative nobody has made yet. If the claims are all specific and comparative, your category rewards comparison. If half fail durability and are still running, frequency is low in your category and you have more room than you assumed.
What you should not do is copy the best-scoring ad. It was built for a brand with a different position, a different price and a different proof available to it. Gartner has been consistent that competitive imitation tends to reproduce the visible layer of a strategy and not the conditions that made it work, which is exactly this failure [Gartner, 2025]. Score for the pattern, build for your own facts.
What to do this week
- Score your last ten ads against the eight criteria, from the finished assets rather than from memory. Note which criterion is absent most often.
- Score twenty long-running competitor ads and compare the two profiles. The gap between them is your creative brief for the quarter.
- Fix one failing component on your best-performing existing ad rather than making a new one. Cheapest available test of the whole method.
- Add the scorecard to the brief, so criteria are decided before production rather than argued after it.
- Get your free score. The Ecommerce Benchmark leaderboard rates your store against real competitors in your category and country across six dimensions, including Creative and Ads, and the paid report breaks the creative dimension down criterion by criterion.
Once you know where you rank, Nexus by Omniconvert identifies which metrics to prioritise and queues the experiments most likely to close the gap, and you approve what goes live. Where the problem the ads are pointing at turns out to be on the site rather than in the creative, the tactical fixes at crobenchmark.com are the cheaper starting point.
What a scorecard cannot do
It cannot rank the survivors. Once several ads pass all eight, the scorecard has done its work and the ranking has to come from spend. Anyone claiming a framework predicts winners is describing hindsight, and Baymard Institute research on why purchases fail is a useful corrective: the reasons are frequently downstream of the ad entirely, in missing information and unexpected cost at checkout [Baymard Institute, 2026].
It cannot supply the idea. Eight criteria describe the shape of a working ad and none of them generates an angle. The angle comes from customer language, review themes and support tickets, and a team using a scorecard as a substitute for that will produce eight-out-of-eight ads with nothing to say.
And it cannot fix a demand problem. An excellent ad for a product the market does not want performs exactly as well as a poor one, slightly faster. Marketing Metrics has long put the probability of selling to an existing customer at roughly sixty to seventy percent against five to twenty percent for a new prospect, which is a standing reminder that creative is rarely the highest-return place to look [Marketing Metrics].
The bottom line
Score before you spend, and score present or absent rather than out of five, because a scale becomes a negotiation and a binary becomes a decision. The first six criteria decide whether an ad can work at all and four of them are cheap to repair, so most failing creative should be rebuilt in one component rather than replaced entirely. The last two, format fit and durability, decide whether an ad survives the moment you increase the budget, which is why an ad that looked strong at low spend so often collapses at high spend. Apply the same eight to the longest-running ads your competitors are paying for, and read the pattern rather than copying the execution. Do that for two quarters and the useful thing you own is not a better ad. It is a file that can tell you why the last twenty worked or did not.
