ZachSearcy
← All frameworks
Framework Testing that produces answers

The Testing Ladder

Most creative testing amounts to building a lot of things, launching them together, and interpreting whatever comes back. That produces a ranking, and a ranking isn't something you can build on.

The premise

A ranking tells you this ad beat that ad. You can't build on it, because you don't know which of the four things that differed was the thing that mattered. Next quarter you're back at zero, guessing again, with a slightly larger folder of assets.

The one rule

Two ads in a test are identical except for the single thing being tested. Offer, audience, landing page, budget, and flight dates are all held. If two things changed, you learned nothing.

The second idea is that tests have an order. Some variables are wide and some are narrow, and a narrow test run before a wide one is unreadable. You can't learn which call to action works while the format underneath it is still moving.

Design for function, not aesthetics.

Karl Blanks & Ben Jesson · Making Websites Win
The founders of Conversion Rate Experts make the case for frequent incremental changes over full redesigns, and their reasoning applies directly to creative. A failed test costs you almost nothing, because you discard it and keep the control. A failed redesign costs you the quarter. They also point out that a 19% lift doesn't require one heroic swing; it can come from small improvements to ten separate elements, which is only possible if you know which element moved.

Why the discipline pays

Creative is the largest controllable lever in a campaign and the one most often left to instinct. Nielsen Catalina Solutions analyzed roughly 500 campaigns for Five Keys to Advertising Effectiveness and attributed 47% of sales contribution to creative, against 9% for targeting, which puts creative ahead by roughly five to one. In digital specifically they put creative higher still, around 56%, because creative quality varies far more there than in TV.

Be honest about what that evidence does and doesn't say

The Nielsen work is CPG, in-store sales, 2017. It establishes that creative is the dominant lever. It does not prove that sequential testing beats batch launching, and it isn't B2B digital. I haven't found a clean published study isolating iterative-versus-batch creative testing. Until one exists, the honest argument is the logical one: if creative is the biggest lever, then knowing which part of it moved the number is the highest-value thing you can learn. Batch launching is the one method that structurally prevents you from learning it.

The ladder

Wide variables first. Each rung's winner becomes the next rung's constant.

RungQuestionPrimary metricDecision it drives
1Campaign: where should money go?Pipeline, cost per opportunityBudget allocation
2Offer: what should we make more of?CPL, lead-to-MQL rateContent roadmap
3Format: how should we build it?CTR, CPC, thumbstopProduction mix
4Creative variable: what should it look and sound like?CTR, landing page CVRDesign and copy direction
5Call to action: what should the button say?CTRMicro-optimization
The comparison rule

You can only compare two things that share every rung above them. A CTA comparison is valid inside one offer, one funnel stage, one format, one audience. It is never valid across a brand's whole ad set, and "CTA X wins overall" is the most common way this rule gets broken, because the winning CTA is usually just the one that ran on the biggest budget at the top of the funnel.

The one exception

Some variables can be pooled across campaigns: visual system questions like light versus dark, people versus abstract, button treatment. The reason is that you deliberately ship both treatments in every campaign, which randomizes the other variables instead of confounding them.

Pooling is valid when the variable is randomized across contexts. Pooling is misleading when the variable is concentrated in one context. A CTA that only ever runs at top of funnel is concentrated. A visual treatment shipped everywhere is randomized.

Watch for placement dependence. A visual variable can flip between platforms, with light beating dark on one and losing on the other. Read each platform separately first and pool only if they agree. Divergence is a finding about placement, not noise.

What counts as one variable

A variable is one decision a maker had to make. Some things that look like one variable are several:

  • "Different spokespeople" varies person, script, setting, and runtime together. Not a test.
  • "A new concept" varies layout, imagery, and message. That's a new control, not a test.
  • "Same message, shorter" varies length and usually word choice. Borderline; write down which one you mean.

If you can't state the variable in five words, it isn't one variable.

Requesting a test

Every request carries seven fields. Two arms in a test share every field except Variable. If any other field differs, the test is invalid, and this way you catch it in the sheet rather than in the data three weeks later.

FieldWhat it does
Test IDGroups arms into one experiment. Two rows sharing an ID are one test.
HypothesisWe believe B beats A on [metric] because [reason].
VariableThe single thing that changes. Five words or fewer.
Held constantEverything else, listed explicitly.
Primary metricOne. Plus a guardrail metric where relevant.
Min impressions per armBelow this the result isn't readable.
Decision ruleWhat we do if A wins, if B wins, if it's flat. Written before launch.
If you can't write the decision rule, don't run the test

A test whose outcome wouldn't change what you do next is a test you're running for the feeling of rigor. Kill it and give the budget to one that matters.

Running and reading

How long a flight needs

The binding constraints are sample size and a full weekday cycle, and they're different things.

Sample size is a budget question, not a time question. At a ~0.5% baseline CTR, detecting a large effect (40%+ relative) needs roughly 18,000 impressions per arm. Whether that takes one week or three depends entirely on daily budget. The same total spend, concentrated, produces the same confidence in a third of the elapsed time.

Seven days is the floor regardless. B2B traffic on a Tuesday behaves nothing like a Friday. Any read shorter than a full weekday cycle is an artifact of when you looked.

What you can and can't detect

At realistic budgets you can reliably read large differences, not small ones. This should shape the test backlog more than it usually does.

  • Readable: person versus product, abstract versus people, question versus stat opening, one message versus another. Big swings.
  • Not readable: button color, minor copy tweaks, small layout shifts. These would need many multiples of typical spend to resolve.

If you can't articulate why the two arms would perform very differently, you're paying to watch a coin flip.

The weekly rhythm

DayWhat happens
TuesdayLaunch the increment. Separate ad sets, locked equal budgets.
Friday, day 4Directional check. Kill anything far behind. Start building the next increment against the leader, without retiring the alternative yet.
Monday, day 7Flight closes. No edits at any point during the week.
TuesdayRead to confidence, lock the variable, launch the next increment.
Separate ad sets, locked equal budgets

Inside a single ad set, the platform's delivery model starves the losing variant within about 48 hours based on early noise. At that point you're reading the algorithm's guess, not the audience's response.

Stop rules

These override any test result, in every sprint.

  • Under the minimum impressions per arm. Don't read it at all. An underpowered result is worse than no result, because it feels like knowledge.
  • Unsubscribes or negative feedback spike. Kill the variant regardless of CTR. A winner that costs list is a loss.
  • CTR up but CPL up with it. The ad is pulling the wrong traffic. Check the audience, not the creative.
  • Two things changed between arms. The test is void. Rebuild it; don't interpret it.

Three outcomes, all useful

Keep. One arm clearly ahead. Lock the variable. The loser retires and isn't rebuilt for another audience as a consolation prize.

Split. The result differs by audience or platform. Stop pooling that variable and assign it by segment. This is a finding, not a failure.

Flat. Within the noise band. The variable isn't a lever, so decide on cost and move the budget to one that is.

What it costs

A weekly sprint is lighter than it sounds once the pre-work exists, because most sprint work is re-rendering and re-cutting rather than originating. Roughly 1.3 FTE spread across four people, not four full-time people.

RolePre-work spikeSteady state, per week
WriterNear full-time, ~1 week~30–40%. Module re-cuts are hours, not days
DesignerNear full-time, ~2 weeks~60% early, dropping toward ~10% as the system locks
Media ownerLight~20%. Build ad sets, hold budgets, pull the read
Decision owner4 sessions~1 hour. The Tuesday read

The falling design curve is the argument to make internally. Six masters, then five, then zero, then two. The same volume of variations built up front would have cost full design effort on every single one, and most of them wouldn't have been testing anything.

Staggered clocks

Most of a sprint's design work (master layout, the render matrix, the size adaptations) is message-agnostic once the visual system exists. It can be built during the flight, before anyone knows the winner. Copy is the last-mile fill, and because the module already exists, it's a two-to-four hour job rather than a new brief.

So the team runs on staggered clocks, not one clock. Design is roughly a sprint ahead; copy is same-week. Trying to synchronize them is what makes a weekly cadence feel impossible.

The operator must not also be the decision owner

Whoever is inside the ad platform all day is subject to its recommendations, and platform recommendations optimize for delivery, not for learning. The person who calls the test needs distance from the interface telling them to consolidate ad sets.

The real limitation, stated plainly

A weekly cadence works for variables read on CTR. It does not work for variables read on conversion, pipeline, or revenue. Those need far more volume and far more time. In practice the top rungs of the ladder move weekly and the bottom rungs move monthly or quarterly. Don't promise a weekly cycle on a bottom-of-funnel offer test.

It breaks when: the audience can't deliver the minimum impressions per arm in a week; one person both makes and decides; there's no pre-work sprint, so every sprint originates from scratch; or someone edits mid-flight "just to help it along."

Next

The cadence only survives if production isn't originating from scratch every week.

AI-Enabled Creative Operations covers the other half: how to encode voice, standards, and editorial judgment into workflows so a lean team can hold this rhythm without adding headcount.