Google Ads Experiments: How to Run an A/B Test That Actually Proves Something

fuse-smo-martin-janecekWritten by Martin J.
Back to blog
Google Ads experiments A/B test setup 2026 — control and variant campaign arms with a traffic split

Your experiment ended on Tuesday. The variant won on conversions, lost on cost per conversion, and moved in a direction you can defend to a client but not to yourself, because you changed the landing page, the headline, and the bidding strategy all at once. So when you last called a test a winner, how many variables had you actually changed? Google will still hand you a winner either way: auto-apply ships winning variants into live campaigns by default, and the confidence level it reads from is 80%, not 95%. A test that fails is a cheap lesson. A test that succeeds at something you never meant to measure is a change you cannot undo.

Here is what happens to your data before you ever see it. The first seven days of every experiment are discarded from the results page, because that is the ramp-up period for a newly created or newly changed campaign. Both arms also have to clear a learning phase that typically takes seven to fourteen days, and a reactive edit inside that window can restart it. You set a start date and assumed the clock began counting. It did not.

An experiment is a control group. One arm keeps doing what it already did, the other carries your change, and the difference between them is your effect. That is the entire mechanism. Everything built around it is packaging.

Which is why an account that improved last month proves nothing on its own. Seasonality moved. A competitor paused. Budget shifted from a campaign you were not watching. Your Quality Score drifted as Google re-evaluated your landing pages. None of that needed your permission, and none of it shows up as a cause in the interface. Without a named change and a group you deliberately did not change, you have a trend, not a result.

The wall most accounts hit is the same one every time: they treat the experiment as a reporting feature rather than a design problem. Setup takes ten minutes. Deciding what the test is even for takes longer, and that is the step that gets skipped.

What You Can and Cannot Test: Custom Experiments vs the Experiments Page

Google renamed "drafts and experiments" to the Experiments page, and the new name hides an old distinction. A draft is a staging area for edits. It splits no traffic and proves nothing. It becomes an experiment only when you link it to a base campaign as the control and set a traffic split. If you have been calling your drafts experiments, you have been running nothing.

What each surface gives you:

  • Custom experiment lets you change a specific element yourself, including bids, keywords, and ads, on a campaign that keeps serving your existing setup as the control.
  • The Experiments page handles the campaigns where Google builds the comparison for you, which in Performance Max means asset group tests and in video means two to four creative arms.
  • A draft inherits everything and tests nothing until you convert it.

An experiment inherits the base campaign's budget, targeting, and bidding strategy, and overrides only what you explicitly change. Search, Display, Demand Gen, video, and Performance Max all support native testing. Shopping does not, and any workaround there is a manual split wearing a different label.

One setting deserves its own warning: sync. Sync propagates changes you make in the base campaign into the experiment while it runs. If you are changing a bid management approach, or rebuilding the account skeleton with a campaign builder, an inherited mid-flight edit invalidates the comparison silently. Turn sync off for the duration of the test and mean it.

Design the Test Before You Launch It: One Variable, One Hypothesis

Write this down before you touch the interface. The worksheet is the deliverable; the experiment is just the receipt.

Google Ads experiment design worksheet 2026 — hypothesis, single variable, traffic split and read date

Element

What you write

Why it decides the outcome

Hypothesis

"Changing X will move metric Y, because Z"

Forces you to name a mechanism instead of a hunch

Single variable

The one thing that differs between arms

Two differences means you cannot attribute the result to either

Traffic split

50/50 unless volume forces otherwise

A small test arm takes longer to accumulate enough data

Read date

A calendar date plus your conversion lag

Stops you from reading the first good week as a result

Decision rule

What result makes you apply it, and what makes you stop

Written in advance, it cannot be bent after the fact

Run the power check first. Since June 9, 2026, Google evaluates an experiment setup before launch and returns an Experiment Power score from campaign volume, the traffic split, and the configured duration:

Power score

Reading

Correction

Low (0-49%)

The test probably cannot reach a conclusion

Move to a higher-volume campaign, extend the duration, or shift to 50/50

Medium (50-79%)

Borderline

Adjust the split or the duration before spending

High (80-99%)

Proceed

Nothing to fix

Available for Performance Max experiments and Broad Match experiments in Search campaigns. Most accounts never open it, then wonder six weeks later why the result is inconclusive.

On the split: 50/50 reaches a conclusion fastest. A 70/30 or 90/10 split protects performance and starves your data, and every practitioner who recommends it tends to skip the power trade-off that comes with it. Google's own floor is four to six weeks, the configured window cannot exceed 85 days, and low-volume accounts routinely need two to twelve weeks to arrive at 95% confidence. The first seven days get thrown out before any of that.

On the money: an experiment does not buy you extra traffic. Both arms are funded from the base campaign's budget, so a test is a decision about how to spend demand you already have. If the change you are testing is a landing page, the conversion side matters as much as the test design, and that belongs in a separate CRO tools conversation.

The Three Ways You Invalidate Your Own Experiment

One: you changed the experiment mid-flight. A paused keyword. A budget bump "just for the weekend". A base-campaign edit that sync carried across. A new ad copy variant added on day nine. Every one of these makes the two arms incomparable from that date forward, and Google does not flag it. The report keeps rendering a percentage as if the test were intact. The most common version is the most innocent-seeming: turning sync on because someone read that keeping campaigns aligned is best practice. Alignment is the opposite of a control group.

Two: you judged on a partial conversion cycle. The interface shows you the result today, which is not the same as the result. High-consideration purchases carry conversion lags of seven to 90 days, so a 21-day read misses the late converters entirely and inverts the winner with depressing regularity. The thresholds worth working from are roughly 30 conversions before an analysis means anything and 50 or more before an A/B decision holds. And the headline number is softer than it looks: the default confidence interval in the Experiments interface is 80%, while the blue asterisk marks 95%, where the underlying variance estimate switches to a two-tailed test on bucketed data. A directional read at 80% and a confident read at 95% are different claims, and the report does not stop you from phrasing the first as the second.

Three: you compared across mismatched periods or audiences. "The experiment arm did better than last month" is not a comparison, because last month had a different auction, a different promo calendar, and possibly a different bidding strategy. The only valid comparison is arm against arm on the same days. A quiet extension of the test window has the same defect: if you keep the window open until the number looks right, you have quietly moved the read date after seeing the data, which is how a result becomes a narrative.

Check the list before you believe anything:

Three ways accounts invalidate their own Google Ads experiment 2026 — mid-flight edits, partial conversion cycles, mismatched periods

Ask yourself

Invalidating answer

What to do

Did anything change in either arm after launch?

Yes, including a synced base-campaign edit

Restart the test, or discard everything before that date

Has the experiment run past your longest conversion lag?

No

Wait, and keep the decision rule untouched

Did both arms clear the seven to fourteen day learning phase?

No

No conclusion yet, at any confidence level

Are you comparing arm to arm on the same days?

No, comparing to last month

Throw the comparison out

Do you have 30 conversions or more in total?

Fewer

Treat it as directional only, never as a winner

Is auto-apply on?

You have not checked

Confirm the confidence level and both success metrics first

Reading a Result You Can Actually Act On

The report gives you conversion rate, cost per conversion, and the difference between arms, with significance flagged. It does not compute whether the experiment had enough power to detect the change in the first place. That question was settled, or not, at launch. This is the honest gap in the tool, and it is the reason the power check belongs in your worksheet rather than in your post-mortem. If you are testing anything built on Smart Bidding, or on the automated bidding layer generally, remember that a bidding change interacts with volume: the smaller arm feeds the system less data while you are testing it.

Auto-apply deserves a sentence of its own. It is on by default, it reads directional results unless you change it, and it will put a winning variant into your live campaign without asking. The safeguard is narrow in a useful way: if your chosen success metric performs significantly worse in the test arm, the change will not roll out. The catch is that an experiment accepts only two success metrics. A third metric you care about can decline for weeks without appearing anywhere in the test.

Two more practical judgments. A "loser" at day 28 deserves a longer window when your conversion lag is longer than the test, when the arm never cleared its learning phase, or when total conversions are under 30. It deserves to be killed when the change is a one-way door, such as a bidding strategy that overwrites historical learning. And whatever you conclude, write it down with the date and the numbers. Six months from now, the only thing standing between you and re-running the same test is that note.

When an Experiment Is the Wrong Tool

If your account cannot generate 30 conversions across both arms in six weeks, the power check will say so before you launch. Three honest alternatives exist.

Split the campaign manually. Create two campaigns with the same campaign structure and split the budget, accepting that you now own the confounds the tool would have handled. Test inside Performance Max asset experiments instead, where the mechanism is native: one experiment per campaign at a time, locked asset groups, a 40-bucket split of 20 control and 20 treatment, an auto-calculated end date, and 95% confidence to call a winner. Ten asset groups run sequentially means 40 to 60 weeks, so prioritisation by traffic volume is the real skill, not setup. And when the question is creative variety across many ad groups rather than one controlled change, that belongs in Google Ads A/B testing, which is a different discipline with a different toolchain.

Automating the Boring Parts of Google Ads Experiments

Setup is where discipline dies. Navigating three separate sections of the interface to stage a variant, remembering the split, and then holding the read date in your head for six weeks is a lot of bookkeeping for a process nobody audits. That bookkeeping is the reason so many Google Ads experiments get launched and never read, and it is the part worth automating.

A chat-driven setup changes what you actually do: you describe the campaign, the hypothesis, and the split, and the draft campaign gets created and linked as an experiment with the split configured. Monitoring then runs on a schedule instead of on your memory, so the read date arrives as data rather than as a surprise.

One honest scope note: automation owns the setup and the watch, not the judgment. It will not tell you whether your test had enough power, and it does not replace reading the confidence level or your conversion lag. Those stay your call, which is exactly where they belong.

Frequently Asked Questions

What are Google Ads experiments?

A Google Ads experiment is a controlled test that runs a changed version of a campaign alongside the original. The unchanged campaign acts as the control, a set share of traffic goes to the variant, and the report shows the difference between the two arms. The tool answers one question: did this specific change cause that result?

How long should a Google Ads experiment run?

Four to six weeks is the practical floor, and Google discards the first seven days of data before showing you anything. Both arms also need seven to fourteen days to clear the learning phase. Reaching 95% confidence takes two to twelve weeks depending on conversion volume, and the configured duration cannot exceed 85 days. If your conversion lag is longer than the test, the test is too short.

Can you run experiments on Performance Max?

Yes. Performance Max supports asset experiments natively, with one experiment per campaign at a time, asset groups locked once the test starts, a 40-bucket split, an auto-calculated end date, and 95% confidence required to declare a winner. The constraint that bites is sequencing: ten asset groups tested one at a time takes 40 to 60 weeks.

Do Google Ads experiments cost extra?

Not directly. Both arms are funded from the base campaign's budget, so you are not buying additional traffic, you are splitting traffic you already pay for. The real cost is the volume your test arm does not get while the experiment runs, plus the six weeks in which the account is not being optimised on the metric you froze.

What is the difference between a draft and an experiment?

A draft is a set of staged edits with no traffic and no control group, so it proves nothing. An experiment links that draft to a base campaign as the control and assigns a traffic split. Every experiment starts life as a draft, but a draft is not yet a test.

A test you can't read is a test you have to run twice

Let the AI set up the experiment, watch the window, and flag the winner when the data is actually ready.

Your competitors are already using AllAble. Are you?

The marketers pulling ahead aren't working harder. They're just working with one tool that does everything — that tool is AllAble. Try it yourself!