Run a Google Ads experiment for at least four full weeks, and longer until each arm has enough conversions to detect the difference you care about. As a rule of thumb, that means a minimum of 1-2 weeks for the trial to settle, plus at least two complete weekly cycles of clean data, plus enough time for late conversions to arrive. A campaign with 400 conversions a month testing a large change can get there in a month; one with 60 conversions a month testing a 10% CPA improvement will not get there in a quarter.

The calendar is a floor. What sets the real length is how many conversions each arm collects, how big a difference you are trying to detect, and how long your customers take to convert after the click.

The four things that set the length

Conversion volume

Statistical tests on CPA and ROAS run on conversions, and conversions are the scarce unit in most accounts. A campaign with 3,000 clicks and 90 conversions a month has the sample size of 90, not 3,000, for any question about cost per conversion. Before you set a date, divide the campaign's monthly conversions by two for a 50/50 split: that is what each arm collects per month.

Conversion lag

A click on day 27 of the test that converts on day 35 counts toward the trial, but only if you wait until day 35 to read it. Check the lag in your own account: add the Conversions column, then Segment > Conversions > Days to conversion over the last 90 days. If 80% of conversions arrive within three days of the click, leave three days after the end date before judging. For B2B lead generation with offline conversion imports, the gap can be weeks.

The learning period

If the trial changes a bid strategy or a target, its bid strategy status shows Learning while Smart Bidding recalibrates. Google's rule of thumb is up to about 7 days, or about 50 conversions, or three conversion cycles, whichever it needs. Any trial also starts with new ads going through review. Treat the first one to two weeks as warm-up and read the result from the period after it.

Weekly cycles

Most accounts convert differently on a Monday than on a Saturday. Run whole weeks so each day of the week is represented equally, and prefer at least two of them after the warm-up so a single odd week does not decide the result.

A worked example: sizing a test in EUR

Say a Search campaign spends EUR 30,000 a month at an average CPC of EUR 1.50. That is 20,000 clicks a month. It converts at 4%, so 800 conversions a month at a CPA of EUR 37.50. You want to test a new landing page and would act on a 15% relative lift in conversion rate, from 4.0% to 4.6%.

For a two-arm test at 95% confidence and 80% power, the standard approximation for the sample per arm is about 16 x p x (1 - p) / d squared, where p is the baseline rate and d is the absolute difference. Here: 16 x 0.04 x 0.96 / 0.006 squared, which comes to roughly 17,000 clicks per arm.

  1. At a 50/50 split, each arm gets 10,000 clicks a month, so 17,000 clicks takes about 7 weeks.
  2. Add one to two weeks of warm-up that you exclude from the reading: 8 to 9 weeks.
  3. Add the conversion lag before you read it: if most conversions land within three days, the verdict comes about 9 weeks after launch.
  4. In that time each arm spends roughly EUR 25,500 and collects about 680 conversions.

Now change one input. If you would only act on a 30% lift (4.0% to 5.2%), the difference doubles and the sample needed drops to a quarter, about 4,300 clicks per arm: under two weeks of data, so three to four weeks in total with warm-up and lag. If you want to detect a 7.5% lift, the sample quadruples to about 68,000 clicks per arm, and the test would run for about seven months. That last case is the answer for many campaigns: the effect you hope for is too small to measure at your volume, so either test a bolder change or test on a larger campaign.

The same logic applies to ROAS, with more noise, because conversion value varies from order to order. A ROAS test usually needs more conversions per arm than a CPA test of the same relative size.

Rules of thumb when you do not want to do the maths

  • Under 30 conversions per arm per month: a CPA or ROAS test is unlikely to reach significance within a quarter on anything smaller than a 40% difference. Test a bolder change or move the test to a bigger campaign.
  • 100 or more conversions per arm over the test: large effects (20% and up) usually become readable.
  • Several hundred conversions per arm: you can start to detect effects in the 10-15% range.
  • Never less than two full weeks of data after warm-up, whatever the volume.

These are rules of thumb. The scorecard's confidence interval is the real test: when it no longer crosses zero on your chosen metric, the result is significant at the confidence level selected, which you can set to 80%, 85% or 95%. How to read Google Ads experiment results covers that step.

When to stop early

Peeking at a test every day and stopping on the first day the result looks significant is the most common way to get a false winner. Random noise crosses the significance line several times in a long test; if you stop on the first crossing, you will ship changes that did nothing. Fix the end date before launch and read the result once.

There are two good reasons to stop before the date:

  • Harm. Agree up front how much you are willing to lose. If the trial's CPA is 40% above the base after two weeks past warm-up, with a confidence interval that sits entirely on the bad side, end it. The downside is known, and waiting only costs money.
  • Breakage. A disapproved ad, a broken landing page, a conversion tag that stopped firing on one arm. The test is measuring the fault, so end it, fix it and start again.

A trial that is winning early is a weaker reason. Early wins shrink as data accumulates, especially while the trial is still out of its learning period, so let a winning test run to its date.

Where GoodLads fits

GoodLads applies its hypotheses as Google Ads experiments where the campaign supports one, and tracks each on a board from live to a verdict on conversion value or CPA. An experiment Google does not mark significant is never counted as a win; the demo shows a win, a loss and a result with no difference.

It does not change the arithmetic above: a campaign with too few conversions for a readable test is still too small, whatever tool launches the experiment. If you would rather query your own experiment data from Claude or ChatGPT, every plan includes the read-only MCP endpoint.

Board with Not live yet, Live and Concluded lanes. Live cards show days elapsed out of the planned test length; concluded cards show what was predicted, what was measured and what happened.
Live tests on the board show how far they are through their planned length, day 14 of 21 for example, so nobody calls one early by accident. From the demo account

Questions people ask

What is the minimum length for a Google Ads experiment?

About four weeks as a rule of thumb: one to two weeks of warm-up, at least two full weeks of data, and time for late conversions. Low-volume campaigns need longer.

How many conversions do I need for a Google Ads experiment?

It depends on the size of the effect. Large effects of 20% or more often become readable at around 100 conversions per arm; effects of 10-15% usually need several hundred per arm.

Should I exclude the learning period from experiment results?

Yes. If the trial changes a bid strategy or target, set the scorecard date range to start after the bid strategy status leaves Learning, usually one to two weeks in.

Can I stop a Google Ads experiment early if it is winning?

You can, but early wins often shrink. Set the end date before launch and stop early only if the trial is causing clear harm or something in the test is broken.