Home / Blog / Paid Ads
Paid Ads

Performance Max Creative A/B Testing: How to Design Experiments That Actually Inform Decisions

Google's new asset-group A/B testing beta gives Performance Max advertisers a more controlled way to compare creative approaches. The real opportunity is to build experiments that distinguish meaningful performance improvements from misleading creative signals.

Sima Damar Beygu
Sima Damar Beygu
Founder, ConversionNest · 10 min read
Performance Max Creative A/B Testing: How to Design Experiments That Actually Inform Decisions

Performance Max advertisers can now test different creative asset sets within the same asset group using Google’s new A/B testing beta.

The feature allows advertisers to compare existing creative assets against alternative headlines, descriptions, images or videos without immediately replacing the original assets.

That creates a more controlled way to evaluate creative decisions inside Performance Max.

But there is an important limitation.

A successful creative A/B test tells you which creative configuration performed better under the conditions of the experiment. It does not automatically tell you whether the advertising generated additional customers or revenue that would not otherwise have existed.

For growth teams, understanding that distinction is essential.

The opportunity is not simply to run more tests. It is to make better creative investment decisions using evidence that matches the business question.

What changed in Performance Max creative testing?

On October 2, 2026, Search Engine Land reported that Google Ads was rolling out A/B testing for Performance Max asset groups.

Google’s official documentation describes the capability as a beta that allows advertisers to compare two different sets of creative assets within the same asset group.

Previously, Performance Max advertisers could use various optimisation experiments, including testing the addition of video assets or creative assets to feed-only campaigns.

The newer asset comparison beta introduces a more flexible approach. Advertisers can compare alternative combinations of existing or newly uploaded creative assets.

For example, a retailer could test product-focused imagery against lifestyle imagery. A subscription business could compare functional product messaging against benefit-led messaging. A travel advertiser could evaluate destination photography against experience-focused creative.

These are potential applications, not demonstrated performance outcomes.

Importantly, the feature remains labelled Beta in Google’s documentation as of October 10, 2026. Availability should therefore be verified within individual Google Ads accounts rather than assumed globally.

How the experiment is structured

Google divides assets into three categories:

  • Control assets (A): Existing assets selected as the reference creative set.
  • Treatment assets (B): Alternative existing or newly uploaded assets being tested.
  • Common assets: Assets not assigned to either test set that remain eligible to serve alongside both variations.
Performance Max experiment comparing control, treatment and common creative assets.


The third category deserves particular attention.

Google states that common assets continue serving to 100% of campaign traffic alongside the respective control and treatment assets.

This means the experiment does not necessarily create two completely isolated advertising experiences. Instead, it compares different creative sets within an environment that may also contain shared assets.

That distinction matters when interpreting results.

Why creative testing matters more in automated campaigns

Performance Max relies on automated systems to determine how advertising assets are combined and delivered across Google’s advertising inventory.

Advertisers provide creative inputs, conversion goals, audience signals and other campaign settings. Google’s systems then determine which eligible combinations to serve.

This creates an important challenge for creative analysis.

An advertiser might upload new images and observe that conversion value increases the following week. But the increase could reflect several factors:

  • Changes in demand or seasonality.
  • Differences in audience composition.
  • Changes in auction competition.
  • Campaign learning and delivery adjustments.
  • The creative change itself.

A simple before-and-after comparison cannot reliably separate these explanations.

A controlled experiment provides stronger evidence because it compares alternative configurations during the same experimental period.

However, it is still necessary to define precisely what is being tested.

The overlooked issue: Common assets can complicate interpretation

One of the most important details in Google’s new testing framework is how common assets are handled.

Consider a hypothetical ecommerce campaign containing:

  • Six product images.
  • Three lifestyle images.
  • Multiple headlines and descriptions.
  • Two videos.

The advertiser wants to understand whether lifestyle imagery performs better than product-focused imagery.

A sensible experiment might assign product-focused images to the control group and lifestyle images to the treatment group. Headlines, descriptions and videos could remain common assets.

This creates a comparison between two image sets while the other creative inputs remain available in both arms.

However, the result should be interpreted carefully.

The experiment evaluates the performance of the image sets within the broader creative environment.

It does not necessarily identify the independent contribution of each image. Nor does it establish that every user saw only the image assigned to their experimental arm.

Google’s automated asset selection and the presence of shared assets remain part of the delivery environment.

Why this changes the testing strategy

Suppose a new lifestyle image set produces better results.

The advertiser can reasonably investigate whether adopting that image set improves the campaign’s performance. But the result does not establish that a particular image, visual element or emotional message caused the improvement.

If the treatment group introduced several changes simultaneously, the experiment evaluates the combined treatment.

This is why creative testing should begin with a clearly defined hypothesis.

Conversion Nest’s recommendation is to test coherent creative hypotheses rather than collections of unrelated changes.

If the question concerns imagery, avoid simultaneously changing headlines, videos and promotional messaging unless the objective is explicitly to compare two complete creative concepts.

The narrower the hypothesis, the easier the result is to interpret.

How to design a useful Performance Max creative A/B test

A practical testing process begins before opening Google Ads.

Step 1: Define the business question

A weak hypothesis might be:

“New creatives will improve performance.”

That statement is too broad to guide a meaningful experiment.

A stronger hypothesis would be:

“Lifestyle imagery showing the product in use will generate higher conversion value at an acceptable return on ad spend than product-only imagery.”

This defines the proposed change, the expected outcome and the commercial constraint. It also provides a basis for deciding what should remain unchanged.

Step 2: Choose one primary creative variable

The most interpretable experiments isolate a meaningful creative difference.

Examples include product-focused versus lifestyle imagery, functional versus emotional messaging, or alternative video concepts.

If multiple asset types change together, treat the test as a comparison between creative packages. Do not present the result as proof that one individual asset caused the difference.

Step 3: Define the control, treatment and common assets

Document exactly which assets belong to each category.

For an imagery experiment, the configuration could be:

  • Images — Control A: Product-only photography.
  • Images — Treatment B: Lifestyle photography.
  • Headlines: Shared across both treatments.
  • Descriptions: Shared across both treatments.
  • Videos: Shared across both treatments.
  • Landing page: Unchanged.
  • Conversion goal: Unchanged.

This is an illustrative test design.

Google’s beta allows advertisers to select existing assets or upload new ones for the treatment group.

Both sets count toward the asset group’s applicable asset limits. The experiment can test assets within only one asset group at a time.

Step 4: Select commercially meaningful success metrics

The correct primary metric depends on the campaign’s objective.

For ecommerce campaigns, conversion value and ROAS may be appropriate.

For lead-generation campaigns, conversions and cost per acquisition may be useful, provided the conversion action represents meaningful business value.

But the evaluation should not stop at platform metrics.

A creative variation that increases revenue while substantially increasing costs may not improve profitability.

Likewise, a variation that generates more leads may create little value if those leads are poorly qualified.

Where possible, connect the experiment to downstream commercial outcomes.

Step 5: Allow sufficient time

Google’s documentation recommends running the new asset comparison experiment for approximately four to six weeks.

The experiment’s end date is initially determined by Google’s Experiment Guidance System, which estimates the duration needed for statistically meaningful results.

Advertisers can modify the end date, but shortening a test without considering conversion volume and conversion lag may reduce its usefulness.

Google also warns against editing, adding or removing assets in the tested asset group while the experiment is running. Those assets become view-only during the experiment.

This restriction protects the comparison from changes that could undermine its interpretation.

For the complete setup requirements, consult Google’s Performance Max asset A/B testing documentation.

How to set up the experiment in Google Ads

For eligible accounts, Google’s documented workflow is:

  1. Open Campaigns and navigate to Experiments.
  2. Create a new experiment.
  3. Select Assets as the variable to test.
  4. Choose Assets provided by you.
  5. Select Performance Max as the campaign type.
  6. Choose Any assets as the experiment type.
  7. Select the campaign and asset group.
  8. Define the control and treatment assets.
  9. Set the traffic split.
  10. Review the proposed duration and schedule the experiment.

Google states that the experiment starts the following day.

The system then compares the alternative creative configurations within the selected asset group.

Availability and interface details may vary because the feature remains in beta.

Important restrictions

Advertisers should check eligibility before committing to a testing plan.

Google identifies several potential obstacles, including incompatible campaign configurations, overlapping experiments, unsupported features and asset-limit issues.

New treatment assets must also pass the standard advertising policy review.

These operational requirements matter because a creative testing roadmap can fail before it begins if the account or campaign is not eligible.

How should you interpret the results?

A creative experiment should be evaluated using both statistical evidence and commercial relevance.

Imagine an ecommerce advertiser testing two creative approaches. The following figures are entirely hypothetical and assume equal advertising spend for illustration.

  • Advertising spend — Control A: $10,000; Treatment B: $10,000.
  • Conversion value — Control A: $30,000; Treatment B: $36,000.
  • ROAS — Control A: 3.0x; Treatment B: 3.6x.
  • Assumed contribution margin before advertising — Control A: 40%; Treatment B: 40%.
  • Contribution before advertising — Control A: $12,000; Treatment B: $14,400.
  • Contribution after advertising — Control A: $2,000; Treatment B: $4,400.

Under these assumptions, Treatment B generates $2,400 more contribution after advertising.

That is a commercially meaningful difference.

However, it would be premature to declare Treatment B the winner without considering statistical uncertainty and whether the experiment was properly conducted.

The numbers alone do not establish that the difference is statistically reliable.

A 50/50 traffic split does not guarantee equal spend

Another important consideration is how experiment traffic is allocated.

Google’s Experiments FAQ explains that traffic splits and budget splits are different.

In Performance Max experiments, traffic allocation does not necessarily produce identical advertising spend across both arms.

Auction dynamics, bidding behaviour and delivery differences can influence spending.

Therefore, advertisers should avoid comparing raw revenue totals without accounting for differences in spend and experimental reporting.

Use the experiment’s reported comparison, uncertainty information and relevant efficiency metrics rather than assuming equal traffic allocation produces equal commercial exposure.

Statistical significance is not the same as business significance

A small performance improvement can be statistically convincing but commercially unimportant.

Conversely, a potentially valuable improvement may remain inconclusive when conversion volume is insufficient.

Growth teams should define the smallest improvement worth implementing before reviewing results.

That threshold should reflect business economics.

For example, a modest ROAS improvement might be valuable in a high-spend account. The same improvement might not justify substantial creative production costs in a smaller account.

The experiment should inform an investment decision, not merely identify a statistically favourable variation.

Creative A/B testing is not Conversion Lift

This distinction is especially important for advertisers already using incrementality measurement.

Creative A/B testing and Conversion Lift answer different questions:

  • Performance Max creative A/B testing asks which creative configuration performs better.
  • Conversion Lift asks how many additional conversions the measured advertising generated.
  • Landing-page A/B testing asks which landing-page experience performs better.
  • Geographic incrementality testing asks what changes when advertising is withheld or altered across selected markets.
Decision framework comparing creative A/B testing, Conversion Lift, landing-page testing and geographic incrementality testing.


A creative experiment compares alternative advertising treatments.

Conversion Lift compares outcomes between users eligible for the measured advertising and a holdout group.

Both can use controlled experimental methods, but they estimate different effects.

A successful creative A/B test does not establish that the campaign itself generates incremental demand.

For a detailed explanation of that distinction, see Conversion Nest’s guide to Google Ads Conversion Lift and incremental ROAS.

This distinction also reflects a broader experimentation principle: the control condition must match the causal question being investigated.

As CXL explains in its discussion of holdout groups, comparing two treatments is not equivalent to evaluating the impact of introducing a treatment in the first place.

Five mistakes that can undermine Performance Max creative experiments

1. Testing too many creative ideas simultaneously

If the treatment changes imagery, headlines, videos and promotional messaging, the result evaluates the combined package.

It cannot reliably identify which individual change drove the outcome.

That may be acceptable when comparing complete creative concepts. It is less useful when the goal is to identify a specific performance driver.

2. Ignoring common assets

Shared assets can continue serving across both experimental arms.

Advertisers should document these assets and consider how they affect the interpretation of the test.

The experiment is not necessarily a clean comparison between two entirely separate advertising experiences.

3. Changing the campaign during the experiment

Material changes to creative assets, conversion goals, bidding or other campaign conditions can make results harder to interpret.

Google explicitly prevents asset editing within the tested asset group while the experiment is running.

Other unavoidable operational changes should be documented and considered during analysis.

4. Declaring a winner too early

Early performance differences can be misleading.

Conversion delays, limited data and normal variation may produce results that change as the experiment continues.

Google recommends allowing sufficient time and considering the experiment’s statistical guidance.

5. Applying the winning treatment without reviewing the final asset configuration

Google’s beta offers an important implementation choice.

When advertisers apply the experiment, they can add the treatment assets to the original asset group.

By default, the option to keep the control assets is selected.

That means applying a successful experiment does not necessarily replace the old creative assets with the winning set. Both sets may remain in the asset group.

This matters because the resulting configuration may differ from the treatment that was evaluated.

Review the final asset mix before applying the result and decide deliberately whether the control assets should remain.

A practical decision framework

Before launching a Performance Max creative experiment, growth teams should be able to answer five questions:

  1. What decision will this experiment inform? Define the action you will take for a positive, negative or inconclusive result.
  2. What is the primary hypothesis? State the creative difference, expected outcome and commercial constraint.
  3. What remains constant? Document common assets, landing pages, conversion goals, bidding settings and other relevant campaign conditions.
  4. What result would be commercially meaningful? Set a decision threshold before seeing the data.
  5. Does this test answer the actual business question? Use an incrementality method if the question is whether advertising caused additional outcomes rather than which creative configuration performed better.

Final takeaway

Performance Max creative A/B testing is a useful step toward more disciplined creative decision-making inside an automated campaign environment.

Its value depends less on the existence of the feature than on the quality of the experiment.

Strong tests begin with a specific commercial question, isolate a coherent creative difference, document common assets, run for sufficient time and distinguish statistical evidence from business value.

Most importantly, teams should not confuse a better-performing creative treatment with proof of incremental advertising impact.

Use the experiment to answer the question it was designed to answer, and use a different measurement method when the decision requires different evidence.

Free audit, powered by Nestor

Is your landing page wasting the clicks you paid for?

Negatives fix who arrives. Nestor reviews what happens next: enter your URL, create a free account, and read what to fix first.

Free account required. No card, no sales call.
Sima Damar Beygu

Sima Damar Beygu

Founder of ConversionNest. 10+ years in growth marketing, managing 300K euro monthly media budgets and scaling acquisition across 15+ markets. Google and Meta certified.

Sima on LinkedIn
Keep reading