A campaign experiment can have accurate reports while comparing the wrong experiences. The control group might receive the same offer through an email, a visitor might encounter both landing pages, or a sales team might extend the new discount to every customer. Contamination means that the intended difference between groups has been weakened or changed. For a ChatGPT ads test, the operational question is what the comparison can still tell you.
The workflow below is our recommended internal method. It assumes no native randomization, user holdout or overlap report in ChatGPT Ads. Establish what your actual tools can support before describing an experiment. Two campaigns with different names do not automatically create separate populations.
Map assignment, exposure and outcome separately
Start with assignment: which unit was meant to enter which group, and under what rule? Then describe exposure: which experience did that unit actually receive? Finally, specify the outcome and observation period. Each stage needs different evidence. A campaign identifier does not establish that a person never encountered another campaign.
Where technically and operationally feasible, your own website could maintain a visitor’s assigned landing page after an ad click. That design tests the page among the defined visitor population. It does not establish the incremental effect of advertising among all prospective customers. When ads influence who clicks, a comparison restricted to clickers may also answer a different question from the original advertising question.
OpenAI’s campaign management documentation describes campaigns, ad groups and ads. Use their stable identifiers to track campaign objects. Keep any person or region assignment required by your design in a separate, appropriate internal register. Review holdout feasibility before promising that exposure can be separated.
Inspect where the experiences can meet
Follow the offer through the whole customer journey. A copied campaign might retain the wrong destination. A shared page template might change both variants. Customers might share a discount code. A regional control period might coincide with a national promotion in another channel. The relevant question is whether an event changes the treatment difference that the groups were supposed to represent.
Distinguish this from a shared external influence. A holiday affecting both groups equally is not automatically contamination. A website change that introduces the test offer into the control experience is a direct leak. A change affecting only one group may instead provide a competing explanation. Name the suspected mechanism rather than describing every issue as noisy data.
Outcome measurement also needs a common definition. If one group records a purchase at payment and another at checkout initiation, the outcomes differ even when exposure is cleanly separated. OpenAI’s conversion tracking documentation explains how events connect to campaigns. A technically functioning connection cannot repair an invalid experimental comparison.
Quantify only what you can establish
Consider a hypothetical experiment with a separately verified assignment register. The control group contains 2,000 people. You know that 160 also received the treatment. Another 250 have unknown exposure status. Known crossover is 160 / 2,000 = 8%. The unknown share is 250 / 2,000 = 12.5%.
Those percentages use people assigned to control as the denominator. Do not divide 160 people by ad impressions. Do not add the unknown cases to confirmed crossover and label the total observed contamination. In this simplified example, the possible contaminated share lies between 8% and 20.5%, provided those 250 are the only uncertain cases. These are logical bounds on missing information, not a statistical confidence interval.
Perform the equivalent check for the treatment group, where some assigned units may never have received the treatment. Record when the leak began and whether it affected a particular variant or channel. When available data cannot establish person overlap, label exposure unknown. An impressive volume of reporting data is not evidence that the groups remained separate.
Trace the leak before interpreting the difference
- Intended control
2,000 people were assigned to control under the external test plan.
- Known crossover
160 also received the treatment: 160 / 2,000 = 8%.
- Unknown exposure
250 have unknown exposure status: 12.5%. Report separately from the 160.
- Decision
Retain assigned groups in the primary analysis and assess what the leak changes.
Resist cleaning the problem out of the analysis
For a properly randomized experiment, analysis by original assignment is generally a central starting point. Moving people into whichever group matches the offer they used can break randomization: that choice may relate to their likelihood of buying. Deleting everyone who encountered both experiences can create the same problem. A tidier table may be a less credible comparison.
Our recommendation is to preserve the primary analysis specified in advance and disclose the deviation. A separate sensitivity analysis can explore alternative assumptions, clearly labelled as such. Do not estimate a supposedly uncontaminated effect by simply dividing the observed effect by the uncontaminated share. That shortcut needs strong assumptions about who was affected and how.
NIST’s design guidance ties the choice of experiment to its objective and factors. The practical implication here is that more observations alone cannot restore a lost comparison. If both groups now receive the same treatment, continuing to collect data does not recreate the intended contrast.
Match the next action to the damage
A short, bounded deviation with a documented start and end may be manageable under an exclusion rule specified before results were examined. An ongoing leak that makes group experiences ambiguous argues for correcting implementation and planning a new comparison. Extensive, unknown overlap may limit the output to a description of observed delivery and attributed outcomes.
Do not assume contamination always makes the effect smaller. Crossovers, selective exposure and unequal measurement can change the result in different directions. If the direction is unknown, say so. A credible limitation is more useful to the next decision than an unsupported claim that the estimate is conservative.
Close the review with four facts: the intended treatment difference, the deviation discovered, the affected units and the decision the remaining evidence supports. Preserve them in the experiment log. If the remaining uncertainty still spans worthwhile benefit and meaningful harm, use the workflow for an inconclusive result. The purpose of this review is to make the contamination an explicit part of the decision rather than a hidden assumption.
Sources and scope
Assess how known crossover and unknown exposure between intended groups affect the interpretation and continuation of a campaign experiment.
- OpenAI Ads: Campaign ManagementRead
- OpenAI Ads: Conversion TrackingRead
- NIST: Selecting an experimental designRead
Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.
