Experiments and evaluation

When is a switchback test suitable for ChatGPT ads?

Evaluate time-based testing for ChatGPT ads. Plan period lengths, carryover assumptions, calendar balance and evidence that the intended switches actually occurred.

Start reading
Editorial illustration: A rippling woven strip has teal and rust sections, with threads continuing across pale transitions.
Editorial illustrationThreads crossing the color boundaries represent carryover when a test switches between conditions.
The working guide

What you can work through.

Experiments and evaluation
  • Calendar alternation does not automatically create a randomized experiment.
  • Carryover determines which period lengths are defensible.
  • Respect assigned time periods instead of treating every click as an independent trial.

Running ChatGPT ads one week and pausing them the next sounds like a straightforward effect test. The problem is that the second week may still be influenced by the first. Customers remember an offer, sales conversations continue and purchases arrive after the initial ad contact. Demand can also differ between weeks for reasons unrelated to advertising.

A switchback design exposes the same market or operation to different conditions in successive periods. This article provides our recommended internal feasibility process. It does not assume a native ChatGPT Ads experiment tool, automatic randomization or guaranteed delivery according to an experimental schedule.

Specify the states the operator can actually distinguish

Define two conditions that can be implemented and verified. One might be a particular campaign active versus paused, where the account and business permit that change. Another might compare two explicit advertising configurations. If message, budget, destination and objective all change together, the test concerns that package rather than any individual component.

OpenAI documents campaign activation and pausing. Those operations do not establish that activation produces the intended exposure during the next period. Separate intended assignment, observed configuration and delivered advertising. An active campaign with no delivery has not implemented the intended exposure contrast.

Define business outcomes for both conditions. If the decision concerns additional completed purchases, count relevant purchases during the comparison periods too. A report containing only conversions credited to advertising cannot establish how many purchases would have happened without it.

Investigate how long an intervention can matter

List plausible mechanisms for carryover: delayed purchases, quotation processes, return visits and offers saved for later. Review the advertiser’s own time from first contact to outcome. A short median is insufficient if a meaningful share of business arrives much later.

Separate business carryover from reporting delay. Waiting for reports to mature can address incomplete data, but it cannot remove the influence of condition A on outcomes observed during condition B. An attribution window does not automatically establish how long the actual advertising influence lasts either.

Research on switchback experiments treats carryover as a central design issue. Our recommendation is to document a justified duration assumption before starting. If influence plausibly persists beyond operationally feasible periods, a short switchback becomes difficult to defend. Consider alternatives such as geographic testing, subject to their own feasibility requirements.

Freeze a schedule before seeing outcomes

Assigning A to weekdays and B to weekends confounds condition with demand patterns. A simple A, B, A, B sequence is not automatically balanced either: a recurring business event may follow the same rhythm. Review weekdays, paydays, promotions and planned operational changes before settling on a calendar.

Where a genuine experiment is feasible, determine assignment externally using a documented randomization rule and appropriate balance constraints. Preserve the original schedule. Replacing upcoming conditions when performance looks disappointing changes the experiment and can compromise its interpretation. NIST’s discussion of blocking provides a methodological basis for handling known nuisance factors.

A transition interval, sometimes called a washout, can be excluded from the primary analysis under a prespecified rule. Exclusion does not guarantee the disappearance of carryover. Justify the interval using the business process and check that advertising and outcome data can support the required time boundaries.

Calculate usable observation time

Here is a hypothetical planning example, not a recommended test duration. Twelve seven-day periods require 84 calendar days. The assignment schedule gives six periods to each condition. If the first day of every period is excluded under the same rule, 72 analysis days remain: 36 for each condition.

There are still only twelve assigned time periods, not 72 independent experiments. Clicks, purchases and adjacent days can be dependent. Analysis must respect the time structure and actual assignment mechanism; otherwise, reported uncertainty can appear substantially smaller than warranted.

The one-day exclusion is included only to demonstrate the time budget. It may be entirely inadequate for a longer purchase journey. A longer transition reduces usable observation time and may require more periods. That tradeoff should be evaluated before the business commits its test budget.

Also account for the cost of operating the schedule. Someone must verify changes, investigate failed switches and preserve records. A theoretically attractive number of short intervals can become impractical if the team cannot reliably execute them. Operational failure affects the intervention the study actually evaluates.

Workflow

From calendar time to analysable time

  1. 12 seven-day periods

    84 calendar days with six periods per condition under a predefined assignment.

  2. 12 transition days

    One prespecified day per period is excluded from the primary analysis.

  3. 72 analysis days

    36 days per condition, but still only 12 assigned time periods.

Hypothetical time budget, not a recommended duration or platform feature.

Design data collection around the analysis

Use an agreed time zone and unambiguous period boundaries. OpenAI Reporting describes daily reports. If the verified source only supports daily resolution, do not pretend it separates exposure before and after a midday switch. Adapt the design to available precision or establish an independently verified supplementary measurement source.

For each period, retain assignment, start and end, transition interval, business outcome, actual delivery and deviations. Keep unsuccessful periods visible. Predefine their treatment in the analysis and distinguish any sensitivity analysis from the primary result. Retaining only periods with strong delivery can select observations according to something affected by treatment.

Review the calendar again after execution, using recorded facts rather than remembered explanations. A warehouse outage may explain an unusual period, but excluding it because its result is inconvenient is different from applying a previously stated rule. Document both the event and how much the conclusion depends on its handling.

Decide what the study can settle

Ask the analyst to assess power for the minimum decision-relevant effect using period count, outcome variability and temporal dependence. A small p-value from an analysis that ignores this structure does not repair a weak schedule. A wide uncertainty interval may leave both meaningful benefit and meaningful harm plausible.

Connect the experiment log to a predefined reporting format: observed difference, uncertainty, deviations and the carryover assumption used. The result then concerns the switching policy in the environment studied. It does not automatically estimate the effect of uninterrupted advertising in another season or under a different operating pattern.

Sources and scope

Decide whether alternating advertising conditions over time is defensible and specify periods, transitions and appropriate units of analysis.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

See what campaign reporting covers.

Explore reporting in AthillyAds and how it fits alongside your own measurement of enquiries, purchases and other business outcomes.