Experiments and evaluation

Choose an effect size worth detecting in an ad test

Choose a meaningful effect size for a ChatGPT ad experiment. Translate business value into a threshold and distinguish it from planned statistical sensitivity.

Start reading
Editorial illustration: A copper ring is held over blue water with large and small ripples.
Editorial illustrationThe different ripple sizes and fixed copper ring evoke the distinction between a change and the chosen frame for assessing it.
The working guide

What you can work through.

Experiments and evaluation
  • Separate the smallest worthwhile effect from planned detection sensitivity.
  • Write absolute percentage points and relative lift explicitly.
  • Do not enlarge the MDE without acknowledging the changed decision scope.

Before asking how many observations a ChatGPT advertising test needs, decide what size of improvement would change the decision. A test that can only detect a doubling may be poorly suited to a choice where a modest increase already pays for the change. Conversely, detecting a tiny difference may not justify an expensive experiment.

Keep two concepts separate. The minimum worthwhile effect is the smallest change that matters commercially. The minimum detectable effect, or MDE, describes the effect size a specified design is planned to detect with chosen error and power assumptions. They should be discussed together, but they are not automatically identical.

Translate the decision into an outcome scale

Start with the primary outcome and its denominator. A purchase probability per assigned visitor differs from purchases per ad click, and both differ from contribution per eligible customer. The effect size must use the same unit that the planned analysis will estimate.

Suppose a hypothetical landing-page experiment for traffic from ChatGPT starts at a 2.0 percent purchase rate. An increase to 2.4 percent is 0.4 percentage points in absolute terms and 20 percent in relative terms. A relative improvement of only 0.4 percent would instead give 2.008 percent: 2.0 multiplied by 1.004. Its absolute gain is 0.008 percentage points, fifty times smaller than the intended 0.4-point gain.

Record the baseline, proposed alternative, absolute difference and relative difference together. When a calculator asks for lift, verify which scale it expects. A correct calculation with a wrongly interpreted input is still the wrong experiment plan.

Compare the alternatives

The same change on two scales

  1. Baseline

    2.0 percent becomes 2.4 percent.

  2. Absolute difference

    Increase of 0.4 percentage points.

  3. Relative difference

    Increase of 20 percent relative to 2.0.

  4. Business question

    Is that increase worth the cost of the change?

Hypothetical conversion rate, not a ChatGPT benchmark.

Derive practical value from the next action

Estimate the cost of adopting the change and the volume over which the result would matter. If the change requires ongoing staff time, that recurring cost belongs in the decision. If it is a small reversible copy edit, the adoption burden may be lower, but the evidence still needs to answer the intended question.

In a hypothetical business case, an additional qualified enquiry has expected contribution after downstream sales costs. Multiplying that value by the extra enquiries associated with a candidate effect can show whether the change is worth pursuing. Keep the assumptions visible, especially when qualification or sales rates are uncertain.

Do not confuse this exercise with assigning a monetary value to every tracked action. The business threshold should use evidence relevant to the outcome. A form-start event cannot inherit the value of a won customer merely because both appear in the same funnel.

Ask what the planned study can detect

NIST’s sample-size guidance shows that detection planning depends on effect size, variability and chosen error assumptions. The appropriate calculation must match the actual design and outcome; a formula for a population mean is not automatically a two-arm conversion-test calculator.

Have the analysis owner translate the worthwhile effect into a suitable power calculation or simulation. The result may show that the available campaign volume cannot resolve such a small difference within a useful period. That is a feasibility finding, not evidence that the smaller difference is economically irrelevant.

Use sample-size input preparation for the calculation handoff. Keep this page’s decision distinct: choosing the effect worth resolving comes before assembling every numerical design input.

Make a sensitivity table instead of one optimistic target

Evaluate several plausible effects around the business threshold. In the hypothetical 2.0 percent baseline, the team might examine 2.2, 2.4 and 2.8 percent alternatives. These are scenarios, not predicted outcomes or published ChatGPT performance benchmarks.

For each alternative, show its economic implication and the information required by the chosen design. This reveals whether the plan is useful only if the treatment performs exceptionally well. It also helps the budget owner decide whether learning about a large effect alone would still be worth the experiment cost.

Do not increase the declared MDE simply to make the sample-size estimate fit the available budget while continuing to claim the original decision is resolved. A larger MDE changes what the study is designed to detect. Record that tradeoff explicitly.

Preserve the threshold when results arrive

Compare the estimated effect and uncertainty with the original practical threshold. A statistically distinguishable change may still be too small to matter. An uncertain estimate may include both worthwhile improvement and negligible effect. Neither situation is captured by a winner label alone.

The confidence-interval guide explains how to read that range against the business decision. For sparse outcomes, low-volume experiment planning considers whether to simplify the question, extend the plan or choose a different learning method.

OpenAI reporting defines the campaign measures available for analysis; it does not choose the advertiser’s economically worthwhile effect. End planning with the outcome scale, practical threshold, planned MDE, calculation assumptions and the decision that remains unresolved if the experiment is less sensitive than the business needs. That record prevents a convenient statistical target from quietly replacing the commercial question.

Sources and scope

Translate a campaign decision into an effect-size target and distinguish practical value from planned statistical sensitivity.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

See what campaign reporting covers.

Explore reporting in AthillyAds and how it fits alongside your own measurement of enquiries, purchases and other business outcomes.