Experiments and evaluation

Read confidence intervals for ChatGPT ads tests

Interpret the estimated difference in a ChatGPT ads conversion test, compare uncertainty with business thresholds, and recognize what an interval cannot prove.

Start reading
Editorial illustration: An elongated green glass block meets a dark and a coral-colored fin on a pale surface.
Editorial illustrationThe glass span and two fins represent judging uncertainty against both zero difference and the business requirement.
The working guide

What you can work through.

Experiments and evaluation
  • Read the interval for the difference, not only separate intervals for each alternative.
  • Compare the interval with zero and with the improvement the business needs.
  • A narrow interval cannot repair assignment problems, measurement errors or immature data.

A report shows a 6.5% booking rate for a new page and 5% for the existing page after clicks from ChatGPT ads. The new page looks promising. Yet the 1.5 percentage point difference is an estimate from limited observations, not a guaranteed improvement next month. A confidence interval describes the precision of that estimate under a particular statistical model.

The useful question is which effect sizes remain compatible with the data and model. That is more informative than placing a green label beside the word significant. Our recommended reporting practice is to show the estimate and interval together, then state precisely what was compared and which decision the uncertainty can support.

Identify the parameter before reading the endpoints

An interval around page B’s conversion rate concerns the level of B’s rate. An interval around B minus A concerns the difference between alternatives. These are different parameters. Overlap between two separate confidence intervals is not itself the correct test of a difference, and subtracting their endpoints does not generally produce the appropriate interval for that difference.

Name the parameter explicitly: for example, the difference in the proportion of assigned visitors who book within a fixed observation window. State whether the effect uses percentage points, relative percentage change or another scale. Moving from 5% to 6.5% is a 1.5 percentage point increase and a 30% relative increase. Those describe the same point estimate but require different interval calculations.

In OpenAI reporting, post-click CVR is conversions divided by clicks. That ratio should not automatically be treated as independent people’s binary booking outcomes. When an export lacks the information required by an inference method, change the method or limit the conclusion rather than feeding the reported ratio into an unsuitable calculator.

Work through one clearly bounded example

Consider a hypothetical website experiment with valid randomization independently implemented by the advertiser after an ad click. Assume 2,000 independent, unique assigned visitors per page, at most one positive booking outcome per person, and an equal, completed observation window. Page A has 100 booking visitors and B has 130. Their observed rates are 100/2,000 = 0.05 and 130/2,000 = 0.065.

For an illustrative normal approximation, the standard error of the difference is the square root of 0.05 × 0.95 / 2,000 + 0.065 × 0.935 / 2,000. This is approximately 0.00736, or 0.736 percentage points. An approximate two-sided 95% interval is 1.5 ± 1.96 × 0.736 percentage points: about 0.06 to 2.94 percentage points.

This calculation illustrates scale and assumes one planned analysis with an appropriate model. NIST’s comparison of two proportions provides statistical background. Sparse outcomes, extreme rates or different observation structures can require another method chosen before examining results. NIST also describes Wilson intervals for a single binomial proportion. That is not automatically the method for the difference between two proportions.

Understand what the confidence level describes

The confidence level concerns the procedure’s coverage over repeated samples under its assumptions. A 95% interval does not assign a 95% probability to the fixed unknown effect being inside this particular completed interval. NIST’s explanation of confidence explicitly distinguishes those interpretations.

The interval is also not a prediction range for the next campaign week. Future visitors, offers, delivery and competitive conditions may change. The calculation addresses sampling uncertainty within the model; it does not automatically cover broken tracking, incorrect assignment or every change in the advertising environment. Keep that limitation visible when the report reaches someone who will not read the technical appendix.

Use two reference points for a business decision

The first reference point is zero difference. In the example, the approximate interval sits just above zero. Under the corresponding test and assumptions, it supports a positive difference at the chosen level. The rounded lower endpoint of 0.06 also shows how close the result is to that boundary. It does not establish that every plausible future outcome will be positive.

The second reference point is the improvement the business needs. Suppose, hypothetically, that adoption requires at least a 1 percentage point increase to justify additional operating cost. An interval from 0.06 to 2.94 includes improvements below and above that threshold. The evidence therefore does not establish that the required commercial improvement has been achieved, even though zero is outside the interval.

Now imagine a different illustrative interval from minus 0.4 to plus 0.6 percentage points. Large gains are less compatible with that data and model than with an interval from minus 5 to plus 8. Both contain zero, but they provide very different decision support. The guide to inconclusive results explains how that distinction should affect the next step.

Decision guide

Compare the interval with two thresholds

  1. Estimate

    6.5% minus 5% = 1.5 percentage points.

  2. Uncertainty

    Approximate 95% interval: 0.06 to 2.94 percentage points.

  3. Zero threshold

    Zero lies just outside the interval under the example model.

  4. Business threshold

    A required +1 percentage point lies inside: sufficient benefit remains uncertain.

Hypothetical normal approximation: 100/2,000 versus 130/2,000 independent visitors.

Check the assumptions that the interval cannot check for you

Repeat visitors, shared company accounts or assignment by region create dependencies that an independent-visitor model does not capture. Ignoring them can make standard errors too small. Repeatedly looking for a favorable interval or selecting the best among many tested variants also changes the inference problem. A report should not present those procedures as one uncomplicated planned comparison.

Return to the definitions and analysis method in the preregistered success metric. Verify assignment, outcome maturity and which units are actually independent. If the records are aggregated by campaign day, a large number of impressions does not magically provide that many independently assigned experimental units. The analysis must follow the design that produced the observations.

A correctly calculated interval can describe an observed difference without proving that advertising caused it. Close the decision note with the estimated effect, its uncertainty and the action that uncertainty permits. Keeping those three statements together makes the interval useful to the advertiser instead of turning it into a decorative technical number.

Sources and scope

Interpret an estimated interval for a conversion outcome and identify which conclusions its uncertainty permits.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

See what campaign reporting covers.

Explore reporting in AthillyAds and how it fits alongside your own measurement of enquiries, purchases and other business outcomes.