Four purchases in a week may be valuable business, but they provide limited evidence for distinguishing two similar campaign approaches. When ChatGPT ads generate few conversions, the first experimental question is whether the intended comparison fits the available traffic. A considered decision to postpone can save both spending and weeks of reporting that could never resolve the proposed choice.
We recommend a feasibility calculation before launching the test. The method below is internal planning guidance, not a description of a native advertising experiment feature. It assumes the team can establish a valid comparison with an appropriate assignment process. If that cannot be done, describe the investigation as observational and limit causal conclusions accordingly.
Keep five different quantities separate
Sample size is the number of relevant analysis units. Those could be unique assigned visitors in an advertiser’s website experiment, or regions or time blocks in a different design. Conversions are positive outcomes, not automatically the sample size. Impressions and clicks can also include repeated exposure from the same people.
The baseline is the conversion probability used for planning. Minimum detectable effect, or MDE, is the effect size the selected design is sized to detect with the chosen power. Alpha is the allowed false-positive probability under the null hypothesis and the model’s conditions. Statistical power is the probability of rejecting the null at a specified true effect under those conditions.
An 80% power target is not an 80% probability that the hypothesis is true, and it is not a quality score for a campaign. NIST’s sample-size guidance explains the relationship between effect size, alpha and power. Check that the calculation matches the actual design: comparing one proportion with a fixed reference differs from comparing two contemporaneous groups.
Translate a meaningful improvement into visitors
Suppose, hypothetically, that an advertiser wants to compare two booking pages after ChatGPT ad clicks and can independently randomize unique visitors correctly. The planning baseline is 2%. The desired detectable increase is to 3%: an absolute improvement of 1 percentage point and a relative improvement of 50%. For this example, choose two-sided alpha of 0.05, 80% power and equal group sizes.
A conventional normal approximation for two independent proportions gives approximately 3,826 visitors per group, or 7,652 in total. The calculation uses 1.96 for alpha, approximately 0.842 for power, and an average proportion of 0.025. The per-group expression squares the sum of 1.96 × √(2 × 0.025 × 0.975) and 0.842 × √(0.02 × 0.98 + 0.03 × 0.97), then divides by 0.01².
This is a rounded planning approximation without continuity correction, not an OpenAI requirement. The exact analysis method, dependencies, missing outcomes and uncertainty in the baseline can change the requirement. Have the analysis owner verify the calculation using the intended method. Assemble the sample-size inputs before accepting a calculator’s output as a campaign plan.
Put the calendar beside the statistical calculation
At 80 eligible new visitors per day in total, collecting 7,652 visitors takes at least approximately 96 days at an unchanged arrival rate. That means 40 visitors per group per day, not 80 in each group. With a seven-day outcome window, the final visitors need another seven days to complete observation. This hypothetical schedule ignores weekends, missing data and technical interruptions.
If a commercial decision is required after 28 days, only about 2,240 assignments are available, or 1,120 per group, before any loss of measurable outcomes. Under the planning probabilities, that corresponds to an average of 22.4 and 33.6 booking visitors respectively. These are expected counts for planning, not promised bookings. Extra decimal places cannot make the calendar more realistic.
Next assess whether the offer, staffing and customer need will remain sufficiently comparable throughout collection. A long experiment could cross a product replacement or a seasonal shift that makes the original question less relevant. Extending the calendar is therefore a design choice, not an automatic remedy for a small number of conversions.
Traffic forecasts need the same discipline. Use eligible, measurable new assignments rather than every reported click. OpenAI’s reporting documentation defines the platform metrics; your experiment may require a narrower population. A change in measurement coverage should be reflected in the feasibility estimate before it becomes a surprise at the results meeting.
Do traffic and calendar fit?
- Requirement
About 3,826 unique visitors per group = 7,652 total.
- Available in 28 days
80 per day × 28 = 2,240 total, only 1,120 per group.
- Collection time
7,652 / 80 rounded up = 96 days, plus outcome follow-up.
- Decision
Check whether the question survives 96 days or choose a clearly scoped measurement pilot.
Choose an alternative that still answers a useful question
One option is to test a larger, commercially meaningful change. This does not allow the team to assume that the effect will be larger. Specify which improvement the experiment needs to distinguish from zero and ask whether that improvement would change the decision. The guide to minimum detectable effect keeps that planning threshold separate from a forecast.
Another option is a measurement pilot. Verify that eligible visitors enter the study, bookings are recorded and outcomes mature as expected. A pilot can improve inputs for future planning without answering which page increases sales. Explicitly rename the objective so that a successful technical check is not later presented as evidence of commercial impact.
An earlier journey event, such as starting a booking, may provide more observations. It can support a separate question about behavior at that stage. Improvement there does not establish more completed bookings. Do not silently replace the primary outcome after purchase volume turns out lower than expected.
Report uncertainty after the test
After completion, we recommend the estimated effect and an appropriate confidence interval, rather than calculating observed power from the same measured effect. A wide interval already shows that commercially different possibilities remain. If another experiment is proposed, plan it prospectively using a relevant effect size and updated assumptions.
Record the final choice as a feasible experiment, a limited pilot or a postponed comparison. Any of these can be sensible. The essential requirement is that the team understands which knowledge the available budget can buy and which question will remain unanswered.
Sources and scope
Assess whether a conversion experiment is feasible at low volume and choose a useful alternative learning objective when it is not.
- OpenAI Ads: ReportingRead
- NIST: Sample sizes requiredRead
- NIST: Comparing two proportionsRead
- NIST: Selecting an experimental designRead
Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.
