An advertising test can produce more clicks, fewer purchases and larger average orders at the same time. If the team chooses its definition of success after seeing those results, the conclusion depends on whichever number looks most attractive. For ChatGPT ads, we recommend a short, dated measurement protocol before outcomes become available. Another analyst should be able to read it and understand the decision that the test was intended to support.
This is a recommended internal working method, not a description of a native preregistration or randomization feature. The Center for Open Science explains preregistration as a way to distinguish planned analyses from analyses developed after examining data. The following workflow applies that distinction to an advertiser’s practical measurement choices.
Give the outcome a specific decision to support
Start with what the team might change. An advertiser choosing between two booking pages after a ChatGPT ad click can reasonably focus on completed bookings among the visitors assigned to those pages. A test of the advertisement itself has a different scope. The message may change who clicks, so conversion among clickers answers a narrower question than the effect across the original assigned audience.
Write down the action a favorable result would support. Keeping a small copy change is different from rebuilding the sales journey, replacing an offer or substantially increasing spending. Record the smallest improvement that would matter to the business. A commercial threshold and a statistical threshold serve different purposes: a precisely estimated improvement can still be too small to justify the work required.
Replace the metric name with an executable definition
“Qualified leads” is not a complete specification. State what qualifies a lead, who makes that judgment, when the judgment is made and how an existing customer is treated. Include the denominator. The percentage of assigned visitors with at least one qualified inquiry differs from the number of qualified inquiries per ad click.
OpenAI defines post-click CVR as conversions divided by clicks in its reporting documentation. That does not make it a unique-person proportion. Repeated clicks or multiple counted events require an analysis that respects their structure. Choose a definition that the available, appropriately collected data can support, rather than assuming every report column represents an independent observation.
Specify the event identity, duplicate handling, exclusions and data source used for the final calculation. OpenAI’s conversion tracking documentation describes connecting conversion events to campaigns. Your internal definition must additionally explain what constitutes a valid business outcome. Use a shared metric dictionary so that the analyst and sales team apply the same qualification rules.
A hypothetical booking-page protocol
Suppose an advertiser can independently run a properly randomized experiment on its own website after an ad click. That is an explicit assumption for this example, not a promised advertising feature. Each eligible new visitor receives a persistent assignment to page A or B. The primary outcome is the proportion of assigned visitors making at least one confirmed booking within seven days of first assignment.
The protocol excludes internal test visits using a rule established before launch. Returning visitors retain their original assignment. Each person contributes at most one positive outcome. The analysis estimates B minus A in percentage points, uses a prespecified two-sided interval, and waits until the final included visitor has completed the observation window. Seven days is a hypothetical internal choice, not a claim about OpenAI’s attribution window.
If A eventually records 32 booking visitors among 800 assigned visitors, its rate is 4%. If B records 40 among 800, its rate is 5%. The difference is 1 percentage point; the relative increase is 25%. Those calculations alone do not establish that changing the page is justified. The protocol specifies the uncertainty method and commercial threshold, as explained in the guide to confidence intervals for ad tests.
The protocol should also name the person responsible for extracting the data and the point at which it becomes final for this analysis. A result calculated while bookings are still arriving is not the same analysis as a mature result, even when the campaign dates on the report are identical. Preserve the extraction date alongside the eligibility dates.
From booking goal to a fixed definition
- Choose the unit
Each assigned visitor can contribute at most one positive booking outcome.
- Set the window
Seven days from assignment is the example, not a platform window.
- Calculate the same outcome
32/800 = 4%. 40/800 = 5%. Difference: 1 percentage point.
- Apply the planned rule
Compare the uncertainty interval and business threshold before adoption.
Decide what the other measurements are allowed to do
Clicks, abandoned forms and page load times can help explain an observed change. Label them diagnostic when they do not decide the main question. Identify any guardrails that can prevent adoption, such as invalid bookings or unacceptable workload for the sales team. Give each guardrail a defined role instead of treating it as another opportunity to declare a winner.
Avoid naming several primary outcomes without specifying how they will be assessed together. If bookings, booking value and qualified bookings can each independently justify success, the analysis needs to address multiple tests. For a small team, our practical recommendation is one primary outcome with clearly labeled supporting information. That makes the eventual tradeoff intelligible to people who did not design the experiment.
Preserve the version while allowing honest amendments
Save the protocol with a timestamp, owner, supporting assumptions and a durable version identifier. An editable file without history cannot establish what was decided before outcomes were known. Ask a colleague outside the creative team to calculate a small synthetic example from the written definitions. Conflicting answers expose ambiguity while it is still inexpensive to fix.
A tracking bug should be corrected. Record what failed, what results had already been seen and how the correction changes the main analysis. Keep the original specification visible instead of overwriting it. An unexpected pattern can still be useful, but treat it as a new hypothesis requiring a test with fresh data.
At the results meeting, place the protocol beside the report. Present the planned outcome even if another metric tells a more appealing story. Then explain deviations, additional findings and the commercial decision. This gives the team a useful record of both what it learned from ChatGPT advertising and which conclusions were specified before the outcome was known.
Sources and scope
Choose and record one primary outcome before a specific test of advertising in ChatGPT begins.
- OpenAI Ads: ReportingRead
- OpenAI Ads: Conversion TrackingRead
- Center for Open Science: PreregistrationRead
Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.
