Experiments and evaluation

Choose the success metric before testing ChatGPT ads

Define the primary outcome for a ChatGPT ads test before results arrive, including the analysis unit, outcome window and rules for changing the plan.

Start reading
Editorial illustration: An unmarked outside caliper frames a pale ceramic vessel, with two other vessels behind it.
Editorial illustrationThe caliper's fixed opening represents a success criterion chosen before the test results are assessed.
The working guide

What you can work through.

Experiments and evaluation
  • Record the formula, analysis unit and observation window with the metric name.
  • Separate the primary outcome from diagnostics and discoveries made later.
  • Preserve the original plan when measurement problems require an amendment.

An advertising test can produce more clicks, fewer purchases and larger average orders at the same time. If the team chooses its definition of success after seeing those results, the conclusion depends on whichever number looks most attractive. For ChatGPT ads, we recommend a short, dated measurement protocol before outcomes become available. Another analyst should be able to read it and understand the decision that the test was intended to support.

This is a recommended internal working method, not a description of a native preregistration or randomization feature. The Center for Open Science explains preregistration as a way to distinguish planned analyses from analyses developed after examining data. The following workflow applies that distinction to an advertiser’s practical measurement choices.

Give the outcome a specific decision to support

Start with what the team might change. An advertiser choosing between two booking pages after a ChatGPT ad click can reasonably focus on completed bookings among the visitors assigned to those pages. A test of the advertisement itself has a different scope. The message may change who clicks, so conversion among clickers answers a narrower question than the effect across the original assigned audience.

Write down the action a favorable result would support. Keeping a small copy change is different from rebuilding the sales journey, replacing an offer or substantially increasing spending. Record the smallest improvement that would matter to the business. A commercial threshold and a statistical threshold serve different purposes: a precisely estimated improvement can still be too small to justify the work required.

Replace the metric name with an executable definition

“Qualified leads” is not a complete specification. State what qualifies a lead, who makes that judgment, when the judgment is made and how an existing customer is treated. Include the denominator. The percentage of assigned visitors with at least one qualified inquiry differs from the number of qualified inquiries per ad click.

OpenAI defines post-click CVR as conversions divided by clicks in its reporting documentation. That does not make it a unique-person proportion. Repeated clicks or multiple counted events require an analysis that respects their structure. Choose a definition that the available, appropriately collected data can support, rather than assuming every report column represents an independent observation.

Specify the event identity, duplicate handling, exclusions and data source used for the final calculation. OpenAI’s conversion tracking documentation describes connecting conversion events to campaigns. Your internal definition must additionally explain what constitutes a valid business outcome. Use a shared metric dictionary so that the analyst and sales team apply the same qualification rules.

A hypothetical booking-page protocol

Suppose an advertiser can independently run a properly randomized experiment on its own website after an ad click. That is an explicit assumption for this example, not a promised advertising feature. Each eligible new visitor receives a persistent assignment to page A or B. The primary outcome is the proportion of assigned visitors making at least one confirmed booking within seven days of first assignment.

The protocol excludes internal test visits using a rule established before launch. Returning visitors retain their original assignment. Each person contributes at most one positive outcome. The analysis estimates B minus A in percentage points, uses a prespecified two-sided interval, and waits until the final included visitor has completed the observation window. Seven days is a hypothetical internal choice, not a claim about OpenAI’s attribution window.

If A eventually records 32 booking visitors among 800 assigned visitors, its rate is 4%. If B records 40 among 800, its rate is 5%. The difference is 1 percentage point; the relative increase is 25%. Those calculations alone do not establish that changing the page is justified. The protocol specifies the uncertainty method and commercial threshold, as explained in the guide to confidence intervals for ad tests.

The protocol should also name the person responsible for extracting the data and the point at which it becomes final for this analysis. A result calculated while bookings are still arriving is not the same analysis as a mature result, even when the campaign dates on the report are identical. Preserve the extraction date alongside the eligibility dates.

Workflow

From booking goal to a fixed definition

  1. Choose the unit

    Each assigned visitor can contribute at most one positive booking outcome.

  2. Set the window

    Seven days from assignment is the example, not a platform window.

  3. Calculate the same outcome

    32/800 = 4%. 40/800 = 5%. Difference: 1 percentage point.

  4. Apply the planned rule

    Compare the uncertainty interval and business threshold before adoption.

Hypothetical website test after an ad click. The method requires valid independent assignment and measurement.

Decide what the other measurements are allowed to do

Clicks, abandoned forms and page load times can help explain an observed change. Label them diagnostic when they do not decide the main question. Identify any guardrails that can prevent adoption, such as invalid bookings or unacceptable workload for the sales team. Give each guardrail a defined role instead of treating it as another opportunity to declare a winner.

Avoid naming several primary outcomes without specifying how they will be assessed together. If bookings, booking value and qualified bookings can each independently justify success, the analysis needs to address multiple tests. For a small team, our practical recommendation is one primary outcome with clearly labeled supporting information. That makes the eventual tradeoff intelligible to people who did not design the experiment.

Preserve the version while allowing honest amendments

Save the protocol with a timestamp, owner, supporting assumptions and a durable version identifier. An editable file without history cannot establish what was decided before outcomes were known. Ask a colleague outside the creative team to calculate a small synthetic example from the written definitions. Conflicting answers expose ambiguity while it is still inexpensive to fix.

A tracking bug should be corrected. Record what failed, what results had already been seen and how the correction changes the main analysis. Keep the original specification visible instead of overwriting it. An unexpected pattern can still be useful, but treat it as a new hypothesis requiring a test with fresh data.

At the results meeting, place the protocol beside the report. Present the planned outcome even if another metric tells a more appealing story. Then explain deviations, additional findings and the commercial decision. This gives the team a useful record of both what it learned from ChatGPT advertising and which conclusions were specified before the outcome was known.

Sources and scope

Choose and record one primary outcome before a specific test of advertising in ChatGPT begins.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

See what campaign reporting covers.

Explore reporting in AthillyAds and how it fits alongside your own measurement of enquiries, purchases and other business outcomes.