Experiments and evaluation

Write a campaign hypothesis that can be proved wrong

Turn a ChatGPT advertising idea into a falsifiable test hypothesis. Define the decision, mechanism, outcome and evidence that would count against it.

Start reading
Editorial illustration: A weight hangs above an intact brick arch with loose clay pieces beside it.
Editorial illustrationThe arch awaiting a load evokes a hypothesis that must be open to testing and possible contradiction.
The working guide

What you can work through.

Experiments and evaluation
  • Name the decision, population, change and comparison.
  • Specify an observable mechanism and a result that would contradict it.
  • Keep business thresholds and design feasibility visible before launch.

“We should test a clearer advertisement” is an idea, not yet a hypothesis. For a ChatGPT campaign, turn that idea into a statement about a specific change, a plausible mechanism and an observable business outcome. The statement should permit an unfavorable result without being rewritten to declare success.

A hypothesis register is a small working document that keeps these decisions visible before results arrive. It need not be a complex research system. Its value comes from separating what the team expected, what it changed and what the evidence later showed.

Start with the decision the test would change

Write the operational choice first. Perhaps the advertiser must decide whether to include a price range in an offer, adopt a different landing-page explanation or continue a particular campaign approach. If either result would lead to exactly the same action, the proposed test may have little immediate decision value.

Specify the intended population and context. A hypothesis about qualified enquiries for a business service should not quietly become a claim about all ChatGPT advertisers. Market, offer, eligibility and the tested campaign period define where the conclusion could reasonably apply.

Also record the current alternative. A treatment described as a better version cannot be reproduced. Preserve the actual message, destination or policy that serves as the comparison, with a version identifier where practical.

Make the mechanism explicit

In a hypothetical service campaign, the team believes that showing a realistic price range will discourage unsuitable enquiries and help suitable prospects decide sooner. That mechanism predicts a quality change, not necessarily a higher click-through rate. The hypothesis should therefore name an outcome that can observe the claimed improvement.

One possible statement is: within the eligible campaign population, the price-range version will increase qualified enquiries per assigned comparison unit enough to justify its implementation cost. The exact unit and allocation method depend on the feasible design and must be specified before launch.

Microsoft Research’s pre-experiment guidance supports forming clear hypotheses and selecting metrics capable of testing them. The register proposed here applies that principle to an advertiser’s own planning. It does not imply an automatic experiment feature in ChatGPT Ads.

Write down what would count against the idea

Ask how the proposed mechanism could fail. The price range may deter suitable prospects as well as unsuitable ones. It may attract clicks without improving qualification. It may reduce lead volume so much that better average quality does not compensate for the lost opportunities.

Record these possibilities before looking at results. Otherwise the team can switch from praising volume to praising quality whenever one of them happens to improve. A useful hypothesis makes the tradeoff explicit enough that an unfavorable combination remains unfavorable after the data arrives.

Separate the primary outcome from diagnostic measures. Clicks can help explain where behavior changed, but they should not replace qualified enquiries as the success measure halfway through the test. Use metric preregistration to fix that distinction in the analysis plan.

Compare the alternatives

From aspiration to a test

  1. Aspiration

    Clearer ads should perform better.

  2. Mechanism

    An early price range may reduce mismatched expectations.

  3. Testable consequence

    More qualified enquiries per assigned comparison unit.

  4. Possible contradiction

    Fewer enquiries without enough improvement in quality.

Hypothetical example for a ChatGPT quote campaign.

Attach a meaningful decision threshold

The smallest worthwhile improvement is a business judgment, not whatever change happens to achieve a favorable statistical result. Estimate the implementation, operating and opportunity costs that the change must justify. Then translate that judgment into an effect on the chosen outcome.

The minimum detectable effect guide distinguishes the worthwhile effect from the study’s planned sensitivity. Keep both in the register. A design capable only of detecting a very large change may not resolve a decision where a modest improvement would already matter.

Avoid inserting an unsupported percentage simply to complete a template. If the economic threshold is uncertain, write down the range and the assumption responsible. Resolving that assumption may be a more valuable next task than launching the experiment immediately.

Check that the hypothesis is testable in practice

Confirm that the change can be delivered distinctly, the comparison can be maintained and the outcome can be observed at the required maturity. Two separately configured campaigns do not automatically create a randomized comparison. Platform delivery choices and different audiences can affect who sees each version.

Use holdout feasibility before claiming an incremental-effect experiment. For a descriptive campaign comparison, label the inference accordingly. OpenAI reporting supplies definitions for campaign metrics, while the advertiser remains responsible for the design and interpretation of its own test.

Finally, assign an owner and a revision rule. The hypothesis can change before the test if new information warrants it, but preserve the earlier version and the reason. After results arrive, record the outcome against the original decision statement. The register then becomes a useful record of what was learned rather than a collection of ideas that all appear successful in hindsight.

Sources and scope

Turn an advertising idea into a falsifiable, decision-linked hypothesis before selecting an experimental design.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

See what campaign reporting covers.

Explore reporting in AthillyAds and how it fits alongside your own measurement of enquiries, purchases and other business outcomes.