Experiments and evaluation

Design a geographic test for ChatGPT ads

Assess whether a geographic test can measure ChatGPT advertising impact. Define market boundaries, business outcomes, comparison logic and uncertainty before launch.

Start reading
Editorial illustration: A stitched landscape with colored regions, small houses, rivers and bridges between areas.
Editorial illustrationThe bounded areas and bridges illustrate regional comparability and possible routes for influence across boundaries.
The working guide

What you can work through.

Experiments and evaluation
  • Geographic areas are the experimental units, not impressions.
  • Verify geographic controls and outcome measurement before selecting markets.
  • A calculated difference needs a credible counterfactual and uncertainty estimate.

A geographic test starts with a map of what the advertiser can actually control and measure. Splitting a sales report into two areas does not create a control group for ChatGPT ads. The advertising intervention must differ according to an explicit plan, while business outcomes remain measurable on comparable terms in the areas where the intervention does not change.

The approach below is our recommended internal planning method. It does not describe a native geographic experimentation feature in ChatGPT Ads or establish that a particular account supports the required boundaries. Rejecting an infeasible design is a valid outcome of this work.

Define the intervention before drawing boundaries

Write the question in business terms: how does a specified increase in advertising change completed purchases during the study? This differs from asking how many purchases a report attributes to an ad. Attribution applies a crediting rule. An effect study estimates what would have happened without the intervention.

Decide whether the outcome is orders, new paying customers or contribution after returns. If the question concerns total business impact, measure the relevant business outcome across each area. Restricting the outcome to people who clicked can omit other responses and selects a population partly created by the advertising itself.

Distinguish adding a channel from increasing an existing investment. A test of additional spending estimates the effect of that increment under the tested conditions. It cannot automatically establish the contribution of every existing campaign. Put that boundary in the original brief so the eventual result answers a question the business intended to ask.

Verify that the proposed map is operational

OpenAI documents campaign settings, but settings alone do not constitute a geographic experiment. Verify which boundaries your actual implementation supports and how other campaigns could overlap. Do not assume that cities, postal codes or custom geographic lists are available.

Choose a separate rule for assigning business outcomes to areas. A delivery address might suit physical orders; a contractual business location might suit corporate customers. Apply the same rule before and during the study. Keep unknown locations visible and check whether their share changes rather than quietly dropping them.

Identify commuting, travel, national offers and other paths for spillover. Someone may encounter advertising in one area and buy in another. Substantial spillover can weaken the contrast or change what the study estimates. The test contamination review helps turn that concern into specific checks for the campaign and business teams.

Select areas without seeing their test results

Assemble historical periods with consistent definitions and a record of local events. Compare levels, trends, weekly patterns and volatility. Areas with similar sales totals may behave differently around paydays or changes in delivery service. One enormous area can dominate the analysis even when the map appears balanced.

Our recommendation is to freeze selection and matching rules before observing test outcomes. If independent areas can genuinely receive separate interventions, the external experiment plan may specify random assignment, potentially within comparable pairs. The advertiser must implement and document that assignment; it should never be presented as an assumed platform feature.

Google Research describes geographic experiments with area-level treatment assignment. NIST explains blocking as a way to account for known nuisance factors in a design. These are methodological references, not evidence that a given ChatGPT account offers the required controls.

A calculation that deliberately stops short of a verdict

Consider a hypothetical example. The treatment areas record 1,000 orders before the study and 1,150 during it. Comparison areas move from 800 to 880 orders over equally long periods. Suppose historical evidence supports a proportional comparison. The comparison areas grew by 10%, implying a counterfactual of 1,100 orders for the treatment areas.

The estimated difference is then 1,150 minus 1,100, or 50 orders. Relative to the counterfactual, that is approximately 4.5%. It is different from the treatment areas’ observed 15% growth. These invented numbers demonstrate the arithmetic; they are not campaign results or expected performance.

That calculation assumes the comparison area’s growth represents the treatment area’s alternative path. A local store expansion or different price changes could invalidate the assumption. Four aggregate totals also provide insufficient information for a credible uncertainty assessment. Preserve observations for each actual area and period alongside selection and assignment records.

Count independent areas, not impressive traffic totals

Thousands of ad impressions do not replace independent geographic units. When intervention is assigned by area, analysis must respect that assignment and dependence over time. Ask the analyst what minimum detectable effect, or MDE, the proposed design can detect at the chosen statistical power. Use historical variation and the intended advertising contrast rather than an arbitrary click target.

An uncertainty interval covering both meaningful gain and meaningful loss produces an inconclusive result. It establishes neither success nor absence of an effect. If the available areas and affordable duration cannot deliver useful precision, narrow the decision before committing a larger budget to the study.

Inspect historical fit without selecting only the prettiest period. A relationship that works in an ordinary month may fail during a promotion. Document such limitations and choose a study window relevant to the future budget decision. A test during an exceptional sales event may be useful specifically for that event, while remaining weak evidence for ordinary trading weeks.

Decision guide

Three gates before a geographic test

  1. Can treatment be separated?

    Verify available geographic controls and how spillover into comparison areas will be detected.

  2. Can outcomes be measured consistently?

    Use the same order definition, location rule and observation period in every area.

  3. Are there enough independent areas?

    Evaluate between-area variation and the target effect size before committing the budget.

Internal assessment model. A failed gate requires another design or a narrower conclusion.

Preserve assignment, execution and outcome separately

Keep the planned assignment, actual campaign changes and execution deviations as separate records. OpenAI Reporting supplies delivery and attributed metrics; add the advertiser’s regional business outcomes. A planned budget difference does not prove an equivalent difference in exposure or expenditure.

Predefine how local promotions, stock shortages and failed delivery will affect interpretation. Review holdout feasibility before reserving markets or promising a causal answer. End planning with a concrete decision: run the specified design, revise the question or decline the study. A documented feasibility limit is more useful than a precise number from a comparison that cannot support the intended conclusion.

Sources and scope

Design or reject a geographic advertising effect study based on actual delivery controls, comparable markets and measurable business outcomes.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

See what campaign reporting covers.

Explore reporting in AthillyAds and how it fits alongside your own measurement of enquiries, purchases and other business outcomes.