Experiments and evaluation

Is the response to ChatGPT ads a novelty effect?

Assess whether early response to ChatGPT ads can last. Compare mature cohorts and business outcomes before turning launch engagement into an ongoing budget forecast.

Start reading
Editorial illustration: A teal teapot and cup on a pale plinth beside large coral gift ribbons and torn wrapping paper.
Editorial illustrationThe lively unwrapping and quiet steam represent the distinction between initial attention and usefulness that needs observation over time.
The working guide

What you can work through.

Experiments and evaluation
  • Falling CTR proves neither novelty nor declining business value.
  • Compare cohorts with equal follow-up and qualification rules.
  • Test durability under the conditions relevant to the budget decision.

An unusually strong first week makes a new ChatGPT campaign difficult to judge. Click volume looks promising and the business wants to decide how much budget the channel deserves. Yet the initial response may combine curiosity about the ad, interest in a temporary offer and purchase demand that existed before launch.

Novelty is one possible explanation for changing response as something becomes familiar. It is not a diagnosis available directly from a falling CTR chart. The process below is our recommended internal approach to identifying what further evidence is needed before early performance becomes a long-term forecast.

Identify what is new to the customer

The novelty could concern the channel, message, product or offer. If all four arrive together, the initial response cannot tell you which component mattered. Record what customers could encounter before launch and what the campaign introduced.

Define the intended lifetime of the activity too. An ad for a limited launch offer may succeed even when response disappears after the offer ends. A decision about continuing advertising requires evidence from more representative conditions. These are different business questions and deserve different evaluation windows.

Microsoft Research discusses changes in experimental results over time. The general issue applies here: an observed result during one period is not automatically a reliable prediction for future customers. That research does not establish performance or novelty effects for ChatGPT advertising.

Put business response beside the click curve

Choose an outcome aligned with the budget decision, such as qualified inquiries or completed purchases after returns. Clicks help explain behavior, but greater activity need not mean that more relevant customers progress. An unclear message can attract curiosity without generating useful business.

Consider a hypothetical campaign with 20,000 impressions in each of three weeks. Click counts are 600, 400 and 300, producing CTRs of 3.0%, 2.0% and 1.5%. After identical follow-up windows, the defined click cohorts contain 30, 28 and 27 qualified leads respectively.

Click response halves between the first and third weeks, while qualified lead count falls by 10%. That warrants a different investigation from saying the campaign lost half its business value. Equally, three weeks and small counts cannot establish a stable effect. These invented figures illustrate the distinction between metrics; they are not observed campaign results.

Report each denominator explicitly. Qualified leads per reported click is a descriptive ratio, not necessarily the probability that a unique person becomes qualified. Repeat clicks, duplicate submissions and cross-period activity require consistent handling before comparisons become interpretable.

Compare the alternatives

Curiosity and business response can move differently

  1. Week 1

    600 clicks produce 3% CTR. 30 qualified leads in the defined cohort.

  2. Week 2

    400 clicks produce 2% CTR. 28 qualified leads under the same assessment rule.

  3. Week 3

    300 clicks produce 1.5% CTR. 27 qualified leads, with uncertainty still unresolved.

Hypothetical weeks with 20,000 impressions each and equally mature lead assessment. The comparison does not establish causality.

Compare cohorts of the same age

Allow every cohort equivalent time to purchase or complete qualification. If first-week inquiries receive two weeks of sales follow-up and recent inquiries receive two days, the operating process creates part of the apparent decline. Preserve the initial contact date and the date of the outcome being counted.

Define which cohort owns a person or order and how duplicates are handled. A returning visitor does not become a new customer because the team retrieves another report. The guide to conversion-rate cohorts helps establish that boundary.

OpenAI documents conversion tracking, but a tracked event does not itself establish durable influence. Our recommendation is to supplement attributed outcomes with consistent advertiser-owned quality and customer-status definitions. Identify assessments performed in the CRM or order system rather than implying they are native advertising metrics.

Investigate competing explanations for the curve

Review budget, bids, offer, destination, stock and actual delivery. Increasing investment can bring a different visitor mix. Changing a form can change lead quality. A slower website can reduce conversion without reducing interest in the ad. Each explanation needs evidence of its own.

Inspect seasonality and simultaneous events as well. A decline after payday differs from the same people gradually responding less to a familiar message. An aggregate weekly report often cannot distinguish those mechanisms.

Avoid statements about individual exposure when the required information is unavailable. Calendar weeks since launch are not the same as the number of times a particular person saw an ad. This workflow assumes no access to individual exposure histories, frequency controls or native segments based on first advertising contact.

Distinguish creative fatigue from novelty cautiously. Both labels can describe falling engagement, but naming a mechanism is stronger than observing a pattern. If the available evidence only establishes that aggregate CTR declined while delivery changed, state that narrower conclusion.

Set observation windows before the initial peak

Specify which initial period is reviewed for operational issues and which period should support the budget decision. Justify those windows using purchase cycles, outcome maturity and relevant calendar patterns. There is no universal number of days after which novelty is exhausted or stable response is guaranteed.

A planned separate analysis of the initial period can be useful. Removing early days afterwards because the remaining curve looks more convincing changes the analysis. Preserve the full series and label retrospective alternatives exploratory. A positive remaining average does not automatically become causal evidence.

For an actual experiment, assess precision for the later period too. A study powered for the whole duration may lack power for a smaller subset. A broad interval around the later effect leaves durability unresolved even if the full-period result is positive.

Choose the next decision with explicit limits

Save comparable report extracts and record evidence for and against sustained response. If delivery is stable but business outcomes remain immature, further follow-up may be appropriate. If conditions changed substantially, define a new study for the new situation.

A planned replication can investigate whether the conclusion recurs under relevant conditions. Changing every component and obtaining similar totals is not a convincing replication. The useful decision record identifies the response observed, the period it represents and what must still be demonstrated before the ongoing budget depends on it.

Sources and scope

Distinguish a possible temporary launch response from sustained business response and determine what evidence is needed before expanding investment.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

See what campaign reporting covers.

Explore reporting in AthillyAds and how it fits alongside your own measurement of enquiries, purchases and other business outcomes.