Agency work and campaign operations

Build an incident runbook for active ChatGPT campaigns

Prepare an incident runbook for active ChatGPT campaigns with clear ownership, evidence capture, containment decisions and recovery checks.

Start reading
Editorial illustration: A miniature track has a yellow gate at a branch and a repaired rail joint visible through a magnifying glass.
Editorial illustrationThe gated branch and inspected joint illustrate how incident response needs both containment of impact and verification of the original fault.
The working guide

What you can work through.

Agency work and campaign operations
  • Name one response owner before parallel work begins.
  • Keep observations and attempted actions in a shared record.
  • Define recovery through the original failure, not a completed task label.

An incident runbook should be usable when the person who normally understands the campaign is unavailable. It needs to tell the next operator where to look, who can decide and how to avoid making the situation harder to reconstruct. A long policy document that nobody can navigate during an incident does not meet that need.

For ChatGPT advertising, build the runbook around a small number of recurring failure types: an incorrect destination, unexpected campaign activity, a measurement problem or a client-reporting defect. Keep detailed diagnosis in separate guides. The runbook coordinates the response across those cases.

Prepare the contact and authority map

Identify the operational owner, the person who can approve consequential campaign changes and the contacts for the website and measurement implementation. Include an agreed alternative when the usual contact is unavailable. Use role-based references where practical so staff changes do not silently invalidate the document.

Write down which actions the operator may take under the existing working arrangement. An incident should not require guessing whether a previously authorized containment action is allowed. Equally, the runbook should not grant new authority merely because an issue is urgent.

Keep the contact map accessible to the relevant team without embedding passwords or API keys. Test that the on-duty operator can find the actual account and supporting records. Access preparation is part of readiness; improvising credential sharing during a live problem is a poor substitute.

Open one incident record

At the start of a response, assign an incident reference and one person responsible for coordinating it. Record the affected account, known resource identifiers, the observed symptom and the time it was noticed. If the beginning of the problem is unknown, preserve that uncertainty.

Use a shared sequence of observations and actions. Each entry should say what was checked or changed, by whom, when and what the result showed. A proposed action, an attempted action and a verified effect are three different states. Label them clearly.

This record helps parallel work stay coherent. One colleague can inspect the destination while another checks campaign state, provided both know the current scope and report their observations in the same place. Avoid several people making overlapping corrective edits without a coordinating owner.

Scope before choosing containment

Identify what is affected and whether exposure appears to be continuing. The problem might concern one ad, several resources sharing a destination or a report that has not yet left the agency. Use defect severity assessment to connect the observed scope with an appropriate response.

Do not assume that a symptom reveals its cause. A missing report row does not by itself establish that a campaign stopped serving. An incorrect destination does not establish that every ad uses it. Record the checks that distinguish these possibilities before presenting a diagnosis as certain.

Choose a proportionate containment action through the authorized workflow. Explain what it is intended to prevent and what it may affect. Preserve the relevant earlier state before making changes where that is practical. Do not promise that a proposed action has stopped exposure until the resulting state has been checked.

Workflow

One coordinated response

  1. Owner and scope

    One owner connects the account, symptom and observations.

  2. Authorized containment

    The action has an intended purpose and a check of its actual effect.

  3. Verified recovery

    Check the original failure and assign the follow-up work.

    Checkpoint
The coordinating procedure points to diagnosis for the observed failure.

Route the technical work to a focused procedure

The runbook should link to specific diagnosis rather than contain every possible API response or website error. For a destination incident, use the broken destination guide. For another failure type, identify the relevant system owner and current technical documentation. If spend appears to exceed the approved allowance, follow the overspend-response procedure to verify the comparison and coordinate authorised containment.

Keep the platform’s procedures separate from internal coordination. OpenAI campaign management describes available operations. The runbook explains who uses an applicable operation in this organization, with what decision and what subsequent check.

If several changes are necessary, sequence them so their effects remain understandable. Fixing the destination, changing creative and modifying measurement at the same time can make it difficult to identify what restored the service. Sometimes parallel corrections are necessary, but the record should show that choice and its limits.

Communicate what is known

Use a brief status format: observed impact, current scope, action taken, unresolved question and next update trigger. A trigger can be the completion of a diagnostic check or a material change in the situation. Avoid inventing a recovery time when the evidence does not support one.

Separate internal technical detail from the information the client needs for a decision. The client may need to know that an offer is temporarily affected and who is handling it, while the engineering team needs request identifiers and diagnostic results. Both messages should describe the same known state.

Record who is responsible for communication so it does not disappear between technical tasks. Also record corrections to earlier statements when new evidence changes the picture. An honest update is more useful than maintaining a confident but outdated explanation.

Verify recovery and leave follow-up work visible

Define recovery against the original failure. If the wrong offer was displayed, check the actual destination and content. If reporting was incorrect, reproduce the affected output. A task marked complete is not sufficient evidence that the user’s problem has been resolved.

Close the immediate response when the agreed operating condition is verified and ongoing ownership is clear. Keep remaining work in named follow-up items rather than treating every improvement as a reason to leave the incident indefinitely open.

Schedule a proportionate post-incident review. Update the runbook when a missing contact, ambiguous authority or ineffective check was part of the difficulty. The document should become easier to use after real experience, with each revision connected to a demonstrated operational need.

Sources and scope

Prepare a concise operating procedure for scoping an incident, assigning ownership, containing impact and verifying recovery.

Working methods and examples are editorial suggestions. Check current platform requirements and available features before implementation.

Your next chapter

Explore the workflow for your clients.

See the agency workflow in AthillyAds, from a client's website to campaign material your team can review.