Minute 0: hold the ceiling and open one record
Create one incident record with a UTC start time, incident lead, business contact, current limit state, recent changes, and next review time. Preserve the absolute spending ceiling already approved; do not raise it merely because a campaign appears successful or lower it merely because traffic looks unfamiliar.
Freeze unrelated releases and export the last normal hour. Record attempts per minute, tokens per minute, orders, payment failures, p95 AI latency, repeated fingerprints, and classification confidence. The baseline makes later changes interpretable.
Minute 5: verify that the event is real
Check whether edge traffic, WordPress AI attempts, and provider usage all rose in the same window. Confirm that alerts are not duplicated and clocks align. A dashboard refresh problem or delayed provider chart should not trigger a production control change.
Ask the business contact whether campaign placements, influencer posts, or email sends began as planned. Look for completed orders and valid payment attempts, not page views alone. Genuine revenue does not prove every request is human, but it changes the cost of a broad shutdown.
| Checkpoint | Evidence required | Permitted action | Next review |
|---|---|---|---|
| 5 min | signal validity and campaign confirmation | hold ceiling; freeze unrelated changes | 15 min |
| 15 min | source/path concentration and orders | narrow one affected source or repetition pattern | 30 min |
| 30 min | customer outcome after change | retain, revert, or tighten the one rule | 60 min |
| 60 min | usage forecast and sustained revenue | approve time-bounded operating mode | named time |
Minute 15: find the narrowest concentration
Segment by route, source, request fingerprint, repetition, model, and context. Compare each segment with orders or inquiries. A single product question repeated hundreds of times may be constrained without slowing cart help, while a broad mix of unique questions that precede purchases may warrant preserved capacity.
Treat Human, Bot, and Unknown as signals with confidence, not verdicts. An Unknown surge during a privacy or proxy change needs corroboration. Prefer a rule tied to proven velocity and repetition on one path, with an expiry, over a site-wide classification block.
Minute 30: judge the change by customer outcomes
After one narrow change, compare AI completion, checkout conversion, payment failures, repeated traffic, and tokens per order with the preceding slice. Revert if customer outcomes worsen without material exposure reduction. Do not stack a second control before the first has an observable window unless the hard ceiling is at risk.
Update the customer-facing status only with confirmed impact. If checkout remains available but assistance is slower, say exactly that and offer an alternative. Set the next update time even if there is no new finding; silence during a visible campaign defect creates unnecessary support load.
Minute 60: leave a bounded operating decision
Forecast remaining campaign usage from the latest stable interval and compare it with the ceiling and revenue path. Approve a temporary rule only with an owner, affected path, threshold, start, expiry, rollback condition, and next review. Hand off if the incident will outlast the current operator.
The response sheet is complete when every checkpoint contains observed numbers, decision, reason, and reviewer. Closure requires stable customer outcomes, controlled usage, expired emergency rules, and an after-action owner; the guide can then turn the timeline into a rehearsal scenario.
Use the midnight demand response sheet from “A campaign went viral at midnight: the first hour without a blunt shutdown” on a real first installation. Download AI Cost Circuit Breaker for free, begin in Monitoring, and move to enforcement only after the expected signals and rollback are verified.