Name the affected feature and the preserved path
Lead with what customers can observe: the product assistant is slower, returning a limited template, or temporarily unavailable. In the next sentence, state that cart and checkout remain available only if a real end-to-end test confirms it. Avoid the broad phrase 'the site is operational' when a visible feature is degraded.
Give a usable alternative such as product search, a size chart, email, phone, or staffed chat. Test the alternative before publishing it and state its hours or expected response. An alternative that feeds the same failed AI path is not a fallback.
Separate confirmed facts from investigation
Include the incident start time in the customer's timezone and the current impact. If the cause is not confirmed, say that the team is investigating elevated demand or a service issue. Do not attribute the event to bots, a provider, or an attack based only on a traffic label or early correlation.
State data status only from evidence. If there is no indication of a data incident, use that narrow wording; do not promise that no data was affected before the relevant logs and systems have been reviewed. Escalate any actual privacy concern to the separate response process.
- Affected feature: [specific assistant or page]
- Unaffected path, after test: [cart/checkout or other]
- Customer alternative: [tested route and hours]
- Started: [date, time, timezone]
- Data status: [confirmed statement only]
- Next update: [exact time, no more than 30 minutes during active impact]
- Closure criterion: [observable customer result]
Use update time instead of a guessed recovery time
Promise the next communication, not an unproven restoration time. A 30-minute update can say that the impact is unchanged, what has been tested, and when the next update will arrive. This keeps trust without forcing operators to rush an unsafe change to meet a guess.
Use one canonical status location and timestamp every revision. Support, social, and store banners should link or copy from that source so customers do not receive contradictory claims. Preserve prior versions for the incident record.
Prepare three stages before the incident
Draft an initial notice, an update, and a resolved notice with placeholders. The initial notice defines impact and alternatives; updates add verified actions and current state; resolution states the customer-facing recovery time and any monitoring period without turning the notice into a technical postmortem.
Assign an operational fact owner and a communications approver. The writer should not infer facts from raw dashboards alone, and the incident lead should not spend the entire response polishing customer copy. Pre-approve the tone and channels before launch days.
Close on a customer test
Publish 'resolved' only after the affected assistant succeeds through the public path, checkout still completes, and key metrics remain stable for the stated observation window. Provider recovery alone is insufficient if a stale cache or queue still degrades customers.
The retained notice kit should contain final texts, publication times, approvers, channels, screenshots, and the evidence supporting each claim. Use the guide to rehearse this communication alongside technical rollback so the first draft is not written under pressure.
Use the degraded-mode notice kit from “Tell customers why the AI assistant is degraded while checkout remains open” on a real first installation. Download AI Cost Circuit Breaker for free, begin in Monitoring, and move to enforcement only after the expected signals and rollback are verified.