
agents.incident-response·version 1.0.0·draft
The alert is acknowledged, the blast radius is known, users are protected, service is restored, and the run is recorded for the postmortem.
An alert fired or a person reports the production service is broken now. Not for a slow-burn degradation with no user impact — open a task and let the next working session pick it up. Not for the analysis afterwards — that is ops.incident-postmortem. If the cause is a bad release you already know about, this playbook and agents.ship-release (rollback) run together. If customers must be told publicly, management.crisis-communication takes over the messaging.
Any step can wait until a date and reopens by itself; every closed step leaves evidence (a note, a link, a number).
Mark the alert acknowledged in the monitoring page and open an incident record with the start time.
Done when the alert shows "acknowledged" and the incident record exists with a timestamp.
Check the service map and dashboards: which tenants, which endpoints, since when, is it total or partial? Compare against the last change (release, migration, config, certificate expiry).
Done when the incident record states affected services, tenants, and start time — or "unknown, investigating".
⛔ Do not touch anything yet. A change made during scoping destroys the evidence.
In order of preference: roll back the last release, disable the offending feature flag, scale up capacity, reroute traffic. Each step is one change, then re-check the dashboard.
Done when the error rate or the failing check visibly improves, or every reversible option is exhausted.
Probe the affected endpoints from outside the network (public checker or second location), not from the server itself.
Done when the outside probe returns correct responses for every affected endpoint.
⛔ A green check from inside the network proves nothing about what users see.
Approval · S5 · owner — the run stops until a named person records the decision
Present impact, expected duration, and a draft status-page note. The owner approves, edits, or declines.
Done when the status page is updated or the owner has recorded "no public note".
⛔ The agent never posts to a status page or mails customers without this approval.
Approval · S6 · owner — the run stops until a named person records the decision
Irreversible · S6 — an agent never closes it alone
If mitigation needs something irreversible — deleting corrupted data, forcing a schema change, dropping a queue — present the exact command, the data at risk, and the alternative. Execute only on approval.
Done when the irreversible action is recorded in the incident record with the approver's name, or skipped.
Dashboards normal for 15 consecutive minutes; the outside probe passes three times in a row; no new alerts.
Done when the incident record shows the recovery window and the monitoring page is green.
Close the incident record with the full timeline: alert time, each action with its time, the cause as currently understood, open questions. Create the postmortem task (ops.incident-postmortem) and notify the owner with a three-line summary: what broke, what fixed it, what is still unknown.
Done when the task exists and the owner has the summary.
| Symptom | Response |
|---|---|
| No on-call owner reachable in 15 min | Escalate to the company owner by the rota's fallback; keep mitigating reversible steps only, record the delay. |
| Mitigation makes it worse | Roll back that one change immediately, note it in the timeline, return to S2 with the new information. |
| Cause unclear after all reversible options | Stabilise with a workaround (maintenance page, degraded mode), get the owner's decision in S5, do not guess at irreversible fixes. |
| Data may be lost or corrupted | Stop. Nothing is deleted. S6 decides with the owner; export what exists before any write. |
| Second alert fires during the incident | Scope it separately in the same record — shared cause or independent? Two independent incidents get two records. |
Incident record id · alert text and firing time · dashboard screenshots at scoping and at recovery · each mitigation command or action with its timestamp · the outside-probe results · the owner's S5 and S6 decisions (who, when) · the final timeline.
After every 5 runs ask: minutes from alert to acknowledgement and to first mitigation — which step waited? Did any mitigation get reverted (that step needs a pre-check)? Was any irreversible action taken without S6 approval, or any S5 communication sent late? Did the outside probe in S4 pass while users still saw the outage? A new version changes the step that caused the wait or the miss, and says so in its change note.
Obligātās sīkdatnes tur sarunu kopā. Analītika ir izslēgta, līdz tu atļauj — tā neliek nevienu sīkdatni un neglabā ierīces identifikatoru.Necessary storage keeps your conversation together. Analytics is off until you allow it — it sets no cookie and stores no device identifier. Ko mēs glabājamWhat we store