
agents.incident-response·versija 1.0.0·melnraksts
Trauksme ir apstiprināta, ietekmes apjoms zināms, lietotāji aizsargāti, pakalpojums atjaunots, un gaita ierakstīta incidenta pēcanalīzei.
An alert fired or a person reports the production service is broken now. Not for a slow-burn degradation with no user impact — open a task and let the next working session pick it up. Not for the analysis afterwards — that is ops.incident-postmortem. If the cause is a bad release you already know about, this playbook and agents.ship-release (rollback) run together. If customers must be told publicly, management.crisis-communication takes over the messaging.
Jebkurš solis var gaidīt līdz datumam un atveras pats; katrs noslēgts solis atstāj pierādījumu (piezīmi, saiti, skaitli).
Mark the alert acknowledged in the monitoring page and open an incident record with the start time.
Izdarīts, kad the alert shows "acknowledged" and the incident record exists with a timestamp.
Check the service map and dashboards: which tenants, which endpoints, since when, is it total or partial? Compare against the last change (release, migration, config, certificate expiry).
Izdarīts, kad the incident record states affected services, tenants, and start time — or "unknown, investigating".
⛔ Do not touch anything yet. A change made during scoping destroys the evidence.
In order of preference: roll back the last release, disable the offending feature flag, scale up capacity, reroute traffic. Each step is one change, then re-check the dashboard.
Izdarīts, kad the error rate or the failing check visibly improves, or every reversible option is exhausted.
Probe the affected endpoints from outside the network (public checker or second location), not from the server itself.
Izdarīts, kad the outside probe returns correct responses for every affected endpoint.
⛔ A green check from inside the network proves nothing about what users see.
Apstiprinājums · S5 · īpašnieks — izpilde apstājas, līdz nosaukts cilvēks ieraksta lēmumu
Present impact, expected duration, and a draft status-page note. The owner approves, edits, or declines.
Izdarīts, kad the status page is updated or the owner has recorded "no public note".
⛔ The agent never posts to a status page or mails customers without this approval.
Apstiprinājums · S6 · īpašnieks — izpilde apstājas, līdz nosaukts cilvēks ieraksta lēmumu
Neatgriezenisks solis · S6 — aģents to nekad nenoslēdz viens
If mitigation needs something irreversible — deleting corrupted data, forcing a schema change, dropping a queue — present the exact command, the data at risk, and the alternative. Execute only on approval.
Izdarīts, kad the irreversible action is recorded in the incident record with the approver's name, or skipped.
Dashboards normal for 15 consecutive minutes; the outside probe passes three times in a row; no new alerts.
Izdarīts, kad the incident record shows the recovery window and the monitoring page is green.
Close the incident record with the full timeline: alert time, each action with its time, the cause as currently understood, open questions. Create the postmortem task (ops.incident-postmortem) and notify the owner with a three-line summary: what broke, what fixed it, what is still unknown.
Izdarīts, kad the task exists and the owner has the summary.
| Pazīme | Rīcība |
|---|---|
| No on-call owner reachable in 15 min | Escalate to the company owner by the rota's fallback; keep mitigating reversible steps only, record the delay. |
| Mitigation makes it worse | Roll back that one change immediately, note it in the timeline, return to S2 with the new information. |
| Cause unclear after all reversible options | Stabilise with a workaround (maintenance page, degraded mode), get the owner's decision in S5, do not guess at irreversible fixes. |
| Data may be lost or corrupted | Stop. Nothing is deleted. S6 decides with the owner; export what exists before any write. |
| Second alert fires during the incident | Scope it separately in the same record — shared cause or independent? Two independent incidents get two records. |
Incident record id · alert text and firing time · dashboard screenshots at scoping and at recovery · each mitigation command or action with its timestamp · the outside-probe results · the owner's S5 and S6 decisions (who, when) · the final timeline.
After every 5 runs ask: minutes from alert to acknowledgement and to first mitigation — which step waited? Did any mitigation get reverted (that step needs a pre-check)? Was any irreversible action taken without S6 approval, or any S5 communication sent late? Did the outside probe in S4 pass while users still saw the outage? A new version changes the step that caused the wait or the miss, and says so in its change note.
Nosaukums un kopsavilkums ir latviski. Detalizētā izpildes kārtība pagaidām ir kanoniskajā angļu valodas versijā; juridiskos un finanšu soļus publicēsim latviski tikai pēc cilvēka pārbaudes.
Obligātās sīkdatnes tur sarunu kopā. Analītika ir izslēgta, līdz tu atļauj — tā neliek nevienu sīkdatni un neglabā ierīces identifikatoru.Necessary storage keeps your conversation together. Analytics is off until you allow it — it sets no cookie and stores no device identifier. Ko mēs glabājamWhat we store