Izveidot kontuCreate account
‹ Visas procedūras
Novērst darbības vides incidentu

agents.incident-response·versija 1.0.0·melnraksts

Novērst darbības vides incidentu

Trauksme ir apstiprināta, ietekmes apjoms zināms, lietotāji aizsargāti, pakalpojums atjaunots, un gaita ierakstīta incidenta pēcanalīzei.

TomsTehnoloģiju vadītājsvadaProfils ›
Kadpēc notikuma — an alert fires — monitoring page, error-rate threshold, or a person reports the service is down
Kas rīkojasaģents rīkojas pēc apstiprinājuma
Laiks10 min to first mitigation; 30–120 min to restored
Valstsjebkura valsts
Pieraksties, lai izpildītuŠī procedūra atveras Brain Club iekšpusē. Pieraksties, lai to lasītu un izpildītu.

Kad izmantot

An alert fired or a person reports the production service is broken now. Not for a slow-burn degradation with no user impact — open a task and let the next working session pick it up. Not for the analysis afterwards — that is ops.incident-postmortem. If the cause is a bad release you already know about, this playbook and agents.ship-release (rollback) run together. If customers must be told publicly, management.crisis-communication takes over the messaging.

Pirms sāc

  • The on-call owner is named and reachable; the agent works, the owner decides.
  • You know the current release version of the affected service and when it last changed.
  • Access exists: dashboards, logs, and the infra module for the affected tenant.

Ko izpilde pieprasa3

  • Apstiprinājums · S5 · īpašnieksizpilde apstājas, līdz nosaukts cilvēks ieraksta lēmumu
  • Apstiprinājums · S6 · īpašnieksizpilde apstājas, līdz nosaukts cilvēks ieraksta lēmumu
  • Neatgriezenisks solis · S6aģents to nekad nenoslēdz viens

Jebkurš solis var gaidīt līdz datumam un atveras pats; katrs noslēgts solis atstāj pierādījumu (piezīmi, saiti, skaitli).

Ceļš8 soļi

  1. Acknowledge and start the clockaģents

    Mark the alert acknowledged in the monitoring page and open an incident record with the start time.

    Izdarīts, kad the alert shows "acknowledged" and the incident record exists with a timestamp.

  2. Scope the blast radiusaģents

    Check the service map and dashboards: which tenants, which endpoints, since when, is it total or partial? Compare against the last change (release, migration, config, certificate expiry).

    Izdarīts, kad the incident record states affected services, tenants, and start time — or "unknown, investigating".

    ⛔ Do not touch anything yet. A change made during scoping destroys the evidence.

  3. Mitigate with reversible steps onlyaģents

    In order of preference: roll back the last release, disable the offending feature flag, scale up capacity, reroute traffic. Each step is one change, then re-check the dashboard.

    Izdarīts, kad the error rate or the failing check visibly improves, or every reversible option is exhausted.

  4. Verify from outsideaģents

    Probe the affected endpoints from outside the network (public checker or second location), not from the server itself.

    Izdarīts, kad the outside probe returns correct responses for every affected endpoint.

    ⛔ A green check from inside the network proves nothing about what users see.

  5. Decide on customer communicationīpašnieksvajag apstiprinājumu · īpašnieks

    Apstiprinājums · S5 · īpašnieks — izpilde apstājas, līdz nosaukts cilvēks ieraksta lēmumu

    Present impact, expected duration, and a draft status-page note. The owner approves, edits, or declines.

    Izdarīts, kad the status page is updated or the owner has recorded "no public note".

    ⛔ The agent never posts to a status page or mails customers without this approval.

  6. Approve any irreversible stepīpašnieksvajag apstiprinājumu · īpašnieksneatgriezenisks

    Apstiprinājums · S6 · īpašnieks — izpilde apstājas, līdz nosaukts cilvēks ieraksta lēmumu

    Neatgriezenisks solis · S6 — aģents to nekad nenoslēdz viens

    If mitigation needs something irreversible — deleting corrupted data, forcing a schema change, dropping a queue — present the exact command, the data at risk, and the alternative. Execute only on approval.

    Izdarīts, kad the irreversible action is recorded in the incident record with the approver's name, or skipped.

  7. Confirm steady stateaģents

    Dashboards normal for 15 consecutive minutes; the outside probe passes three times in a row; no new alerts.

    Izdarīts, kad the incident record shows the recovery window and the monitoring page is green.

  8. Record and hand overaģents

    Close the incident record with the full timeline: alert time, each action with its time, the cause as currently understood, open questions. Create the postmortem task (ops.incident-postmortem) and notify the owner with a three-line summary: what broke, what fixed it, what is still unknown.

    Izdarīts, kad the task exists and the owner has the summary.

Pārbaudes — kā zinām, ka izdevās

  • The outside probe answers correctly for every endpoint named in S2 — not only the one that alerted.
  • Every action in the timeline has a timestamp and an actor (agent or named person).
  • No irreversible step exists in the record without an approval name next to it.
  • The postmortem task is open and linked to the incident record.

Ja noiet greizi

PazīmeRīcība
No on-call owner reachable in 15 minEscalate to the company owner by the rota's fallback; keep mitigating reversible steps only, record the delay.
Mitigation makes it worseRoll back that one change immediately, note it in the timeline, return to S2 with the new information.
Cause unclear after all reversible optionsStabilise with a workaround (maintenance page, degraded mode), get the owner's decision in S5, do not guess at irreversible fixes.
Data may be lost or corruptedStop. Nothing is deleted. S6 decides with the owner; export what exists before any write.
Second alert fires during the incidentScope it separately in the same record — shared cause or independent? Two independent incidents get two records.

Ko atstāj katrs solis

  1. S1the alert shows "acknowledged" and the incident record exists with a timestamp.
  2. S2the incident record states affected services, tenants, and start time — or "unknown, investigating".
  3. S3the error rate or the failing check visibly improves, or every reversible option is exhausted.
  4. S4the outside probe returns correct responses for every affected endpoint.
  5. S5the status page is updated or the owner has recorded "no public note".
  6. S6the irreversible action is recorded in the incident record with the approver's name, or skipped.
  7. S7the incident record shows the recovery window and the monitoring page is green.
  8. S8the task exists and the owner has the summary.

Ko saglabāt

Incident record id · alert text and firing time · dashboard screenshots at scoping and at recovery · each mitigation command or action with its timestamp · the outside-probe results · the owner's S5 and S6 decisions (who, when) · the final timeline.

Kā šī procedūra uzlabojas

After every 5 runs ask: minutes from alert to acknowledgement and to first mitigation — which step waited? Did any mitigation get reverted (that step needs a pre-check)? Was any irreversible action taken without S6 approval, or any S5 communication sent late? Did the outside probe in S4 pass while users still saw the outage? A new version changes the step that caused the wait or the miss, and says so in its change note.

Nosaukums un kopsavilkums ir latviski. Detalizētā izpildes kārtība pagaidām ir kanoniskajā angļu valodas versijā; juridiskos un finanšu soļus publicēsim latviski tikai pēc cilvēka pārbaudes.