Docs

How 3AM works

Signals in, proven fixes out. What happens between an alert firing and a fix being verified.

When an alert fires, 3AM opens an incident and works it the way a careful engineer would: find out which service is affected, read the runbooks and code, prove what is wrong with read-only checks, propose the fix, ask a person, make the change, and check that it worked.

Detect

3AM reads alerts from the tools you already run: Alertmanager, PagerDuty, Kubernetes, Splunk, Datadog. Related alerts on the same service join one incident. Alerts that were already firing when 3AM started are treated as chronic.

Understand

It works out which service the alert concerns, who owns it and what depends on it. The map comes from your CMDB, monitoring labels, repositories and clusters. It also searches your runbooks and code for that service.

Prove

Each possible root cause is a class with read-only checks: a PromQL query, a SQL query, a Kubernetes read, a config read. A cause is confirmed only when its required checks pass on live data. Look-alike causes are told apart. If 3AM can't prove anything, it says so and hands over the evidence. It never guesses.

Approve

If the confirmed cause has a known fix, 3AM rehearses it first (a dry run that changes nothing), then asks for approval in Slack, Symphony, PagerDuty, by phone or e-mail, with the evidence and the rehearsal attached.

Fix and verify

After approval, 3AM makes the change and re-runs the checks until the cause has cleared. The incident closes as mitigated only when the evidence says so.

You stay in control

RailWhat it means
Shadow mode firstA new install proposes and records but changes nothing. The daily digest compares what 3AM would have done with what your team did
Autonomy per serviceYou let 3AM act on one service at a time (L1: one approval per action). Raising a service takes two approvers; lowering is instant
PolicyRules by action, risk, environment and time of day; quorum for risky changes; actions per hour; a halt switch that stops everything at once
RehearsalEvery change is dry-run first: SQL in a rolled-back transaction, Kubernetes with server-side dry run
Audit logEvery signal, check, model call, decision, approval and action, hash-chained and signed. Your auditors can verify it

Where models fit in

Models never decide anything in 3AM. Diagnosis is checks against live systems, fixes come from vetted recipes, and every change needs a person. An optional on-prem model writes the incident note for the people on call. If the model is missing or gets it wrong, a template is used instead.

On this page