Packs
How a pack defines a cause, how 3AM proves it, and how the fix is chosen and carried out.
3AM's knowledge of a technology comes in a pack. Packs are the strongest proof 3AM has, and the source of its exact fixes. They aren't the only way it finds a cause: when no pack knows the failure, 3AM still looks at what changed, what is abnormal and what the errors say (How 3AM finds the cause).
A pack is a list of root causes, called classes. Each class says:
- which alerts it can explain (
signals); - the read-only checks that prove or rule it out (
checks); - which checks must pass to confirm it (
confirm), and which only add context (support); - what to do about it (
remediations): an exact fix, or advice for a person when the fix needs judgement.
Anatomy of a class
This is a real class from the VictoriaMetrics pack, trimmed. A storage node is unreachable: confirm it from VictoriaMetrics, find the node's pod through Kubernetes, and restart that pod.
{
"class": "vm_storage_node_unreachable",
"signals": ["VMStorageNodeUnreachable", "RPCErrors", "VMRemoteWriteFailing"],
"checks": [
{ "id": "unreachable", "kind": "read", "connector": "victoriametrics", "action": "cluster_status",
"field": "unreachable_nodes", "expect": { "op": ">", "value": 0 } },
{ "id": "where", "kind": "read", "connector": "kubernetes", "action": "locate",
"params": { "address": "{unreachable.unreachable_node}" },
"field": "found", "expect": { "op": "==", "value": true } },
{ "id": "one_pod", "kind": "read", "connector": "kubernetes", "action": "locate",
"params": { "address": "{unreachable.unreachable_node}" },
"field": "restart_target", "expect": { "op": "matches", "value": "^pod$" } }
],
"confirm": ["unreachable"],
"support": ["where"],
"remediations": [
{ "id": "restart_pod", "action": "restart_pod", "connector": "kubernetes",
"params": { "namespace": "{where.namespace}", "pod": "{where.pod}" },
"bind": { "where.namespace": "[a-z0-9]([-a-z0-9]*[a-z0-9])?", "where.pod": "[a-z0-9]([-a-z0-9.]{0,251}[a-z0-9])?" },
"when": { "check": "one_pod", "passed": true },
"risk": "medium", "reversible": false },
{ "id": "restore_storage_node", "action": "human",
"summary": "Bring the vmstorage node back (its address is in the evidence)." }
]
}What to notice:
{unreachable.unreachable_node}. A check can use another check's output. The node to locate is the onecluster_statusreported, not whatever alert opened the incident.bind. Every value filled into a fix must fully match its pattern. A pod name with a space, a quote or a;stops the plan, so odd values in the evidence can never become part of a command.when. Picks between fixes from the evidence: restart one pod if it's a single StatefulSet pod, otherwise restart the workload.- The
humanremediation. It's the fallback when the executable fix can't run, for example because there's no Kubernetes connector.
From alert to fix
- Select. The classes whose
signalsinclude the incident's alerts are candidates. A pack applies when 3AM finds its technology in your estate, and a firing alert that names one of its signals turns it on too. - Check. All candidates' checks run in parallel, in rounds when one check needs another's output.
- Decide. A class is confirmed when its
confirmchecks pass, and refuted when they definitely fail. When a check has no data, the class is inconclusive. A confirmed class whose checks already held half an hour before the alert is already true before the incident, not its cause. If nothing is confirmed yet, the investigation goes on (How 3AM finds the cause). - Plan. For the first confirmed class with a fix that can run, the parameters are filled in from the evidence and validated. System 1 can veto a fix the evidence doesn't support.
- Policy. Shadow records it. L1 asks for approval: one person, or two for high risk. Rules can refuse it outright. See Safety rails.
- Act and verify. After approval the fix runs, and the class's checks run again until it's no longer confirmed.
Fixes that cross connectors
The pack that proves a cause and the connector that carries out the fix don't have to be the same: above, VictoriaMetrics proves it and Kubernetes fixes it. Every fix follows the same rules:
- The target comes from the evidence, never from the alert.
- Values are validated against their patterns.
- A server-side dry run comes before anyone is asked.
- High-risk changes, such as growing a volume (which can't be undone), need two approvers.
- The cause's checks must clear afterwards.
An action that reports success but changes nothing is reported as failed. For example, a storage driver that accepts a resize and grows nothing.
Packs available today
| Pack | Causes |
|---|---|
| Databases (MySQL, MariaDB, PostgreSQL) | Not answering, nothing listening, out of disk, won't start, recovering, refusing credentials, connection slots full, lock waits, read-only primary, replication stopped (with or without an error), a replication slot holding WAL, transaction-ID wraparound |
| MySQL / MariaDB | Table lock blocking, row-lock contention, connection leak, unreachable, read-only primary, replica lag, the application's SQL failing against the schema |
| Servers (over SSH) | Application out of memory, disk full, won't start, database down, died while its database was away, failing after a deploy, stopped, running but stuck; a server's disk nearly full, not answering SSH, its host key changed, refusing 3AM's key |
| Kubernetes | Bad release, dependency down, memory limit too low, missing config, image unavailable, unschedulable, volume filling up |
| Argo CD | Bad sync, degraded by a dependency, manifests don't render, Git unreachable, Git credentials, repo-server down, component down, controller stalled, stuck sync, failed sync, missing resources, drift |
| Prometheus / Alertmanager | Target unreachable, refusing credentials or answering garbage; config reload failed; rules failing; no Alertmanager; notifications failing; component down; storage full |
| VictoriaMetrics | Targets; storage read-only or filling up; remote write unreachable, refused, rejected or failing; series limit; rows rejected; churn; config reload; rules; datasource; notifier; storage node unreachable; component down |
| nginx edge | Rate limit too strict, timeouts too low, expired TLS certificate, authentication failures, authorisation failures (a permission removed from a role) |
| ActiveMQ | Broker down, consumer stopped or slow, dead letters growing |
| JVM / Tomcat | Instance crashing, heap pressure, JDBC pool exhaustion, instances on different releases |
| Apache Fineract | Stuck batch job, failing batch job, GL mapping gaps, unusable GL accounts, unbalanced journals, duplicate postings, event hook disabled, maker-checker backlog |
| ServiceNow | Not answering, refusing credentials or a table, rate-limited, MID Server down, fulfilment stuck or failing |
| 3AM's own connections | A signal source 3AM can't reach |
Packs ship sealed with your licence. 3am-core packs list shows which ones yours opens.