Packs

How a pack defines a cause, how 3AM proves it, and how the fix is chosen and carried out.

3AM's knowledge of a technology comes in a pack. Packs are the strongest proof 3AM has, and the source of its exact fixes. They aren't the only way it finds a cause: when no pack knows the failure, 3AM still looks at what changed, what is abnormal and what the errors say (How 3AM finds the cause).

A pack is a list of root causes, called classes. Each class says:

  • which alerts it can explain (signals);
  • the read-only checks that prove or rule it out (checks);
  • which checks must pass to confirm it (confirm), and which only add context (support);
  • what to do about it (remediations): an exact fix, or advice for a person when the fix needs judgement.

Anatomy of a class

This is a real class from the VictoriaMetrics pack, trimmed. A storage node is unreachable: confirm it from VictoriaMetrics, find the node's pod through Kubernetes, and restart that pod.

packs/victoriametrics · vm_storage_node_unreachable
{
  "class": "vm_storage_node_unreachable",
  "signals": ["VMStorageNodeUnreachable", "RPCErrors", "VMRemoteWriteFailing"],
  "checks": [
    { "id": "unreachable", "kind": "read", "connector": "victoriametrics", "action": "cluster_status",
      "field": "unreachable_nodes", "expect": { "op": ">", "value": 0 } },
    { "id": "where", "kind": "read", "connector": "kubernetes", "action": "locate",
      "params": { "address": "{unreachable.unreachable_node}" },
      "field": "found", "expect": { "op": "==", "value": true } },
    { "id": "one_pod", "kind": "read", "connector": "kubernetes", "action": "locate",
      "params": { "address": "{unreachable.unreachable_node}" },
      "field": "restart_target", "expect": { "op": "matches", "value": "^pod$" } }
  ],
  "confirm": ["unreachable"],
  "support": ["where"],
  "remediations": [
    { "id": "restart_pod", "action": "restart_pod", "connector": "kubernetes",
      "params": { "namespace": "{where.namespace}", "pod": "{where.pod}" },
      "bind": { "where.namespace": "[a-z0-9]([-a-z0-9]*[a-z0-9])?", "where.pod": "[a-z0-9]([-a-z0-9.]{0,251}[a-z0-9])?" },
      "when": { "check": "one_pod", "passed": true },
      "risk": "medium", "reversible": false },
    { "id": "restore_storage_node", "action": "human",
      "summary": "Bring the vmstorage node back (its address is in the evidence)." }
  ]
}

What to notice:

  • {unreachable.unreachable_node}. A check can use another check's output. The node to locate is the one cluster_status reported, not whatever alert opened the incident.
  • bind. Every value filled into a fix must fully match its pattern. A pod name with a space, a quote or a ; stops the plan, so odd values in the evidence can never become part of a command.
  • when. Picks between fixes from the evidence: restart one pod if it's a single StatefulSet pod, otherwise restart the workload.
  • The human remediation. It's the fallback when the executable fix can't run, for example because there's no Kubernetes connector.

From alert to fix

  1. Select. The classes whose signals include the incident's alerts are candidates. A pack applies when 3AM finds its technology in your estate, and a firing alert that names one of its signals turns it on too.
  2. Check. All candidates' checks run in parallel, in rounds when one check needs another's output.
  3. Decide. A class is confirmed when its confirm checks pass, and refuted when they definitely fail. When a check has no data, the class is inconclusive. A confirmed class whose checks already held half an hour before the alert is already true before the incident, not its cause. If nothing is confirmed yet, the investigation goes on (How 3AM finds the cause).
  4. Plan. For the first confirmed class with a fix that can run, the parameters are filled in from the evidence and validated. System 1 can veto a fix the evidence doesn't support.
  5. Policy. Shadow records it. L1 asks for approval: one person, or two for high risk. Rules can refuse it outright. See Safety rails.
  6. Act and verify. After approval the fix runs, and the class's checks run again until it's no longer confirmed.

Fixes that cross connectors

The pack that proves a cause and the connector that carries out the fix don't have to be the same: above, VictoriaMetrics proves it and Kubernetes fixes it. Every fix follows the same rules:

  • The target comes from the evidence, never from the alert.
  • Values are validated against their patterns.
  • A server-side dry run comes before anyone is asked.
  • High-risk changes, such as growing a volume (which can't be undone), need two approvers.
  • The cause's checks must clear afterwards.

An action that reports success but changes nothing is reported as failed. For example, a storage driver that accepts a resize and grows nothing.

Packs available today

PackCauses
Databases (MySQL, MariaDB, PostgreSQL)Not answering, nothing listening, out of disk, won't start, recovering, refusing credentials, connection slots full, lock waits, read-only primary, replication stopped (with or without an error), a replication slot holding WAL, transaction-ID wraparound
MySQL / MariaDBTable lock blocking, row-lock contention, connection leak, unreachable, read-only primary, replica lag, the application's SQL failing against the schema
Servers (over SSH)Application out of memory, disk full, won't start, database down, died while its database was away, failing after a deploy, stopped, running but stuck; a server's disk nearly full, not answering SSH, its host key changed, refusing 3AM's key
KubernetesBad release, dependency down, memory limit too low, missing config, image unavailable, unschedulable, volume filling up
Argo CDBad sync, degraded by a dependency, manifests don't render, Git unreachable, Git credentials, repo-server down, component down, controller stalled, stuck sync, failed sync, missing resources, drift
Prometheus / AlertmanagerTarget unreachable, refusing credentials or answering garbage; config reload failed; rules failing; no Alertmanager; notifications failing; component down; storage full
VictoriaMetricsTargets; storage read-only or filling up; remote write unreachable, refused, rejected or failing; series limit; rows rejected; churn; config reload; rules; datasource; notifier; storage node unreachable; component down
nginx edgeRate limit too strict, timeouts too low, expired TLS certificate, authentication failures, authorisation failures (a permission removed from a role)
ActiveMQBroker down, consumer stopped or slow, dead letters growing
JVM / TomcatInstance crashing, heap pressure, JDBC pool exhaustion, instances on different releases
Apache FineractStuck batch job, failing batch job, GL mapping gaps, unusable GL accounts, unbalanced journals, duplicate postings, event hook disabled, maker-checker backlog
ServiceNowNot answering, refusing credentials or a table, rate-limited, MID Server down, fulfilment stuck or failing
3AM's own connectionsA signal source 3AM can't reach

Packs ship sealed with your licence. 3am-core packs list shows which ones yours opens.

On this page