# Packs

> How a pack defines a cause, how 3AM proves it, and how the fix is chosen and carried out.

Source: https://docs.3am.si/how-3am-decides/packs · Markdown: https://docs.3am.si/how-3am-decides/packs.md · All docs: https://docs.3am.si/llms.txt
3AM is an on-prem AI on-call engineer for regulated enterprises (https://3am.si).

3AM's knowledge of a technology comes in a **pack**. Packs are the strongest proof 3AM has, and the source of its
exact fixes. They aren't the only way it finds a cause: when no pack knows the failure, 3AM still looks at what
changed, what is abnormal and what the errors say ([How 3AM finds the cause](https://docs.3am.si/how-3am-decides)).

A pack is a list of root causes, called classes. Each class says:

* which alerts it can explain (`signals`);
* the read-only checks that prove or rule it out (`checks`);
* which checks must pass to confirm it (`confirm`), and which only add context (`support`);
* what to do about it (`remediations`): an exact fix, or advice for a person when the fix needs judgement.

## Anatomy of a class

This is a real class from the VictoriaMetrics pack, trimmed. A storage node is unreachable: confirm it from
VictoriaMetrics, find the node's pod through Kubernetes, and restart that pod.

```json title="packs/victoriametrics · vm_storage_node_unreachable"
{
  "class": "vm_storage_node_unreachable",
  "signals": ["VMStorageNodeUnreachable", "RPCErrors", "VMRemoteWriteFailing"],
  "checks": [
    { "id": "unreachable", "kind": "read", "connector": "victoriametrics", "action": "cluster_status",
      "field": "unreachable_nodes", "expect": { "op": ">", "value": 0 } },
    { "id": "where", "kind": "read", "connector": "kubernetes", "action": "locate",
      "params": { "address": "{unreachable.unreachable_node}" },
      "field": "found", "expect": { "op": "==", "value": true } },
    { "id": "one_pod", "kind": "read", "connector": "kubernetes", "action": "locate",
      "params": { "address": "{unreachable.unreachable_node}" },
      "field": "restart_target", "expect": { "op": "matches", "value": "^pod$" } }
  ],
  "confirm": ["unreachable"],
  "support": ["where"],
  "remediations": [
    { "id": "restart_pod", "action": "restart_pod", "connector": "kubernetes",
      "params": { "namespace": "{where.namespace}", "pod": "{where.pod}" },
      "bind": { "where.namespace": "[a-z0-9]([-a-z0-9]*[a-z0-9])?", "where.pod": "[a-z0-9]([-a-z0-9.]{0,251}[a-z0-9])?" },
      "when": { "check": "one_pod", "passed": true },
      "risk": "medium", "reversible": false },
    { "id": "restore_storage_node", "action": "human",
      "summary": "Bring the vmstorage node back (its address is in the evidence)." }
  ]
}
```

What to notice:

* **`{unreachable.unreachable_node}`.** A check can use another check's output. The node to locate is the one
  `cluster_status` reported, not whatever alert opened the incident.
* **`bind`.** Every value filled into a fix must fully match its pattern. A pod name with a space, a quote or a `;`
  stops the plan, so odd values in the evidence can never become part of a command.
* **`when`.** Picks between fixes from the evidence: restart one pod if it's a single StatefulSet pod, otherwise
  restart the workload.
* **The `human` remediation.** It's the fallback when the executable fix can't run, for example because there's no
  Kubernetes connector.

## From alert to fix

1. **Select.** The classes whose `signals` include the incident's alerts are candidates. A pack applies when 3AM finds
   its technology in your estate, and a firing alert that names one of its signals turns it on too.
2. **Check.** All candidates' checks run in parallel, in rounds when one check needs another's output.
3. **Decide.** A class is *confirmed* when its `confirm` checks pass, and *refuted* when they definitely fail. When a
   check has no data, the class is *inconclusive*. A confirmed class whose checks already held half an hour before
   the alert is *already true before the incident*, not its cause. If nothing is confirmed yet, the investigation
   goes on ([How 3AM finds the cause](https://docs.3am.si/how-3am-decides)).
4. **Plan.** For the first confirmed class with a fix that can run, the parameters are filled in from the evidence and
   validated. System 1 can veto a fix the evidence doesn't support.
5. **Policy.** Shadow records it. L1 asks for approval: one person, or two for high risk. Rules can refuse it
   outright. See [Safety rails](https://docs.3am.si/how-3am-decides/safety).
6. **Act and verify.** After approval the fix runs, and the class's checks run again until it's no longer confirmed.

## Fixes that cross connectors

The pack that proves a cause and the connector that carries out the fix don't have to be the same: above,
VictoriaMetrics proves it and Kubernetes fixes it. Every fix follows the same rules:

* The target comes from the evidence, never from the alert.
* Values are validated against their patterns.
* A server-side dry run comes before anyone is asked.
* High-risk changes, such as growing a volume (which can't be undone), need two approvers.
* The cause's checks must clear afterwards.

An action that reports success but changes nothing is reported as failed. For example, a storage driver that accepts
a resize and grows nothing.

## Packs available today

| Pack                                   | Causes                                                                                                                                                                                                                                                        |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Databases (MySQL, MariaDB, PostgreSQL) | Not answering, nothing listening, out of disk, won't start, recovering, refusing credentials, connection slots full, lock waits, read-only primary, replication stopped (with or without an error), a replication slot holding WAL, transaction-ID wraparound |
| MySQL / MariaDB                        | Table lock blocking, row-lock contention, connection leak, unreachable, read-only primary, replica lag, the application's SQL failing against the schema                                                                                                      |
| Servers (over SSH)                     | Application out of memory, disk full, won't start, database down, died while its database was away, failing after a deploy, stopped, running but stuck; a server's disk nearly full, not answering SSH, its host key changed, refusing 3AM's key              |
| Kubernetes                             | Bad release, dependency down, memory limit too low, missing config, image unavailable, unschedulable, volume filling up                                                                                                                                       |
| Argo CD                                | Bad sync, degraded by a dependency, manifests don't render, Git unreachable, Git credentials, repo-server down, component down, controller stalled, stuck sync, failed sync, missing resources, drift                                                         |
| Prometheus / Alertmanager              | Target unreachable, refusing credentials or answering garbage; config reload failed; rules failing; no Alertmanager; notifications failing; component down; storage full                                                                                      |
| VictoriaMetrics                        | Targets; storage read-only or filling up; remote write unreachable, refused, rejected or failing; series limit; rows rejected; churn; config reload; rules; datasource; notifier; storage node unreachable; component down                                    |
| nginx edge                             | Rate limit too strict, timeouts too low, expired TLS certificate, authentication failures, authorisation failures (a permission removed from a role)                                                                                                          |
| ActiveMQ                               | Broker down, consumer stopped or slow, dead letters growing                                                                                                                                                                                                   |
| JVM / Tomcat                           | Instance crashing, heap pressure, JDBC pool exhaustion, instances on different releases                                                                                                                                                                       |
| Apache Fineract                        | Stuck batch job, failing batch job, GL mapping gaps, unusable GL accounts, unbalanced journals, duplicate postings, event hook disabled, maker-checker backlog                                                                                                |
| ServiceNow                             | Not answering, refusing credentials or a table, rate-limited, MID Server down, fulfilment stuck or failing                                                                                                                                                    |
| 3AM's own connections                  | A signal source 3AM can't reach                                                                                                                                                                                                                               |

Packs ship sealed with your licence. `3am-core packs list` shows which ones yours opens.
