Kubernetes
3AM watches your namespaces for failing workloads, proves why they fail, and rolls back bad releases with your approval.
What 3AM does with Kubernetes
| In short | |
|---|---|
| Failures in | Crash loops, OOM kills, image pull failures, missing config, unschedulable pods and stuck rollouts become 3AM incidents. No metrics stack is needed: 3AM reads pod and workload state from the API server |
| Understands | Six root causes, each proven from the cluster: a bad release, a dependency that is down, a memory limit that is too low, missing config, an unavailable image, unschedulable pods |
| Actions | With approval: roll back a bad release, restart or scale a workload, restart one pod, grow a volume. Every change is rehearsed first with a server-side dry run, then verified |
| Fixes for other connectors | When VictoriaMetrics, Prometheus or another pack proves a component is hung or its disk is full, 3AM finds the workload it runs in and fixes it here |
| Notes | 3AM's diagnosis appears as an Event on the workload (kubectl describe shows it) |
1. Give 3AM least-privilege access
3AM needs one ServiceAccount and, in each namespace it watches, a Role allowing only:
- read pods, pod logs, Services, Deployments, StatefulSets, DaemonSets and ReplicaSets;
- patch Deployments, StatefulSets and DaemonSets, including their scale (approved restarts, scaling and rollbacks);
- delete a single pod, so its controller recreates it (approved
restart_pod: controller-managed pods only); - read and patch PersistentVolumeClaims (approved
expand_volume: only where the StorageClass allows it); - create Events (optional: 3AM's notes).
Cluster-wide, 3AM may only read StorageClasses, to know whether a volume may grow. No exec, no Secrets, no
deleting workloads or volumes.
You can take any action away: delete its rule. The connector's health then lists it as not allowed, and 3AM hands
that fix to a person instead. The manifests are in the bundle under kubernetes/. Apply them with a cluster admin's
credentials:
kubectl apply -f 3am-rbac.yaml # namespace 3am, the ServiceAccount, its token, StorageClass readThen once per namespace 3AM should watch:
for NS in bank-prod payments-prod; do
sed "s/WATCHED_NAMESPACE/$NS/" 3am-role.yaml | kubectl apply -f -
doneapiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata: {name: 3am-connector, namespace: WATCHED_NAMESPACE}
rules:
- apiGroups: [""]
resources: [pods, pods/log, services]
verbs: [get, list, watch]
- apiGroups: [apps]
resources: [deployments, statefulsets, daemonsets, replicasets]
verbs: [get, list, watch]
- apiGroups: [apps]
resources: [deployments, statefulsets, daemonsets, deployments/scale, statefulsets/scale]
verbs: [patch]
- apiGroups: [""]
resources: [pods]
verbs: [delete] # restart_pod
- apiGroups: [""]
resources: [persistentvolumeclaims]
verbs: [get, list, patch] # expand_volume
- apiGroups: [""]
resources: [events]
verbs: [create]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata: {name: 3am-connector, namespace: WATCHED_NAMESPACE}
subjects: [{kind: ServiceAccount, name: 3am-connector, namespace: 3am}]
roleRef: {apiGroup: rbac.authorization.k8s.io, kind: Role, name: 3am-connector}2. Collect the token, the CA and the API address
# the token
kubectl -n 3am get secret 3am-connector-token -o jsonpath='{.data.token}' | base64 -d > 3am-k8s-token
# the cluster's CA certificate
kubectl config view --raw --minify -o jsonpath='{.clusters[0].cluster.certificate-authority-data}' | base64 -d > k8s-ca.pem
# the API server address
kubectl config view --minify -o jsonpath='{.clusters[0].cluster.server}'; echoManaged clusters
On EKS, AKS and GKE, the address and CA in your kubeconfig are the ones to use. If your kubeconfig authenticates with a cloud plugin, that's fine: 3AM uses the ServiceAccount token, not your kubeconfig.
3. Connect it in 3AM
Settings → Add a connection → Kubernetes (or the wizard's Systems step). Paste the token, choose the CA file:
kind: kubernetesalerts inactions| Setting | What to enter | Required | Default |
|---|---|---|---|
api_url | API server URL | yes | — |
tokensecret | Service-account token | yes | — |
ca_file | API server CA certificate (PEM file; empty: system CAs) | no | — |
verify_tls | Verify TLS | yes | true |
namespaces | Namespaces 3AM watches and may act in | yes | — |
Test connection asks the API server, for each namespace, whether 3AM's token can do every thing it needs, and lists anything missing. A healthy result:
Connected. Kubernetes v1.30.4; all permissions present in bank-prod, payments-prodNetwork: outbound HTTPS from 3AM to the API server. If 3AM runs in the same cluster, that's the
kubernetes.default.svc address.
What 3AM understands
Every cause is proven from the cluster before anything is proposed:
| Root cause | Proven when | 3AM does |
|---|---|---|
| Bad release | A rollout in the last 6 hours; the previous revision's pods differ; the new revision's pods fail; the crashed run's logs show no dependency errors | Rolls back to the last revision whose pods actually differ (approved, rehearsed, then verified until the pods recover) |
| Dependency down | Pods crash and the crashed run's logs show they can't reach something (connection refused, timeouts, unknown host) | Does not restart or roll back. It names the dependency; the pods recover once it is back |
| Memory limit too low | The last exit was OOMKilled, in the last 30 minutes | Advice: raise the limit (the current one is in the evidence). If a release caused it, it rolls back instead |
| Missing config | CreateContainerConfigError naming the ConfigMap or Secret | Advice: create it |
| Image unavailable | Image pull failing | Advice: tag, registry or pull secret. If the image came with a recent release, it rolls back |
| Unschedulable | The scheduler reports why | Advice: capacity, node selector, affinity or taints |
| Volume filling up | KubePersistentVolumeFillingUp fires, the kubelet's volume stats confirm under 15% free, and the PVC is real | Grows the volume by half where its StorageClass allows it (two approvers: a volume never shrinks back), then checks it has room. Otherwise advice |
Bad release and dependency down exclude each other on purpose. Restarting a pod whose database is down is the classic wrong move, and 3AM won't make it.
Signals
| Signal | When |
|---|---|
KubePodCrashLooping | A container is waiting in CrashLoopBackOff |
KubeContainerOOMKilled | A container's last exit was OOMKilled, in the last 30 minutes |
KubeImagePullBackOff | ImagePullBackOff or ErrImagePull |
KubeContainerConfigError | CreateContainerConfigError |
KubePodUnschedulable | Pending and unschedulable for more than 2 minutes |
KubeDeploymentRolloutStuck | A rollout exceeded its progress deadline |
The names match the kube-prometheus alerts, so the same root causes apply whichever source reports the failure. One incident per problem per workload: three crash-looping replicas are one incident.
Actions
| Action | What it does |
|---|---|
rollback | Puts back the last revision whose pods differ. Unlike kubectl rollout undo, a restart-only revision is skipped |
restart | A rolling restart, as kubectl rollout restart |
scale | Sets the replica count (0–100). A DaemonSet is never scaled |
restart_pod | Deletes one pod so its controller (StatefulSet, ReplicaSet, DaemonSet) recreates it, then waits until the replacement is Ready. Refuses a pod nothing would bring back |
expand_volume | Grows a PVC. Refuses when the StorageClass doesn't allow expansion, the size isn't larger, or it's more than 4x the current size. Succeeds only when the volume has actually grown: a driver that accepts the request and grows nothing is reported as a failure |
locate | Read-only. Finds the pod and workload behind an address (URL, host:port, pod IP, Service or pod DNS name), with its volume. This is how other packs' fixes find their target |
logs | Read-only. previous=true reads the crashed run |
Every change is rehearsed with dryRun=All: the API server validates and admits it, including your admission
webhooks, and nothing changes. Two fences protect other teams: 3AM refuses namespaces outside its list, and the Role
refuses them even if that list were wrong.
Troubleshooting
| Test connection says | Fix |
|---|---|
| the token is invalid or expired | Re-read the token (step 2); if you used a short-lived token, use the 3am-connector-token Secret |
| the API server's certificate is not trusted: set ca_file | Choose the CA file from step 2 |
| the token cannot: list pods in X | The Role isn't applied in namespace X; run the step 1 loop for X |
| An action fails with the 3AM Role does not allow this | The namespace's Role is missing a rule; re-apply 3am-role.yaml |
| No incidents for a failing pod | Its namespace isn't in Namespaces |
Certification
Certified on 2026-10-04 on a real cluster (Kubernetes 1.34) with exactly this RBAC and only 3AM's token: 32 live connector checks, and 30 end-to-end checks. Two broken releases (a config panic and a missing image) were proven, rolled back and verified healthy; a release whose database was unreachable was correctly left alone; an OOM and a missing ConfigMap were diagnosed with advice and nothing executed.
Fixes, certified live on 2026-10-04 (54 of 54 checks, on a real Kubernetes cluster, with 3AM inside it under the client RBAC). Components were hung mid-run, meaning their pod was Running and answered nothing. Each was detected, proven, located, rehearsed, approved, restarted and verified:
- vmagent, a vmstorage node (only that pod), a scrape target of vmagent and one of Prometheus, Alertmanager (seen by both monitoring connectors, restarted once) and Prometheus.
A read-only storage node got a proposal to grow its volume, approved by two people. The test storage driver accepts the request but can't grow volumes, and 3AM reported the fix as failed rather than claiming success.