Docs
Connectors

Kubernetes

3AM watches your namespaces for failing workloads, proves why they fail, and rolls back bad releases with your approval.

What 3AM does with Kubernetes

In short
Failures inCrash loops, OOM kills, image pull failures, missing config, unschedulable pods and stuck rollouts become 3AM incidents. No metrics stack is needed: 3AM reads pod and workload state from the API server
UnderstandsSix root causes, each proven from the cluster: a bad release, a dependency that is down, a memory limit that is too low, missing config, an unavailable image, unschedulable pods
ActionsWith approval: roll back a bad release, restart or scale a workload, restart one pod, grow a volume. Every change is rehearsed first with a server-side dry run, then verified
Fixes for other connectorsWhen VictoriaMetrics, Prometheus or another pack proves a component is hung or its disk is full, 3AM finds the workload it runs in and fixes it here
Notes3AM's diagnosis appears as an Event on the workload (kubectl describe shows it)

1. Give 3AM least-privilege access

3AM needs one ServiceAccount and, in each namespace it watches, a Role allowing only:

  • read pods, pod logs, Services, Deployments, StatefulSets, DaemonSets and ReplicaSets;
  • patch Deployments, StatefulSets and DaemonSets, including their scale (approved restarts, scaling and rollbacks);
  • delete a single pod, so its controller recreates it (approved restart_pod: controller-managed pods only);
  • read and patch PersistentVolumeClaims (approved expand_volume: only where the StorageClass allows it);
  • create Events (optional: 3AM's notes).

Cluster-wide, 3AM may only read StorageClasses, to know whether a volume may grow. No exec, no Secrets, no deleting workloads or volumes.

You can take any action away: delete its rule. The connector's health then lists it as not allowed, and 3AM hands that fix to a person instead. The manifests are in the bundle under kubernetes/. Apply them with a cluster admin's credentials:

kubectl apply -f 3am-rbac.yaml                    # namespace 3am, the ServiceAccount, its token, StorageClass read

Then once per namespace 3AM should watch:

for NS in bank-prod payments-prod; do
  sed "s/WATCHED_NAMESPACE/$NS/" 3am-role.yaml | kubectl apply -f -
done

2. Collect the token, the CA and the API address

# the token
kubectl -n 3am get secret 3am-connector-token -o jsonpath='{.data.token}' | base64 -d > 3am-k8s-token
# the cluster's CA certificate
kubectl config view --raw --minify -o jsonpath='{.clusters[0].cluster.certificate-authority-data}' | base64 -d > k8s-ca.pem
# the API server address
kubectl config view --minify -o jsonpath='{.clusters[0].cluster.server}'; echo

Managed clusters

On EKS, AKS and GKE, the address and CA in your kubeconfig are the ones to use. If your kubeconfig authenticates with a cloud plugin, that's fine: 3AM uses the ServiceAccount token, not your kubeconfig.

3. Connect it in 3AM

Settings → Add a connection → Kubernetes (or the wizard's Systems step). Paste the token, choose the CA file:

kind: kubernetesalerts inactions
SettingWhat to enterRequiredDefault
api_urlAPI server URLyes—
tokensecretService-account tokenyes—
ca_fileAPI server CA certificate (PEM file; empty: system CAs)no—
verify_tlsVerify TLSyestrue
namespacesNamespaces 3AM watches and may act inyes—

Test connection asks the API server, for each namespace, whether 3AM's token can do every thing it needs, and lists anything missing. A healthy result:

Connected. Kubernetes v1.30.4; all permissions present in bank-prod, payments-prod

Network: outbound HTTPS from 3AM to the API server. If 3AM runs in the same cluster, that's the kubernetes.default.svc address.

What 3AM understands

Every cause is proven from the cluster before anything is proposed:

Root causeProven when3AM does
Bad releaseA rollout in the last 6 hours; the previous revision's pods differ; the new revision's pods fail; the crashed run's logs show no dependency errorsRolls back to the last revision whose pods actually differ (approved, rehearsed, then verified until the pods recover)
Dependency downPods crash and the crashed run's logs show they can't reach something (connection refused, timeouts, unknown host)Does not restart or roll back. It names the dependency; the pods recover once it is back
Memory limit too lowThe last exit was OOMKilled, in the last 30 minutesAdvice: raise the limit (the current one is in the evidence). If a release caused it, it rolls back instead
Missing configCreateContainerConfigError naming the ConfigMap or SecretAdvice: create it
Image unavailableImage pull failingAdvice: tag, registry or pull secret. If the image came with a recent release, it rolls back
UnschedulableThe scheduler reports whyAdvice: capacity, node selector, affinity or taints
Volume filling upKubePersistentVolumeFillingUp fires, the kubelet's volume stats confirm under 15% free, and the PVC is realGrows the volume by half where its StorageClass allows it (two approvers: a volume never shrinks back), then checks it has room. Otherwise advice

Bad release and dependency down exclude each other on purpose. Restarting a pod whose database is down is the classic wrong move, and 3AM won't make it.

Signals

SignalWhen
KubePodCrashLoopingA container is waiting in CrashLoopBackOff
KubeContainerOOMKilledA container's last exit was OOMKilled, in the last 30 minutes
KubeImagePullBackOffImagePullBackOff or ErrImagePull
KubeContainerConfigErrorCreateContainerConfigError
KubePodUnschedulablePending and unschedulable for more than 2 minutes
KubeDeploymentRolloutStuckA rollout exceeded its progress deadline

The names match the kube-prometheus alerts, so the same root causes apply whichever source reports the failure. One incident per problem per workload: three crash-looping replicas are one incident.

Actions

ActionWhat it does
rollbackPuts back the last revision whose pods differ. Unlike kubectl rollout undo, a restart-only revision is skipped
restartA rolling restart, as kubectl rollout restart
scaleSets the replica count (0–100). A DaemonSet is never scaled
restart_podDeletes one pod so its controller (StatefulSet, ReplicaSet, DaemonSet) recreates it, then waits until the replacement is Ready. Refuses a pod nothing would bring back
expand_volumeGrows a PVC. Refuses when the StorageClass doesn't allow expansion, the size isn't larger, or it's more than 4x the current size. Succeeds only when the volume has actually grown: a driver that accepts the request and grows nothing is reported as a failure
locateRead-only. Finds the pod and workload behind an address (URL, host:port, pod IP, Service or pod DNS name), with its volume. This is how other packs' fixes find their target
logsRead-only. previous=true reads the crashed run

Every change is rehearsed with dryRun=All: the API server validates and admits it, including your admission webhooks, and nothing changes. Two fences protect other teams: 3AM refuses namespaces outside its list, and the Role refuses them even if that list were wrong.

Troubleshooting

Test connection saysFix
the token is invalid or expiredRe-read the token (step 2); if you used a short-lived token, use the 3am-connector-token Secret
the API server's certificate is not trusted: set ca_fileChoose the CA file from step 2
the token cannot: list pods in XThe Role isn't applied in namespace X; run the step 1 loop for X
An action fails with the 3AM Role does not allow thisThe namespace's Role is missing a rule; re-apply 3am-role.yaml
No incidents for a failing podIts namespace isn't in Namespaces

Certification

Certified on 2026-10-04 on a real cluster (Kubernetes 1.34) with exactly this RBAC and only 3AM's token: 32 live connector checks, and 30 end-to-end checks. Two broken releases (a config panic and a missing image) were proven, rolled back and verified healthy; a release whose database was unreachable was correctly left alone; an OOM and a missing ConfigMap were diagnosed with advice and nothing executed.

Fixes, certified live on 2026-10-04 (54 of 54 checks, on a real Kubernetes cluster, with 3AM inside it under the client RBAC). Components were hung mid-run, meaning their pod was Running and answered nothing. Each was detected, proven, located, rehearsed, approved, restarted and verified:

  • vmagent, a vmstorage node (only that pod), a scrape target of vmagent and one of Prometheus, Alertmanager (seen by both monitoring connectors, restarted once) and Prometheus.

A read-only storage node got a proposal to grow its volume, approved by two people. The test storage driver accepts the request but can't grow volumes, and 3AM reported the fix as failed rather than claiming success.

On this page