Argo CD
Deploy history as evidence, the GitOps-correct rollback, and fixes when Argo CD itself breaks.
With Argo CD connected, every incident on an Argo-managed workload starts with the question an on-call engineer asks
first: what changed? 3AM shows the revision that went out, its author and commit message, and when it synced. Argo
CD also changes how a bad release is undone, because its self-heal re-applies Git: a kubectl rollback of an
Argo-managed Deployment is undone within seconds (we measured 15). With this connector, 3AM rolls back through Argo CD
instead. It also watches Argo CD itself: a repo-server or controller that hangs, a sync stuck on a hook, a Git host
that's unreachable.
Before you start
- Admin access to Argo CD, once, to create a local account for 3AM.
- Network access from 3AM to the Argo CD API server, and to the components' metrics ports if you want 3AM to notice them hanging.
- The Kubernetes connector, so 3AM can prove why a workload fails and restart Argo CD's own
components. Add the
argocdnamespace to its namespaces and apply the 3AM Role there.
1. Create the 3AM account
Add a local account with an API key, and a role that can read applications, sync them, and turn auto-sync off for a rollback:
data:
accounts.threeam: apiKeydata:
policy.csv: |
p, role:threeam, applications, get, */*, allow
p, role:threeam, applications, sync, */*, allow # sync, roll back, stop a stuck sync
p, role:threeam, applications, update, */*, allow # turn auto-sync off before a rollback
p, role:threeam, projects, get, *, allow
g, threeam, role:threeamLeave out sync and update for a read-only account: 3AM then diagnoses and hands fixes to a person. Scope */* to
<project>/* to limit 3AM to some projects.
Create the token:
argocd account generate-token --account threeam > 3am-argocd-token2. Add the connector
In the console: Settings → Add a connection → Argo CD. Or in connectors.json:
{ "name": "argocd", "kind": "argocd",
"settings": { "url": "https://argocd-server.argocd.svc",
"token": { "file": "/etc/3am/secrets/3am-argocd-token" },
"ca_file": "/etc/3am/secrets/argocd-ca.pem",
"component_urls": ["http://argocd-metrics.argocd.svc:8082",
"http://argocd-repo-server.argocd.svc:8084",
"http://argocd-server-metrics.argocd.svc:8083"] } }Argo CD serves a self-signed certificate unless you gave it one. Point ca_file at the CA that signed it, rather than
turning verification off.
kind: argocdalerts inactions| Setting | What to enter | Required | Default |
|---|---|---|---|
url | Argo CD API URL, e.g. https://argocd.bank.internal | yes | — |
tokensecret | API token of the 3AM account | yes | — |
projects | Projects to watch (empty: all) | no | — |
component_urls | Component metrics URLs: application-controller (8082), repo-server (8084), server (8083) | no | — |
ca_file | CA certificate for a private CA (PEM file) | no | — |
verify_tls | Verify TLS | yes | true |
degraded_after_s | Seconds Degraded before it's an incident | yes | 120 |
stuck_after_s | Seconds a sync may run before it's stuck | yes | 900 |
drift_after_s | Seconds OutOfSync (no sync running) before it's drift | yes | 1800 |
stall_after_s | Seconds without reconciling before the controller is stalled | yes | 600 |
3. Verify
Argo CD v3.5.3; 7 applicationsIf a component listed in component_urls doesn't answer, Test connection names it.
When an Argo-managed workload has a bad release, the incident's proposed fix is an Argo CD rollback, shown as
rollback id=0, name=payments (via argocd), where argocd is your connector's name. It is never a kubectl
rollback.
What it detects
| Signal | When |
|---|---|
ArgoCDAppDegraded | An app has been Degraded for degraded_after_s (120 s). Carries the failing workload, so the Kubernetes pack can prove why |
ArgoCDSyncFailed | A sync failed while applying resources |
ArgoCDSyncStuck | A sync has run longer than stuck_after_s (15 min), usually a hook that never finishes |
ArgoCDComparisonError | Manifests can't be generated: they don't render, Git is unreachable or refuses credentials, or the repo-server is down |
ArgoCDAppMissing | Resources the app should manage are missing from the cluster |
ArgoCDAppOutOfSync | Live resources differ from Git for drift_after_s (30 min) with self-heal off |
ArgoCDControllerStalled | No app reconciled for stall_after_s (10 min) |
ArgoCDComponentDown | The API server or a component in component_urls doesn't answer |
Signals carry the app in the name label, like argocd_app_info. So if you already alert with the community rules
(ArgoCdAppUnhealthy, ArgoCdAppSyncFailed, ArgoCdAppOutOfSync) through Prometheus, the same causes apply.
What it proves and fixes
Fixed with approval:
| Cause | How 3AM knows | Fix |
|---|---|---|
| Bad release | A sync in the last 6 hours, the app Degraded, its pods failing, and the crash logs show no dependency errors | Roll back through Argo CD: auto-sync off, back to the previous deployment, verified healthy |
| Sync stuck on a hook | The operation has run too long, waiting on a hook | Stop the sync, so later syncs can run. If Git has moved on, auto-sync starts a new sync that waits on the same hook, and the note says so: the hook needs fixing |
| Resources missing, auto-sync off | Health Missing and nothing syncing | Sync from Git (dry run first, nothing pruned) |
| Repo-server down or hung | Manifest generation fails with the repo-server's connection error, or it doesn't answer | Restart its Deployment (Kubernetes connector) |
| Application controller stalled | Nothing has reconciled for too long | Restart its pod (Kubernetes connector) |
| A component down | It doesn't answer | Restart its pod or workload (Kubernetes connector) |
Handed to a person:
| Cause | How 3AM knows |
|---|---|
| Degraded because a dependency is down | The crash logs show it can't reach something. Rolling back wouldn't help, so 3AM doesn't |
| Manifests don't render | The Helm, Kustomize or YAML error, with file and line. The running version is untouched |
| Git unreachable | The Git host's DNS, network or TLS error |
| Git refuses credentials | authentication required and similar |
| Sync failed applying resources | Each resource's error: validation, an admission webhook, an immutable field, a quota |
| Drift | What differs from Git. It may be a deliberate hotfix, and syncing would undo it |
How a rollback works
- Rehearse. Argo CD renders the target revision. Argo refuses a dry-run rollback while auto-sync is on, so rendering is the rehearsal: it proves the old revision still produces valid manifests.
- Approve. The request shows the commit that went out ("payments 1.1: read ledger url from config", Jane Doe) and the revision it goes back to.
- Turn auto-sync off, so self-heal can't re-apply the bad commit.
- Roll back to the previous deployment, and wait until the app is healthy.
- Hand back. The incident says which commit to revert in Git, and to turn auto-sync back on (
set_auto_sync, or in Argo CD).
The Kubernetes pack follows the same rule. A bad release of an Argo-managed workload is
rolled back through Argo CD. If 3AM has no Argo CD connection, it hands the rollback to a person rather than using
kubectl. It knows which app manages a workload from the argocd.argoproj.io/tracking-id annotation.
Actions
| Action | Parameters | Notes |
|---|---|---|
rollback | name, id | History id to go back to. Turns auto-sync off first |
sync | name, prune | Dry run first |
terminate_operation | name | Stops a running sync and waits until Argo CD lets go of its hooks. Warns when auto-sync is about to start a new one |
set_auto_sync | name, enabled | Turn automated sync with self-heal back on after a rollback |
app_status, managing_app, controller_status, components_status, repo_status | Read-only evidence |
Certification
Live, on Argo CD v3.5.3 and v3.3.14 (the oldest supported line), with the least-privilege account above and 3AM running inside the cluster with the client RBAC (2026-10-05). 53 of 53 checks passed on each version:
- The connector. Health, discovery, the deploy behind each app (commit, author, message), and every error explained: a wrong token, an untrusted certificate, an app the account can't see.
- A bad release on an auto-syncing app. Rolled back through Argo CD with auto-sync off, still healthy 45 seconds
later. A
kubectlrollback is undone by self-heal within 15 seconds. The approval showed which commit went out, and the note said which commit to revert in Git. - A release whose database was down. Left alone, with advice.
- Argo CD and Git problems. Manifests that don't render, an unreachable Git host and a resource changed by hand got advice, and nothing ran. A sync stuck on a hook was stopped. Resources deleted from a manual app were re-synced from Git after a dry run.
- Argo CD itself, hung mid-run. The repo-server and the application controller were each frozen. 3AM located them, restarted them through the Kubernetes connector and verified they were back.
- The audit trail. No token reached the ledger, and the ledger verified.
Troubleshooting
the token is invalid or expired
Generate a new token for the threeam account (step 1).
the certificate is not trusted: set ca_file
Argo CD's certificate is self-signed or from your internal CA. Set CA certificate to the CA that signed it.
the 3AM account's Argo CD role does not allow this
The account lacks sync or update in argocd-rbac-cm. That's deliberate for a read-only setup; add them for fixes.
Argo CD's components aren't restarted
The Kubernetes connector must watch the argocd namespace, with the 3AM Role applied there.
VictoriaMetrics
Alerts from vmalert or Alertmanager, MetricsQL as evidence, and fixes when VictoriaMetrics itself is in trouble. Single-node and cluster.
ServiceNow
Incidents and change records kept current, recent changes as evidence, knowledge articles in every diagnosis, fixes ordered through the catalogue, and ServiceNow's own failures handled.