Argo CD

Deploy history as evidence, the GitOps-correct rollback, and fixes when Argo CD itself breaks.

With Argo CD connected, every incident on an Argo-managed workload starts with the question an on-call engineer asks first: what changed? 3AM shows the revision that went out, its author and commit message, and when it synced. Argo CD also changes how a bad release is undone, because its self-heal re-applies Git: a kubectl rollback of an Argo-managed Deployment is undone within seconds (we measured 15). With this connector, 3AM rolls back through Argo CD instead. It also watches Argo CD itself: a repo-server or controller that hangs, a sync stuck on a hook, a Git host that's unreachable.

Before you start

  • Admin access to Argo CD, once, to create a local account for 3AM.
  • Network access from 3AM to the Argo CD API server, and to the components' metrics ports if you want 3AM to notice them hanging.
  • The Kubernetes connector, so 3AM can prove why a workload fails and restart Argo CD's own components. Add the argocd namespace to its namespaces and apply the 3AM Role there.

1. Create the 3AM account

Add a local account with an API key, and a role that can read applications, sync them, and turn auto-sync off for a rollback:

argocd-cm
data:
  accounts.threeam: apiKey
argocd-rbac-cm
data:
  policy.csv: |
    p, role:threeam, applications, get, */*, allow
    p, role:threeam, applications, sync, */*, allow      # sync, roll back, stop a stuck sync
    p, role:threeam, applications, update, */*, allow    # turn auto-sync off before a rollback
    p, role:threeam, projects, get, *, allow
    g, threeam, role:threeam

Leave out sync and update for a read-only account: 3AM then diagnoses and hands fixes to a person. Scope */* to <project>/* to limit 3AM to some projects.

Create the token:

argocd account generate-token --account threeam > 3am-argocd-token

2. Add the connector

In the console: Settings → Add a connection → Argo CD. Or in connectors.json:

/etc/3am/connectors.json
{ "name": "argocd", "kind": "argocd",
  "settings": { "url": "https://argocd-server.argocd.svc",
                "token": { "file": "/etc/3am/secrets/3am-argocd-token" },
                "ca_file": "/etc/3am/secrets/argocd-ca.pem",
                "component_urls": ["http://argocd-metrics.argocd.svc:8082",
                                   "http://argocd-repo-server.argocd.svc:8084",
                                   "http://argocd-server-metrics.argocd.svc:8083"] } }

Argo CD serves a self-signed certificate unless you gave it one. Point ca_file at the CA that signed it, rather than turning verification off.

kind: argocdalerts inactions
SettingWhat to enterRequiredDefault
urlArgo CD API URL, e.g. https://argocd.bank.internalyes—
tokensecretAPI token of the 3AM accountyes—
projectsProjects to watch (empty: all)no—
component_urlsComponent metrics URLs: application-controller (8082), repo-server (8084), server (8083)no—
ca_fileCA certificate for a private CA (PEM file)no—
verify_tlsVerify TLSyestrue
degraded_after_sSeconds Degraded before it's an incidentyes120
stuck_after_sSeconds a sync may run before it's stuckyes900
drift_after_sSeconds OutOfSync (no sync running) before it's driftyes1800
stall_after_sSeconds without reconciling before the controller is stalledyes600

3. Verify

Test connection
Argo CD v3.5.3; 7 applications

If a component listed in component_urls doesn't answer, Test connection names it.

When an Argo-managed workload has a bad release, the incident's proposed fix is an Argo CD rollback, shown as rollback id=0, name=payments (via argocd), where argocd is your connector's name. It is never a kubectl rollback.

What it detects

SignalWhen
ArgoCDAppDegradedAn app has been Degraded for degraded_after_s (120 s). Carries the failing workload, so the Kubernetes pack can prove why
ArgoCDSyncFailedA sync failed while applying resources
ArgoCDSyncStuckA sync has run longer than stuck_after_s (15 min), usually a hook that never finishes
ArgoCDComparisonErrorManifests can't be generated: they don't render, Git is unreachable or refuses credentials, or the repo-server is down
ArgoCDAppMissingResources the app should manage are missing from the cluster
ArgoCDAppOutOfSyncLive resources differ from Git for drift_after_s (30 min) with self-heal off
ArgoCDControllerStalledNo app reconciled for stall_after_s (10 min)
ArgoCDComponentDownThe API server or a component in component_urls doesn't answer

Signals carry the app in the name label, like argocd_app_info. So if you already alert with the community rules (ArgoCdAppUnhealthy, ArgoCdAppSyncFailed, ArgoCdAppOutOfSync) through Prometheus, the same causes apply.

What it proves and fixes

Fixed with approval:

CauseHow 3AM knowsFix
Bad releaseA sync in the last 6 hours, the app Degraded, its pods failing, and the crash logs show no dependency errorsRoll back through Argo CD: auto-sync off, back to the previous deployment, verified healthy
Sync stuck on a hookThe operation has run too long, waiting on a hookStop the sync, so later syncs can run. If Git has moved on, auto-sync starts a new sync that waits on the same hook, and the note says so: the hook needs fixing
Resources missing, auto-sync offHealth Missing and nothing syncingSync from Git (dry run first, nothing pruned)
Repo-server down or hungManifest generation fails with the repo-server's connection error, or it doesn't answerRestart its Deployment (Kubernetes connector)
Application controller stalledNothing has reconciled for too longRestart its pod (Kubernetes connector)
A component downIt doesn't answerRestart its pod or workload (Kubernetes connector)

Handed to a person:

CauseHow 3AM knows
Degraded because a dependency is downThe crash logs show it can't reach something. Rolling back wouldn't help, so 3AM doesn't
Manifests don't renderThe Helm, Kustomize or YAML error, with file and line. The running version is untouched
Git unreachableThe Git host's DNS, network or TLS error
Git refuses credentialsauthentication required and similar
Sync failed applying resourcesEach resource's error: validation, an admission webhook, an immutable field, a quota
DriftWhat differs from Git. It may be a deliberate hotfix, and syncing would undo it

How a rollback works

  1. Rehearse. Argo CD renders the target revision. Argo refuses a dry-run rollback while auto-sync is on, so rendering is the rehearsal: it proves the old revision still produces valid manifests.
  2. Approve. The request shows the commit that went out ("payments 1.1: read ledger url from config", Jane Doe) and the revision it goes back to.
  3. Turn auto-sync off, so self-heal can't re-apply the bad commit.
  4. Roll back to the previous deployment, and wait until the app is healthy.
  5. Hand back. The incident says which commit to revert in Git, and to turn auto-sync back on (set_auto_sync, or in Argo CD).

The Kubernetes pack follows the same rule. A bad release of an Argo-managed workload is rolled back through Argo CD. If 3AM has no Argo CD connection, it hands the rollback to a person rather than using kubectl. It knows which app manages a workload from the argocd.argoproj.io/tracking-id annotation.

Actions

ActionParametersNotes
rollbackname, idHistory id to go back to. Turns auto-sync off first
syncname, pruneDry run first
terminate_operationnameStops a running sync and waits until Argo CD lets go of its hooks. Warns when auto-sync is about to start a new one
set_auto_syncname, enabledTurn automated sync with self-heal back on after a rollback
app_status, managing_app, controller_status, components_status, repo_statusRead-only evidence

Certification

Live, on Argo CD v3.5.3 and v3.3.14 (the oldest supported line), with the least-privilege account above and 3AM running inside the cluster with the client RBAC (2026-10-05). 53 of 53 checks passed on each version:

  • The connector. Health, discovery, the deploy behind each app (commit, author, message), and every error explained: a wrong token, an untrusted certificate, an app the account can't see.
  • A bad release on an auto-syncing app. Rolled back through Argo CD with auto-sync off, still healthy 45 seconds later. A kubectl rollback is undone by self-heal within 15 seconds. The approval showed which commit went out, and the note said which commit to revert in Git.
  • A release whose database was down. Left alone, with advice.
  • Argo CD and Git problems. Manifests that don't render, an unreachable Git host and a resource changed by hand got advice, and nothing ran. A sync stuck on a hook was stopped. Resources deleted from a manual app were re-synced from Git after a dry run.
  • Argo CD itself, hung mid-run. The repo-server and the application controller were each frozen. 3AM located them, restarted them through the Kubernetes connector and verified they were back.
  • The audit trail. No token reached the ledger, and the ledger verified.

Troubleshooting

the token is invalid or expired

Generate a new token for the threeam account (step 1).

the certificate is not trusted: set ca_file

Argo CD's certificate is self-signed or from your internal CA. Set CA certificate to the CA that signed it.

the 3AM account's Argo CD role does not allow this

The account lacks sync or update in argocd-rbac-cm. That's deliberate for a read-only setup; add them for fixes.

Argo CD's components aren't restarted

The Kubernetes connector must watch the argocd namespace, with the 3AM Role applied there.

On this page