# Argo CD

> Deploy history as evidence, the GitOps-correct rollback, and fixes when Argo CD itself breaks.

Source: https://docs.3am.si/connectors/argocd · Markdown: https://docs.3am.si/connectors/argocd.md · All docs: https://docs.3am.si/llms.txt
3AM is an on-prem AI on-call engineer for regulated enterprises (https://3am.si).

With Argo CD connected, every incident on an Argo-managed workload starts with the question an on-call engineer asks
first: what changed? 3AM shows the revision that went out, its author and commit message, and when it synced. Argo
CD also changes how a bad release is undone, because its self-heal re-applies Git: a `kubectl` rollback of an
Argo-managed Deployment is undone within seconds (we measured 15). With this connector, 3AM rolls back through Argo CD
instead. It also watches Argo CD itself: a repo-server or controller that hangs, a sync stuck on a hook, a Git host
that's unreachable.

**Before you start**

* Admin access to Argo CD, once, to create a local account for 3AM.
* Network access from 3AM to the Argo CD API server, and to the components' metrics ports if you want 3AM to notice
  them hanging.
* The [Kubernetes connector](https://docs.3am.si/connectors/kubernetes), so 3AM can prove why a workload fails and restart Argo CD's own
  components. Add the `argocd` namespace to its namespaces and apply the 3AM Role there.

## 1. Create the 3AM account

Add a local account with an API key, and a role that can read applications, sync them, and turn auto-sync off for a
rollback:

```yaml title="argocd-cm"
data:
  accounts.threeam: apiKey
```

```yaml title="argocd-rbac-cm"
data:
  policy.csv: |
    p, role:threeam, applications, get, */*, allow
    p, role:threeam, applications, sync, */*, allow      # sync, roll back, stop a stuck sync
    p, role:threeam, applications, update, */*, allow    # turn auto-sync off before a rollback
    p, role:threeam, projects, get, *, allow
    g, threeam, role:threeam
```

Leave out `sync` and `update` for a read-only account: 3AM then diagnoses and hands fixes to a person. Scope `*/*` to
`<project>/*` to limit 3AM to some projects.

Create the token:

```bash
argocd account generate-token --account threeam > 3am-argocd-token
```

## 2. Add the connector

In the console: **Settings → Add a connection → Argo CD**. Or in `connectors.json`:

```json title="/etc/3am/connectors.json"
{ "name": "argocd", "kind": "argocd",
  "settings": { "url": "https://argocd-server.argocd.svc",
                "token": { "file": "/etc/3am/secrets/3am-argocd-token" },
                "ca_file": "/etc/3am/secrets/argocd-ca.pem",
                "component_urls": ["http://argocd-metrics.argocd.svc:8082",
                                   "http://argocd-repo-server.argocd.svc:8084",
                                   "http://argocd-server-metrics.argocd.svc:8083"] } }
```

Argo CD serves a self-signed certificate unless you gave it one. Point `ca_file` at the CA that signed it, rather than
turning verification off.

Connector settings (kind `argocd`; alerts in, actions):

| Setting | What to enter | Required | Default |
| --- | --- | --- | --- |
| `url` | Argo CD API URL, e.g. https://argocd.bank.internal | yes | — |
| `token` (secret) | API token of the 3AM account | yes | — |
| `projects` | Projects to watch (empty: all) | no | — |
| `component_urls` | Component metrics URLs: application-controller (8082), repo-server (8084), server (8083) | no | — |
| `ca_file` | CA certificate for a private CA (PEM file) | no | — |
| `verify_tls` | Verify TLS | yes | `true` |
| `degraded_after_s` | Seconds Degraded before it's an incident | yes | `120` |
| `stuck_after_s` | Seconds a sync may run before it's stuck | yes | `900` |
| `drift_after_s` | Seconds OutOfSync (no sync running) before it's drift | yes | `1800` |
| `stall_after_s` | Seconds without reconciling before the controller is stalled | yes | `600` |

## 3. Verify

```text title="Test connection"
Argo CD v3.5.3; 7 applications
```

If a component listed in `component_urls` doesn't answer, Test connection names it.

When an Argo-managed workload has a bad release, the incident's proposed fix is an Argo CD rollback, shown as
`rollback id=0, name=payments (via argocd)`, where `argocd` is your connector's name. It is never a `kubectl`
rollback.

## What it detects

| Signal                    | When                                                                                                                        |
| ------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| `ArgoCDAppDegraded`       | An app has been Degraded for `degraded_after_s` (120 s). Carries the failing workload, so the Kubernetes pack can prove why |
| `ArgoCDSyncFailed`        | A sync failed while applying resources                                                                                      |
| `ArgoCDSyncStuck`         | A sync has run longer than `stuck_after_s` (15 min), usually a hook that never finishes                                     |
| `ArgoCDComparisonError`   | Manifests can't be generated: they don't render, Git is unreachable or refuses credentials, or the repo-server is down      |
| `ArgoCDAppMissing`        | Resources the app should manage are missing from the cluster                                                                |
| `ArgoCDAppOutOfSync`      | Live resources differ from Git for `drift_after_s` (30 min) with self-heal off                                              |
| `ArgoCDControllerStalled` | No app reconciled for `stall_after_s` (10 min)                                                                              |
| `ArgoCDComponentDown`     | The API server or a component in `component_urls` doesn't answer                                                            |

Signals carry the app in the `name` label, like `argocd_app_info`. So if you already alert with the community rules
(`ArgoCdAppUnhealthy`, `ArgoCdAppSyncFailed`, `ArgoCdAppOutOfSync`) through Prometheus, the same causes apply.

## What it proves and fixes

**Fixed with approval:**

| Cause                            | How 3AM knows                                                                                                | Fix                                                                                                                                                              |
| -------------------------------- | ------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Bad release                      | A sync in the last 6 hours, the app Degraded, its pods failing, and the crash logs show no dependency errors | Roll back through Argo CD: auto-sync off, back to the previous deployment, verified healthy                                                                      |
| Sync stuck on a hook             | The operation has run too long, waiting on a hook                                                            | Stop the sync, so later syncs can run. If Git has moved on, auto-sync starts a new sync that waits on the same hook, and the note says so: the hook needs fixing |
| Resources missing, auto-sync off | Health `Missing` and nothing syncing                                                                         | Sync from Git (dry run first, nothing pruned)                                                                                                                    |
| Repo-server down or hung         | Manifest generation fails with the repo-server's connection error, or it doesn't answer                      | Restart its Deployment (Kubernetes connector)                                                                                                                    |
| Application controller stalled   | Nothing has reconciled for too long                                                                          | Restart its pod (Kubernetes connector)                                                                                                                           |
| A component down                 | It doesn't answer                                                                                            | Restart its pod or workload (Kubernetes connector)                                                                                                               |

**Handed to a person:**

| Cause                                 | How 3AM knows                                                                            |
| ------------------------------------- | ---------------------------------------------------------------------------------------- |
| Degraded because a dependency is down | The crash logs show it can't reach something. Rolling back wouldn't help, so 3AM doesn't |
| Manifests don't render                | The Helm, Kustomize or YAML error, with file and line. The running version is untouched  |
| Git unreachable                       | The Git host's DNS, network or TLS error                                                 |
| Git refuses credentials               | `authentication required` and similar                                                    |
| Sync failed applying resources        | Each resource's error: validation, an admission webhook, an immutable field, a quota     |
| Drift                                 | What differs from Git. It may be a deliberate hotfix, and syncing would undo it          |

### How a rollback works

1. **Rehearse.** Argo CD renders the target revision. Argo refuses a dry-run rollback while auto-sync is on, so
   rendering is the rehearsal: it proves the old revision still produces valid manifests.
2. **Approve.** The request shows the commit that went out ("payments 1.1: read ledger url from config", Jane Doe) and
   the revision it goes back to.
3. **Turn auto-sync off,** so self-heal can't re-apply the bad commit.
4. **Roll back** to the previous deployment, and wait until the app is healthy.
5. **Hand back.** The incident says which commit to revert in Git, and to turn auto-sync back on (`set_auto_sync`, or
   in Argo CD).

The [Kubernetes pack](https://docs.3am.si/connectors/kubernetes) follows the same rule. A bad release of an Argo-managed workload is
rolled back through Argo CD. If 3AM has no Argo CD connection, it hands the rollback to a person rather than using
`kubectl`. It knows which app manages a workload from the `argocd.argoproj.io/tracking-id` annotation.

## Actions

| Action                                                                                | Parameters        | Notes                                                                                                               |
| ------------------------------------------------------------------------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------- |
| `rollback`                                                                            | `name`, `id`      | History id to go back to. Turns auto-sync off first                                                                 |
| `sync`                                                                                | `name`, `prune`   | Dry run first                                                                                                       |
| `terminate_operation`                                                                 | `name`            | Stops a running sync and waits until Argo CD lets go of its hooks. Warns when auto-sync is about to start a new one |
| `set_auto_sync`                                                                       | `name`, `enabled` | Turn automated sync with self-heal back on after a rollback                                                         |
| `app_status`, `managing_app`, `controller_status`, `components_status`, `repo_status` |                   | Read-only evidence                                                                                                  |

## Certification

Live, on Argo CD v3.5.3 and v3.3.14 (the oldest supported line), with the least-privilege account above and 3AM
running inside the cluster with the client RBAC (2026-10-05). 53 of 53 checks passed on each version:

* **The connector.** Health, discovery, the deploy behind each app (commit, author, message), and every error
  explained: a wrong token, an untrusted certificate, an app the account can't see.
* **A bad release on an auto-syncing app.** Rolled back through Argo CD with auto-sync off, still healthy 45 seconds
  later. A `kubectl` rollback is undone by self-heal within 15 seconds. The approval showed which commit went out,
  and the note said which commit to revert in Git.
* **A release whose database was down.** Left alone, with advice.
* **Argo CD and Git problems.** Manifests that don't render, an unreachable Git host and a resource changed by hand got
  advice, and nothing ran. A sync stuck on a hook was stopped. Resources deleted from a manual app were re-synced
  from Git after a dry run.
* **Argo CD itself, hung mid-run.** The repo-server and the application controller were each frozen. 3AM located them,
  restarted them through the Kubernetes connector and verified they were back.
* **The audit trail.** No token reached the ledger, and the ledger verified.

## Troubleshooting

### `the token is invalid or expired`

Generate a new token for the `threeam` account (step 1).

### `the certificate is not trusted: set ca_file`

Argo CD's certificate is self-signed or from your internal CA. Set **CA certificate** to the CA that signed it.

### `the 3AM account's Argo CD role does not allow this`

The account lacks `sync` or `update` in `argocd-rbac-cm`. That's deliberate for a read-only setup; add them for fixes.

### Argo CD's components aren't restarted

The Kubernetes connector must watch the `argocd` namespace, with the 3AM Role applied there.
