Docs
Connectors

Prometheus and Alertmanager

Alerts in, PromQL evidence for every check, approved silences, and 3AM knows when your monitoring itself is broken.

What 3AM does with Prometheus

In short
Alerts inFiring alerts from Alertmanager become 3AM incidents. Silenced and inhibited alerts don't
Evidence3AM runs PromQL to prove root causes for its other packs (MySQL, Kubernetes, JVM, nginx…)
Monitoring itself3AM notices when monitoring is broken (targets down, rules failing, config not reloading, notifications failing), because then its own evidence is unreliable
FixesWith approval, through the Kubernetes connector: restart a hung Prometheus, Alertmanager or scrape target; grow Prometheus' volume when its storage is full. Each is rehearsed, then verified
ActionsWith approval: silence alerts in Alertmanager for a while, and expire silences

1. Prepare access

3AM needs read access to Prometheus' HTTP API, and to Alertmanager's. It needs write access to Alertmanager only if it may create silences (always approved).

Your setupWhat to give 3AM
No authenticationJust the URLs
Prometheus web.yml with basic authA user and password (add a user for 3AM to basic_auth_users)
A reverse proxy / SSO in front (bearer tokens)A token for 3AM
Alertmanager behind a different proxySeparate Alertmanager credentials (fields below)
HTTPS with an internal CAThe CA certificate (PEM)

Adding a basic-auth user for 3AM to Prometheus' web.yml (the password must be a bcrypt hash):

htpasswd -nbBC 10 3am 'a-long-random-password' | cut -d: -f2
# web.yml
basic_auth_users:
  3am: $2y$10$…          # the hash from above

2. Exporters 3AM's checks use

3AM's checks read standard exporter metrics. A stock exporter covers the core checks; a few need an extra flag:

TechnologyExporterCovered by defaultEnable for more
MySQL / MariaDBmysqld_exporterup, read-only, connection usage, row-lock waits--collect.info_schema.processlist (per-user connections); replication lag is reported on replicas
Alertmanagerits own /metrics (scrape it)notification failures
Kuberneteskube-state-metricscontainer restarts, OOM kills

When a metric is missing, the check is inconclusive, never passed. The MySQL pack confirms through SQL where a metric isn't standard.

3. Connect it in 3AM

Settings → Add a connection → Prometheus / Alertmanager (or the wizard's Monitoring step):

kind: prometheusalerts inmetrics and queriesactions
SettingWhat to enterRequiredDefault
urlPrometheus base URLyes—
alertmanager_urlAlertmanager base URLno—
bearer_tokensecretBearer token, if requiredno—
usernameBasic-auth user, if requiredno—
passwordsecretBasic-auth passwordno—
alertmanager_bearer_tokensecretAlertmanager bearer token (if different)no—
alertmanager_usernameAlertmanager basic-auth user (if different)no—
alertmanager_passwordsecretAlertmanager basic-auth password (if different)no—
ca_fileCA certificate for a private CA (PEM file)no—
verify_tlsVerify TLSyestrue

Test connection runs a real query (not just a ping) and checks Alertmanager with its own credentials. A healthy result shows what 3AM found:

Connected. Prometheus 3.15.0 + Alertmanager 0.34.1
jobs: 6 (mysql, prometheus, …)   services: …   alert rules: 95   targets: 32

Network: outbound from 3AM to Prometheus and Alertmanager. Nothing connects to 3AM.

4. Check it end to end

Create a test rule that always fires, and watch it reach 3AM:

# a rule file Prometheus loads
groups:
  - name: 3am-test
    rules:
      - alert: ThreeAMPipelineTest
        expr: vector(1)
        labels: {severity: critical, service: monitoring}

Reload Prometheus. Within a minute the console shows an incident for ThreeAMPipelineTest (3AM will find nothing to prove, and say so). Remove the rule afterwards.

What 3AM understands about monitoring

3AM raises these itself from Prometheus' status, even when you have no alert rule for them. If your own rule for the same thing is already firing, 3AM uses yours instead:

Root causeProven by3AM does
Prometheus or Alertmanager down or hungIt doesn't answer (3AM says so even when Prometheus is the one down)Restarts its pod or workload, then checks it answers
Prometheus' storage fullThe TSDB fails to write its WAL or to compactGrows Prometheus' volume by half (two approvers), then checks the failures stop
Target unreachableThe scrape error is connection refused, DNS or timeoutRestarts the target's workload when it runs in a namespace 3AM watches, then checks the target is up. Otherwise advice
Target refuses credentialsThe scrape error is 401/403Advice: update the scrape job's credentials
Target answers garbageThe scrape error is a parse or content-type failureAdvice: check the exporter
Config reload failedPrometheus' runtime statusAdvice: promtool check config, fix, reload
Rules failingRule health, with each failing rule's errorAdvice: fix the named rules (their alerts can't fire until then)
No AlertmanagerPrometheus' Alertmanager discoveryAdvice: alerts reach nobody until it is fixed
Notifications failingalertmanager_notifications_failed_total rising, and which integrationAdvice: fix the receiver's endpoint or credentials

Restarts and volume growth go through the Kubernetes connector, which finds the workload behind Prometheus', Alertmanager's or the target's address. Without it the same causes are handed to a person. Configuration causes (credentials, reload, rules, receivers) stay with a person, with the error attached.

Silences

ActionWhat it does
silenceSilences matching alerts for 1–1440 minutes, with a comment and createdBy: 3AM. The rehearsal lists which alerts it would silence
unsilenceExpires a silence

Troubleshooting

Test connection saysFix
Prometheus: authentication failedCheck the user and password, or the bearer token
Alertmanager: authentication failedAlertmanager has its own credentials: fill in the Alertmanager fields
the certificate is not trusted: set ca_fileChoose your internal CA's certificate (PEM)
Prometheus refused the query (bad_data)Shown for a malformed PromQL expression, with Prometheus' own message
Silenced alerts still appearWithout an Alertmanager URL, 3AM reads Prometheus' alerts, which know nothing of silences: add the Alertmanager URL

Certification

Certified on 2026-10-04 against Prometheus 3.15 (HTTPS with an internal CA, basic auth) and Alertmanager 0.34 behind a bearer-token proxy, with the real mysqld_exporter 0.20: 46 live checks. End to end: the database was set read-only, this Prometheus fired MySQLReadOnly through Alertmanager, 3AM proved it with this Prometheus' metric, fixed it (approved), and the alert resolved.

Fixes, certified live on 2026-10-04 (54 of 54 checks, on a real Kubernetes cluster, with 3AM inside it under the client RBAC). Components were hung mid-run, meaning their pod was Running and answered nothing. Each was detected, proven, located, rehearsed, approved, restarted and verified:

  • vmagent, a vmstorage node (only that pod), a scrape target of vmagent and one of Prometheus, Alertmanager (seen by both monitoring connectors, restarted once) and Prometheus.

A read-only storage node got a proposal to grow its volume, approved by two people. The test storage driver accepts the request but can't grow volumes, and 3AM reported the fix as failed rather than claiming success.

On this page