Prometheus and Alertmanager
Alerts in, PromQL evidence for every check, approved silences, and 3AM knows when your monitoring itself is broken.
What 3AM does with Prometheus
| In short | |
|---|---|
| Alerts in | Firing alerts from Alertmanager become 3AM incidents. Silenced and inhibited alerts don't |
| Evidence | 3AM runs PromQL to prove root causes for its other packs (MySQL, Kubernetes, JVM, nginx…) |
| Monitoring itself | 3AM notices when monitoring is broken (targets down, rules failing, config not reloading, notifications failing), because then its own evidence is unreliable |
| Fixes | With approval, through the Kubernetes connector: restart a hung Prometheus, Alertmanager or scrape target; grow Prometheus' volume when its storage is full. Each is rehearsed, then verified |
| Actions | With approval: silence alerts in Alertmanager for a while, and expire silences |
1. Prepare access
3AM needs read access to Prometheus' HTTP API, and to Alertmanager's. It needs write access to Alertmanager only if it may create silences (always approved).
| Your setup | What to give 3AM |
|---|---|
| No authentication | Just the URLs |
Prometheus web.yml with basic auth | A user and password (add a user for 3AM to basic_auth_users) |
| A reverse proxy / SSO in front (bearer tokens) | A token for 3AM |
| Alertmanager behind a different proxy | Separate Alertmanager credentials (fields below) |
| HTTPS with an internal CA | The CA certificate (PEM) |
Adding a basic-auth user for 3AM to Prometheus' web.yml (the password must be a bcrypt hash):
htpasswd -nbBC 10 3am 'a-long-random-password' | cut -d: -f2# web.yml
basic_auth_users:
3am: $2y$10$… # the hash from above2. Exporters 3AM's checks use
3AM's checks read standard exporter metrics. A stock exporter covers the core checks; a few need an extra flag:
| Technology | Exporter | Covered by default | Enable for more |
|---|---|---|---|
| MySQL / MariaDB | mysqld_exporter | up, read-only, connection usage, row-lock waits | --collect.info_schema.processlist (per-user connections); replication lag is reported on replicas |
| Alertmanager | its own /metrics (scrape it) | notification failures | |
| Kubernetes | kube-state-metrics | container restarts, OOM kills |
When a metric is missing, the check is inconclusive, never passed. The MySQL pack confirms through SQL where a metric isn't standard.
3. Connect it in 3AM
Settings → Add a connection → Prometheus / Alertmanager (or the wizard's Monitoring step):
kind: prometheusalerts inmetrics and queriesactions| Setting | What to enter | Required | Default |
|---|---|---|---|
url | Prometheus base URL | yes | — |
alertmanager_url | Alertmanager base URL | no | — |
bearer_tokensecret | Bearer token, if required | no | — |
username | Basic-auth user, if required | no | — |
passwordsecret | Basic-auth password | no | — |
alertmanager_bearer_tokensecret | Alertmanager bearer token (if different) | no | — |
alertmanager_username | Alertmanager basic-auth user (if different) | no | — |
alertmanager_passwordsecret | Alertmanager basic-auth password (if different) | no | — |
ca_file | CA certificate for a private CA (PEM file) | no | — |
verify_tls | Verify TLS | yes | true |
Test connection runs a real query (not just a ping) and checks Alertmanager with its own credentials. A healthy result shows what 3AM found:
Connected. Prometheus 3.15.0 + Alertmanager 0.34.1
jobs: 6 (mysql, prometheus, …) services: … alert rules: 95 targets: 32Network: outbound from 3AM to Prometheus and Alertmanager. Nothing connects to 3AM.
4. Check it end to end
Create a test rule that always fires, and watch it reach 3AM:
# a rule file Prometheus loads
groups:
- name: 3am-test
rules:
- alert: ThreeAMPipelineTest
expr: vector(1)
labels: {severity: critical, service: monitoring}Reload Prometheus. Within a minute the console shows an incident for ThreeAMPipelineTest (3AM will find nothing to
prove, and say so). Remove the rule afterwards.
What 3AM understands about monitoring
3AM raises these itself from Prometheus' status, even when you have no alert rule for them. If your own rule for the same thing is already firing, 3AM uses yours instead:
| Root cause | Proven by | 3AM does |
|---|---|---|
| Prometheus or Alertmanager down or hung | It doesn't answer (3AM says so even when Prometheus is the one down) | Restarts its pod or workload, then checks it answers |
| Prometheus' storage full | The TSDB fails to write its WAL or to compact | Grows Prometheus' volume by half (two approvers), then checks the failures stop |
| Target unreachable | The scrape error is connection refused, DNS or timeout | Restarts the target's workload when it runs in a namespace 3AM watches, then checks the target is up. Otherwise advice |
| Target refuses credentials | The scrape error is 401/403 | Advice: update the scrape job's credentials |
| Target answers garbage | The scrape error is a parse or content-type failure | Advice: check the exporter |
| Config reload failed | Prometheus' runtime status | Advice: promtool check config, fix, reload |
| Rules failing | Rule health, with each failing rule's error | Advice: fix the named rules (their alerts can't fire until then) |
| No Alertmanager | Prometheus' Alertmanager discovery | Advice: alerts reach nobody until it is fixed |
| Notifications failing | alertmanager_notifications_failed_total rising, and which integration | Advice: fix the receiver's endpoint or credentials |
Restarts and volume growth go through the Kubernetes connector, which finds the workload behind Prometheus', Alertmanager's or the target's address. Without it the same causes are handed to a person. Configuration causes (credentials, reload, rules, receivers) stay with a person, with the error attached.
Silences
| Action | What it does |
|---|---|
silence | Silences matching alerts for 1–1440 minutes, with a comment and createdBy: 3AM. The rehearsal lists which alerts it would silence |
unsilence | Expires a silence |
Troubleshooting
| Test connection says | Fix |
|---|---|
| Prometheus: authentication failed | Check the user and password, or the bearer token |
| Alertmanager: authentication failed | Alertmanager has its own credentials: fill in the Alertmanager fields |
| the certificate is not trusted: set ca_file | Choose your internal CA's certificate (PEM) |
| Prometheus refused the query (bad_data) | Shown for a malformed PromQL expression, with Prometheus' own message |
| Silenced alerts still appear | Without an Alertmanager URL, 3AM reads Prometheus' alerts, which know nothing of silences: add the Alertmanager URL |
Certification
Certified on 2026-10-04 against Prometheus 3.15 (HTTPS with an internal CA, basic auth) and Alertmanager 0.34 behind
a bearer-token proxy, with the real mysqld_exporter 0.20: 46 live checks. End to end: the database was set
read-only, this Prometheus fired MySQLReadOnly through Alertmanager, 3AM proved it with this Prometheus' metric,
fixed it (approved), and the alert resolved.
Fixes, certified live on 2026-10-04 (54 of 54 checks, on a real Kubernetes cluster, with 3AM inside it under the client RBAC). Components were hung mid-run, meaning their pod was Running and answered nothing. Each was detected, proven, located, rehearsed, approved, restarted and verified:
- vmagent, a vmstorage node (only that pod), a scrape target of vmagent and one of Prometheus, Alertmanager (seen by both monitoring connectors, restarted once) and Prometheus.
A read-only storage node got a proposal to grow its volume, approved by two people. The test storage driver accepts the request but can't grow volumes, and 3AM reported the fix as failed rather than claiming success.
Kubernetes
3AM watches your namespaces for failing workloads, proves why they fail, and rolls back bad releases with your approval.
VictoriaMetrics
Alerts from vmalert or Alertmanager, MetricsQL evidence for every check, approved silences and reloads, and 3AM knows when VictoriaMetrics itself is in trouble.