VictoriaMetrics
Alerts from vmalert or Alertmanager, MetricsQL as evidence, and fixes when VictoriaMetrics itself is in trouble. Single-node and cluster.
3AM reads firing alerts from Alertmanager (silence-aware) or straight from vmalert, and runs every pack's PromQL on
VictoriaMetrics unchanged. It also watches VictoriaMetrics itself, through each component's own /metrics. When
storage turns read-only, vmagent can't deliver, a storage node drops out or a component hangs, 3AM proves which, and
fixes it through the Kubernetes connector where it can.
Before you start
- Read access to the query endpoint and to each component's
/metrics. - Write access to Alertmanager silences and to
/-/reloadon vmagent and vmalert, only if 3AM may use them. Both are approved every time. - Outbound access from 3AM to every URL below.
1. Create credentials
| Your setup | Give 3AM |
|---|---|
| No authentication | The URLs |
-httpAuth.username / -httpAuth.password | That user and password. 3AM uses them for every component |
| vmauth in front | A vmauth user with read access to the query paths and /metrics |
| Alertmanager behind a different proxy | Its own credentials (the alertmanager_* settings) |
HTTPS (-tls) with an internal CA | The CA certificate, as a PEM file |
A read-only vmauth user for a cluster:
users:
- username: 3am
password: a-long-random-password
url_map:
- src_paths: ["/select/.*"]
url_prefix: "http://vmselect:8481"2. Decide what to connect
The query endpoint alone can't show disk, queues or limits. Those live in each component's /metrics, so give 3AM
the components' own addresses too.
| Setting | Single-node | Cluster |
|---|---|---|
url | http://victoriametrics:8428 | http://vmselect:8481/select/0/prometheus (your tenant) |
vmagent_urls | each vmagent, e.g. http://vmagent:8429 | same |
vmalert_url | http://vmalert:8880 | same |
component_urls | each vminsert, vmselect and vmstorage | |
alertmanager_url | http://alertmanager:9093 | same |
On Kubernetes, list each vmstorage pod through its headless Service (vmstorage-0.vmstorage.monitoring.svc:8482, …)
so 3AM can tell which node is in trouble, and restart only that one.
If you leave out the vmagents, 3AM can't see scrape targets or remote write. Leave out the cluster components and it can't see storage nodes. Checks that need them report inconclusive, never passed.
3. Add the connector
In the console: Settings → Add a connection → VictoriaMetrics. Or in connectors.json:
{ "name": "vm", "kind": "victoriametrics",
"settings": { "url": "http://vmselect.monitoring.svc:8481/select/0/prometheus",
"vmagent_urls": ["http://vmagent.monitoring.svc:8429"],
"vmalert_url": "http://vmalert.monitoring.svc:8880",
"component_urls": ["http://vminsert.monitoring.svc:8480",
"http://vmstorage-0.vmstorage.monitoring.svc:8482",
"http://vmstorage-1.vmstorage.monitoring.svc:8482"],
"alertmanager_url": "http://alertmanager.monitoring.svc:9093",
"username": "3am", "password": { "file": "/etc/3am/secrets/vm-password" } } }kind: victoriametricsalerts inmetrics and queriesactions| Setting | What to enter | Required | Default |
|---|---|---|---|
url | Query URL: single-node (http://victoriametrics:8428), vmselect with tenant (http://vmselect:8481/select/0/prometheus), or vmauth | yes | — |
vmagent_urls | vmagent base URLs (scrape targets and remote write) | no | — |
vmalert_url | vmalert base URL (rules and alerts) | no | — |
component_urls | Cluster components 3AM watches: vmstorage, vminsert, vmselect base URLs | no | — |
alertmanager_url | Alertmanager base URL (routed alerts, silences) | no | — |
bearer_tokensecret | Bearer token, if required | no | — |
username | Basic-auth user, if required | no | — |
passwordsecret | Basic-auth password | no | — |
alertmanager_bearer_tokensecret | Alertmanager bearer token (if different) | no | — |
alertmanager_username | Alertmanager basic-auth user (if different) | no | — |
alertmanager_passwordsecret | Alertmanager basic-auth password (if different) | no | — |
ca_file | CA certificate for a private CA (PEM file) | no | — |
verify_tls | Verify TLS | yes | true |
4. Verify
Test connection runs a query and reads every component's /metrics. Each component is named with its version:
Connected. vmselect v1.153.0-cluster + vmagent v1.153.0 + vmalert v1.153.0 + vminsert v1.153.0-cluster
+ vmstorage v1.153.0-cluster + vmstorage v1.153.0-cluster + Alertmanager 0.34.1For the whole path, load a rule that always fires into vmalert and reload it:
groups:
- name: 3am-test
rules:
- alert: ThreeAMPipelineTest
expr: vector(1)
labels: { severity: critical, service: monitoring }curl -X POST http://vmalert:8880/-/reloadWithin a minute the console shows an incident for ThreeAMPipelineTest. Remove the rule afterwards.
What it proves and fixes
3AM raises these itself, even with no rule for them. If you run VictoriaMetrics' own alerting rules
(RowsRejectedOnIngestion, TooManyRemoteWriteErrors, ConfigurationReloadFailure, …), 3AM uses yours and proves
the cause behind them.
Fixed with approval (through the Kubernetes connector):
| Cause | How 3AM knows | Fix |
|---|---|---|
| Storage read-only (disk full) | vm_storage_is_read_only: free disk below -storage.minFreeDiskSpaceBytes | Grow that node's volume by half. Two approvers |
| Disk filling up | Over 80% used, not read-only yet | Grow the fullest node's volume. Two approvers |
| Component down or hung | It doesn't answer | Restart its pod (StatefulSet) or its workload |
| Storage node unreachable | vminsert's per-node reachability, vmselect's connection errors | Restart that vmstorage pod only |
| Scrape target unreachable | The scrape error, per job | Restart the target's workload |
Handed to a person with the evidence (the right fix is a judgement call):
| Cause | How 3AM knows | Your data |
|---|---|---|
| vmagent can't reach storage | Retries and no answer at all | Kept in vmagent's queue |
| Storage refuses vmagent's credentials | 401 / 403 | Kept in vmagent's queue |
| Storage rejects vmagent's data | 400 / 409, blocks dropped | Lost |
| Storage fails writes | 5xx / 429, while storage is writable and every node reachable | Kept in vmagent's queue |
| Series limit reached | *_series_limit_rows_dropped_total rising, and the metrics with the most series | New series lost |
| Rows rejected | vm_rows_ignored_total rising, with the reason | Those rows lost |
| High churn | Over 10% of ingested rows create new series; the labels with the most values | |
| Target refuses credentials, or answers garbage | The scrape error. A garbage target stays up in vmagent with 0 samples, so up == 0 never fires | No metrics from it |
| Config reload failed | *_config_last_reload_successful = 0 | Runs on the last good config |
| Rules broken / datasource unreachable | vmalert rule health: a rule error versus a connection error | Alerts can't fire |
| vmalert can't notify | vmalert_alerts_send_errors_total rising | Nobody is paged |
A few things worth knowing about how it decides:
- The fix targets the proven culprit. If vmagent's alert opens the incident but a storage node is read-only, 3AM grows that node's volume, not vmagent's.
- Lost vs delayed. vmagent records each storage answer's exact status code, so 3AM can tell whether data is dropped or only queued.
- No guessing. "Storage fails writes" is only concluded when storage is visible, writable and complete. Otherwise the read-only disk or the lost node is the cause.
- Only what's happening now. Counters are compared with 3AM's earlier readings, so last week's dropped rows aren't today's incident.
- Partial results are never evidence. With a storage node down, vmselect still answers, flagged partial; 3AM refuses those answers.
Actions
| Action | Parameters | Notes |
|---|---|---|
silence | alertname, minutes, comment, matchers | 1–1440 minutes in Alertmanager. The rehearsal lists what it would silence |
unsilence | silence_id | |
reload | component (vmagent or vmalert), instance | vmagent answers 200 even when it refuses a broken file, so 3AM checks the reload status afterwards |
targets, storage_status, remote_write_status, ingestion_status, cardinality, components_status, vmalert_status, cluster_status, volume_usage | Read-only evidence |
Troubleshooting
VictoriaMetrics: authentication failed
Check the user and password, or the bearer token.
Alertmanager: authentication failed
Alertmanager has its own credentials. Fill in the alertmanager_* settings.
the certificate is not trusted: set ca_file
Set CA certificate to your internal CA, in PEM format.
answers, but partially: a storage node is unreachable
A vmstorage node is down. 3AM refuses partial results as evidence until it's back.
vmagent (…): HTTP 404
That URL isn't a vmagent, or a proxy hides /metrics. Give 3AM the vmagent's own address.
Certification
Live, on VictoriaMetrics v1.153.0 and the long-term-support v1.136.0, 97 checks on each (2026-10-04).
The estate:
- single-node over HTTPS with an internal CA and basic auth;
- a cluster (vminsert, vmselect, two vmstorage);
- 7 vmagents and 2 vmalerts;
- Alertmanager behind a bearer-token proxy, and the real
mysqld_exporter.
One thing was broken per component. Targets:
- refusing connections, rejecting credentials or answering garbage;
- carrying too many labels, or churning series.
vmagents writing:
- to a dead port, or with wrong credentials;
- through a proxy that rejects the data;
- into read-only storage, or past a series limit.
Alerting and storage:
- a dead datasource, a broken rule, a dead notifier and a broken reload;
- a storage node and a vmagent stopped mid-run.
Every cause was proven from live evidence. Read-only storage was never taken for failing storage, and partial answers were refused. On Kubernetes, hung vmagent, vmstorage and exporter pods were restarted and verified (details).