VictoriaMetrics

Alerts from vmalert or Alertmanager, MetricsQL as evidence, and fixes when VictoriaMetrics itself is in trouble. Single-node and cluster.

3AM reads firing alerts from Alertmanager (silence-aware) or straight from vmalert, and runs every pack's PromQL on VictoriaMetrics unchanged. It also watches VictoriaMetrics itself, through each component's own /metrics. When storage turns read-only, vmagent can't deliver, a storage node drops out or a component hangs, 3AM proves which, and fixes it through the Kubernetes connector where it can.

Before you start

  • Read access to the query endpoint and to each component's /metrics.
  • Write access to Alertmanager silences and to /-/reload on vmagent and vmalert, only if 3AM may use them. Both are approved every time.
  • Outbound access from 3AM to every URL below.

1. Create credentials

Your setupGive 3AM
No authenticationThe URLs
-httpAuth.username / -httpAuth.passwordThat user and password. 3AM uses them for every component
vmauth in frontA vmauth user with read access to the query paths and /metrics
Alertmanager behind a different proxyIts own credentials (the alertmanager_* settings)
HTTPS (-tls) with an internal CAThe CA certificate, as a PEM file

A read-only vmauth user for a cluster:

vmauth config
users:
  - username: 3am
    password: a-long-random-password
    url_map:
      - src_paths: ["/select/.*"]
        url_prefix: "http://vmselect:8481"

2. Decide what to connect

The query endpoint alone can't show disk, queues or limits. Those live in each component's /metrics, so give 3AM the components' own addresses too.

SettingSingle-nodeCluster
urlhttp://victoriametrics:8428http://vmselect:8481/select/0/prometheus (your tenant)
vmagent_urlseach vmagent, e.g. http://vmagent:8429same
vmalert_urlhttp://vmalert:8880same
component_urlseach vminsert, vmselect and vmstorage
alertmanager_urlhttp://alertmanager:9093same

On Kubernetes, list each vmstorage pod through its headless Service (vmstorage-0.vmstorage.monitoring.svc:8482, …) so 3AM can tell which node is in trouble, and restart only that one.

If you leave out the vmagents, 3AM can't see scrape targets or remote write. Leave out the cluster components and it can't see storage nodes. Checks that need them report inconclusive, never passed.

3. Add the connector

In the console: Settings → Add a connection → VictoriaMetrics. Or in connectors.json:

/etc/3am/connectors.json
{ "name": "vm", "kind": "victoriametrics",
  "settings": { "url": "http://vmselect.monitoring.svc:8481/select/0/prometheus",
                "vmagent_urls": ["http://vmagent.monitoring.svc:8429"],
                "vmalert_url": "http://vmalert.monitoring.svc:8880",
                "component_urls": ["http://vminsert.monitoring.svc:8480",
                                   "http://vmstorage-0.vmstorage.monitoring.svc:8482",
                                   "http://vmstorage-1.vmstorage.monitoring.svc:8482"],
                "alertmanager_url": "http://alertmanager.monitoring.svc:9093",
                "username": "3am", "password": { "file": "/etc/3am/secrets/vm-password" } } }
kind: victoriametricsalerts inmetrics and queriesactions
SettingWhat to enterRequiredDefault
urlQuery URL: single-node (http://victoriametrics:8428), vmselect with tenant (http://vmselect:8481/select/0/prometheus), or vmauthyes—
vmagent_urlsvmagent base URLs (scrape targets and remote write)no—
vmalert_urlvmalert base URL (rules and alerts)no—
component_urlsCluster components 3AM watches: vmstorage, vminsert, vmselect base URLsno—
alertmanager_urlAlertmanager base URL (routed alerts, silences)no—
bearer_tokensecretBearer token, if requiredno—
usernameBasic-auth user, if requiredno—
passwordsecretBasic-auth passwordno—
alertmanager_bearer_tokensecretAlertmanager bearer token (if different)no—
alertmanager_usernameAlertmanager basic-auth user (if different)no—
alertmanager_passwordsecretAlertmanager basic-auth password (if different)no—
ca_fileCA certificate for a private CA (PEM file)no—
verify_tlsVerify TLSyestrue

4. Verify

Test connection runs a query and reads every component's /metrics. Each component is named with its version:

Test connection
Connected. vmselect v1.153.0-cluster + vmagent v1.153.0 + vmalert v1.153.0 + vminsert v1.153.0-cluster
+ vmstorage v1.153.0-cluster + vmstorage v1.153.0-cluster + Alertmanager 0.34.1

For the whole path, load a rule that always fires into vmalert and reload it:

groups:
  - name: 3am-test
    rules:
      - alert: ThreeAMPipelineTest
        expr: vector(1)
        labels: { severity: critical, service: monitoring }
curl -X POST http://vmalert:8880/-/reload

Within a minute the console shows an incident for ThreeAMPipelineTest. Remove the rule afterwards.

What it proves and fixes

3AM raises these itself, even with no rule for them. If you run VictoriaMetrics' own alerting rules (RowsRejectedOnIngestion, TooManyRemoteWriteErrors, ConfigurationReloadFailure, …), 3AM uses yours and proves the cause behind them.

Fixed with approval (through the Kubernetes connector):

CauseHow 3AM knowsFix
Storage read-only (disk full)vm_storage_is_read_only: free disk below -storage.minFreeDiskSpaceBytesGrow that node's volume by half. Two approvers
Disk filling upOver 80% used, not read-only yetGrow the fullest node's volume. Two approvers
Component down or hungIt doesn't answerRestart its pod (StatefulSet) or its workload
Storage node unreachablevminsert's per-node reachability, vmselect's connection errorsRestart that vmstorage pod only
Scrape target unreachableThe scrape error, per jobRestart the target's workload

Handed to a person with the evidence (the right fix is a judgement call):

CauseHow 3AM knowsYour data
vmagent can't reach storageRetries and no answer at allKept in vmagent's queue
Storage refuses vmagent's credentials401 / 403Kept in vmagent's queue
Storage rejects vmagent's data400 / 409, blocks droppedLost
Storage fails writes5xx / 429, while storage is writable and every node reachableKept in vmagent's queue
Series limit reached*_series_limit_rows_dropped_total rising, and the metrics with the most seriesNew series lost
Rows rejectedvm_rows_ignored_total rising, with the reasonThose rows lost
High churnOver 10% of ingested rows create new series; the labels with the most values
Target refuses credentials, or answers garbageThe scrape error. A garbage target stays up in vmagent with 0 samples, so up == 0 never firesNo metrics from it
Config reload failed*_config_last_reload_successful = 0Runs on the last good config
Rules broken / datasource unreachablevmalert rule health: a rule error versus a connection errorAlerts can't fire
vmalert can't notifyvmalert_alerts_send_errors_total risingNobody is paged

A few things worth knowing about how it decides:

  • The fix targets the proven culprit. If vmagent's alert opens the incident but a storage node is read-only, 3AM grows that node's volume, not vmagent's.
  • Lost vs delayed. vmagent records each storage answer's exact status code, so 3AM can tell whether data is dropped or only queued.
  • No guessing. "Storage fails writes" is only concluded when storage is visible, writable and complete. Otherwise the read-only disk or the lost node is the cause.
  • Only what's happening now. Counters are compared with 3AM's earlier readings, so last week's dropped rows aren't today's incident.
  • Partial results are never evidence. With a storage node down, vmselect still answers, flagged partial; 3AM refuses those answers.

Actions

ActionParametersNotes
silencealertname, minutes, comment, matchers1–1440 minutes in Alertmanager. The rehearsal lists what it would silence
unsilencesilence_id
reloadcomponent (vmagent or vmalert), instancevmagent answers 200 even when it refuses a broken file, so 3AM checks the reload status afterwards
targets, storage_status, remote_write_status, ingestion_status, cardinality, components_status, vmalert_status, cluster_status, volume_usageRead-only evidence

Troubleshooting

VictoriaMetrics: authentication failed

Check the user and password, or the bearer token.

Alertmanager: authentication failed

Alertmanager has its own credentials. Fill in the alertmanager_* settings.

the certificate is not trusted: set ca_file

Set CA certificate to your internal CA, in PEM format.

answers, but partially: a storage node is unreachable

A vmstorage node is down. 3AM refuses partial results as evidence until it's back.

vmagent (…): HTTP 404

That URL isn't a vmagent, or a proxy hides /metrics. Give 3AM the vmagent's own address.

Certification

Live, on VictoriaMetrics v1.153.0 and the long-term-support v1.136.0, 97 checks on each (2026-10-04).

The estate:

  • single-node over HTTPS with an internal CA and basic auth;
  • a cluster (vminsert, vmselect, two vmstorage);
  • 7 vmagents and 2 vmalerts;
  • Alertmanager behind a bearer-token proxy, and the real mysqld_exporter.

One thing was broken per component. Targets:

  • refusing connections, rejecting credentials or answering garbage;
  • carrying too many labels, or churning series.

vmagents writing:

  • to a dead port, or with wrong credentials;
  • through a proxy that rejects the data;
  • into read-only storage, or past a series limit.

Alerting and storage:

  • a dead datasource, a broken rule, a dead notifier and a broken reload;
  • a storage node and a vmagent stopped mid-run.

Every cause was proven from live evidence. Read-only storage was never taken for failing storage, and partial answers were refused. On Kubernetes, hung vmagent, vmstorage and exporter pods were restarted and verified (details).

On this page