Skip to content
SupaCovedocs

12 / 13

Monitoring & metrics

/metrics semantics, dashboard denominators, alerting suggestions

Access

GET /metrics requires an authenticated session (Prometheus scrapes must carry session credentials) and answers with Cache-Control: no-store. Unauthenticated requests receive a 401 JSON error.

Families

All families are also served under their pre-rename supabackup_* names as equal-valued deprecated aliases during the transition (removed two tagged releases after the rename); switch scrapers to supacove_*. Exception: the pre-Phase-7 shortcut gauges supabackup_jobs_succeeded / supabackup_jobs_failed exist under their legacy names only — query supacove_jobs{status="succeeded"} / {status="failed"} instead.

MetricSemantics
supacove_jobs{status=…}current distribution of task states (gauge, not a monotonic counter; the historical _total-suffixed families were renamed)
supacove_verification{status=…}verification-state distribution over succeeded jobs
supacove_last_success_timestamp{database=…}newest success per database (unix seconds)
supacove_databases_protection{state=…}protection distribution (fresh/expired/never)
supacove_remote_commits / supacove_remote_upload_failurescurrent counts of remote-committed (incl. deleted-after-commit) and upload-failed jobs — gauges queried from the jobs table, NOT standalone counters; never apply rate()
supacove_outbox_pending / supacove_outbox_deadnotification queue health (in-flight count has no metric — use the console delivery log)
supacove_staging_bytesstaging usage (a collector failure still emits a best-effort value AND sets scrape_errors)
supacove_uptime_seconds / supacove_go_goroutines / supacove_heap_alloc_bytesprocess health: uptime, goroutine count, heap allocation
supacove_scrape_errors{collector=…}collector-fault signal: nonzero means some families of that scrape may be missing

Labels carry only low-cardinality values (states, database names, collector names) — never job IDs, hosts or credentials.

Console statistics denominators

  • Export success rate: over jobs that began executing, those whose export completed (including "export fine, upload failed"); jobs canceled while queued are excluded on both sides.
  • Archive total: a subtotal over RECORDED samples (backup_stats rows); historical rows may hold ciphertext sizes (old semantics), so the sum can be slightly high on upgraded instances.
  • Average duration: export-through-remote-commit wall time over succeeded jobs, positive samples only; no samples renders —, never a fake 0.
  • The export rate denominator covers terminal jobs only; running tasks do not participate.

Alerting suggestions

  1. time() - supacove_last_success_timestamp older than your configured freshness threshold (the timestamp is the job's started_at, an approximation of the success snapshot) → backups stalled. Databases that never succeeded are absent from this family; cover them with supacove_databases_protection{state="never"}.
  2. supacove_outbox_dead > 0 → an alert never reached its webhook.
  3. supacove_staging_bytes growing → destination outage or a long reclamation grace.
  4. Dead-man switch silence (external) → process death or fleet-wide failure.

Last updated

On this page