diff --git a/content/docs/server/meta.json b/content/docs/server/meta.json index a80c2e9d80..e95dd70277 100644 --- a/content/docs/server/meta.json +++ b/content/docs/server/meta.json @@ -1,4 +1,4 @@ { "title": "Server", - "pages": ["introduction", "configuration", "topic-options", "storage-engine", "durability", "networking", "security", "docker", "linux-tuning", "benchmarking"] + "pages": ["introduction", "configuration", "topic-options", "storage-engine", "durability", "networking", "security", "monitoring", "docker", "linux-tuning", "benchmarking"] } diff --git a/content/docs/server/monitoring.mdx b/content/docs/server/monitoring.mdx new file mode 100644 index 0000000000..8c7e79e0e1 --- /dev/null +++ b/content/docs/server/monitoring.mdx @@ -0,0 +1,51 @@ +--- +title: Monitoring +description: "Scrape the server's Prometheus metrics: the token the endpoint needs, what it exports, and which metrics are worth alerting on." +--- + +The server exports Prometheus metrics on its HTTP API at `/metrics`. Both the HTTP API and the metrics endpoint are on by default. See [`[http.metrics]`](/docs/server/configuration#httpmetrics) to change the path or turn it off. + +This page covers the server's own metrics. The connectors runtime has a separate endpoint with its own metric names; see [Connectors observability](/docs/connectors/observability). + +## Get a token for the scraper + +`/metrics` needs a bearer token, like every other read on the HTTP API. Any logged-in user can scrape it; there's no separate monitoring permission. Prometheus can't log in, so create a personal access token and put it in the scrape config: + +```bash +iggy pat create prometheus 90d +``` + +The token is printed once. The Prometheus scrape config is in [`[http.metrics]`](/docs/server/configuration#httpmetrics). + +The token can read every metric the server exports, so store it like any other secret: in a Kubernetes Secret or your secrets manager, not in a file you commit. With the Helm chart, set it through `server.serviceMonitor.authorization`. See [Docker & Helm](/docs/server/docker). + +## What's exported + +The core metrics are below. The gauges are read fresh on every scrape. + +| Metric | Type | What it counts | +|--------|------|----------------| +| `http_requests_total` | counter | HTTP API requests | +| `streams`, `topics`, `partitions`, `segments` | gauge | Current number of each | +| `messages` | gauge | Messages stored | +| `users` | gauge | Users | +| `clients` | gauge | Connected clients | + +The names have no `iggy_` prefix, so watch for clashes if Prometheus scrapes other exporters with similar names. + +Each shard also exports its own lower-level metrics, labelled `shard=""`. These cover the write-ahead log (the `partition_wal_*` series), partition housekeeping (`partitions_*`), and requests or frames the server refused or dropped. To see the full list with descriptions, scrape a running server: + +```bash +curl -s -H "Authorization: Bearer " http://localhost:3000/metrics | grep '^# HELP' +``` + +## What to alert on + +- **`frame_drops_total{variant="forward_client_send"}`.** Any increase means a client never got a reply to a request. Most other `frame_drops_total` variants are retried or repaired automatically. This series only appears after the first drop, so on a healthy server it won't be in the scrape at all. Write the alert so that a missing series counts as no drops. +- **`stats_rollup_underflows_total`.** It goes up when a stream, topic or partition total would have dropped below zero and was held at zero instead. A few during deletes, purges or restores are expected. A steady rise at other times is worth a look. + +## Grafana and OpenTelemetry + +Add Prometheus as a data source in Grafana, then build panels from the metric names above. + +The server doesn't export metrics over OTLP. To bring them into an OpenTelemetry pipeline, scrape `/metrics` with the collector's Prometheus receiver, using the same bearer token. Server export of logs and traces over OTLP isn't available yet; see [`[telemetry]`](/docs/server/configuration#telemetry).