Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion content/docs/server/meta.json
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
{
"title": "Server",
"pages": ["introduction", "configuration", "topic-options", "storage-engine", "durability", "networking", "security", "docker", "linux-tuning", "benchmarking"]
"pages": ["introduction", "configuration", "topic-options", "storage-engine", "durability", "networking", "security", "monitoring", "docker", "linux-tuning", "benchmarking"]
}
51 changes: 51 additions & 0 deletions content/docs/server/monitoring.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
---
title: Monitoring
description: "Scrape the server's Prometheus metrics: the token the endpoint needs, what it exports, and which metrics are worth alerting on."
---

The server exports Prometheus metrics on its HTTP API at `/metrics`. Both the HTTP API and the metrics endpoint are on by default. See [`[http.metrics]`](/docs/server/configuration#httpmetrics) to change the path or turn it off.

This page covers the server's own metrics. The connectors runtime has a separate endpoint with its own metric names; see [Connectors observability](/docs/connectors/observability).

## Get a token for the scraper

`/metrics` needs a bearer token, like every other read on the HTTP API. Any logged-in user can scrape it; there's no separate monitoring permission. Prometheus can't log in, so create a personal access token and put it in the scrape config:

```bash
iggy pat create prometheus 90d
```

The token is printed once. The Prometheus scrape config is in [`[http.metrics]`](/docs/server/configuration#httpmetrics).

The token can read every metric the server exports, so store it like any other secret: in a Kubernetes Secret or your secrets manager, not in a file you commit. With the Helm chart, set it through `server.serviceMonitor.authorization`. See [Docker & Helm](/docs/server/docker).

## What's exported

The core metrics are below. The gauges are read fresh on every scrape.

| Metric | Type | What it counts |
|--------|------|----------------|
| `http_requests_total` | counter | HTTP API requests |
| `streams`, `topics`, `partitions`, `segments` | gauge | Current number of each |
| `messages` | gauge | Messages stored |
| `users` | gauge | Users |
| `clients` | gauge | Connected clients |

The names have no `iggy_` prefix, so watch for clashes if Prometheus scrapes other exporters with similar names.

Each shard also exports its own lower-level metrics, labelled `shard="<id>"`. These cover the write-ahead log (the `partition_wal_*` series), partition housekeeping (`partitions_*`), and requests or frames the server refused or dropped. To see the full list with descriptions, scrape a running server:

```bash
curl -s -H "Authorization: Bearer <token>" http://localhost:3000/metrics | grep '^# HELP'
```

## What to alert on

- **`frame_drops_total{variant="forward_client_send"}`.** Any increase means a client never got a reply to a request. Most other `frame_drops_total` variants are retried or repaired automatically. This series only appears after the first drop, so on a healthy server it won't be in the scrape at all. Write the alert so that a missing series counts as no drops.
- **`stats_rollup_underflows_total`.** It goes up when a stream, topic or partition total would have dropped below zero and was held at zero instead. A few during deletes, purges or restores are expected. A steady rise at other times is worth a look.

## Grafana and OpenTelemetry

Add Prometheus as a data source in Grafana, then build panels from the metric names above.

The server doesn't export metrics over OTLP. To bring them into an OpenTelemetry pipeline, scrape `/metrics` with the collector's Prometheus receiver, using the same bearer token. Server export of logs and traces over OTLP isn't available yet; see [`[telemetry]`](/docs/server/configuration#telemetry).
Loading