Dashboards get a new schema — and Prometheus parity

Alongside the Prometheus exporter source covered in the v6 announcement, pgwatch's bundled dashboards were rebuilt from scratch on Grafana's v13 schema (dashboard.grafana.app/v2). The visible change is a unified row layout and a shared "colophon" footer; the more consequential one is that Prometheus sources now reach first-class dashboard parity with Postgres instead of trailing behind it.

Four new Prometheus-side dashboards ship in this beta:

  • Lock Details, mirroring the Postgres v13 "Locks (PG19+)" view: lock waits per second by type, average wait time per acquisition, fast-path lock pool exceeded counts, and time elapsed since the last pg_stat_lock reset.
  • Server Log Events, breaking events down by database and by instance, with "Top ERROR generating DBs" and "Top FATAL generating instances" panels linking into the per-database overview.
  • Postgres Version Overview, a single "Monitored DBs by version" panel driven by the settings metric.
  • Stored Procedures, matching the existing Postgres sproc views.

The Patroni Cluster Overview is new as well, fed by the patroni Prometheus preset. It adds per-cluster health summaries, node role and leader-lock panels, DCS-last-seen age, and WAL write, received and replayed locations. Two panels in particular — "Clusters Without Leader" and "Paused Clusters" — reduce split-brain exposure to a number that should stay at zero.

The reaper: fixing what a real outage exposed

The second half of the release comes out of a production incident. Roughly 400 sources stopped collecting entirely once or twice a month and only recovered after a full restart. The trigger was a shared-infrastructure brownout affecting DNS and the network, but pgwatch amplified the damage in three ways that are now closed.

  1. Deadlines on every database round-trip. A metric fetch, a Ping or a discovery query could previously stall on a half-open TCP connection until the kernel timed it out — 15 to 30 minutes on Linux. Each of those calls now runs under a bounded context.
  2. A parallel sweep behind a fixed ceiling. A single hanging source used to serialize every source queued behind it in the per-refresh sweep; at 400 sources and 10-20 seconds of failure cost each, one sweep could stretch to an hour or two. Collection now runs through a bounded goroutine group capped at 32 sources in flight. The cap is a fixed constant, deliberately not configurable and independent of fleet size, so a brownout cannot escalate into a reconnection storm. During an outage, a 400-source sweep costs about [400/32] multiplied by the per-source timeout rather than 400 multiplied by it.
  3. Discovery remembers its last good answer. Sources using postgres-continuous-discovery periodically re-run a discovery query to pick up newly created databases. A transient failure — a DNS hiccup, a brief permission error — used to tear down monitoring for every already-discovered database and rebuild it on the next successful resolution. Now a failed resolution with a populated cache logs a warning and keeps serving the last known-good list; teardown happens only when no successful resolution has ever been cached. This is the same caching approach pgwatch already applied to Patroni cluster membership.

Configuration is unchanged. What changes is how monitoring behaves when the network underneath it falters.

Other changes in v6.0.0-beta

  • Four new metrics.
  • JWT secret hardening: Web UI and API tokens are now signed with a random 32-byte secret generated once per process at startup, replacing a hardcoded key compiled into the source. Tokens forged with the old key are rejected.
  • Fixes for a SourceReaper deadlock, a sink outage that previously took collection down rather than degrading gracefully, NaN/Inf values reaching stored measurements unsanitized, and a broken v12 "Replication lag" panel.
  • WebUI updated to React 19 and MUI v9.
  • PostgreSQL 19 joins the supported targets, plus a Grafana provisioning fix for Grafana 12 and later.
  • Task build tool support for contributors: task build, task test, task lint and related targets, cross-platform.

Beta caveats

The Prometheus source and the v13 dashboard schema are both new ground. Test them outside production first, and report problems through the project's issue tracker. The complete changelog, including dependency bumps not listed above, is on the v6.0.0-beta release page.

A token forged with the old hardcoded JWT key is rejected outright.