Skip to content

Metrics API

GET /metrics
Authorization: Bearer <Z4J_METRICS_AUTH_TOKEN>

Not a REST-versioned endpoint (Prometheus convention), and closed by default. A fresh packaged SQLite install auto-mints and persists a Z4J_METRICS_AUTH_TOKEN during z4j serve; every other deployment must configure one explicitly. Scrapes must present it. If neither a token nor Z4J_METRICS_PUBLIC=1 is configured, the endpoint returns 401. Setting Z4J_METRICS_PUBLIC=1 deliberately opens it and emits a startup warning. Z4J_METRICS_ENABLED=false removes the route entirely (requests return 404); the token and public switches only control who may scrape an enabled endpoint.

  • z4j_events_ingested_total{project,engine,kind} (counter): wire-level event ingest count by project, engine adapter, and event kind.
  • z4j_tasks_total{project,task_name,state} (counter): observed tasks by final state. High-cardinality task names are bounded by the server's overflow policy: the label is truncated to 128 characters, and once a project has 1000 distinct names every further new name is folded into __overflow__.
  • z4j_task_duration_seconds{project,task_name} (histogram): task wall-clock duration buckets.
  • z4j_commands_total{project,action,status} (counter): commands dispatched to agents.
  • z4j_command_late_results_total{status} (counter): command results received after the row was already terminal.
  • z4j_agents_online{project} (gauge): currently connected agent count.
  • z4j_workers_online{project} (gauge): currently online worker count.
  • z4j_queue_depth{project,queue,engine} (gauge): pending messages in a queue.
  • z4j_ws_connections (gauge): registered as "live WebSocket connections held by this worker", but nothing in the gateway sets it, so it carries no signal.
  • z4j_db_pool_size (gauge): configured SQLAlchemy pool size. Populated at scrape time. Rises only when pool_size setting changes; useful baseline.
  • z4j_db_pool_checked_out (gauge): pool connections currently checked out (in active use). Steady-state under burst load is a key contention indicator; should be much less than pool_size.
  • z4j_brain_rss_bytes (gauge): brain process RSS in bytes, sampled at scrape time from /proc/self/status. 0 on non-Linux. The slope under sustained load is the headline leak signal; flat or slow growth is healthy, rapid growth indicates either an unbounded cache (tune Z4J_DATABASE_STATEMENT_CACHE_SIZE) or a new retention path. See Brain memory tuning for the playbook.
  • z4j_postgres_deadlocks_total (counter): total Postgres DeadlockDetectedError instances observed via the asyncpg/SQLAlchemy handle_error event listener. Should hover at or near zero in steady state; sustained non-zero rates on INSERT INTO workers / UPDATE agents / UPDATE queues point at lock-order contention.
  • z4j_notifications_sent_total{project,channel_type,status} (counter): notification deliveries attempted.
  • z4j_notifications_cooldown_skipped_total{project,trigger} (counter): dispatches skipped because the cooldown window had not elapsed.
  • z4j_scheduler_misfires_detected_total{project} (counter): schedules the brain detected as misfired (expected fire past the grace window). A sustained non-zero rate means a scheduler is dead, partitioned, or badly behind.
  • z4j_agents_offline_detected_total{project} (counter): agent-offline episodes the brain confirmed (no heartbeat past the offline timeout plus alert grace). Each episode counts once, however long the agent stays down.
  • z4j_automation_rule_fires_total{project,action,outcome} (counter): automation rule action executions by action type and outcome, including dry-run and failsafe-skipped ones. Labelled by project, not by rule; per-rule detail is in the automation.rule.fired audit rows.
  • z4j_automation_circuit_trips_total{project} (counter): rule circuit-breaker trips into notify-only failsafe mode.
  • z4j_automation_firings_dropped_total{project,reason} (counter): firings dropped before dispatch because the pending queue was full. A drop is permanent; a sustained rate means automation is shedding load.
  • z4j_automation_notify_coalesced_total{project} (counter): notify actions suppressed by the per-rule coalesce window; the first alert in each window still goes out.
  • z4j_automation_outbox_enqueued_total{project,trigger} (counter): firings persisted to the durable outbox for later replay. Pairs with the dropped counter: enqueued is recoverable, dropped is lost.
  • z4j_schedule_fires_partition_failures_total{op,reason} (counter): schedule_fires partition create/drop operations that failed. reason="default_blocked" means a day's rows sit in the DEFAULT partition, so its daily partition cannot be created and retention-by-DROP is blocked until DEFAULT is cleared.
  • z4j_inmemory_state_items{subsystem} (gauge): registered process-local state, currently longpoll_sessions, activity_rate_limit_users, otel_dynamic_sensitive_hosts and metric_task_name_projects. It is refreshed by the worker handling the scrape; multiprocess aggregation sums live workers, whose other samples may be one scrape behind.
  • z4j_swallowed_exceptions_total{module,site} (counter): intentional exception swallows at I/O boundaries (metric updates, WebSocket close during shutdown, asyncpg teardown). A sustained non-zero rate signals a subsystem in trouble even when no error-level log fires.
  • z4j_background_task_error_active{task} (gauge): covers audit_retention and wal_checkpoint; 1 means that worker's most recent pass failed, and 0 means a later pass succeeded. In multiprocess mode the maximum is retained, including a failed worker that later died, until the serve run restarts.
  • z4j_audit_retention_pruned_total (gauge despite the _total suffix): cumulative audit rows pruned during the current serve run; it resets on restart. Do not apply counter-only functions such as rate() or increase().
  • z4j_audit_retention_last_run_timestamp (gauge): Unix timestamp of the most recent retention pass.
  • z4j_audit_retention_last_deleted (gauge): rows deleted by the worker whose pass was most recently published. In a multiworker deployment this is not a brain-wide deletion total and can be zero even when another worker pruned rows; use z4j_audit_retention_pruned_total to confirm cumulative work for the current serve run.
  • z4j_audit_chain_verifications_total{outcome} (counter): scheduled audit-chain verification runs by outcome. clean, failed (the chain did not verify and needs investigation), or error (the run could not complete, which is a different problem and must not be read as evidence of tampering). Only emitted while Z4J_AUDIT_CHAIN_VERIFY_ENABLED is on.
  • z4j_audit_chain_rows_verified (gauge): rows walked by the most recent successful or integrity-failed verification that walked at least one row. An error or zero-row run leaves the prior value in place, so pair it with z4j_audit_chain_verifications_total{outcome=...}.
  • z4j_audit_forward_lag_rows (gauge): audit rows written but not yet acknowledged by the audit webhook receiver, read from the database at the start of every forwarder pass. Only emitted while Z4J_AUDIT_WEBHOOK_URL is set; the leader is the only process that sets it.
  • z4j_audit_forward_failures_total{reason} (counter): audit webhook attempts that did not get a 2xx, by non_2xx, post_raised, ssrf_or_dns or cursor_conflict. Each is a retry from the durable cursor, not a drop. See audit webhook forwarding.
  • z4j_auth_ip_denied_total{surface} (counter): requests refused by a source-address allowlist, by surface (dashboard, api, agent). Every denial also writes an auth.ip_denied audit row, so a count that climbs while the rows do not means the audit write is failing, not that the denials stopped. See threat model.
  • z4j_wal_checkpoint_last_run_timestamp (gauge): Unix timestamp of the most recent SQLite WAL checkpoint (0 on Postgres).
  • z4j_wal_checkpoint_pages_last (gauge): pages checkpointed in the most recent successful SQLite WAL pass; -1 means no useful result yet and is also the permanent value on PostgreSQL.

Shareable Grafana dashboards live in the repository checkout's deploy/grafana/ directory (overview, tasks, agents, scheduler, notifications). They are not included in the installed wheel or the z4j package sdist (the sdist's deploy/ is the package's own directory, which holds the Caddyfile, the Helm chart, the Kubernetes manifests and the systemd units, not these dashboards). Import in Grafana via Dashboards, New, Import, then upload the JSON. The overview dashboard covers:

  1. Health stats: agents online, brain RSS, DB pool utilisation, deadlocks per minute.
  2. Task throughput and outcome: tasks per minute by final state, task duration p50 / p95 / p99.
  3. Agents and queues: agents online by project, queue depth.
  4. Background tasks (self-watch): background task error flags, swallowed exceptions per minute by module.

Configure your Prometheus datasource to scrape the brain's /metrics with the Z4J_METRICS_AUTH_TOKEN bearer:

prometheus.yml
scrape_configs:
- job_name: z4j
metrics_path: /metrics
authorization:
type: Bearer
credentials: $Z4J_METRICS_AUTH_TOKEN
static_configs:
- targets: ["z4j.internal:7700"]