Metrics API
Endpoint
Section titled “Endpoint”GET /metricsAuthorization: Bearer <Z4J_METRICS_AUTH_TOKEN>Not a REST-versioned endpoint (Prometheus convention), and closed by default.
A fresh packaged SQLite install auto-mints and persists a
Z4J_METRICS_AUTH_TOKEN during z4j serve; every other deployment must
configure one explicitly. Scrapes must present it. If neither a token nor
Z4J_METRICS_PUBLIC=1 is configured, the endpoint returns 401. Setting
Z4J_METRICS_PUBLIC=1 deliberately opens it and emits a startup warning.
Z4J_METRICS_ENABLED=false removes the route entirely (requests return 404);
the token and public switches only control who may scrape an enabled endpoint.
Catalog
Section titled “Catalog”Ingest and lifecycle
Section titled “Ingest and lifecycle”z4j_events_ingested_total{project,engine,kind}(counter): wire-level event ingest count by project, engine adapter, and event kind.z4j_tasks_total{project,task_name,state}(counter): observed tasks by final state. High-cardinality task names are bounded by the server's overflow policy: the label is truncated to 128 characters, and once a project has 1000 distinct names every further new name is folded into__overflow__.z4j_task_duration_seconds{project,task_name}(histogram): task wall-clock duration buckets.z4j_commands_total{project,action,status}(counter): commands dispatched to agents.z4j_command_late_results_total{status}(counter): command results received after the row was already terminal.
Agents, workers, queues
Section titled “Agents, workers, queues”z4j_agents_online{project}(gauge): currently connected agent count.z4j_workers_online{project}(gauge): currently online worker count.z4j_queue_depth{project,queue,engine}(gauge): pending messages in a queue.z4j_ws_connections(gauge): registered as "live WebSocket connections held by this worker", but nothing in the gateway sets it, so it carries no signal.
Database
Section titled “Database”z4j_db_pool_size(gauge): configured SQLAlchemy pool size. Populated at scrape time. Rises only whenpool_sizesetting changes; useful baseline.z4j_db_pool_checked_out(gauge): pool connections currently checked out (in active use). Steady-state under burst load is a key contention indicator; should be much less thanpool_size.z4j_brain_rss_bytes(gauge): brain process RSS in bytes, sampled at scrape time from/proc/self/status.0on non-Linux. The slope under sustained load is the headline leak signal; flat or slow growth is healthy, rapid growth indicates either an unbounded cache (tuneZ4J_DATABASE_STATEMENT_CACHE_SIZE) or a new retention path. See Brain memory tuning for the playbook.z4j_postgres_deadlocks_total(counter): total PostgresDeadlockDetectedErrorinstances observed via the asyncpg/SQLAlchemyhandle_errorevent listener. Should hover at or near zero in steady state; sustained non-zero rates onINSERT INTO workers/UPDATE agents/UPDATE queuespoint at lock-order contention.
Notifications
Section titled “Notifications”z4j_notifications_sent_total{project,channel_type,status}(counter): notification deliveries attempted.z4j_notifications_cooldown_skipped_total{project,trigger}(counter): dispatches skipped because the cooldown window had not elapsed.
Automation and scheduler reliability
Section titled “Automation and scheduler reliability”z4j_scheduler_misfires_detected_total{project}(counter): schedules the brain detected as misfired (expected fire past the grace window). A sustained non-zero rate means a scheduler is dead, partitioned, or badly behind.z4j_agents_offline_detected_total{project}(counter): agent-offline episodes the brain confirmed (no heartbeat past the offline timeout plus alert grace). Each episode counts once, however long the agent stays down.z4j_automation_rule_fires_total{project,action,outcome}(counter): automation rule action executions by action type and outcome, including dry-run and failsafe-skipped ones. Labelled by project, not by rule; per-rule detail is in theautomation.rule.firedaudit rows.z4j_automation_circuit_trips_total{project}(counter): rule circuit-breaker trips into notify-only failsafe mode.z4j_automation_firings_dropped_total{project,reason}(counter): firings dropped before dispatch because the pending queue was full. A drop is permanent; a sustained rate means automation is shedding load.z4j_automation_notify_coalesced_total{project}(counter): notify actions suppressed by the per-rule coalesce window; the first alert in each window still goes out.z4j_automation_outbox_enqueued_total{project,trigger}(counter): firings persisted to the durable outbox for later replay. Pairs with the dropped counter: enqueued is recoverable, dropped is lost.z4j_schedule_fires_partition_failures_total{op,reason}(counter):schedule_firespartition create/drop operations that failed.reason="default_blocked"means a day's rows sit in the DEFAULT partition, so its daily partition cannot be created and retention-by-DROP is blocked until DEFAULT is cleared.
In-memory state
Section titled “In-memory state”z4j_inmemory_state_items{subsystem}(gauge): registered process-local state, currentlylongpoll_sessions,activity_rate_limit_users,otel_dynamic_sensitive_hostsandmetric_task_name_projects. It is refreshed by the worker handling the scrape; multiprocess aggregation sums live workers, whose other samples may be one scrape behind.
Reliability and self-watch
Section titled “Reliability and self-watch”z4j_swallowed_exceptions_total{module,site}(counter): intentional exception swallows at I/O boundaries (metric updates, WebSocket close during shutdown, asyncpg teardown). A sustained non-zero rate signals a subsystem in trouble even when no error-level log fires.z4j_background_task_error_active{task}(gauge): coversaudit_retentionandwal_checkpoint; 1 means that worker's most recent pass failed, and 0 means a later pass succeeded. In multiprocess mode the maximum is retained, including a failed worker that later died, until the serve run restarts.
Audit retention and WAL
Section titled “Audit retention and WAL”z4j_audit_retention_pruned_total(gauge despite the_totalsuffix): cumulative audit rows pruned during the current serve run; it resets on restart. Do not apply counter-only functions such asrate()orincrease().z4j_audit_retention_last_run_timestamp(gauge): Unix timestamp of the most recent retention pass.z4j_audit_retention_last_deleted(gauge): rows deleted by the worker whose pass was most recently published. In a multiworker deployment this is not a brain-wide deletion total and can be zero even when another worker pruned rows; usez4j_audit_retention_pruned_totalto confirm cumulative work for the current serve run.z4j_audit_chain_verifications_total{outcome}(counter): scheduled audit-chain verification runs by outcome.clean,failed(the chain did not verify and needs investigation), orerror(the run could not complete, which is a different problem and must not be read as evidence of tampering). Only emitted whileZ4J_AUDIT_CHAIN_VERIFY_ENABLEDis on.z4j_audit_chain_rows_verified(gauge): rows walked by the most recent successful or integrity-failed verification that walked at least one row. An error or zero-row run leaves the prior value in place, so pair it withz4j_audit_chain_verifications_total{outcome=...}.z4j_audit_forward_lag_rows(gauge): audit rows written but not yet acknowledged by the audit webhook receiver, read from the database at the start of every forwarder pass. Only emitted whileZ4J_AUDIT_WEBHOOK_URLis set; the leader is the only process that sets it.z4j_audit_forward_failures_total{reason}(counter): audit webhook attempts that did not get a 2xx, bynon_2xx,post_raised,ssrf_or_dnsorcursor_conflict. Each is a retry from the durable cursor, not a drop. See audit webhook forwarding.z4j_auth_ip_denied_total{surface}(counter): requests refused by a source-address allowlist, by surface (dashboard,api,agent). Every denial also writes anauth.ip_deniedaudit row, so a count that climbs while the rows do not means the audit write is failing, not that the denials stopped. See threat model.z4j_wal_checkpoint_last_run_timestamp(gauge): Unix timestamp of the most recent SQLite WAL checkpoint (0on Postgres).z4j_wal_checkpoint_pages_last(gauge): pages checkpointed in the most recent successful SQLite WAL pass;-1means no useful result yet and is also the permanent value on PostgreSQL.
Grafana dashboard
Section titled “Grafana dashboard”Shareable Grafana dashboards live in the repository checkout's
deploy/grafana/ directory (overview, tasks, agents, scheduler,
notifications). They are not included in the installed wheel or the z4j
package sdist (the sdist's deploy/ is the package's own directory, which
holds the Caddyfile, the Helm chart, the Kubernetes manifests and the systemd
units, not these dashboards). Import in Grafana via Dashboards,
New, Import, then upload the
JSON. The overview dashboard covers:
- Health stats: agents online, brain RSS, DB pool utilisation, deadlocks per minute.
- Task throughput and outcome: tasks per minute by final state, task duration p50 / p95 / p99.
- Agents and queues: agents online by project, queue depth.
- Background tasks (self-watch): background task error flags, swallowed exceptions per minute by module.
Configure your Prometheus datasource to scrape the brain's /metrics with the Z4J_METRICS_AUTH_TOKEN bearer:
scrape_configs: - job_name: z4j metrics_path: /metrics authorization: type: Bearer credentials: $Z4J_METRICS_AUTH_TOKEN static_configs: - targets: ["z4j.internal:7700"]