Skip to content

Dashboard data

These routes feed the dashboard's Home, Overview, Trends, Queues, Workers, Issues and task-detail pages. They are read-only except the personal saved views. Every path below is under /api/v1.

Every route except the two public health probes and the first-boot setup surface resolves the caller the same way: a valid session cookie wins when both a cookie and an Authorization: Bearer z4k_... header are present; otherwise the API key is checked strictly, and a revoked, expired or inactive-user key is refused with 401. Cookie sessions also pass the MFA enrollment and re-verification gates described in MFA enforcement.

Project-scoped routes resolve {slug} first and then require membership at viewer or above. A global admin passes without a membership row. An unknown slug, an archived project and a project the caller is not a member of all return the same 404, so slugs cannot be enumerated.

API keys are authorized by the route's tag. The home, stats, events, queues and workers routes require home:read, tasks:read, tasks:read, queues:read and workers:read. The trends, issues and saved-views routes have no scope mapping, so every API key receives 403 on them (admin:* included); they are session-cookie only. A project-bound key may call the home routes (the results are filtered to its bound project) but is refused with 403 on /health/system and /health/deep, because those paths carry no project slug.

Status Meaning
401 No valid session cookie or API key.
403 The API key lacks the scope, is bound to another project, or the route has no scope mapping; a cookie session is blocked by an MFA gate; or the CSRF check failed on a saved-view write.
404 Unknown or archived project, non-member caller, unknown worker, or unknown saved view.
409 Saved-view name already used in the project, or the cap of 100 personal views reached.
422 Query or body validation failure.
GET /api/v1/home/summary

Any authenticated user. Returns HomeSummaryPublic:

{
"user": {"id": "...", "email": "...", "display_name": null, "is_admin": false},
"aggregate": {
"tasks_24h": 0, "failures_24h": 0, "failure_rate_24h": 0.0,
"workers_online": 0, "workers_total": 0,
"agents_online": 0, "agents_total": 0, "stuck_commands": 0
},
"projects": [
{
"id": "...", "slug": "billing", "name": "Billing", "environment": "production",
"role": "operator",
"tasks_24h": 0, "failures_24h": 0, "failure_rate_24h": 0.0,
"workers_online": 0, "workers_total": 0,
"agents_online": 0, "agents_total": 0, "stuck_commands": 0,
"last_activity_at": null,
"health": "healthy"
}
],
"attention": [
{"kind": "agent_offline", "severity": "critical", "project_id": "...",
"project_slug": "billing", "project_name": "Billing",
"message": "Agent offline", "count": 1}
]
}

Visibility: a non-admin sees the active projects they hold a membership on. A global admin sees the first Z4J_ADMIN_PROJECT_LIST_CAP projects (default 500), active ones only, with role null where they hold no membership row; this is a bounded first page, not an exhaustive tenant-wide list. A project-bound API key sees only its bound project. Archived projects never appear. With nothing visible the response carries zeros and empty lists.

Field Source
tasks_24h, failures_24h Count of task.received and task.failed events in the last 24 hours.
failure_rate_24h failures_24h / tasks_24h, capped at 1.0; 1.0 when there are failures but no received events; 0.0 when there are no failures.
workers_online, workers_total Worker rows in state online, and all worker rows.
agents_online, agents_total Unrevoked agent rows in state online, and all unrevoked agent rows.
stuck_commands Commands in state dispatched whose timeout_at is in the past.
last_activity_at Latest event occurred_at within the last 30 days; null when the project has no event in that window.
aggregate Sums of the card counters over the visible projects, with the same bounded failure rate.

health is decided in order: offline when at least one agent exists and none is online; degraded when the failure rate is above 5%, or any command is stuck, or some but not all agents are offline; idle when there were no tasks in 24 hours and no worker is online; otherwise healthy.

Attention items, one per applicable rule per project:

kind When severity count
agent_offline Some agents are offline. critical when none is online, otherwise warning. Offline agents.
high_failure_rate Failures with zero received events, or more than 20 tasks with a failure rate above 5%. critical for the zero-receive case and for a rate above 20%, otherwise warning. failures_24h.
stuck_commands At least one stuck command. warning. Stuck commands.

Attention items are sorted critical first, then warning, then by project name. Cards are sorted by project name.

GET /api/v1/home/recent-failures?limit=50&cursor=...

Any authenticated user; the same visibility rules as the summary. limit is 1 to 200 (default 50). Returns the most recent task.failed events across the visible projects:

{
"items": [
{
"id": "...", "occurred_at": "...",
"project_id": "...", "project_slug": "billing", "project_name": "Billing",
"engine": "celery", "task_id": "...",
"task_name": "billing.charge", "worker": "celery@web-01",
"exception": "...", "priority": "normal"
}
],
"next_cursor": "2026-01-01T00:00:00+00:00|<uuid hex>"
}

Pagination is keyset on (occurred_at, id); the cursor is <iso8601>|<uuid hex> and next_cursor is null on the last page. An unparseable cursor is treated as no cursor and the feed starts from the top. task_name, worker and exception are read from the event payload; exception is truncated to 500 characters and priority defaults to normal when the payload carries none.

GET /api/v1/projects/{slug}/stats?hours=24

Role: viewer. hours selects the window for the rate-based counters and accepts 1, 6, 24, 72 or 168; any other integer silently falls back to 24, and window_hours reports the effective value. The _24h field names are kept for compatibility and follow the selected window. Returns StatsResponse:

{
"window_hours": 24,
"tasks_by_state": {"pending": 0, "received": 0, "started": 0, "success": 0,
"failure": 0, "retry": 0, "revoked": 0, "rejected": 0, "unknown": 0},
"tasks_total": 0,
"tasks_failed_24h": 0, "tasks_succeeded_24h": 0, "failure_rate_24h": 0.0,
"agents_online": 0, "agents_offline": 0,
"workers_online": 0, "workers_offline": 0,
"commands_pending": 0, "commands_completed_24h": 0,
"commands_failed_24h": 0, "commands_timeout_24h": 0,
"queue_depths": [
{"name": "default", "engine": "celery", "pending_count": 0,
"broker_type": "redis", "last_seen_at": "..."}
],
"system_health": {"status": "healthy", "agents_all_online": true,
"queue_depth_ok": true, "failure_rate_ok": true, "brain_db_ok": true}
}

tasks_by_state and tasks_total count every task in the project regardless of the window. tasks_failed_24h and tasks_succeeded_24h count tasks whose finished_at falls inside the window; failure_rate_24h is failed / (failed + succeeded) over that window and 0.0 when nothing finished. Agent counters exclude revoked agents. commands_pending counts commands in pending or dispatched; the three command outcome counters use completed_at inside the window. queue_depths lists every queue row with its reported pending_count.

system_health: agents_all_online is true when no agent is offline and at least one is online; queue_depth_ok when the summed pending_count is below 10000; failure_rate_ok when the rate is below 0.1; brain_db_ok is always true. status is degraded when any of the first three is false and critical when no agent is online.

GET /api/v1/projects/{slug}/trends?window=24h&bucket=1h

Role: viewer. window is one of 1h, 6h, 24h, 72h, 7d (default 24h); bucket is one of 1m, 5m, 15m, 1h, 1d (default 1h). Any other value is 422. A pair that would produce more than 500 buckets (for example 7d with 1m) returns 400; that check runs before the project is resolved. Returns TrendsResponse:

{
"window": "24h",
"bucket": "1h",
"series": [
{"t": "2026-04-15T10:00:00Z", "success": 120, "failure": 2, "retry": 0,
"revoked": 0, "total": 122, "avg_runtime_ms": 843}
]
}

Only tasks with a non-null finished_at inside the window and a state of success, failure, retry or revoked are counted; in-flight tasks are excluded. series holds one entry per bucket that has at least one such task, ascending by t, which is the UTC bucket start. avg_runtime_ms is the integer mean of the non-null runtime_ms values in the bucket, or null.

GET /api/v1/projects/{slug}/queues

Role: viewer. Returns a bare list of QueuePublic, no parameters and no pagination:

[
{"id": "...", "project_id": "...", "name": "default", "engine": "celery",
"broker_type": "redis", "broker_url_hint": "...",
"last_seen_at": "...", "created_at": "..."}
]

The queue depth is not on this shape; read queue_depths on the stats route.

GET /api/v1/projects/{slug}/workers

Role: viewer. Returns a bare list of WorkerPublic:

[
{
"id": "...", "project_id": "...", "engine": "celery", "name": "celery@web-01",
"hostname": "web-01", "pid": 4242, "concurrency": 8, "queues": ["default"],
"state": "online", "last_heartbeat": "...", "load_average": [0.1, 0.2, 0.3],
"active_tasks": 0, "processed": 0, "failed": 0, "succeeded": 0, "retried": 0,
"created_at": "..."
}
]

processed, succeeded, failed and retried are aggregated from the events table per worker name, so they survive worker restarts; they are not the engine's own inspect counters.

GET /api/v1/projects/{slug}/workers/{worker_id}

Role: viewer. worker_id is the worker's UUID (a non-UUID is 422). An unknown id, or a worker that belongs to another project, returns 404 worker not found. The response is WorkerDetailPublic: every WorkerPublic field plus metadata, the mapping the worker reported on its heartbeat ({} when it reported none).

GET /api/v1/projects/{slug}/workers/lint

Role: viewer. Evaluates the configuration each worker already reports and flags settings that lose or leak work; see worker configuration lint for the rules. It is read-only and advisory. Returns ProjectLintPublic:

{
"workers_evaluated": 1,
"workers_not_evaluated": 0,
"findings_by_severity": {"critical": 1},
"workers": [
{
"worker_id": "...", "worker_name": "celery@web-01", "engine": "celery",
"hostname": "web-01", "evaluated": true,
"findings": [
{"rule_id": "...", "severity": "critical", "setting": "...",
"title": "...", "detail": "...", "remedy": "..."}
]
}
]
}

A worker counts as evaluated only when its engine has lint rules and it reported a configuration; an empty configuration is still a report. Workers are ordered worst first (most severe finding, then more findings, then name), with unevaluated workers last. workers_not_evaluated is reported separately because an unjudged worker is not a clean one.

GET /api/v1/projects/{slug}/issues?engine=celery&status=ongoing&hours=24&cursor=...&limit=50

Role: viewer. An issue is one failure fingerprint grouped across runs and engines; see issues. All parameters are optional: engine filters by engine, status is ongoing or recovered, hours is a look-back of 1 to 8760 hours, limit is 1 to 200 (default 50), cursor is the opaque cursor from the previous page. Returns IssueListResponse:

{
"items": [
{
"fingerprint": "...", "status": "ongoing",
"occurrences": 12, "open_count": 3, "recovered_count": 9,
"first_seen": "...", "last_seen": "...",
"engine_count": 1, "engines": ["celery"],
"sample_exception": "...", "sample_task_name": "billing.charge"
}
],
"next_cursor": null
}

status is recovered when open_count is zero and ongoing otherwise.

GET /api/v1/projects/{slug}/events?engine=celery&task_id=...&cursor=...&limit=50

Role: viewer. engine (1 to 40 characters) and task_id (1 to 200 characters) are required. limit accepts 1 to 5000; when omitted the page size is Z4J_REST_DEFAULT_PAGE_SIZE (default 50), and any value above Z4J_REST_MAX_PAGE_SIZE (default 500) is capped to it. Returns EventListResponse:

{
"items": [
{"id": "...", "project_id": "...", "agent_id": "...", "engine": "celery",
"task_id": "...", "kind": "task.failed", "occurred_at": "...",
"payload": {}}
],
"next_cursor": "..."
}

Pagination is keyset on (occurred_at, id) with an opaque cursor. A full page always carries a next_cursor, so the final request can return an empty items list. A malformed cursor is treated as no cursor.

Personal task filter presets, scoped to the calling user and the project. Project membership never grants access to another user's views. All four routes require viewer membership and are session-cookie only (the tag has no API-key scope mapping). The writes are CSRF-protected.

GET /api/v1/projects/{slug}/saved-views
POST /api/v1/projects/{slug}/saved-views
PUT /api/v1/projects/{slug}/saved-views/{view_id}
DELETE /api/v1/projects/{slug}/saved-views/{view_id}

GET returns the caller's views in this project as a bare list ordered by name. POST creates one and returns 201; PUT replaces the name and filters of one of the caller's views and returns 200; DELETE returns 204. A view_id that is not one of the caller's views returns 404. A name that another of the caller's views in the project already uses (case-insensitive) returns 409, as does creating a 101st view.

Write body (SavedViewWrite; unknown keys are 422):

{
"name": "Failed critical",
"filters": {
"state": "failure",
"priority": ["critical", "high"],
"search": "billing"
}
}

name is 1 to 80 characters after trimming. state is one of pending, received, started, success, failure, retry, revoked, rejected, unknown, or null. priority holds up to four of critical, high, normal, low; duplicates are dropped and the list is returned in that canonical order. search is at most 200 characters and defaults to "". Responses are SavedViewPublic: the write fields plus id, created_at and updated_at.

GET /api/v1/health # liveness, public
GET /api/v1/health/ready # readiness, public
GET /api/v1/health/system # build and platform detail, authenticated
GET /api/v1/health/deep # per-subsystem report, authenticated

/health does no I/O and returns {"status": "ok"}. It carries no version field on purpose.

/health/ready returns 503 with {"status": "unready", "reason": "starting"} until lifespan startup has finished, 503 with reason database when SELECT 1 fails or takes longer than two seconds, and otherwise 200 with {"status": "ready"}.

/health/system requires any authenticated user. It returns z4j_version, python_version, python_implementation, os, architecture, pid, database_type, database_version and packages (the installed versions of fastapi, uvicorn, sqlalchemy, pydantic and celery, where present). On PostgreSQL it adds database_size_mb and database_connections. If the database query fails, database_type is unknown and database_error names the failure.

/health/deep requires any authenticated user and runs one probe per subsystem, each under a three-second deadline, on the request's own database session. It returns DeepHealthResponse with 200 when no check failed (degraded is a warning and stays 200) and 503 when any check reports failed; the body has the same shape either way:

{
"status": "ok",
"coverage": {"startup_equivalent": false, "detail": "..."},
"checks": {
"database": {"status": "ok", "latency_ms": 1.2},
"migrations": {"status": "ok", "revision": "..."},
"audit_chain": {"status": "ok", "activated": true, "scope": "state-only"}
}
}

A failed check carries detail; migrations adds expected when the database revision is not the head this build boots on. A probe that exceeds its deadline or raises is reported as failed rather than omitted. Fields a probe had no answer for are left out of its entry. coverage states that the verdict is not the one the boot path would give the same database, in either direction; see monitoring for how to read it.

GET /api/v1/setup/status
GET /setup
POST /api/v1/setup/complete

This is the first-boot surface that creates the bootstrap admin and the default project; see first admin. It is gated by the setup token printed in the brain logs, not by a session or CSRF: GET /setup (at the root, not under /api/v1) serves the inline HTML form only while no user exists and an active token row exists, and returns 404 otherwise; POST /api/v1/setup/complete takes {"token", "email", "display_name", "password"}, consumes the single-use token, creates the admin and project in one transaction, sets the session cookie, and returns {"user", "project_id"} under a per-IP throttle of 5 attempts per 15 minutes (429). GET /api/v1/setup/status returns {"first_boot": true} anonymously while no user exists; after provisioning it returns 401 to anonymous callers and {"first_boot": false} to any authenticated session or API key.