Skip to content

Tasks API

Task reads (list, detail, tree) live under /projects/{slug}/tasks. Most agent-executed mutations (retry, cancel, bulk retry, queue purge, and worker control) go through /projects/{slug}/commands so the brain can mint an audit-chained command record. Bulk delete is the brain-local exception at /projects/{slug}/tasks/bulk-delete.

GET /api/v1/projects/{slug}/tasks

Role: viewer.

Query param Type Notes
state string One of pending, received, started, success, failure, retry, revoked, rejected, unknown. Invalid and empty values return 422.
priority string Comma-separated, e.g. critical,high.
name string Literal task-name substring.
search string Literal, case-insensitive substring across name, queue, worker, and task ID.
queue string Exact match.
worker string Exact match on worker_name.
since / until RFC 3339 Bound on received_at.
cursor opaque Pagination cursor from the prior page.
limit int Request validation accepts 1..5000, then normal JSON listing clamps it to Z4J_REST_MAX_PAGE_SIZE (default 500).
include_total bool Default false. When true the response also carries total_count, the exact filtered count independent of cursor / limit (one extra SQL aggregate; exports ignore it). Live data can change between reads, so it is not a sealed bulk-action target.
format string csv, xlsx, or json. Export mode ignores cursor pagination and serves at most Z4J_TASKS_EXPORT_MAX_ROWS rows (default 50,000; 100 to 100,000). A filter that matches more than the cap is refused, never truncated: 422 validation_error with details.cap, details.format and details.setting, and a message pointing at the narrowing filters (state, priority, queue, worker, since, until, name, search). XLSX additionally fails above its separate 25,000-row in-memory cap.
fields string Comma-separated projection for export.

Response (JSON mode):

{
"items": [
{
"id": "01H...",
"project_id": "...",
"engine": "celery",
"task_id": "9d2c...",
"name": "email.send",
"queue": "default",
"state": "failure",
"priority": "normal",
"args": ["<redacted>"],
"kwargs": {"to": "<email>"},
"result": null,
"exception": "ValueError: ...",
"traceback": "...",
"retry_count": 0,
"eta": null,
"received_at": "...",
"started_at": "...",
"finished_at": "...",
"runtime_ms": 142,
"worker_name": "worker-1@host",
"parent_task_id": null,
"root_task_id": null,
"tags": [],
"created_at": "...",
"updated_at": "..."
}
],
"next_cursor": "...",
"total_count": null
}
GET /api/v1/projects/{slug}/tasks/{engine}/{task_id}

Role: viewer. Same TaskPublic shape as the list endpoint.

GET /api/v1/projects/{slug}/tasks/{engine}/{task_id}/tree

Role: viewer. Walks to the task's root via root_task_id (or treats the task itself as root if the field is null) and returns every task in the project sharing that root. Capped at 500 nodes; the response carries truncated: true when the cap kicks in.

POST /api/v1/projects/{slug}/tasks/bulk-delete

Role: admin. CSRF-protected. Throttled by the shared bulk-action bucket used by other destructive and expensive routes.

{
"filter_state": "failure",
"filter_priority": ["critical", "high"],
"filter_search": "email_%"
}

The request must select exactly one mode: a non-empty task_ids list of at most 1000 task-row UUIDs, or at least one validated filter_* field. Empty or all-null bodies, empty ID/priority lists, blank text filters, unknown state or priority values, duplicate priorities, and mixed ID/filter requests return 422. Every mode is scoped to the path project.

Filter mode accepts filter_state, filter_priority, filter_search, filter_name, filter_queue, filter_worker, filter_since, and filter_until. Terms are intersected. filter_search mirrors list-search semantics across task name, queue, worker, and task ID; filter_name searches only task names. Both are literal substring filters, so %, _, and escape characters do not broaden the selection. Filter mode deletes at most 10,000 rows in stable task-UUID order. Returns {"deleted_count": N}.

Commands (retry, cancel, bulk-retry, purge-queue)

Section titled “Commands (retry, cancel, bulk-retry, purge-queue)”

Every action that needs an agent round-trip is issued as a command. The brain inserts a row in commands, signs it, and delivers it over WebSocket or the agent long-poll transport. The agent returns signed command_ack and command_result frames over that same transport; there is no result POST endpoint.

GET /api/v1/projects/{slug}/commands?status=&cursor=&limit=

Role: viewer. status is one of pending, dispatched, completed, failed, timeout, or cancelled. An unknown filter value is ignored and therefore returns all statuses.

GET /api/v1/projects/{slug}/commands/{command_id}

Role: viewer.

Both read routes return the same command shape, but a viewer gets the operational view only: issued_by is null, payload is {}, and result is null. Operators and admins see all three. A purge confirm_token inside the payload is redacted for every role.

POST /api/v1/projects/{slug}/commands/retry-task

Role: operator.

{
"agent_id": "...",
"engine": "celery",
"task_id": "9d2c...",
"idempotency_key": "optional-string",
"override_args": [],
"override_kwargs": {"foo": "bar"},
"eta_seconds": 60
}

Supply neither override field to retry by reference, or supply both override_args and override_kwargs to replace the original inputs; a partial override is rejected. Each field has a 64 KiB serialized cap, but the whole request is also subject to the lower default HTTP body limit. eta_seconds (optional, 0 to 86400) delays the re-enqueue; the brain resolves it to an absolute ETA when the request arrives, so queueing or an offline agent does not restart the delay. override_kwargs keys beginning with __z4j_ are reserved control metadata and are rejected with 422.

engine is not checked against a list. The target agent's current session must advertise the engine in its engine list with the retry_task capability and the adapter's attested safe retry contract; otherwise the request is 422 and the error names the engine. The brain never substitutes another engine for the one sent. An engine whose adapter retries by re-submitting (Huey) must be given both override fields, or the request is 409.

POST /api/v1/projects/{slug}/commands/cancel-task

Role: operator.

{
"agent_id": "...",
"engine": "celery",
"task_id": "9d2c...",
"idempotency_key": "optional-string"
}

engine follows the same rule as retry, with the cancel_task capability: an engine or action the agent's session does not advertise is 422.

GET /api/v1/projects/{slug}/dead-letters?engine=rq&queue=default&limit=100&cursor=

Role: viewer. API-key scope: tasks:read.

Query param Type Notes
engine string Required. The engine whose dead-letter store to read. Checked for shape only; an online agent must advertise it with list_dead_letters.
queue string Restrict the page to one queue. Absent, every queue the adapter knows.
limit int Page size, 1..200, default 100.
cursor opaque The next_cursor of the previous page, passed back verbatim.

The brain keeps no copy of an engine's dead letters. Each request issues one dlq.list command to an online agent whose session advertises the engine with list_dead_letters, waits for the agent's result and returns the page it sent, newest first. The command and its result stay in the commands list like any other command, and every listing writes a dead_letters.list audit row with the engine, queue and limit. Because each listing fans a command out to an agent and holds the request for up to fifteen seconds, it draws on the same per-address bulk-action bucket as bulk retry, purge and trigger-now (rate limits): past the budget the answer is 429 with a Retry-After, before any command is issued. An adapter that refuses the cursor (invalid dead-letter cursor) is 422; any other adapter failure, including a broker error that happens to mention a cursor, is 502.

{
"entries": [
{
"task_id": "9d2c...",
"task_name": "email.send",
"queue": "default",
"failed_at": "2026-10-02T12:00:00Z",
"error_excerpt": "ValueError: ...",
"attempts": 3
}
],
"next_cursor": "100",
"total": 212,
"engine": "rq"
}

task_id is the id the requeue command below takes. task_name, failed_at, attempts and total are null or empty when the engine does not record them; error_excerpt is the redacted tail of the stored failure text, at most 512 characters. A Dramatiq RabbitMQ broker reports total only, with an empty entries list, because AMQP has no non-destructive read of a queue.

Errors: 409 when no online agent advertises list_dead_letters for the engine (the body lists the online agents and what each advertises for it), 504 when the agent does not answer within the wait bound of 15 seconds, or Z4J_COMMAND_TIMEOUT_SECONDS when that is shorter (the body names the command), 422 for a malformed engine, limit or cursor, and 502 when the agent refused the listing or returned something that is not a page.

POST /api/v1/projects/{slug}/commands/requeue-dead-letter

Role: operator. CSRF-protected. Returns 202 with the command.

{
"agent_id": "...",
"engine": "rq",
"task_id": "9d2c...",
"idempotency_key": "optional-string"
}

Moves one dead-lettered task back onto its queue. The route is gated on what the target agent advertises, not on an engine list: whether a requeue is safe is a property of the engine's own dead-letter primitive, which only the adapter knows, so an adapter advertises requeue_dead_letter only when it has one. RQ's failed job registry is such a primitive. An agent whose session does not advertise the capability for the engine (the Celery and Dramatiq adapters among them) is refused with 422 and nothing touches the broker. The task_id comes from the dead-letter listing above.

POST /api/v1/projects/{slug}/commands/bulk-retry

Role: operator. CSRF-protected and throttled by the shared bulk-action bucket (10 requests per minute per IP).

{
"agent_id": "...",
"filter": {"task_ids": ["9d2c..."], "engine": "celery"},
"max": 1000,
"idempotency_key": "optional-string"
}

This compatibility command route accepts only an explicit, non-empty task_ids selection and requires an engine that the target agent's session advertises with the bulk_retry capability (422 otherwise); every selected ID must belong to this project and engine. max is bounded 1..10000 and truncates the deduplicated selection to its first max IDs. Filter-only/all-matching retries return 410 and must use the durable bulk retry requests resource.

POST /api/v1/projects/{slug}/commands/purge-queue

Role: admin. CSRF-protected and throttled by the shared bulk-action bucket (10 requests per minute per IP).

{
"agent_id": "...",
"queue": "default",
"confirm_token": "<hmac of (queue_name, observed_depth)>",
"force": false,
"idempotency_key": "optional-string"
}

Clients may pass a precomputed confirm_token, or observed_depth so the brain computes the keyed token, or force=true. The HTTP route also accepts none of them and returns 202; in that case the agent refuses the command asynchronously. force=true bypasses the token/depth guards. The Celery adapter logs that bypass at CRITICAL; the RQ and Dramatiq paths do not make the same logging guarantee.

The durable resource behind filter-driven bulk retries. Base path: /api/v1/projects/{slug}/bulk-retry-requests. Every route requires operator; the three POST routes are CSRF-protected, and create is also throttled by the shared bulk-action bucket. These routes are session-only: the collection has no API-key scope mapping, so a bearer key receives 403 (scope_missing) even with admin:*.

POST /api/v1/projects/{slug}/bulk-retry-requests
GET /api/v1/projects/{slug}/bulk-retry-requests/{request_id}
POST /api/v1/projects/{slug}/bulk-retry-requests/{request_id}/pause
POST /api/v1/projects/{slug}/bulk-retry-requests/{request_id}/resume

Create body:

{
"idempotency_key": "retry-billing-failures-1",
"filter": {"engine": "celery", "state": "failure", "name": "billing"},
"max": 1000
}

idempotency_key is required (1 to 200 chars) and identifies the request: replaying the same key with the same body returns the existing request, and the same key with a different body is 409. Create returns 202 when the request is in_progress and 200 for any other status, always with a Location header pointing at the request. agent_id is optional; when it is absent each sealed child binds to a compatible session in the project at send time. filter accepts only task_ids, engine, state (alias status), queue, name, search, priority, since, and until; any other key is 400. Explicit task_ids require an engine. A filter engine must be one the named agent_id, or without one some agent in the project, advertises with the bulk_retry capability; otherwise the request is 400 and the error names the engine. priority cannot be combined with task_ids. max is 1 to 10000 (default 1000).

The request object reports status (in_progress, paused, blocked, succeeded, failed, partial, indeterminate, or no_match), control_state (running, paused, or blocked), the sealed canonical_digest and plan_digest, target_agent_id, per-child counts (total, pending, claimed, unobserved, succeeded, failed, unknown), and created_at, sealed_at, deadline_at, and last_progress_at. pause and resume flip control_state and write bulk_retry_request.paused / bulk_retry_request.resumed audit rows.

The same /commands router exposes five worker-control routes. Each accepted command is audit-logged and signed, then delivered to the named agent over WebSocket or long-poll.

Worker control is an operator capability, not an admin one. Only queue purge is raised to admin, because it destroys queued work. Assign operator with that in mind.

POST /api/v1/projects/{slug}/commands/restart-worker

Role: operator.

{
"agent_id": "...",
"worker_name": "celery@worker-1",
"idempotency_key": "optional-string"
}

For Celery this broadcasts pool_restart(reload=True): the parent worker stays alive while its child pool is respawned. The broadcast is fire-and-forget and requires Celery's pool-restart support; it is not proof that in-flight work was preserved. Other adapters may refuse the operation or use the guarded supervisor-based self-exit fallback.

POST /api/v1/projects/{slug}/commands/pool-resize

Role: operator.

{
"agent_id": "...",
"worker_name": "celery@worker-1",
"delta": 2,
"idempotency_key": "optional-string"
}

delta is bounded -100..100 and must be non-zero. Positive grows the pool, negative shrinks it. The agent's engine adapter translates the delta into the engine-native call (pool_grow / pool_shrink). Unsupported adapters still leave the HTTP call at 202; the command later reaches status="failed" with an error string. The wire result has status, not an ok field.

POST /api/v1/projects/{slug}/commands/add-consumer

Role: operator.

{
"agent_id": "...",
"worker_name": "celery@worker-1",
"queue": "billing",
"idempotency_key": "optional-string"
}

Tells the worker to start consuming from the named queue.

POST /api/v1/projects/{slug}/commands/cancel-consumer

Role: operator. Same body shape as add-consumer. The worker stops consuming from queue but keeps running on its remaining queues.

POST /api/v1/projects/{slug}/commands/rate-limit

Role: operator.

{
"agent_id": "...",
"task_name": "myapp.tasks.send_email",
"rate": "100/m",
"worker_name": "celery@worker-1",
"idempotency_key": "optional-string"
}

rate follows Celery's grammar: integer optionally suffixed with /s, /m, or /h; "0" clears the limit. The pattern is enforced server-side. worker_name is optional in the request schema, but omission does not broadcast: the dispatcher falls back to the command target id, which for this route is the task name, and sends the control command to a worker with that name. The command can therefore report success without changing the intended fleet. No CRITICAL global-throttle log or audit flag is produced on this path; pass an explicit worker name.

Every worker-control command requires an agent_id, which the agents API returns. Discover engine-native worker identifiers (for example Celery's worker@host) from GET /api/v1/projects/{slug}/workers. The separate /api/v1/projects/{slug}/agent-workers collection describes z4j agent processes, not engine-native workers; neither collection is nested under /agents. agent-workers is session-only: it has no API-key scope mapping, so a bearer key receives 403 even with admin:*.