Skip to content

Unified action surface

Every engine implements "retry" and "cancel" differently:

  • Celery - task.retry() from inside the task, or re-apply with app.send_task.
  • RQ - Queue.requeue(job_id), or job.requeue().
  • Dramatiq - no first-class retry API; middleware-based.
  • Huey - retries are configured per-task; no ad-hoc retry surface.
  • arq - no retry API; re-enqueue manually.
  • taskiq - per-task retry decorators.

z4j exposes one command vocabulary, but only when an adapter can honor the specific command without reconstructing redacted inputs or pretending a broker operation succeeded.

Six verbs exist, gated by each adapter's advertised capabilities:

Verb Semantics
retry_task Re-enqueue only when the engine can recover the authoritative inputs or the operator supplies complete replacements.
cancel_task Invoke an engine primitive that covers the adapter's documented pending/running contract.
bulk_retry Retry an explicit brain-selected set only on adapters that implement it safely.
purge_queue Drop pending messages only where the adapter can measure and target the requested queue.
requeue_dead_letter Put one parked task back on its queue, only where the engine has a dead-letter store the adapter can address by task id.
dlq.list Read one page of the engine's dead-letter store without consuming it; gated by the list_dead_letters capability.

Each engine adapter advertises its capability tokens in the hello frame, per engine, and the brain reads that map before it issues an engine command. retry_task, cancel_task, requeue_dead_letter, bulk_retry and dlq.list go to an agent only when its recorded hello lists the engine and carries the action's capability token for it; the retry family (retry_task, bulk_retry) also needs retry_by_reference_v1, the adapter's attested safe retry contract. The brain holds no engine list of its own: the engine string is checked for shape only, an engine no agent advertises is refused by name (422 on the command routes, 409 on the dead-letter listing, each naming what the agents do advertise), and it is never rewritten to another engine. Every other verb (purge_queue and the worker controls) is written and delivered without a brain-side check, and the agent's dispatcher fails closed: a verb the adapter does not advertise is answered with a command_result of status: "failed" and the error adapter '<engine>' does not support action '<verb>', and that failed result is audited like any other outcome.

The dashboard hides a control the caller's role does not allow, and disables Retry and Cancel unless the selected agent's hello advertised that token for the task's engine.

An agent on the long-poll transport sends no hello, so its capability map is unknown. The dashboard still offers the control for it, the brain-side check does not refuse it (its claim admits each command against the contracts it declares on that poll, and its dispatcher fails closed on an engine it has not loaded), and the brain delivers a retry only to a poll whose declared retry contracts cover the task's engine.

The brain stores redacted task inputs, so it does not rebuild an executable payload from its task or event rows. Retry-with-different-inputs supplies both complete replacement collections explicitly.

Every action writes an audit row with the requested command and outcome.

See API § tasks for exact endpoints.

  • If the agent is offline, the command row is still written and stays pending. A Brain on SQLite, which uses the in-process command registry, answers 503 agent_offline; a PostgreSQL Brain answers with the pending command. Either way the row is delivered if the agent reconnects before its deadline (Z4J_COMMAND_TIMEOUT_SECONDS, default 60), and the timeout sweeper closes it otherwise. Check the command on the Commands page before issuing the action again: a second command sent while the first is pending can run the action twice.
  • An agent on the long-poll transport holds no WebSocket session but is not offline while it polls: the request returns the pending command and the agent claims it on its next poll.
  • If the command times out (60s), the timeout sweeper sets the command row to status timeout with the error command timed out before agent responded, and the audit log captures the failure. The task state is not modified.
  • A mutation timeout can be indeterminate when the broker call may still finish; the adapter reports that state instead of claiming a clean failure.
  • No blanket retry policy - the unified action surface is operator-driven. Automatic retries happen only where you arm an automation rule with a retry action, which is admin-gated, circuit-broken, and audited. For per-task backoff and attempt limits, use the engine's native retry configuration.
  • No side-effect-safety guarantees - retrying a task that already half-ran is the user's call. z4j does not introspect idempotency.