Skip to content

Wire protocol

Protocol version: 2. v2 adds a per-frame HMAC envelope plus replay protection; v1 is not accepted on the wire. The canonical definitions live in z4j_core.transport.frames; see the WebSocket protocol reference for the full schema.

  • WebSocket, TLS in production (wss://), plaintext for local dev (ws://).
  • One JSON object per WebSocket frame, UTF-8. There is no WS-layer batching; the event_batch frame carries the application-level batch.
  • Authentication: Authorization: Bearer <agent-token> on the WebSocket handshake.
  • HTTP longpoll fallback for networks that block WebSockets: POST /api/v1/agent/events to upload signed frames and GET /api/v1/agent/commands to poll for inbound commands. Same z4j_core frame types, same HMAC envelope.

Every frame carries the version, an id, and a type. Stateful frames also carry the HMAC envelope (nonce, seq, hmac); handshake frames (hello / hello_ack) are unsigned because the agent and brain are still negotiating which key to use.

{
"v": 2,
"type": "event_batch",
"id": "<agent-generated, 1..64 chars>",
"ts": "2026-04-16T12:34:56.789Z",
"nonce": "<urlsafe random>",
"seq": 4281,
"hmac": "<base64 HMAC over the canonical envelope>",
"payload": { "events": [] }
}

First frame the agent sends. Declares the agent's protocol version, framework, engines, schedulers, and host info. The brain validates compatibility before accepting the connection.

The hot path. payload.events is capped at 5000 entries; the agent's batcher caps itself at 500. Signed in v2 so a stolen bearer token alone cannot forge events.

Default cadence is 10 seconds; the brain returns its preferred heartbeat_interval_seconds in hello_ack and the agent honours that.

Periodic self-report sent alongside the heartbeat: consecutive failure counts per error class, last success, session age, buffer depth and version metadata. The brain stores eligible frames in agent_status_history on a best-effort basis (rate limiting or a persistence failure may drop one).

Reply to a brain-initiated command. Correlated by command_id.

The agent sends this first-stage receipt immediately after accepting a brain-initiated command and before executing it. The later command_result reports the execution outcome.

A task-definition delta for one engine (engine, added and updated definitions, removed names). The brain accepts the frame and discards it; nothing on the brain consumes registry deltas, and the agent runtime does not send them.

Carries agent_id, project_id, session_id, plus tuning parameters (heartbeat interval, max frame size).

Round-trip ack so the agent knows which buffered batch it can drop. Includes the original frame id (acked_id) so the agent's in-flight map can match precisely.

Brain-initiated work, correlated by command_id. The REST routes under /api/v1/projects/{slug}/commands/... issue retry_task, cancel_task, requeue_dead_letter, bulk_retry, purge_queue, restart_worker, pool_grow, pool_shrink, add_consumer, cancel_consumer and rate_limit; the dead-letter listing route (GET /api/v1/projects/{slug}/dead-letters) issues dlq.list and returns the agent's command_result as its HTTP response. The brain's own workers and schedule routes issue reconcile_task (the reconciliation sweep), submit_task (manual fire of a schedule owned by z4j-scheduler), schedule.fire, schedule.enable, schedule.disable, schedule.trigger_now, schedule.resync, schedule.external.activate and schedule.external.control. There is no schedule create or update on the wire: a schedule owned by z4j-scheduler is stored in the brain and reaches the agent as schedule.fire; a schedule owned by an external scheduler is inventoried by the agent's scheduler adapter, which receives only enable, disable, trigger_now and resync. target always carries type and id; a command the brain has resolved to one engine also carries engine in target (bulk_retry from its filter, dlq.list from the queue it reads), so a host running several adapters binds the right one. purge_queue carries no engine.

Carries code, message and fatal (default false). The brain's only emitter is scheduler_upgrade_required with fatal: true, sent when a schedule event needs a current scheduler adapter, and the brain does not close the socket after it. The agent closes on any fatal error frame, and treats the codes scheduler_upgrade_required, agent_incompatible and protocol_incompatible as terminal for its build (slow reconnect schedule until something is upgraded).

  • Agent to brain: the agent's seq is monotonic per signed-frame stream. Replay protection rejects out-of-order or repeated nonces.
  • The brain dedupes accepted events by a deterministic event id: for a task event, uuid5 of project_id:engine:task_id:kind:occurred_at (at second precision); for any other event, uuid5 of project_id:agent_event_id. A retried batch after a brain-side restart is therefore idempotent; seq is replay protection, not the dedupe key.
  • The event_batch_ack carries received, accepted, rejected so the agent can confirm exactly which rows landed.
  • Brain to agent: command_id is a UUID. The agent replies with the same command_id in command_result. Commands time out server-side: a sweeper (every 5 seconds by default) moves a command whose timeout_at has passed to status timeout with the error command timed out before agent responded, and the REST side reads that from the command row.

hello.payload.protocol_version must be "2". A mismatch closes the WebSocket with code 4426. The agent treats that close as terminal for its build: it keeps reconnecting, but on a slow schedule (120 seconds, backing off to 3600) instead of the normal one, so an upgraded deployment is picked up without restarting the host process.

Forward-compat: agents and brain ignore unknown fields. A verb an agent's adapter does not advertise is answered with command_result status: "failed" and the error adapter '<name>' does not support action '<verb>' (a verb the dispatcher has no branch for at all answers unrecognized action '<verb>'); the brain records that as a failed command rather than a protocol error.