Skip to content

Architecture

z4j is built as brain + many agents:

  • Brain - one per organization. FastAPI + TanStack Router v1 dashboard. PostgreSQL persistence (SQLite is accepted and forces a single-worker brain). Serves the UI, the REST API, and the agent WebSocket endpoint.
  • Agent - one per application process. Thin pip package. Opens an outbound WebSocket to z4j. Captures events from the local queue engine; executes actions z4j sends back.

This separation buys:

  1. Persistence independence - history lives in Postgres, not in the broker. Broker crashes don't erase history.
  2. Engine abstraction - z4j speaks one wire protocol; agents translate to/from the engine's API.
  3. Network safety - agents initiate all connections. z4j never needs to reach into your application VPC.
your code: celery_app.send_task("email.send", ...)
│
▼
Celery broker (Redis / RabbitMQ / SQS)
│
▼
Celery worker picks up
│
▼
z4j-celery signal hooks capture task_received / task_prerun / task_postrun / task_failure / task_retry / task_revoked
│
▼
z4j-bare dispatcher: redacts secrets, buffers, flushes on batch/time window
│
▼
WebSocket frame: {type: "event_batch", payload: {events: [...]}}
│
▼
brain: validates, persists to `events` table, fans out to connected dashboards over their /ws/dashboard WebSocket
dashboard: user clicks "Retry"
│
▼
REST: POST /api/v1/projects/{slug}/commands/retry-task
│
▼
brain: authorize -> mint signed command row -> look up target agent -> push WebSocket frame
│
▼
WebSocket frame: {type: "command", payload: {action: "retry_task", ...}}
│
▼
agent: dispatch to engine adapter
│
▼
adapter: call the engine operation only when the capability is advertised;
otherwise the command is refused
│
▼
response frame: {type: "command_result", payload: {status: "success", ...}}
│
▼
brain: write audit log entry, return the command row to the dashboard
Axis Examples What it adapts
Framework django / flask / fastapi / bare Process boot, settings parsing, ASGI/WSGI teardown
Engine celery / rq / dramatiq / huey / arq / taskiq Task enqueue, event capture, retry/cancel semantics
Scheduler celery-beat / rq-scheduler / apscheduler / huey-periodic / arq-cron / taskiq-scheduler Periodic task CRUD

They compose freely. A Django + Celery + Beat app uses three adapters; a Flask + RQ + rq-scheduler app uses three different ones. Any combination is supported.

Table Purpose Retention
projects Tenants Unlimited
agents Registered agents Unlimited; a daily hygiene sweep soft-revokes agents silent for Z4J_AGENT_STALE_PRUNE_DAYS (default 30, 0 disables), leaving the row as a tombstone
tasks Task identity (one row per task_id) Unlimited (admin bulk-delete API only)
events Per-state event stream (sent/started/.../finished) Z4J_EVENT_RETENTION_DAYS, default 30
schedules Scheduler entries Unlimited
audit_log Admin actions, auth events. HMAC-chained Z4J_AUDIT_RETENTION_DAYS, default 90

See database schema for full field docs.

Why WebSocket (not polling / not gRPC / not message queue)

Section titled “Why WebSocket (not polling / not gRPC / not message queue)”
  • Agent-initiated outbound only - no inbound firewall holes. This is a security win.
  • Bidirectional with low overhead - one socket carries events and commands. An HTTP longpoll fallback (POST /api/v1/agent/events, GET /api/v1/agent/commands) is available for networks that block WebSockets.
  • gRPC - would add a heavy dep (grpcio) to every agent. WebSocket is in stdlib-adjacent space via websockets.
  • Message queue - adding Redis/RabbitMQ as a dependency for the control plane (when the thing we're observing often is those brokers) creates a circular operational dep.
Failure Behavior
Agent to brain network partition Agent keeps appending events to its on-disk SQLite buffer, which is written first in any case; at the size or byte cap the oldest entries are evicted and counted (capacity_evicted_frames). Reconnects with exponential backoff. On reconnect, drains the buffer.
Brain crash Postgres has the data. Agents reconnect. No events are acknowledged until persisted - at-least-once delivery.
Postgres crash API routes answer with a generic 500 while the database is unreachable; /api/v1/health/ready runs SELECT 1 with a timeout and answers 503, so a load balancer stops routing to the brain. Agents buffer as above.
Worker / queue backlog Doesn't affect z4j - we observe it, we don't participate in it.

See reconciliation for how we detect "task said it started, never said it finished."