Skip to content

Scaling

A single brain replica is the simplest topology. Validate capacity for your workload before choosing it. The main scaling dimensions are:

  1. Postgres -- query, write, and retention load; see the self-hosting guide.
  2. WebSocket connections -- one socket per agent; measure process memory at your target connection count.
  3. Event persistence -- batched into Postgres; benchmark representative event shape, rate, retention, and query load on the intended hosts.

Within one replica, z4j serve starts min(4, cpu) uvicorn worker processes unless --workers says otherwise. SQLite, the local registry backend and Z4J_EMBEDDED_SCHEDULER each force --workers=1, because an in-memory registry and an embedded scheduler child are process-local. Background workers whose ticks have cross-process side effects run under a per-worker advisory lock, so one process per interval runs them and the others no-op.

Every worker verifies the audit chain before it serves, and the walk holds the chain lock from its first page to its last, so the workers of one z4j serve verify one after another: the walk is linear in the trail (roughly thirty thousand rows a second on ordinary hardware), and the fourth worker waits for the three before it. That wait is bounded by Z4J_STARTUP_VERIFY_LOCK_TIMEOUT_MS (default 600000, ten minutes), not by the per-request Z4J_DB_LOCK_TIMEOUT_MS; a worker that waited more than a second logs that it waited for a sibling worker's walk, with the seconds. On a large trail budget startup as workers x walk time, raise the bound if your trail and worker count exceed ten minutes together, and give the readiness probe an initial delay that covers it. The same applies to several replicas starting against one database at once.

Multiple brain replicas are supported on Postgres. The brain selects its registry and dashboard-fan-out backend from Z4J_REGISTRY_BACKEND:

  • postgres_notify (the default on Postgres) -- agent commands and dashboard updates fan out across replicas via Postgres LISTEN/NOTIFY. Each agent's WebSocket lives on whichever replica it happened to connect to; commands minted on any replica route to the right one through the registry. Dashboard subscribers connected to one replica still see events captured by another.
  • local (forced on SQLite, since SQLite has no LISTEN/NOTIFY) -- single-process only.

What you still need to provide yourself:

  • Sticky session routing on /ws -- each agent's WebSocket must pin to one brain pod. Configure your load balancer's session affinity (e.g. nginx-ingress's nginx.ingress.kubernetes.io/affinity: cookie, an ALB target-group's stickiness, or your service-mesh equivalent).
  • TLS termination in front of the brain. The brain itself speaks plaintext WebSocket on its bind port; production deployments put a reverse proxy in front.

The dashboard fan-out is over WebSocket (/ws/dashboard), not SSE; cross-replica delivery is the postgres_notify DashboardHub.

The standalone scheduler scales on its own axis. One instance handles any number of schedules; more instances buy failover, and with postgres_per_project a division of projects between them, because only a project's leader ticks it. Run them with Z4J_SCHEDULER_LEADER_BACKEND=postgres (one leader, warm standbys) or postgres_per_project, each with Z4J_SCHEDULER_LEADER_PG_DSN; the shipped PostgreSQL Compose stack starts two and --scale scheduler=N adds more. See high availability.

  • Read replicas help dashboards but not the hot event-persist path.
  • Native partitioning on events(occurred_at) is built in; partition retention drops the oldest partition once it ages past Z4J_EVENT_RETENTION_DAYS.
  • The brain runs SET statement_timeout, lock_timeout and idle_in_transaction_session_timeout on every PostgreSQL connection it opens, from Z4J_DB_STATEMENT_TIMEOUT_MS (default 10000), Z4J_DB_LOCK_TIMEOUT_MS (default 3000) and Z4J_DB_IDLE_IN_TX_TIMEOUT_MS (default 30000); a value set on the database role is overridden for the session, so tune the env vars rather than the role.

Agents scale with your app. One agent per app process; the worker-first protocol identifies each worker by (agent_id, worker_id) so multi-worker servers (gunicorn, uwsgi) coexist under a single agent identity. No coordination between agents; deploying more app replicas registers more workers automatically.

If you reach a capacity ceiling after benchmarking, file an issue with the workload shape and measurements. We want the feedback.