Production checklist
-
Z4J_SECRETandZ4J_SESSION_SECRETeach independently set to 64 hex chars of entropy. -
Z4J_AUDIT_CHAIN_SECRETset to a further independent key of at least 32 bytes. Outside development this is required and has no fallback, so the brain will refuse to start without it. Store it where whoever administers the database cannot read it, otherwise the audit chain proves less than it appears to. -
Z4J_PUBLIC_URLset to the public HTTPS URL (used in invitation and reset links, the setup banner, the dashboard WebSocket origin check, and the CSP and HSTS headers). Outsidedevit must start withhttps://unlessZ4J_ALLOW_HTTP_PUBLIC_URL=true. CORS origins come fromZ4J_CORS_ORIGINS, and cookie flags followZ4J_ENVIRONMENT. - TLS terminated in front of the brain; the proxy forwards
X-Forwarded-Forand is listed inZ4J_TRUSTED_PROXIES. - Postgres 17+ with backups (
pg_dumpschedule or managed service snapshots). - At least two admin accounts (avoids lockout if one is lost).
- Email notification channel configured per project if you use invitations / password reset.
-
Z4J_METRICS_AUTH_TOKENset explicitly (only a fresh packaged SQLite install mints one on its own) and/metricsscraped by Prometheus. - Container resource limits set (not unlimited).
- Container healthcheck configured.
- Log aggregation sinks z4j's JSON stdout.
- Audit-chain verification runs on a schedule, one of two ways:
z4j audit verifyin cron / a Kubernetes CronJob, or the built-in worker viaZ4J_AUDIT_CHAIN_VERIFY_ENABLED=true(off by default, daily, leader-gated). -
Z4J_AUDIT_RETENTION_DAYSset to your actual retention obligation. The default is 90 days and rows past the window are deleted, not archived. - If you need audit evidence that survives a compromised database role: chain head exported on a schedule to a sink that role cannot rewrite, and verification run against it with
z4j audit verify --known-head. Without this, a clean verification means "nothing outside z4j's own write path touched this log", not "nothing touched this log".
Scheduler
Section titled “Scheduler”- A scheduler is running. The shipped Compose stacks start one next to the brain (the
schedulerservice runsz4j-scheduler servefrom the brain image); outside Compose, run it yourself, because nothing fires a schedule otherwise./readyon port 7800 of the scheduler answers 200. - The scheduler channel is mutual TLS:
Z4J_SCHEDULER_GRPC_TLS_CERT,Z4J_SCHEDULER_GRPC_TLS_KEYandZ4J_SCHEDULER_GRPC_TLS_CAset on the brain,Z4J_SCHEDULER_GRPC_ALLOWED_CNSnaming the client identity,Z4J_SCHEDULER_GRPC_REQUIRE_ALLOWLIST=true, andZ4J_SCHEDULER_GRPC_INSECUREunset. - The certificate material is backed up or owned. The Compose stacks mint a private CA and both leaf certificates with the
scheduler-certsone-shot into thez4j_scheduler_pkivolume; back that volume up with the data, or replace its contents with certificates from your own PKI (see z4j-scheduler). - Port 7701 on the brain is not published, but it is reachable by any container on the stack network, so mutual TLS plus the CN allowlist is the gate, not the network. The stacks never publish it to the host.
- More than one replica, with
Z4J_SCHEDULER_LEADER_BACKEND=postgresandZ4J_SCHEDULER_LEADER_PG_DSN, if a scheduler outage must not stop fires. The PostgreSQL stack starts two;--scale scheduler=Nadds more. -
Z4J_SCHEDULER_METRICS_AUTH_TOKENset together withZ4J_SCHEDULER_METRICS_ENABLED=trueif Prometheus scrapes the scheduler. The stacks ship scheduler metrics off. - An alert on
z4j_scheduler_misfires_detected_total, which the brain raises when a schedule is late, so a dead scheduler is noticed before its users notice.
Agents
Section titled “Agents”- One agent token per agent, stored in the app's secret manager.
- Token rotation plan documented (mint new, deploy, revoke old).
-
Z4J_BRAIN_URLuses the publichttps://brain base URL in production. The agent derives thewss://WebSocket endpoint itself. -
agent_namedistinguishes web / worker / beat / cron processes. - Egress firewall allows WebSocket to brain.
Security posture
Section titled “Security posture”- MFA enrolled for every admin, and
Z4J_MFA_ENFORCE_FOR_ADMINS(orZ4J_MFA_ENFORCE_FOR_ALL) set. Enforcement is off by default and gates cookie sessions only, so also revoke personal API keys for anyone who misses the enrollment deadline. See MFA. - Recovery codes distributed and stored somewhere the user can reach without their phone, and the
z4j reset-mfaescape hatch written into your runbook. - Redaction rules reviewed for your domain-specific secrets (see redaction).
- Rate limits verified on
/auth/*,/setup,/api/v1/invitations/previewand/api/v1/invitations/accept(/inviteis the dashboard page, not the throttled route). - No
Z4J_BOOTSTRAP_ADMIN_*left in env after first boot. - Password policy appropriate for your org (length, complexity, denylist - all enforced by default).
- Audit retention policy documented. The built-in worker deletes rows older than
Z4J_AUDIT_RETENTION_DAYS(90 days by default); archive externally first if you must retain them longer.
Operations
Section titled “Operations”- Runbook for "brain down" (restart, check Postgres, check disk).
- Runbook for "agent offline" (check token, network, app logs).
- Runbook for "stuck tasks" (reconciliation worker status, manual retry).
- Disaster recovery tested - can you restore Postgres + re-mint agent tokens?
Compliance notes
Section titled “Compliance notes”z4j is not certified against SOC 2, HIPAA, or ISO 27001. If you need compliance, the audit log export + Postgres backups provide most of the raw evidence, but you own the policies, controls, and external audit.
Know the gaps in that evidence before you build a control on top of it:
- The HMAC chain is evidence against everything that writes through z4j. It is not evidence against a role that can write the audit tables directly, which can roll the log back to a shortened history that still verifies. Export the chain head off-box on a schedule if that matters (see HMAC audit chain).
- Denied requests are mostly not recorded. A rejection for insufficient role writes no audit row on most API families; only mutating schedule routes offer one, best-effort.
- There is no SSO, OIDC, SAML, or SCIM. Identity federation and user deprovisioning have to run in a reverse proxy in front of the brain.