Skip to content

Production checklist

  • Z4J_SECRET and Z4J_SESSION_SECRET each independently set to 64 hex chars of entropy.
  • Z4J_AUDIT_CHAIN_SECRET set to a further independent key of at least 32 bytes. Outside development this is required and has no fallback, so the brain will refuse to start without it. Store it where whoever administers the database cannot read it, otherwise the audit chain proves less than it appears to.
  • Z4J_PUBLIC_URL set to the public HTTPS URL (used in invitation and reset links, the setup banner, the dashboard WebSocket origin check, and the CSP and HSTS headers). Outside dev it must start with https:// unless Z4J_ALLOW_HTTP_PUBLIC_URL=true. CORS origins come from Z4J_CORS_ORIGINS, and cookie flags follow Z4J_ENVIRONMENT.
  • TLS terminated in front of the brain; the proxy forwards X-Forwarded-For and is listed in Z4J_TRUSTED_PROXIES.
  • Postgres 17+ with backups (pg_dump schedule or managed service snapshots).
  • At least two admin accounts (avoids lockout if one is lost).
  • Email notification channel configured per project if you use invitations / password reset.
  • Z4J_METRICS_AUTH_TOKEN set explicitly (only a fresh packaged SQLite install mints one on its own) and /metrics scraped by Prometheus.
  • Container resource limits set (not unlimited).
  • Container healthcheck configured.
  • Log aggregation sinks z4j's JSON stdout.
  • Audit-chain verification runs on a schedule, one of two ways: z4j audit verify in cron / a Kubernetes CronJob, or the built-in worker via Z4J_AUDIT_CHAIN_VERIFY_ENABLED=true (off by default, daily, leader-gated).
  • Z4J_AUDIT_RETENTION_DAYS set to your actual retention obligation. The default is 90 days and rows past the window are deleted, not archived.
  • If you need audit evidence that survives a compromised database role: chain head exported on a schedule to a sink that role cannot rewrite, and verification run against it with z4j audit verify --known-head. Without this, a clean verification means "nothing outside z4j's own write path touched this log", not "nothing touched this log".
  • A scheduler is running. The shipped Compose stacks start one next to the brain (the scheduler service runs z4j-scheduler serve from the brain image); outside Compose, run it yourself, because nothing fires a schedule otherwise. /ready on port 7800 of the scheduler answers 200.
  • The scheduler channel is mutual TLS: Z4J_SCHEDULER_GRPC_TLS_CERT, Z4J_SCHEDULER_GRPC_TLS_KEY and Z4J_SCHEDULER_GRPC_TLS_CA set on the brain, Z4J_SCHEDULER_GRPC_ALLOWED_CNS naming the client identity, Z4J_SCHEDULER_GRPC_REQUIRE_ALLOWLIST=true, and Z4J_SCHEDULER_GRPC_INSECURE unset.
  • The certificate material is backed up or owned. The Compose stacks mint a private CA and both leaf certificates with the scheduler-certs one-shot into the z4j_scheduler_pki volume; back that volume up with the data, or replace its contents with certificates from your own PKI (see z4j-scheduler).
  • Port 7701 on the brain is not published, but it is reachable by any container on the stack network, so mutual TLS plus the CN allowlist is the gate, not the network. The stacks never publish it to the host.
  • More than one replica, with Z4J_SCHEDULER_LEADER_BACKEND=postgres and Z4J_SCHEDULER_LEADER_PG_DSN, if a scheduler outage must not stop fires. The PostgreSQL stack starts two; --scale scheduler=N adds more.
  • Z4J_SCHEDULER_METRICS_AUTH_TOKEN set together with Z4J_SCHEDULER_METRICS_ENABLED=true if Prometheus scrapes the scheduler. The stacks ship scheduler metrics off.
  • An alert on z4j_scheduler_misfires_detected_total, which the brain raises when a schedule is late, so a dead scheduler is noticed before its users notice.
  • One agent token per agent, stored in the app's secret manager.
  • Token rotation plan documented (mint new, deploy, revoke old).
  • Z4J_BRAIN_URL uses the public https:// brain base URL in production. The agent derives the wss:// WebSocket endpoint itself.
  • agent_name distinguishes web / worker / beat / cron processes.
  • Egress firewall allows WebSocket to brain.
  • MFA enrolled for every admin, and Z4J_MFA_ENFORCE_FOR_ADMINS (or Z4J_MFA_ENFORCE_FOR_ALL) set. Enforcement is off by default and gates cookie sessions only, so also revoke personal API keys for anyone who misses the enrollment deadline. See MFA.
  • Recovery codes distributed and stored somewhere the user can reach without their phone, and the z4j reset-mfa escape hatch written into your runbook.
  • Redaction rules reviewed for your domain-specific secrets (see redaction).
  • Rate limits verified on /auth/*, /setup, /api/v1/invitations/preview and /api/v1/invitations/accept (/invite is the dashboard page, not the throttled route).
  • No Z4J_BOOTSTRAP_ADMIN_* left in env after first boot.
  • Password policy appropriate for your org (length, complexity, denylist - all enforced by default).
  • Audit retention policy documented. The built-in worker deletes rows older than Z4J_AUDIT_RETENTION_DAYS (90 days by default); archive externally first if you must retain them longer.
  • Runbook for "brain down" (restart, check Postgres, check disk).
  • Runbook for "agent offline" (check token, network, app logs).
  • Runbook for "stuck tasks" (reconciliation worker status, manual retry).
  • Disaster recovery tested - can you restore Postgres + re-mint agent tokens?

z4j is not certified against SOC 2, HIPAA, or ISO 27001. If you need compliance, the audit log export + Postgres backups provide most of the raw evidence, but you own the policies, controls, and external audit.

Know the gaps in that evidence before you build a control on top of it:

  • The HMAC chain is evidence against everything that writes through z4j. It is not evidence against a role that can write the audit tables directly, which can roll the log back to a shortened history that still verifies. Export the chain head off-box on a schedule if that matters (see HMAC audit chain).
  • Denied requests are mostly not recorded. A rejection for insufficient role writes no audit row on most API families; only mutating schedule routes offer one, best-effort.
  • There is no SSO, OIDC, SAML, or SCIM. Identity federation and user deprovisioning have to run in a reverse proxy in front of the brain.