Kubernetes
The z4j chart is published as an OCI artifact and deploys the brain, with the standalone scheduler and an evaluation PostgreSQL as opt-in components. The chart version equals the z4j release it was built for, so one --version pins the chart and the image together.
helm install z4j oci://ghcr.io/z4jdev/charts/z4j \ --namespace z4j --create-namespacekubectl --namespace z4j logs deployment/z4j -f # capture the first-boot setup URLkubectl --namespace z4j port-forward svc/z4j 7700:7700The defaults mirror the evaluation compose stack: one replica, bundled SQLite on a PersistentVolumeClaim, secrets minted on first boot, reachable only through port-forward. A production install supplies the public origin, explicit secrets, a PostgreSQL URL and an ingress:
kubectl --namespace z4j create secret generic z4j-secrets \ --from-literal=app-secret="$(openssl rand -hex 48)" \ --from-literal=session-secret="$(openssl rand -hex 48)" \ --from-literal=audit-chain-secret="$(openssl rand -hex 48)" \ --from-literal=database-url='postgresql+asyncpg://z4j:PASSWORD@db.internal:5432/z4j?sslmode=verify-full&sslrootcert=/etc/ssl/certs/ca-certificates.crt'
helm install z4j oci://ghcr.io/z4jdev/charts/z4j --namespace z4j \ --set brain.publicUrl=https://z4j.example.com \ --set 'brain.allowedHosts={z4j.example.com}' \ --set brain.allowHttpPublicUrl=false \ --set brain.secrets.existingSecret=z4j-secrets \ --set brain.database.existingSecret=z4j-secrets \ --set brain.ingress.enabled=true \ --set 'brain.ingress.hosts[0].host=z4j.example.com' \ --set 'brain.ingress.hosts[0].paths[0].path=/' \ --set 'brain.ingress.hosts[0].paths[0].pathType=Prefix'scheduler.enabled=true adds the standalone scheduler and turns on the brain's mTLS gRPC listener; pki.mintJob.enabled=true lets a hook Job mint the CA and both certificates into Secrets, the way the compose kit's scheduler-certs.sh does, or point scheduler.tls.existingSecret and brain.schedulerGrpc.tls.existingSecret at material you manage (cert-manager is supported for the scheduler side). The scheduler's leader election needs PostgreSQL; a SQLite brain can only run it with scheduler.leader.backend=single and one replica.
The chart refuses to render a release it knows would crash-loop (a PostgreSQL install without secrets, scheduler metrics without a bearer token, more than one SQLite replica) and names the value to set. helm show values oci://ghcr.io/z4jdev/charts/z4j prints every value with its comment; helm show readme prints the full chart guide, including Secret shapes and the rotation procedure. Every published chart digest is signed keylessly with cosign; verify it with cosign verify ghcr.io/z4jdev/charts/z4j@sha256:<digest> and the identity regexp ^https://github.com/z4jdev/z4j/.
The brain Deployment uses strategy: Recreate for the reason explained under the manifest below, and its data claim is kept on helm uninstall.
The data claim arrives from a dynamic provisioner owned by root, fsGroup changes only its group, and the brain's secret store refuses a /data it does not own outright with mode 0700. The chart therefore runs a fix-data-permissions init container from the same image before the brain: root with only CHOWN and FOWNER, it gives /data to the brain's uid and gid with mode 0700 and changes nothing that is already right. It is the one root container in the release, so the namespace needs the PodSecurity baseline profile or an exemption for it. Where policy forbids that, set brain.persistence.fixPermissions.enabled=false and prepare the volume once from a shell that reaches it as root (the node for hostPath-backed storage, or a one-off pod that mounts the claim): chown 10001:10001 <mount> && chmod 0700 <mount>. On a storage driver that applies fsGroup on every mount, also clear brain.podSecurityContext.fsGroup (--set brain.podSecurityContext.fsGroup=null), or the kubelet adds the group bits back on the next start and the brain refuses the directory again.
Minimum manifest
Section titled “Minimum manifest”Without Helm, the same objects ship as plain manifests in the deploy/kubernetes/ directory of the z4j sdist and the source repository (z4j-brain.yaml and z4j-scheduler.yaml, with CHANGE ME markers). The essentials:
apiVersion: apps/v1kind: Deploymentmetadata: name: z4jspec: replicas: 1 # Recreate, not the default. Kubernetes defaults to RollingUpdate, and its # maxSurge of 25% ROUNDS UP, so even at replicas: 1 it starts a second pod # before terminating the first. The new pod migrates the database on boot # while the old one is still writing, which is exactly the mixed-version # window described under upgrades: the old brain writes change-log envelopes # without the columns the new one reads, and a fire that lands in that window # is skipped rather than deferred. "One replica" is a statement about desired # count, not about how many processes are alive during a rollout. strategy: type: Recreate selector: matchLabels: { app: z4j } template: metadata: labels: { app: z4j } spec: # A provisioned claim arrives owned by root, and the secret store # refuses a /data that uid 10001 (the image's user) does not own with # mode 0700; fsGroup would not fix that either, it only sets the # group. This root step runs the same image with CHOWN and FOWNER # only and changes nothing once the volume is right. It needs the # PodSecurity baseline profile; see the note below this manifest. initContainers: - name: fix-data-permissions image: z4jdev/z4j:latest securityContext: runAsUser: 0 runAsNonRoot: false allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: { drop: [ALL], add: [CHOWN, FOWNER] } command: ["/bin/sh", "-ec"] args: - | [ "$(stat -c %u:%g /data)" = "10001:10001" ] || chown 10001:10001 /data [ "$(stat -c %a /data)" = "700" ] || chmod 0700 /data volumeMounts: - name: data mountPath: /data containers: - name: brain image: z4jdev/z4j:latest ports: [{ containerPort: 7700 }] env: # database-url must carry exactly one sslmode set to require, # verify-ca or verify-full (the verify modes also need sslrootcert). # The brain refuses a PostgreSQL URL without one; only # Z4J_REQUIRE_DB_SSL=false, for a database on a private network, # relaxes that. - name: Z4J_DATABASE_URL valueFrom: { secretKeyRef: { name: z4j-secrets, key: database-url } } - name: Z4J_SECRET valueFrom: { secretKeyRef: { name: z4j-secrets, key: app-secret } } - name: Z4J_SESSION_SECRET valueFrom: { secretKeyRef: { name: z4j-secrets, key: session-secret } } - name: Z4J_AUDIT_CHAIN_SECRET valueFrom: { secretKeyRef: { name: z4j-secrets, key: audit-chain-secret } } - name: Z4J_PUBLIC_URL value: https://z4j.example.com - name: Z4J_ALLOWED_HOSTS value: '["z4j.example.com"]' readinessProbe: httpGet: { path: /api/v1/health/ready, port: 7700 } periodSeconds: 10 livenessProbe: httpGet: { path: /api/v1/health, port: 7700 } periodSeconds: 30 resources: requests: { cpu: "200m", memory: "256Mi" } limits: { cpu: "2", memory: "2Gi" } volumeMounts: - name: data mountPath: /data # Z4J_HOME is /data and the image only declares an anonymous volume # there. The brain keeps owner-private state in it (generated secrets, # restore-recovery and activation manifests, allowed hosts, embedded # PKI), so without a claim that state is discarded with the pod. A # ReadWriteOnce claim is enough under the Recreate strategy above. volumes: - name: data persistentVolumeClaim: { claimName: z4j-data }---apiVersion: v1kind: PersistentVolumeClaimmetadata: name: z4j-dataspec: accessModes: [ReadWriteOnce] resources: requests: { storage: 1Gi }Plus a Service + Ingress per your cluster's conventions.
The fix-data-permissions init container is the one root container in the manifest and needs the PodSecurity baseline profile or an exemption on the namespace. Where policy forbids it, delete it and prepare the volume once from a shell that reaches it as root: chown 10001:10001 <mount> && chmod 0700 <mount>. Do not add fsGroup to a manifest without the init container on a storage driver that applies it on mount: the kubelet then adds group bits on every start and the brain refuses the directory again. The shipped z4j-brain.yaml carries the same init container with the same comment.
WebSocket ingress
Section titled “WebSocket ingress”Your Ingress must allow WebSocket upgrades. Example (nginx-ingress):
metadata: annotations: nginx.ingress.kubernetes.io/proxy-read-timeout: "3600" nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"Horizontal scaling
Section titled “Horizontal scaling”Running more than one brain replica requires sticky session routing for /ws/agent, and for /ws/dashboard when browsers use it (each agent pins to one brain pod). z4j provides no affinity-balance helper of its own, so multi-replica deploys rely on the load balancer's own session-affinity setting.
Postgres
Section titled “Postgres”Use a managed Postgres (Cloud SQL, RDS, Crunchy) or an operator (Zalando, CNPG). Do not run Postgres in a StatefulSet with local storage unless you really know what you're doing.
Secrets
Section titled “Secrets”Inject Z4J_*_SECRET via Secret objects or external managers (Vault, AWS Secrets Manager, GCP Secret Manager). Do not hard-code.
Observability
Section titled “Observability”- Scrape
/metricswith Prometheus (BearerZ4J_METRICS_AUTH_TOKEN, or setZ4J_METRICS_PUBLIC=1if the port is firewalled). - Ship stdout JSON logs with Fluent Bit / Vector.