Skip to content

Troubleshooting

Before anything else, run the framework-side doctor as the same user the service runs under. It probes everything below in order and reports the first specific failure.

Stack Command
Django python manage.py z4j_doctor
Flask python -m z4j_flask doctor
FastAPI python -m z4j_fastapi doctor
Bare Python python -m z4j_bare doctor

Add --no-websocket to skip the last probe, which starts and stops the local transport runtime and never reaches the brain. Add --json for scripts. Exits 0 on all-green, 1 on any failure.

  1. Run the doctor - it surfaces the most common failures with a specific reason. The list below covers the rest.
  2. Check the agent's app logs for the z4j.runtime.* loggers at boot (messages begin z4j agent ... or z4j buffer: ...). Look for handshake errors.
  3. Verify Z4J_BRAIN_URL is an https:// URL in production. The agent derives wss:// from it; a plain http:// URL is accepted only for localhost, 127.0.0.1 or ::1, or with dev mode on.
  4. Verify the token is the one shown at mint time (tokens are not recoverable - re-mint if lost).
  5. Egress firewall: the agent's host must reach z4j on TCP 443 (or wherever your proxy is).
  6. Proxy WebSocket passthrough: nginx-ingress needs proxy_set_header Upgrade and proxy_set_header Connection.

Buffer relocated under gunicorn / uvicorn service users

Section titled “Buffer relocated under gunicorn / uvicorn service users”

A WARNING reading z4j buffer: /var/www/.z4j is not writable; falling back to /tmp/z4j-33. Set Z4J_HOME to a persistent writable location to silence this warning. (or any path that's not the running user's writable home) in the service log; no PermissionError reaches the log. The service user has an unwritable $HOME. The agent auto-relocates the buffer to $TMPDIR/z4j-{uid}/buffer-{pid}.sqlite and logs that single WARNING instead of crashing. See service-user deployments.

  1. Agent is online (Agents page shows online)?
  2. Engine is auto-detected (Engines column on the Agents page)?
  3. For Django: INSTALLED_APPS includes z4j_django after any Celery apps?
  4. For Flask: z4j.init_app(app, ...) was called on the app factory?
  5. For FastAPI: agent is inside the lifespan context manager?
  6. Task names registered (Workers, open the worker, Registered tasks)?

"N agents are on an older wire protocol" banner

Section titled “"N agents are on an older wire protocol" banner”

The agent's package major differs from the brain's versions snapshot. Agents on the brain's major keep working and show an update available badge when a newer package exists; a different major shows incompatible. Re-deploy agents with pip install -U z4j-*.

Someone modified the audit_log table directly, or a backup restore is incomplete.

  1. Read the MISMATCHES (<count>) block the command prints: one line per finding, capped at 100 with a closing line naming how many further findings are not shown. Rows outside the authenticated generations are reported as their own finding.
  2. If known-intentional (e.g., planned DB surgery), document the break externally. The chain from there is not recoverable.
  3. If unexpected, treat as a compromise event - preserve the DB, alert security, investigate.

The body carries "error": "conflict" and the message an agent with that name already exists in this project: another agent already has that (project_id, name). Pick a different name or delete the old agent first.

Schedule controls missing or failing for a scheduler

Section titled “Schedule controls missing or failing for a scheduler”

The scheduler adapter advertises the operations it supports (list, create, update, delete, enable, disable, trigger_now) and the dashboard offers only those; there is no read_only flag. celery-beat advertises the full set, but without django-celery-beat (for example a PersistentScheduler backing store) the write operations fail when used. See schedulers overview.

  1. Email channel configured? Use the channel test endpoint (POST /api/v1/projects/{slug}/notifications/channels/{channel_id}/test) or the Test button on the dashboard's Notifications, Channels page.
  2. From: domain has SPF / DKIM / DMARC?
  3. Sender reputation good? Gmail drops many SMTP senders silently.
  1. Check rate(z4j_events_ingested_total[1m]) broken down by project and engine - is one agent emitting millions of events?
  2. Redaction patterns not looping on giant payloads? - the redactor truncates any single value above 8192 bytes before matching.
  3. Hot task - check z4j_task_duration_seconds by task_name. There is no per-endpoint HTTP histogram; get request latency from your reverse proxy.

The target agent was offline when z4j dispatched the command. The command still stays pending and is delivered if the agent reconnects before its deadline (Z4J_COMMAND_TIMEOUT_SECONDS, default 60). Check it on the Commands page before retrying: a second retry sent while the first is pending can run the task twice. Pick a different agent for the same engine only after the first command has timed out.

Include z4j logs (with X-Request-Id), agent logs, and a description of the sequence - file at github.com/z4jdev/z4j/issues.