Troubleshooting
Run the doctor first
Section titled “Run the doctor first”Before anything else, run the framework-side doctor as the same user the service runs under. It probes everything below in order and reports the first specific failure.
| Stack | Command |
|---|---|
| Django | python manage.py z4j_doctor |
| Flask | python -m z4j_flask doctor |
| FastAPI | python -m z4j_fastapi doctor |
| Bare Python | python -m z4j_bare doctor |
Add --no-websocket to skip the last probe, which starts and stops the local transport runtime and never reaches the brain. Add --json for scripts. Exits 0 on all-green, 1 on any failure.
Agent shows offline
Section titled “Agent shows offline”- Run the doctor - it surfaces the most common failures with a specific reason. The list below covers the rest.
- Check the agent's app logs for the
z4j.runtime.*loggers at boot (messages beginz4j agent ...orz4j buffer: ...). Look for handshake errors. - Verify
Z4J_BRAIN_URLis anhttps://URL in production. The agent deriveswss://from it; a plainhttp://URL is accepted only forlocalhost,127.0.0.1or::1, or with dev mode on. - Verify the token is the one shown at mint time (tokens are not recoverable - re-mint if lost).
- Egress firewall: the agent's host must reach z4j on TCP 443 (or wherever your proxy is).
- Proxy WebSocket passthrough: nginx-ingress needs
proxy_set_header Upgradeandproxy_set_header Connection.
Buffer relocated under gunicorn / uvicorn service users
Section titled “Buffer relocated under gunicorn / uvicorn service users”A WARNING reading z4j buffer: /var/www/.z4j is not writable; falling back to /tmp/z4j-33. Set Z4J_HOME to a persistent writable location to silence this warning. (or any path that's not the running user's writable home) in the service log; no PermissionError reaches the log. The service user has an unwritable $HOME. The agent auto-relocates the buffer to $TMPDIR/z4j-{uid}/buffer-{pid}.sqlite and logs that single WARNING instead of crashing. See service-user deployments.
Events don't appear
Section titled “Events don't appear”- Agent is online (Agents page shows online)?
- Engine is auto-detected (Engines column on the Agents page)?
- For Django:
INSTALLED_APPSincludesz4j_djangoafter any Celery apps? - For Flask:
z4j.init_app(app, ...)was called on the app factory? - For FastAPI: agent is inside the lifespan context manager?
- Task names registered (Workers, open the worker, Registered tasks)?
"N agents are on an older wire protocol" banner
Section titled “"N agents are on an older wire protocol" banner”The agent's package major differs from the brain's versions snapshot. Agents on the brain's major keep working and show an update available badge when a newer package exists; a different major shows incompatible. Re-deploy agents with pip install -U z4j-*.
Audit chain verify fails
Section titled “Audit chain verify fails”Someone modified the audit_log table directly, or a backup restore is incomplete.
- Read the
MISMATCHES (<count>)block the command prints: one line per finding, capped at 100 with a closing line naming how many further findings are not shown. Rows outside the authenticated generations are reported as their own finding. - If known-intentional (e.g., planned DB surgery), document the break externally. The chain from there is not recoverable.
- If unexpected, treat as a compromise event - preserve the DB, alert security, investigate.
409 conflict on mint-token
Section titled “409 conflict on mint-token”The body carries "error": "conflict" and the message an agent with that name already exists in this project: another agent already has that (project_id, name). Pick a different name or delete the old agent first.
Schedule controls missing or failing for a scheduler
Section titled “Schedule controls missing or failing for a scheduler”The scheduler adapter advertises the operations it supports (list, create, update, delete, enable, disable, trigger_now) and the dashboard offers only those; there is no read_only flag. celery-beat advertises the full set, but without django-celery-beat (for example a PersistentScheduler backing store) the write operations fail when used. See schedulers overview.
Password reset email not arriving
Section titled “Password reset email not arriving”- Email channel configured? Use the channel test endpoint (
POST /api/v1/projects/{slug}/notifications/channels/{channel_id}/test) or the Test button on the dashboard's Notifications, Channels page. From:domain has SPF / DKIM / DMARC?- Sender reputation good? Gmail drops many SMTP senders silently.
Very high CPU
Section titled “Very high CPU”- Check
rate(z4j_events_ingested_total[1m])broken down byprojectandengine- is one agent emitting millions of events? - Redaction patterns not looping on giant payloads? - the redactor truncates any single value above 8192 bytes before matching.
- Hot task - check
z4j_task_duration_secondsbytask_name. There is no per-endpoint HTTP histogram; get request latency from your reverse proxy.
503 agent_offline on a retry
Section titled “503 agent_offline on a retry”The target agent was offline when z4j dispatched the command. The command still stays pending and is delivered if the agent reconnects before its deadline (Z4J_COMMAND_TIMEOUT_SECONDS, default 60). Check it on the Commands page before retrying: a second retry sent while the first is pending can run the task twice. Pick a different agent for the same engine only after the first command has timed out.
When stuck
Section titled “When stuck”Include z4j logs (with X-Request-Id), agent logs, and a description of the sequence - file at github.com/z4jdev/z4j/issues.