Skip to content

Migrate from Flower to z4j

This guide walks through replacing Flower with z4j on a production Celery deployment. Both tools observe the same Celery broker events, so you can run them side by side during the cut-over and turn Flower off only when you are confident z4j covers your workflow.

End state: persistent task history in Postgres, retry/cancel/bulk actions from the dashboard, Celery Beat schedule CRUD (with django-celery-beat), RBAC with invitations, an HMAC-chained audit log, and the same dashboard for any other engines you add later (RQ, Dramatiq, Huey, arq, taskiq).

Plan a dedicated change window; duration depends on your deployment and validation.

You are running Celery 5.3 or newer (z4j-celery alone accepts Celery 5.2.2 and later) with Flower in front of it, and you have outgrown Flower for at least one of these reasons:

  • Worker crashes leave tasks "started" forever because Flower has no reconciler.
  • Flower forgets everything on restart and you needed to know what failed last week.
  • You want bulk retry, schedule CRUD, or RBAC and Flower does not have them.
  • Your stack is going multi-engine (RQ alongside Celery, for example) and you want one dashboard.
  • An auditor asked for an HMAC-chained audit log of who retried which job and when.

If none of that applies and you are happy with Flower, do not migrate. Flower is a fine tool inside the scope it was built for.

  • Celery 5.3 or newer (z4j-celery alone accepts 5.2.2 and later; z4j-celerybeat needs 5.3).
  • A reachable Postgres or SQLite path for z4j. Production should use Postgres; SQLite is fine for evaluation.
  • The same broker URL and result backend the workers already use. No data migration is required.
  • Permission to add one pip package (z4j-celery) to the worker venv and one container or process for z4j.

Flower keeps running unchanged. z4j is added in parallel. Neither tool interferes with the other.

Terminal window
# Option A: pip + bundled SQLite, fastest path for evaluation
pip install z4j
z4j serve
# Option B: docker compose, recommended for anything other than a laptop
git clone https://github.com/z4jdev/z4j
cd z4j
docker compose up -d

The first boot prints a setup URL to stderr. Open it, create the first admin, and you have an empty dashboard.

Verify the install:

Terminal window
z4j check # config + DB connectivity + the stored migration revision (no head compare)
z4j status # user/project/agent/task counts, version

At this point z4j is running but has no agents and no events. Flower is still your source of truth.

Phase 2: Add the z4j-celery agent to your workers

Section titled “Phase 2: Add the z4j-celery agent to your workers”

Install the adapter into the venv your Celery workers use:

Terminal window
pip install z4j-celery

Set these environment variables on the worker process. Z4J_BRAIN_URL, Z4J_TOKEN, Z4J_HMAC_SECRET and Z4J_PROJECT_ID are required; Z4J_AGENT_NAME is optional and unset by default:

Terminal window
export Z4J_BRAIN_URL="https://z4j.example.com"
export Z4J_TOKEN="<bearer-from-the-dashboard>"
export Z4J_HMAC_SECRET="<hmac-from-the-dashboard>"
export Z4J_PROJECT_ID="default" # required: slug of the project the agent was minted in
export Z4J_AGENT_NAME="celery-prod-1" # optional, unset by default

Mint the token and HMAC secret in the dashboard on the project's Agents page, New agent. The dashboard prints them once; store them in your secrets manager.

Restart one worker with Celery logging at INFO (-l info). You should see this line in the worker log (Celery adds its own timestamp and process prefix; the logger name is z4j.adapter.celery.worker_bootstrap):

INFO: z4j worker bootstrap: agent runtime started (celery_app=..., framework=...)

The message is logged at INFO, so Celery's default WARNING level hides it even while the agent is running. If you log at INFO and still do not see that line, the agent is not running. Common causes:

  • Worker was started as celery beat or celery inspect, not celery worker. The bootstrap signal only fires under celery worker.
  • Z4J_BRAIN_URL is unreachable from the worker. Check DNS and TLS from inside the worker pod or container.
  • Z4J_TOKEN is wrong. z4j logs an unauthorized request when this happens.

Once the first worker reports in, verify its task state in the z4j dashboard, then restart and verify the remaining workers.

Phase 3: Validate parity over a representative business cycle

Section titled “Phase 3: Validate parity over a representative business cycle”

Run Flower and z4j side by side for one full business cycle. The goal is to make sure z4j shows the same operational picture you trusted Flower for, plus the new things you came for.

A reasonable validation checklist:

  • Live tasks appear in z4j as they appear in Flower.
  • Task counts per queue match between the two dashboards.
  • A task you intentionally fail shows up with the full traceback in z4j.
  • A task you intentionally retry from z4j actually re-runs.
  • A scheduled task fires at its expected time.
  • The audit log entry for the retry shows your user, the timestamp, and the action.
  • A task stranded in started by a worker restart is corrected by reconciliation (every 300 s the brain probes tasks stuck longer than 15 min and rewrites its own snapshot to the engine's state; it never re-runs a task).
  • Secret arguments are redacted by default in the z4j detail view.

If any of these fail, file an issue at github.com/z4jdev/z4j/issues with the agent log and z4j log. Do not turn Flower off until they all pass.

Three things to do, in order:

  1. Move any operator playbooks or runbooks that reference flower.example.com to point at z4j.example.com.
  2. Stop the Flower process or remove the flower service from docker-compose.yml. Keep the entry commented out in git for one release cycle in case you need to roll back.
  3. Tell the team. Include a one-paragraph diff so they know what changed and where the new buttons are.

z4j keeps recording from now on. Flower's last in-memory snapshot is gone the moment you stop it, but you no longer care.

If validation finds a problem, use the rollback steps below. z4j and Flower remain independent during the overlap period.

Terminal window
# Re-enable Flower (it never had state to migrate, so it just starts).
docker compose up -d flower
# Stop the z4j agent on workers (uninstall or unset Z4J_BRAIN_URL).
# Workers keep running normally; only the agent thread stops phoning home.
pip uninstall -y z4j-celery

Operators can keep using Flower exactly as before. z4j retains the history it has already recorded; you can come back to it later.

No. Passive capture never enqueues anything: the agent reads broker events, worker signals, and the result backend. Operator actions you trigger from the dashboard (retry, cancel, bulk retry, purge, restart, rate limit) do act on the broker and workers through the agent, and each is audit-logged; a retry re-enqueues the original task via send_task, so it runs again by design. Running z4j alongside Flower or any other observer is safe.

Does z4j handle Celery chords, groups, and chains?

Section titled “Does z4j handle Celery chords, groups, and chains?”

Yes. The z4j-celery adapter understands chord/group/chain parents and renders them as a tidy-tree DAG with runtime badges per node. Flower shows these as a flat parent/child list, which is one of the things you may have come here to fix.

Install pip install z4j-celerybeat on the host that runs celery beat. Full schedule CRUD (create, edit, delete, enable/disable, trigger-now) needs django-celery-beat with beat running its DatabaseScheduler; the adapter writes to the PeriodicTask rows. Static beat_schedule entries are listed read-only. RedBeat and the default file-backed scheduler are not write targets.

Do I need to change my broker, result backend, or worker code?

Section titled “Do I need to change my broker, result backend, or worker code?”

No. Capture is passive; only the operator actions you trigger from the dashboard act on the broker and workers. Your CELERY_BROKER_URL, CELERY_RESULT_BACKEND, task definitions, and worker invocation stay exactly as they are.

Will my secret task arguments leak to the dashboard?

Section titled “Will my secret task arguments leak to the dashboard?”

Secrets are scrubbed by default. The redaction layer recursively walks args and kwargs and replaces values that match common credential patterns (password, token, secret, api_key, authorization, etc.). You extend the patterns on the agent side: set Z4J_REDACTION_EXTRA_KEY_PATTERNS and Z4J_REDACTION_EXTRA_VALUE_PATTERNS in the worker environment, pass the matching install_agent(...) arguments, or use the redaction dict the Django and Flask adapters accept in their Z4J settings. There are no project-level redaction settings in the dashboard.

Does z4j work with django-celery-beat / redbeat?

Section titled “Does z4j work with django-celery-beat / redbeat?”

django-celery-beat yes, redbeat no. The adapter has two sources: django-celery-beat (read/write, when the package is installed and Django apps are ready) and the static beat_schedule (read-only). Entries you create in the z4j dashboard land in the django-celery-beat tables, and entries you create directly in django-celery-beat are visible in the z4j dashboard; beat picks either up only when it runs the DatabaseScheduler.

What happens to historical tasks Flower had cached?

Section titled “What happens to historical tasks Flower had cached?”

They are gone the moment Flower restarts, the same as before. z4j starts a fresh, persistent record from the moment its agent connects. There is no migration step because there is nothing in Flower to migrate from.

The agent buffers events in a crash-safe on-disk SQLite store (~/.z4j/buffer.sqlite by default) and ships them to z4j in batches over a single WebSocket. Measure agent overhead and brain/database throughput with a representative workload before setting production capacity expectations.

Can I still use Flower's HTTP API for scripts I have already written?

Section titled “Can I still use Flower's HTTP API for scripts I have already written?”

z4j ships a REST and WebSocket API documented under API reference. Map each call you scripted against Flower's API to its z4j counterpart in that reference rather than assuming a one-to-one equivalent; operator actions through the z4j API leave an audit-log entry recording who called them.

Is there a paid version with different features?

Section titled “Is there a paid version with different features?”

No. z4j is open source under split licensing: z4j is AGPL v3, the adapters are Apache 2.0. Every feature is in the open-source release. There is no commercial gate, no telemetry, and no phone-home. The only version-check request is the operator-initiated Settings, System, Check for updates fetch of versions.json from GitHub; set Z4J_VERSION_CHECK_URL to an empty string to disable it.

Check these integration pitfalls during validation:

  • Worker started without celery worker. The bootstrap signal only fires under that command. celery beat, celery inspect, celery purge will not start the agent.
  • Reverse proxy strips the WebSocket upgrade. The agent uses a long-lived WebSocket. nginx, Cloudflare, and Traefik all need explicit upgrade handling. See the TLS setup guide for working configs.
  • Token shared across multiple agents. Each worker host should get its own agent token. Sharing one token means you cannot tell which host is reporting which event.
  • HMAC clock skew. The HMAC frames embed a timestamp. If the worker clock drifts by more than 60 seconds from the z4j clock, z4j rejects the frames. Run NTP everywhere.
  • Missing Z4J_PROJECT_ID. The agent refuses to start with missing required Z4J settings: project_id (or Z4J_PROJECT_ID); there is no fallback to a default project. Set it to the slug of the project whose agent token you minted.