Skip to content

z4j-scheduler - the dynamic scheduler service

z4j-scheduler is a separate companion process that fires schedules against any z4j-supported engine (Celery, RQ, Dramatiq, Huey, arq, taskiq). Unlike the per-engine scheduler adapters in this section (which surface celery-beat / rq-scheduler / APScheduler / etc. in the dashboard), z4j-scheduler IS the scheduler - operators who choose it delete celery-beat from their stack.

Each schedule's catch_up policy decides what happens to a missed occurrence: one that falls due while the scheduler's watch of the brain is unhealthy and is older than the on-time grace by the time the watch recovers. An occurrence still inside the grace is late, not missed, and fires under every policy; a brain restart takes seconds and usually fits. For a missed one, skip discards it, logging a warning, counting it in z4j_scheduler_slots_discarded_total and showing it on the schedule's dashboard page as "Skipped in last 24h" (which counts gaps, not occurrences: a three-slot gap is one skip there), while fire_one_missed runs the latest missed occurrence as soon as the watch recovers. New schedules created from the dashboard or through POST /schedules default to fire_one_missed; choose skip for work that must never run late.

A single-binary, leader-elected, gRPC-connected scheduler service. Three responsibilities:

  1. Tick - sleep until the earliest next_fire_at in the in-memory schedule cache (fed by brain over the gRPC watch stream), waking early when the cache changes and at least every 30 seconds.
  2. Dispatch - for each due fire, send a FireSchedule gRPC frame to z4j. Brain dispatches a schedule.fire command to an online agent of the project that connected over WebSocket and advertised the schedule's engine; a long-poll agent never qualifies, so a project served only by long-poll agents sees every fire buffered.
  3. Reconcile - every 15 minutes (configurable), do a full sync from brain so a missed delete event in the watch stream doesn't leave a stale schedule firing forever.

Idle CPU is essentially zero. Memory is bounded by the number of schedules.

How it differs from the per-engine scheduler adapters

Section titled “How it differs from the per-engine scheduler adapters”

The packages on the schedulers overview page (z4j-celerybeat, z4j-rqscheduler, etc.) are adapters - they surface an EXISTING native scheduler (celery-beat, rq-scheduler) in the z4j dashboard so operators can see and edit those schedules in one UI.

z4j-scheduler is a replacement - it owns the schedule storage itself, in z4j's database, and dispatches across any engine. Operators who run z4j-scheduler don't need celery-beat or rq-scheduler at all.

The marketing version of this matrix lives at z4j.com/schedulers/z4j-scheduler broken into six categorized tables. The terse engineering version follows.

Capability celery-beat django-celery-beat rq-scheduler APScheduler 4 system cron z4j-scheduler
Engines supported Celery only Celery only RQ only APScheduler in-process Shell exec All 6: Celery, RQ, Dramatiq, Huey, arq, taskiq
Framework agnostic Yes Django only Yes Yes Yes (host OS) Yes - Django/Flask/FastAPI/bare
Multiple engines, one process No No No No No Yes
Capability celery-beat django-celery-beat rq-scheduler APScheduler 4 system cron z4j-scheduler
Edit live (no restart) No Django admin only No Persistent jobstore only No Dashboard / declarative / REST
Built-in dashboard No Django admin (basic) No No No Yes - fire history, run-now, edit
Fire history per schedule No No No No syslog only Buffered + acked + searchable
Audit log of edits No Django auditlog (3rd party) No No No HMAC-chained
Manual fire-now button No No No API only No Yes - dashboard + REST
RBAC / project scoping No Django auth only No No UNIX permissions Project + role-scoped
Capability celery-beat django-celery-beat rq-scheduler APScheduler 4 system cron z4j-scheduler
HA / leader election Single only Single only Single only Pluggable (manual) Single host Postgres advisory-lock leader, rolling-restart safe
Catch-up on outage All-or-nothing All-or-nothing Default fire-all Coalescing only No Per-schedule: skip / fire-one-missed / fire-all-missed
DST / IANA tz correctness Yes Yes Partial Yes Yes Yes (validated at API)
Solar (sunrise/sunset) Yes Yes No No No Yes
Capability celery-beat django-celery-beat rq-scheduler APScheduler 4 system cron z4j-scheduler
Importer FROM other schedulers N/A N/A N/A N/A N/A celery-beat, django-celery-beat, rq-scheduler, APScheduler, system crontab, Huey, arq, taskiq
Exporter TO other schedulers N/A N/A N/A N/A N/A celery-beat, rq-scheduler, APScheduler, crontab, Huey, arq, taskiq
Coexist with native scheduler - - - - - Yes - z4j-celerybeat coexistence adapter
Declarative-in-source Yes (beat_schedule) No (DB-only) Manual Yes crontab file Yes - z4j_scheduler.declarative reconciler

The importer and exporter cover the schedulers people actually migrate off. Huey periodic tasks, arq cron jobs and taskiq schedule labels live in your source rather than in a store, so the importer reads them from an import path (--huey-app, --arq-settings, --taskiq-broker) and the exporter renders a Python module you paste back. The huey-periodic, arq-cron and taskiq-scheduler adapters still surface them read-only in the dashboard if you keep the native scheduler. Dramatiq has no native scheduler, so --from dramatiq and --to dramatiq print migration guidance instead.

Capability celery-beat django-celery-beat rq-scheduler APScheduler 4 system cron z4j-scheduler
Wire HMAC + replay protection Broker-dependent Broker-dependent No No N/A HMAC-SHA256 + per-session seq+nonce binding
HMAC-chained audit log No No No No No Per-row HMAC, prev-hmac chain
Secret rotation without losing history N/A N/A N/A N/A N/A Multi-key verification window, but agents must be re-credentialed
Capability celery-beat django-celery-beat rq-scheduler APScheduler 4 system cron z4j-scheduler
License BSD-3 BSD-3 MIT MIT BSD Apache-2.0
Process model Long-running daemon Daemon + DB Long-running daemon In-process or daemon init / systemd Standalone OR brain-embedded subprocess
Idle CPU footprint Low Low Low Low Negligible Negligible (sleeps until the next fire; 16-worker dispatch pool)
┌─────────────────────────────────────────────┐
│ Your app (Django/Flask/FastAPI/bare) │
│ + Celery (or RQ / arq / Huey / etc.) │
│ + z4j-X engine adapter │
│ + z4j-bare agent runtime │
└──────────────────┬──────────────────────────┘
│ WebSocket / longpoll
▼
┌─────────────────────────────────────────────┐
│ z4j - server + dashboard │
│ z4j-scheduler - fires schedules at the │
│ right time │
└─────────────────────────────────────────────┘

z4j-scheduler does not run tasks. It only decides WHEN they should run and tells brain to dispatch them. Your existing engine worker (celery, RQ worker, etc.) runs the tasks on its own broker. This means:

  • No new broker. Your existing Redis / RabbitMQ / Postgres / etc. stays as the message bus.
  • No new worker process. Your existing celery worker / rq worker / arq worker continues running tasks unchanged.
  • The only new thing is one scheduler process (or a few, for HA).

Every fire carries an identity derived deterministically from the schedule and its exact scheduled slot. The same logical fire therefore always produces the same identity, no matter which instance produces it or how many times it is retried.

That single property is what makes running more than one instance safe. If an instance dies after dispatching a fire but before recording it, and another instance picks the schedule up and dispatches the same slot, the brain recognises the second dispatch as the same fire rather than a new one and deduplicates it. Without deterministic identity, every failover would be a coin flip between a missed fire and a duplicate one.

The retry rules follow from the same idea. Transient transport failures are retried with capped exponential backoff, because retrying is free when the identity is stable. Permission and precondition failures are not retried, because repeating them cannot change the answer.

Leader election uses PostgreSQL advisory locks, which are session-scoped: when the holder's connection drops the lock is released automatically, and a standby's next poll acquires it.

Two backends are available:

Z4J_SCHEDULER_LEADER_BACKEND Behaviour
single Always leader. Correct for one instance, and the only option without PostgreSQL.
postgres One leader for the whole cluster. Standbys idle until failover.
postgres_per_project One lock per project, so instances divide the projects between them.

postgres_per_project is the one to reach for with more than a couple of instances. Work is shared rather than parked in a hot standby, which also means the failover path is exercised continuously instead of only during an incident.

Takeover has no fixed response-time promise. A clean stop typically fails over in one to three seconds and a hard partition in 30 to 60 seconds, but both figures come from the environment, not from the scheduler. A killed process closes its connection immediately, so the lock frees at once. A partitioned host cannot send anything, so the lock is held until the database's TCP keepalive gives up. Tune the keepalive on the PostgreSQL host, and Z4J_SCHEDULER_LEADER_HEARTBEAT_SECONDS on the scheduler (default of two seconds, anywhere from half a second to a minute; it sets both the leader's liveness probe and the standby's acquire retry), if your schedules are tight enough that a lost minute matters. Z4J_SCHEDULER_LEADER_POLL_INTERVAL_SECONDS is still accepted for compatibility but nothing reads it.

HA does not make task execution exactly-once. The brain deduplicates the dispatch of a slot by its fire identity; the task itself then runs on your engine's worker under that engine's own delivery semantics.

The scheduler fails closed and says so. What each condition does:

  • Watch stream down. The tick engine stops dispatching, withholds any transient retry so a fire never reaches a brain it has not negotiated with, and resumes once the stream reconnects and re-syncs. z4j_scheduler_watch_healthy drops to 0 at once; /ready turns 503 only when the outage outlasts Z4J_SCHEDULER_ON_TIME_GRACE_SECONDS. A slot that came due in the gap and is still inside the grace when the stream is back fires late, under every catch-up policy; only a slot older than the grace is missed and goes to the schedule's catch-up policy.
  • Cadence mismatch mid-flight. A brain restarted during a fire's retry window can answer FIRE_CADENCE_SEMANTICS_MISMATCH from a runtime this instance never negotiated with. The first such answer holds the schedule locally, logs a WARNING and renegotiates the cadence contract; agreement lifts the hold and the slot is retried on the next tick. Only a mismatch that survives an agreed renegotiation is reported to the brain as a durable quarantine (cadence_semantics_mismatch), which disables the schedule until an operator lifts it.
  • Brain and scheduler on different pinned sets. Both sides hash the libraries and timezone data that decide fire times, and the brain refuses a scheduler whose hash differs: the watch stream fails with FAILED_PRECONDITION, /ready answers 503, and a scheduler started against such a brain exits. Nothing fires and nothing is recorded against the schedules; they resume when both run the same release, and the occurrences missed meanwhile follow the catch-up policy. The upgrade guide gives the order that avoids the window.
  • Brain shutdown. A brain that has begun shutting down admits no new fire: it answers FireSchedule with gRPC UNAVAILABLE, which is what a scheduler sees from a brain that is gone, and writes nothing. The scheduler logs a transient error and retries the same slot when a brain is back. A fire accepted before the stop signal arrived is outside that gate: if the process ends before delivering it, the delivery timeout holds the schedule until an operator resolves it, which is why an upgrade stops the scheduler first.
  • Tombstone pressure. A project that accumulates 10,000 deletes since its last stable snapshot is paused: nothing in it fires until a covering snapshot lands. The pause logs a WARNING naming the project and the count and requests an immediate snapshot resync rather than waiting for the periodic one, which remains the backstop if that resync fails.
  • A networked single leader. single is not an election: every replica started with it considers itself the leader. At startup, single bound off loopback outside dev logs one WARNING saying so, because that is the shape of a deployment that grows a second replica. The brain fences the duplicate dispatches (no task runs twice), but the FireSchedule load doubles and the replicas' catch-up decisions can diverge under clock skew. Run one replica, or elect through PostgreSQL.
  • Shutdown. On SIGTERM the process stops admitting slots, waits up to Z4J_SCHEDULER_FIRE_TIMEOUT_SECONDS for fires already sent to the brain to be acknowledged, then releases the leader gate and closes the gRPC channel. A fire still in flight past that bound is logged and abandoned; the brain fences a late acceptance, so a standby that takes over cannot double-fire it.

Verifying a migration before you commit to it

Section titled “Verifying a migration before you commit to it”

The shipped importer can compare the normalized schedule set with the rows already in the brain. It cannot independently predict what celery-beat would fire, so it cannot certify timing equivalence.

Terminal window
# Static schedule-set diff. Nothing is written.
export Z4J_SCHEDULER_BRAIN_API_TOKEN='<project-admin-token-from-secret-manager>'
z4j-scheduler import --from celery --celery-app myapp.celery:app \
--project myproject --brain-url https://brain.internal --verify

--verify prints insert, update, unchanged, and delete counts and implies --dry-run. It does not compare fire times. The reserved --duration option is deliberately fail-closed: --verify --duration 24h runs the static diff, then exits 2 and says there is no independent source-side firing oracle. Do not treat it as cutover certification.

For homelab / single-instance deploys, z4j can spawn z4j-scheduler as a supervised subprocess in its own lifespan. Auto-mints loopback mTLS at boot.

Z4J_EMBEDDED_SCHEDULER=true
Z4J_SCHEDULER_GRPC_ENABLED=true

That's it - z4j spawns z4j-scheduler at startup and supervises it (bounded auto-restart, graceful SIGTERM). One container, no extra ops surface.

For production, run z4j-scheduler as a separate process or container. Multiple instances elect a leader via Postgres advisory lock; only the leader ticks. Followers stay warm.

Terminal window
pip install z4j-scheduler
export Z4J_SCHEDULER_BRAIN_GRPC_URL=brain.internal:7701
export Z4J_SCHEDULER_BRAIN_REST_URL=https://brain.internal
export Z4J_SCHEDULER_TLS_CERT=/etc/z4j/tls/scheduler-1.crt
export Z4J_SCHEDULER_TLS_KEY=/etc/z4j/tls/scheduler-1.key
export Z4J_SCHEDULER_TLS_CA=/etc/z4j/tls/ca.crt
export Z4J_SCHEDULER_METRICS_AUTH_TOKEN='<from-secret-manager>'
# Required for multiple scheduler instances. A single instance may leave
# Z4J_SCHEDULER_LEADER_BACKEND at its default of single.
export Z4J_SCHEDULER_LEADER_BACKEND=postgres
export Z4J_SCHEDULER_LEADER_PG_DSN='postgresql://z4j:password@db.internal/z4j'
z4j-scheduler serve

serve takes no command-line options. It reads Z4J_SCHEDULER_* settings, and also a .env file in the working directory, and fails startup if required endpoints, mTLS material, production metrics authentication, or the selected leader backend are invalid. It exits 0 on a clean stop and 1 on a configuration or startup failure.

The process serves four HTTP endpoints on Z4J_SCHEDULER_BIND_HOST and Z4J_SCHEDULER_BIND_PORT:

  • /health always answers 200 once the process is up.
  • /ready answers 503 while any of brain_client, cache_initial_sync or leader_gate is not ready, and names the ones that are not. It also answers 503 once the watch stream has been unhealthy for longer than Z4J_SCHEDULER_ON_TIME_GRACE_SECONDS, naming watch_unhealthy with the outage length in seconds, and answers 200 again as soon as the stream reconnects. A shorter reconnect is a blip the engine rides out on its own. Reconnect attempts back off exponentially up to Z4J_SCHEDULER_GRPC_RECONNECT_BACKOFF_MAX_SECONDS, and that penalty clears once the stream has stayed healthy for one Z4J_SCHEDULER_RECONCILE_INTERVAL_SECONDS, so a drop after a long quiet stretch reconnects fast while a flapping stream keeps backing off.
  • /info reports version, instance id, uptime, readiness, whether the watch stream is healthy and the number of schedules loaded.
  • /metrics is Prometheus text, behind a bearer token when Z4J_SCHEDULER_METRICS_AUTH_TOKEN is set, and absent entirely when Z4J_SCHEDULER_METRICS_ENABLED is false.

z4j-scheduler info reads that base URL from --url, or from Z4J_SCHEDULER_INFO_URL, defaulting to http://localhost:7800.

z4j side enables the gRPC server with:

Z4J_SCHEDULER_GRPC_ENABLED=true
Z4J_SCHEDULER_GRPC_ALLOWED_CNS=["scheduler-prod","scheduler-staging"]

Mutual TLS is required: z4j's gRPC server presents its server cert; the scheduler presents a client cert whose CN must be in the allow-list. The only plaintext path is a dev-only opt-in, refused unless the environment setting on each side is exactly dev: Z4J_SCHEDULER_GRPC_INSECURE=true on z4j (which also bypasses the CN allow-list) and Z4J_SCHEDULER_INSECURE_GRPC=true on the scheduler.

The Compose files shipped with the brain package start the scheduler next to the brain, so a default deployment fires schedules with no further setup. This holds for docker-compose.yml (SQLite) and docker-compose.postgres.yml (PostgreSQL), with docker-compose.caddy.yml layered on either. Three services take part:

  • scheduler-certs runs first and exits. It is the brain image running docker/scheduler-certs.sh (POSIX sh and openssl, nothing else) as root. It mints a private CA, a server certificate for the brain (DNS names z4j, brain and localhost) and a client certificate for the scheduler, writes them into the z4j_scheduler_pki volume, and hands them to the uid the brain and the scheduler run as. It is idempotent: a valid bundle is left alone, and a certificate is minted again only when it is missing, does not chain to the CA, does not match its key, or is within 30 days of expiry. Restart the brain and the scheduler after a renewal. The CA key stays readable by root only.
  • z4j serves the scheduler gRPC channel on port 7701 of the private Compose network over mutual TLS, with Z4J_SCHEDULER_GRPC_TLS_CERT, Z4J_SCHEDULER_GRPC_TLS_KEY and Z4J_SCHEDULER_GRPC_TLS_CA pointing at the minted files, Z4J_SCHEDULER_GRPC_ALLOWED_CNS naming exactly one client identity, and Z4J_SCHEDULER_GRPC_REQUIRE_ALLOWLIST=true. The port is never published to the host.
  • scheduler is the same image with its entrypoint switched to z4j-scheduler serve. It presents the client certificate (CN z4j-scheduler), starts once the brain reports healthy, and answers /health, /ready and /info on port 7800 inside the network, where the dashboard's Schedulers page reaches it through Z4J_SCHEDULER_INFO_URLS.

The SQLite stack runs one scheduler with Z4J_SCHEDULER_LEADER_BACKEND=single. That scheduler binds every interface of its container outside dev, so it logs the single-leader WARNING described under "Outages, mismatches and shutdown" at every start; for a one-replica stack that is expected and needs no change. The PostgreSQL stack runs two replicas with Z4J_SCHEDULER_LEADER_BACKEND=postgres; Z4J_SCHEDULER_LEADER_PG_DSN defaults to the stack's own database, built from the same POSTGRES_USER, POSTGRES_PASSWORD and POSTGRES_DB values the brain uses, so the pair elects a leader through the advisory lock and the standby takes over when the leader stops. Add replicas with:

Terminal window
docker compose -f docker-compose.postgres.yml up -d --scale scheduler=3

Every replica presents the same client certificate: the allow-list names a role, not an instance, and the instance id defaults to the container hostname. A replica beyond the second appears on the dashboard only once its /info URL is added to Z4J_SCHEDULER_INFO_URLS on the brain.

Settings you may put in .env:

  • Z4J_PKI_SCHEDULER_CN changes the minted CN and the brain's allow-list together.
  • Z4J_SCHEDULER_METRICS_ENABLED=true together with Z4J_SCHEDULER_METRICS_AUTH_TOKEN turns /metrics on. The stacks ship it off because the scheduler binds every interface of its container and refuses to serve unauthenticated metrics outside dev. Set both: an empty token locks the endpoint rather than opening it.
  • Z4J_SCHEDULER_LEADER_PG_DSN, when the PostgreSQL password contains URL delimiters or the lock should live in another database.

To bring your own PKI, replace the files in the z4j_scheduler_pki volume with your CA certificate and the two certificate and key pairs, keeping the file names. To add a client identity signed by the generated CA, for a scheduler that runs outside the stack, mint it from the one-shot's container and allow-list the name on the brain:

Terminal window
docker compose run --rm --user 0:0 --entrypoint z4j scheduler-certs \
mint-scheduler-cert --name scheduler-ext --ca-cert /pki/ca.crt \
--ca-key /pki/ca.key --out-dir /pki

Every setting is read from the environment with the Z4J_SCHEDULER_ prefix, or from a .env file in the working directory.

Brain connection:

Variable Default Effect
Z4J_SCHEDULER_BRAIN_GRPC_URL required Host and port of the brain's scheduler gRPC server.
Z4J_SCHEDULER_BRAIN_REST_URL required Read by check, doctor and status only; the running service never calls it.
Z4J_SCHEDULER_TLS_CERT none Client certificate PEM presented to the brain.
Z4J_SCHEDULER_TLS_KEY none Client private key PEM.
Z4J_SCHEDULER_TLS_CA none CA bundle that signed the brain's server certificate.
Z4J_SCHEDULER_INSECURE_GRPC false Plaintext gRPC. Refused unless Z4J_SCHEDULER_ENVIRONMENT is exactly dev, and the brain must also allow it.
Z4J_SCHEDULER_GRPC_KEEPALIVE_SECONDS 30 (5 to 300) gRPC keepalive ping interval.
Z4J_SCHEDULER_GRPC_RECONNECT_BACKOFF_MAX_SECONDS 30 (1 to 300) Maximum backoff between watch-stream reconnects.

Identity, HTTP surface and logging:

Variable Default Effect
Z4J_SCHEDULER_INSTANCE_ID hostname Label shown in logs and on /info.
Z4J_SCHEDULER_BIND_HOST 0.0.0.0 Interface for the HTTP surface. A loopback bind exempts /metrics from the token requirement.
Z4J_SCHEDULER_BIND_PORT 7800 (1 to 65535) Port for the HTTP surface.
Z4J_SCHEDULER_ENVIRONMENT production Only the exact value dev relaxes the plaintext-gRPC and metrics-token gates. Any other value, including staging, is production posture.
Z4J_SCHEDULER_LOG_LEVEL INFO One of DEBUG, INFO, WARNING, ERROR, CRITICAL.
Z4J_SCHEDULER_LOG_JSON true JSON log lines when true.
Z4J_SCHEDULER_METRICS_ENABLED true Mounts /metrics; false leaves it unmounted.
Z4J_SCHEDULER_METRICS_AUTH_TOKEN none Bearer token for /metrics. Startup refuses a non-loopback bind without it outside dev.

Leader election:

Variable Default Effect
Z4J_SCHEDULER_LEADER_BACKEND single single, postgres or postgres_per_project.
Z4J_SCHEDULER_LEADER_PG_DSN none PostgreSQL DSN for the advisory lock. Both Postgres backends refuse to start without it.
Z4J_SCHEDULER_LEADER_NAMESPACE z4j-scheduler-global Hashed into the advisory-lock key. Two clusters sharing one database must use different namespaces.
Z4J_SCHEDULER_LEADER_HEARTBEAT_SECONDS 2.0 (0.5 to 60) Cadence of the leader's liveness probe and the standby's acquire retry.

Tick and dispatch:

Variable Default Effect
Z4J_SCHEDULER_ON_TIME_GRACE_SECONDS 30.0 (0 to 300) How late a slot may be and still fire as on-time rather than as a missed slot handed to catch-up. A brain restart takes seconds, so a slot that comes due inside this window is late, not missed.
Z4J_SCHEDULER_RECONCILE_INTERVAL_SECONDS 900 (60 to 86400) Seconds between full re-syncs of the schedule cache from the brain. Also how long the watch stream must stay healthy before a drop clears its reconnect penalty.
Z4J_SCHEDULER_FIRE_TIMEOUT_SECONDS 10 (1 to 3600) Deadline of each FireSchedule attempt, and how long shutdown waits for fires already in flight.
Z4J_SCHEDULER_FIRE_RETRY_MAX 3 (0 to 10) Retries after the first attempt, on transient gRPC errors only.
Z4J_SCHEDULER_FIRE_RETRY_BACKOFF_SECONDS 1.0 (0 to 60) Base retry delay; it doubles per attempt up to a 10 second cap, with jitter.

Reverse trigger server, off by default. Turn it on only for a brain that calls out to the scheduler for run-now; a brain with durable schedule control fires operator triggers itself and refuses a trigger routed here:

Variable Default Effect
Z4J_SCHEDULER_TRIGGER_GRPC_ENABLED false Starts the mTLS trigger server.
Z4J_SCHEDULER_TRIGGER_GRPC_BIND_HOST 0.0.0.0 Bind interface of that server.
Z4J_SCHEDULER_TRIGGER_GRPC_BIND_PORT 7802 (0 to 65535) Bind port of that server.
Z4J_SCHEDULER_TRIGGER_GRPC_TLS_CERT none Server certificate PEM. Required when the server is enabled.
Z4J_SCHEDULER_TRIGGER_GRPC_TLS_KEY none Server private key PEM.
Z4J_SCHEDULER_TRIGGER_GRPC_TLS_CA none CA bundle that validates the brain's client certificate.
Z4J_SCHEDULER_TRIGGER_GRPC_ALLOWED_CNS [] JSON list of accepted client CNs. Empty trusts any certificate the CA validates, with a warning.
Z4J_SCHEDULER_TRIGGER_GRPC_REQUIRE_ALLOWLIST false Refuse to start when the allow-list is empty.
Z4J_SCHEDULER_TRIGGER_GRPC_GRACE_SECONDS 5.0 (0.1 to 60) Drain window for in-flight trigger calls at shutdown.

Accepted for configuration compatibility and read by nothing:

Variable Default Note
Z4J_SCHEDULER_DATABASE_URL none Deprecated and unread. Use Z4J_SCHEDULER_LEADER_PG_DSN.
Z4J_SCHEDULER_PROJECTS * Deprecated and unread. Project scope comes from the certificate's project binding.
Z4J_SCHEDULER_LEADER_POLL_INTERVAL_SECONDS 2 Deprecated and unread. Use Z4J_SCHEDULER_LEADER_HEARTBEAT_SECONDS.
Z4J_SCHEDULER_PRO_LICENSE_KEY none Reserved and unread. Setting it enables nothing.

The scheduler's own /metrics endpoint exposes:

Metric Labels What it says
z4j_scheduler_is_leader project 1 on the instance holding the lock. This is the failover signal.
z4j_scheduler_schedules_loaded project, kind Schedules in the in-memory cache. Recounted after every full sync and whenever a schedule is created, deleted, or changes kind or project.
z4j_scheduler_fires_total status Fires dispatched, by outcome.
z4j_scheduler_fire_latency_seconds FireSchedule round-trip time.
z4j_scheduler_fire_variance_seconds schedule_id, engine, project Difference between the due time and the dispatch.
z4j_scheduler_tick_drift_seconds How late the tick loop woke.
z4j_scheduler_slots_discarded_total catch_up Missed slots dropped by the catch-up policy.
z4j_scheduler_grpc_calls_total method, status Calls to the brain, by RPC method and gRPC status name (OK, UNAVAILABLE, DEADLINE_EXCEEDED, ...). A streaming call counts once, when its stream ends.
z4j_scheduler_watch_stream_reconnects_total Watch-stream reconnects.
z4j_scheduler_watch_healthy project 1 while the watch stream is connected, 0 while it is down or reconnecting. The tick engine stops dispatching and /ready turns 503 on the same signal. project is * when one stream covers every project the certificate binds.
z4j_scheduler_engine_iterations_total Tick-engine iterations.
z4j_scheduler_engine_iteration_failures_total Iterations that raised.
z4j_scheduler_per_schedule_fires_total schedule_id, schedule_name, status Per-schedule fire counts.
z4j_scheduler_per_schedule_fire_latency_seconds schedule_id, schedule_name Per-schedule latency.

The brain publishes z4j_scheduler_misfires_detected_total separately; see misfire detection.

You can put schedules into z4j-scheduler's database three ways:

The Schedules page in the dashboard (per-project) has a full CRUD UI: name, engine, kind (cron, interval, solar, clocked), expression, task name, args, kwargs, queue, catch-up policy. An audit row is written on every change. The REST surface backing the UI is the schedules API.

2. Declarative (in your app's startup hook)

Section titled “2. Declarative (in your app's startup hook)”

Commit your schedules in source. Reconciler posts the dict to brain on app startup; same shape across django, flask, fastapi:

from z4j_scheduler.declarative import ScheduleSpec, reconcile
await reconcile(
schedules=[
ScheduleSpec(
name="hourly-cleanup",
engine="celery",
kind="cron",
expression="0 * * * *",
task_name="myapp.tasks.cleanup",
),
],
project="my-app",
source="declarative",
brain_url="http://brain:7700",
api_token=settings.Z4J_BRAIN_API_TOKEN,
)

Re-running reconcile with the same dict changes nothing (every row counts as unchanged), but each call still writes one schedules.import audit row for the batch.

Import from any existing scheduler:

Terminal window
# Celery beat schedule -> z4j
z4j-scheduler import --from celery --celery-app myapp:app \
--project myproject --brain-url https://brain.internal
# rq-scheduler -> z4j
z4j-scheduler import --from rq --redis-url redis://... \
--project myproject --brain-url https://brain.internal
# APScheduler 3.x SQLAlchemy jobstore -> z4j (Redis and Mongo jobstores are not accepted)
z4j-scheduler import --from apscheduler --jobstore-url postgresql://... \
--project myproject --brain-url https://brain.internal
# system crontab -> z4j. --task-prefix names the task that runs each command
# line; --user-column is for /etc/crontab, which has a user field.
z4j-scheduler import --from cron --crontab /etc/crontab --user-column \
--task-prefix myapp.shell.exec_command \
--project myproject --brain-url https://brain.internal
# Huey periodic tasks -> z4j (reads the @huey.periodic_task registry)
z4j-scheduler import --from huey --huey-app myapp.tasks:huey \
--project myproject --brain-url https://brain.internal
# arq cron jobs -> z4j (reads WorkerSettings.cron_jobs)
z4j-scheduler import --from arq --arq-settings myapp.worker:WorkerSettings \
--project myproject --brain-url https://brain.internal
# taskiq schedule labels -> z4j (reads the broker's task registry)
z4j-scheduler import --from taskiq --taskiq-broker myapp.tkq:broker \
--project myproject --brain-url https://brain.internal

Each importer installs its source's engine through an extra: z4j-scheduler[celery-import], [rq-import], [apscheduler-import], [huey-import], [arq-import] or [taskiq-import]. The huey, arq and taskiq importers import your application module, so run them where that module and its engine resolve. --celery-app, --huey-app, --arq-settings and --taskiq-broker import and execute the named module, so the path is trusted code: point them only at a module you would run.

Round-trip: every importer pairs with an exporter (--to celery, rq, apscheduler, cron, huey, arq or taskiq; --to dramatiq prints guidance only). z4j-scheduler export --to celery --project myproject --brain-url https://brain.internal --out beat_schedule.py renders a Python beat_schedule file you can drop back into Celery if you ever decide to leave z4j. --out defaults to stdout, and --source / --scheduler filter which rows are exported.

Per-schedule field that decides what happens to a missed slot: one that is older than Z4J_SCHEDULER_ON_TIME_GRACE_SECONDS by the time the scheduler reaches it, after its own outage or the brain's. A slot still inside the grace is late, not missed, and fires under every policy: a 35 second outage on a 10 second interval, at the default 30 second grace, fires all three slots on recovery. A schedule whose interval is shorter than the grace can therefore burst up to the grace divided by the interval fires when the scheduler catches up; lower the grace if that burst is unwelcome.

  • skip - discard the missed slots; the next regular tick fires once.
  • fire_one_missed - fire once on recovery, then resume normal cadence. Right for "nightly report" semantics.
  • fire_all_missed - fire once for every missed slot. Right for "every-5-minute metric backfill" semantics. Slot computation is bounded at 10,000 missed slots, and the tick loop dispatches at most 256 fires per pass, so a long backlog is spread over several ticks rather than issued in one burst. A 12-hour outage on a per-minute schedule still means 720 fires.

Default: fire_one_missed for a schedule created from the dashboard or through POST /schedules without a catch_up value. Imported and declarative definitions default to skip unless they say otherwise, and an existing schedule keeps whatever policy it has.

A schedule that comes due while its previous run is still going fires anyway. There is no per-schedule overlap control: the overlap_policy field accepts only allow, and any other value is refused at write time rather than accepted and quietly ignored. If a run must not overlap itself, take a lock inside the task.

A misfire is an enabled interval or cron schedule whose expected next fire has come and gone without the scheduler firing it. The usual cause is the scheduler process being down or partitioned - which is exactly why detection runs brain-side, not scheduler-side: a scheduler-side check cannot report its own death. The scheduler's tick engine already covers the alive-but-behind case via each schedule's catch-up policy; the misfire detector covers the case the tick engine structurally cannot.

A periodic brain worker computes each enabled interval/cron schedule's expected next fire from its cadence (anchored on its last run, or its creation time when it has never fired) and flags any schedule whose expected fire is late past a grace window. Each detection, once per misfire episode:

  • writes a scheduler.misfire_detected row to the HMAC-chained audit log, naming the schedule, its expected fire, and how late it is;
  • fires any schedule.misfired automation rule on the project;
  • fans out to any schedule.misfired notification subscription (delivery channels + in-app bell);
  • bumps the z4j_scheduler_misfires_detected_total{project} Prometheus counter - a sustained non-zero rate means a scheduler is dead, partitioned, or badly behind.

When the scheduler recovers and the schedule fires, the episode ends; the next gap is a fresh episode and alerts again. Solar and one-shot (clocked) schedules are not misfire-detected: solar cadence depends on a location the brain does not hold, and a one-shot has no recurring slot to miss.

Recent misfires for one schedule are readable at GET /api/v1/projects/{slug}/schedules/{schedule_id}/misfires (viewer role); the project-wide history across every schedule is at GET /api/v1/projects/{slug}/schedules/misfires - see the schedules API. From a shell, the z4j misfires --project <slug> CLI command (with --json) prints the same project-wide history without the dashboard.

Knobs (detection only runs when Z4J_SCHEDULER_GRPC_ENABLED is on):

Variable Default Description
Z4J_SCHEDULER_MISFIRE_GRACE_SECONDS 60 Lateness past the expected fire before a schedule counts as misfired. The grace absorbs normal fire latency, gRPC jitter, and small clock skew. Range 5..3600.
Z4J_SCHEDULER_MISFIRE_SWEEP_SECONDS 60 Detector cadence. 0 disables misfire detection entirely.

Every fire the scheduler dispatches is recorded per-schedule: status (pending, accepted, delivered, buffered, acked_success, acked_failed, failed, or one of the terminal_* resolutions), scheduled-for vs fired-at, ack latency, and error detail, with manual "run now" fires attributed to the operator who triggered them. See Schedule fire history for the lifecycle, the API, and retention mechanics.

Terminal window
# 1. Install and prepare a disabled systemd unit. Do not start it yet.
pip install z4j-scheduler
# 2. Check the normalized schedule-set diff. This does not compare fire times.
z4j-scheduler import --from celery --celery-app myapp:app \
--project myproject --brain-url https://brain.internal --verify
# 3. Import while celery-beat remains the only process firing schedules.
z4j-scheduler import --from celery --celery-app myapp:app \
--project myproject --brain-url https://brain.internal
# 4. Verify the definitions in the dashboard. They cannot be confirmed as
# firing until z4j-scheduler becomes the one live scheduler.
# 5. Cut over without ever running both schedulers. If the replacement does
# not become ready, stop it and restore celery-beat immediately.
sudo systemctl stop celery-beat
if ! sudo systemctl start z4j-scheduler || \
! curl --fail --retry 10 --retry-connrefused --retry-delay 1 \
http://127.0.0.1:7800/ready; then
sudo systemctl stop z4j-scheduler
sudo systemctl start celery-beat
exit 1
fi
# 6. After observing successful fires, disable celery-beat at boot. Keep its
# configuration available until the rollback window closes.
sudo systemctl disable celery-beat
# 7. (Optional, after the rollback window) uninstall django-celery-beat if
# that separate package is no longer used. Celery workers still need Celery.
pip uninstall django-celery-beat

The example assumes the standalone service has already been configured with the environment shown above and exposes its readiness endpoint on loopback port 7800. For embedded mode, prepare and test the brain configuration first, then use the same stop-old, start-new, verify-ready, or restore-old sequence.

Coexistence with celery-beat - gradual migration

Section titled “Coexistence with celery-beat - gradual migration”

If you can't stop celery-beat in one shot (e.g., shared Postgres schedule table with another team's app), keep both running and use z4j-celerybeat as the coexistence adapter:

  • celery-beat continues firing its schedules.
  • z4j-celerybeat surfaces those celery-beat schedules in the z4j dashboard - read AND write (when django-celery-beat is installed).
  • z4j-scheduler can fire its OWN schedules alongside (different rows in z4j's database).
  • Your dashboard shows both. Every row carries a source label (dashboard by default; the CLI writes cli), and the scheduler field tells them apart: celery-beat for the adapter's rows, z4j-scheduler for z4j's own.

When you're ready to fully migrate, run the importer and disable celery-beat.

  • You only have one engine and one schedule. A single crontab line is simpler.
  • You explicitly want celery-beat's exact behavior (e.g., a custom celery-beat scheduler class your team wrote). z4j-celerybeat surfaces that scheduler's existing schedules; z4j-scheduler replaces it.
  • You've committed to APScheduler 4's persistence model. That model is in-process; z4j-scheduler is out-of-process. Different shape, different trade-offs.
  • Schedulers overview for the per-engine scheduler adapters (different surface; complementary).
  • The standalone CLI: z4j-scheduler --help lists serve, version, check, status, info, doctor, schedules add/list/trigger/disable/edit/history/enable, import, export, and restart (an informational stub that prints the supervisor command for your deploy shape and exits 1).