z4j-scheduler - the dynamic scheduler service
z4j-scheduler is a separate companion process that fires schedules
against any z4j-supported engine (Celery, RQ, Dramatiq, Huey, arq,
taskiq). Unlike the per-engine scheduler adapters in this section
(which surface celery-beat / rq-scheduler / APScheduler / etc. in the
dashboard), z4j-scheduler IS the scheduler - operators who choose it
delete celery-beat from their stack.
Each schedule's catch_up policy decides what happens to a missed
occurrence: one that falls due while the scheduler's watch of the brain is
unhealthy and is older than the on-time grace by the time the watch
recovers. An occurrence still inside the grace is late, not missed, and
fires under every policy; a brain restart takes seconds and usually fits.
For a missed one, skip discards it, logging a warning, counting it in
z4j_scheduler_slots_discarded_total and showing it on the schedule's
dashboard page as "Skipped in last 24h" (which counts gaps, not
occurrences: a three-slot gap is one skip there), while fire_one_missed
runs the latest missed occurrence as soon as the watch recovers. New schedules
created from the dashboard or through POST /schedules default to
fire_one_missed; choose skip for work that must never run late.
What it is
Section titled “What it is”A single-binary, leader-elected, gRPC-connected scheduler service. Three responsibilities:
- Tick - sleep until the earliest
next_fire_atin the in-memory schedule cache (fed by brain over the gRPC watch stream), waking early when the cache changes and at least every 30 seconds. - Dispatch - for each due fire, send a
FireSchedulegRPC frame to z4j. Brain dispatches aschedule.firecommand to an online agent of the project that connected over WebSocket and advertised the schedule's engine; a long-poll agent never qualifies, so a project served only by long-poll agents sees every fire buffered. - Reconcile - every 15 minutes (configurable), do a full sync from brain so a missed delete event in the watch stream doesn't leave a stale schedule firing forever.
Idle CPU is essentially zero. Memory is bounded by the number of schedules.
How it differs from the per-engine scheduler adapters
Section titled “How it differs from the per-engine scheduler adapters”The packages on the schedulers overview page
(z4j-celerybeat, z4j-rqscheduler, etc.) are adapters - they
surface an EXISTING native scheduler (celery-beat, rq-scheduler) in
the z4j dashboard so operators can see and edit those schedules in
one UI.
z4j-scheduler is a replacement - it owns the schedule storage
itself, in z4j's database, and dispatches across any engine.
Operators who run z4j-scheduler don't need celery-beat or
rq-scheduler at all.
Comparison matrix
Section titled “Comparison matrix”The marketing version of this matrix lives at z4j.com/schedulers/z4j-scheduler broken into six categorized tables. The terse engineering version follows.
Engine + framework reach
Section titled “Engine + framework reach”| Capability | celery-beat | django-celery-beat | rq-scheduler | APScheduler 4 | system cron | z4j-scheduler |
|---|---|---|---|---|---|---|
| Engines supported | Celery only | Celery only | RQ only | APScheduler in-process | Shell exec | All 6: Celery, RQ, Dramatiq, Huey, arq, taskiq |
| Framework agnostic | Yes | Django only | Yes | Yes | Yes (host OS) | Yes - Django/Flask/FastAPI/bare |
| Multiple engines, one process | No | No | No | No | No | Yes |
Operations & observability
Section titled “Operations & observability”| Capability | celery-beat | django-celery-beat | rq-scheduler | APScheduler 4 | system cron | z4j-scheduler |
|---|---|---|---|---|---|---|
| Edit live (no restart) | No | Django admin only | No | Persistent jobstore only | No | Dashboard / declarative / REST |
| Built-in dashboard | No | Django admin (basic) | No | No | No | Yes - fire history, run-now, edit |
| Fire history per schedule | No | No | No | No | syslog only | Buffered + acked + searchable |
| Audit log of edits | No | Django auditlog (3rd party) | No | No | No | HMAC-chained |
| Manual fire-now button | No | No | No | API only | No | Yes - dashboard + REST |
| RBAC / project scoping | No | Django auth only | No | No | UNIX permissions | Project + role-scoped |
Reliability & catch-up
Section titled “Reliability & catch-up”| Capability | celery-beat | django-celery-beat | rq-scheduler | APScheduler 4 | system cron | z4j-scheduler |
|---|---|---|---|---|---|---|
| HA / leader election | Single only | Single only | Single only | Pluggable (manual) | Single host | Postgres advisory-lock leader, rolling-restart safe |
| Catch-up on outage | All-or-nothing | All-or-nothing | Default fire-all | Coalescing only | No | Per-schedule: skip / fire-one-missed / fire-all-missed |
| DST / IANA tz correctness | Yes | Yes | Partial | Yes | Yes | Yes (validated at API) |
| Solar (sunrise/sunset) | Yes | Yes | No | No | No | Yes |
Migration & lock-in
Section titled “Migration & lock-in”| Capability | celery-beat | django-celery-beat | rq-scheduler | APScheduler 4 | system cron | z4j-scheduler |
|---|---|---|---|---|---|---|
| Importer FROM other schedulers | N/A | N/A | N/A | N/A | N/A | celery-beat, django-celery-beat, rq-scheduler, APScheduler, system crontab, Huey, arq, taskiq |
| Exporter TO other schedulers | N/A | N/A | N/A | N/A | N/A | celery-beat, rq-scheduler, APScheduler, crontab, Huey, arq, taskiq |
| Coexist with native scheduler | - | - | - | - | - | Yes - z4j-celerybeat coexistence adapter |
| Declarative-in-source | Yes (beat_schedule) | No (DB-only) | Manual | Yes | crontab file | Yes - z4j_scheduler.declarative reconciler |
The importer and exporter cover the schedulers people actually migrate off.
Huey periodic tasks, arq cron jobs and taskiq schedule labels live in your
source rather than in a store, so the importer reads them from an import path
(--huey-app, --arq-settings, --taskiq-broker) and the exporter renders a
Python module you paste back. The huey-periodic, arq-cron and taskiq-scheduler
adapters still surface them read-only in the dashboard if you keep the native
scheduler. Dramatiq has no native scheduler, so --from dramatiq and
--to dramatiq print migration guidance instead.
Security
Section titled “Security”| Capability | celery-beat | django-celery-beat | rq-scheduler | APScheduler 4 | system cron | z4j-scheduler |
|---|---|---|---|---|---|---|
| Wire HMAC + replay protection | Broker-dependent | Broker-dependent | No | No | N/A | HMAC-SHA256 + per-session seq+nonce binding |
| HMAC-chained audit log | No | No | No | No | No | Per-row HMAC, prev-hmac chain |
| Secret rotation without losing history | N/A | N/A | N/A | N/A | N/A | Multi-key verification window, but agents must be re-credentialed |
License + project shape
Section titled “License + project shape”| Capability | celery-beat | django-celery-beat | rq-scheduler | APScheduler 4 | system cron | z4j-scheduler |
|---|---|---|---|---|---|---|
| License | BSD-3 | BSD-3 | MIT | MIT | BSD | Apache-2.0 |
| Process model | Long-running daemon | Daemon + DB | Long-running daemon | In-process or daemon | init / systemd | Standalone OR brain-embedded subprocess |
| Idle CPU footprint | Low | Low | Low | Low | Negligible | Negligible (sleeps until the next fire; 16-worker dispatch pool) |
Architecture
Section titled “Architecture”┌─────────────────────────────────────────────┐│ Your app (Django/Flask/FastAPI/bare) ││ + Celery (or RQ / arq / Huey / etc.) ││ + z4j-X engine adapter ││ + z4j-bare agent runtime │└──────────────────┬──────────────────────────┘ │ WebSocket / longpoll ▼┌─────────────────────────────────────────────┐│ z4j - server + dashboard ││ z4j-scheduler - fires schedules at the ││ right time │└─────────────────────────────────────────────┘z4j-scheduler does not run tasks. It only decides WHEN they
should run and tells brain to dispatch them. Your existing engine
worker (celery, RQ worker, etc.) runs the tasks on its own broker.
This means:
- No new broker. Your existing Redis / RabbitMQ / Postgres / etc. stays as the message bus.
- No new worker process. Your existing celery worker / rq worker / arq worker continues running tasks unchanged.
- The only new thing is one scheduler process (or a few, for HA).
Fire identity and why failover is safe
Section titled “Fire identity and why failover is safe”Every fire carries an identity derived deterministically from the schedule and its exact scheduled slot. The same logical fire therefore always produces the same identity, no matter which instance produces it or how many times it is retried.
That single property is what makes running more than one instance safe. If an instance dies after dispatching a fire but before recording it, and another instance picks the schedule up and dispatches the same slot, the brain recognises the second dispatch as the same fire rather than a new one and deduplicates it. Without deterministic identity, every failover would be a coin flip between a missed fire and a duplicate one.
The retry rules follow from the same idea. Transient transport failures are retried with capped exponential backoff, because retrying is free when the identity is stable. Permission and precondition failures are not retried, because repeating them cannot change the answer.
High availability
Section titled “High availability”Leader election uses PostgreSQL advisory locks, which are session-scoped: when the holder's connection drops the lock is released automatically, and a standby's next poll acquires it.
Two backends are available:
Z4J_SCHEDULER_LEADER_BACKEND |
Behaviour |
|---|---|
single |
Always leader. Correct for one instance, and the only option without PostgreSQL. |
postgres |
One leader for the whole cluster. Standbys idle until failover. |
postgres_per_project |
One lock per project, so instances divide the projects between them. |
postgres_per_project is the one to reach for with more than a couple of
instances. Work is shared rather than parked in a hot standby, which also
means the failover path is exercised continuously instead of only during an
incident.
Takeover has no fixed response-time promise. A clean stop typically
fails over in one to three seconds and a hard partition in 30 to 60 seconds,
but both figures come from the environment, not from the scheduler. A killed
process closes its connection immediately, so the lock frees at once. A
partitioned host cannot send anything, so the lock is held until the
database's TCP keepalive gives up. Tune the keepalive on the PostgreSQL host,
and Z4J_SCHEDULER_LEADER_HEARTBEAT_SECONDS on the scheduler (default of two
seconds, anywhere from half a second to a minute; it sets both the leader's
liveness probe and the standby's acquire retry), if your schedules are tight
enough that a lost minute matters.
Z4J_SCHEDULER_LEADER_POLL_INTERVAL_SECONDS is still accepted for
compatibility but nothing reads it.
HA does not make task execution exactly-once. The brain deduplicates the dispatch of a slot by its fire identity; the task itself then runs on your engine's worker under that engine's own delivery semantics.
Outages, mismatches and shutdown
Section titled “Outages, mismatches and shutdown”The scheduler fails closed and says so. What each condition does:
- Watch stream down. The tick engine stops dispatching, withholds any
transient retry so a fire never reaches a brain it has not negotiated with,
and resumes once the stream reconnects and re-syncs.
z4j_scheduler_watch_healthydrops to 0 at once;/readyturns 503 only when the outage outlastsZ4J_SCHEDULER_ON_TIME_GRACE_SECONDS. A slot that came due in the gap and is still inside the grace when the stream is back fires late, under every catch-up policy; only a slot older than the grace is missed and goes to the schedule's catch-up policy. - Cadence mismatch mid-flight. A brain restarted during a fire's retry
window can answer
FIRE_CADENCE_SEMANTICS_MISMATCHfrom a runtime this instance never negotiated with. The first such answer holds the schedule locally, logs a WARNING and renegotiates the cadence contract; agreement lifts the hold and the slot is retried on the next tick. Only a mismatch that survives an agreed renegotiation is reported to the brain as a durable quarantine (cadence_semantics_mismatch), which disables the schedule until an operator lifts it. - Brain and scheduler on different pinned sets. Both sides hash the
libraries and timezone data that decide fire times, and the brain refuses a
scheduler whose hash differs: the watch stream fails with
FAILED_PRECONDITION,/readyanswers 503, and a scheduler started against such a brain exits. Nothing fires and nothing is recorded against the schedules; they resume when both run the same release, and the occurrences missed meanwhile follow the catch-up policy. The upgrade guide gives the order that avoids the window. - Brain shutdown. A brain that has begun shutting down admits no new
fire: it answers
FireSchedulewith gRPCUNAVAILABLE, which is what a scheduler sees from a brain that is gone, and writes nothing. The scheduler logs a transient error and retries the same slot when a brain is back. A fire accepted before the stop signal arrived is outside that gate: if the process ends before delivering it, the delivery timeout holds the schedule until an operator resolves it, which is why an upgrade stops the scheduler first. - Tombstone pressure. A project that accumulates 10,000 deletes since its last stable snapshot is paused: nothing in it fires until a covering snapshot lands. The pause logs a WARNING naming the project and the count and requests an immediate snapshot resync rather than waiting for the periodic one, which remains the backstop if that resync fails.
- A networked
singleleader.singleis not an election: every replica started with it considers itself the leader. At startup,singlebound off loopback outsidedevlogs one WARNING saying so, because that is the shape of a deployment that grows a second replica. The brain fences the duplicate dispatches (no task runs twice), but theFireScheduleload doubles and the replicas' catch-up decisions can diverge under clock skew. Run one replica, or elect through PostgreSQL. - Shutdown. On SIGTERM the process stops admitting slots, waits up to
Z4J_SCHEDULER_FIRE_TIMEOUT_SECONDSfor fires already sent to the brain to be acknowledged, then releases the leader gate and closes the gRPC channel. A fire still in flight past that bound is logged and abandoned; the brain fences a late acceptance, so a standby that takes over cannot double-fire it.
Verifying a migration before you commit to it
Section titled “Verifying a migration before you commit to it”The shipped importer can compare the normalized schedule set with the rows already in the brain. It cannot independently predict what celery-beat would fire, so it cannot certify timing equivalence.
# Static schedule-set diff. Nothing is written.export Z4J_SCHEDULER_BRAIN_API_TOKEN='<project-admin-token-from-secret-manager>'z4j-scheduler import --from celery --celery-app myapp.celery:app \ --project myproject --brain-url https://brain.internal --verify--verify prints insert, update, unchanged, and delete counts and implies
--dry-run. It does not compare fire times. The reserved --duration option
is deliberately fail-closed: --verify --duration 24h runs the static diff,
then exits 2 and says there is no independent source-side firing oracle. Do
not treat it as cutover certification.
Deploy modes
Section titled “Deploy modes”Embedded mode - single container
Section titled “Embedded mode - single container”For homelab / single-instance deploys, z4j can spawn z4j-scheduler as a supervised subprocess in its own lifespan. Auto-mints loopback mTLS at boot.
Z4J_EMBEDDED_SCHEDULER=trueZ4J_SCHEDULER_GRPC_ENABLED=trueThat's it - z4j spawns z4j-scheduler at startup and supervises it (bounded auto-restart, graceful SIGTERM). One container, no extra ops surface.
Standalone mode - production / HA
Section titled “Standalone mode - production / HA”For production, run z4j-scheduler as a separate process or container. Multiple instances elect a leader via Postgres advisory lock; only the leader ticks. Followers stay warm.
pip install z4j-scheduler
export Z4J_SCHEDULER_BRAIN_GRPC_URL=brain.internal:7701export Z4J_SCHEDULER_BRAIN_REST_URL=https://brain.internalexport Z4J_SCHEDULER_TLS_CERT=/etc/z4j/tls/scheduler-1.crtexport Z4J_SCHEDULER_TLS_KEY=/etc/z4j/tls/scheduler-1.keyexport Z4J_SCHEDULER_TLS_CA=/etc/z4j/tls/ca.crtexport Z4J_SCHEDULER_METRICS_AUTH_TOKEN='<from-secret-manager>'
# Required for multiple scheduler instances. A single instance may leave# Z4J_SCHEDULER_LEADER_BACKEND at its default of single.export Z4J_SCHEDULER_LEADER_BACKEND=postgresexport Z4J_SCHEDULER_LEADER_PG_DSN='postgresql://z4j:password@db.internal/z4j'
z4j-scheduler serveserve takes no command-line options. It reads Z4J_SCHEDULER_* settings, and
also a .env file in the working directory, and fails startup if required
endpoints, mTLS material, production metrics authentication, or the selected
leader backend are invalid. It exits 0 on a clean stop and 1 on a
configuration or startup failure.
The process serves four HTTP endpoints on Z4J_SCHEDULER_BIND_HOST and
Z4J_SCHEDULER_BIND_PORT:
/healthalways answers 200 once the process is up./readyanswers 503 while any ofbrain_client,cache_initial_syncorleader_gateis not ready, and names the ones that are not. It also answers 503 once the watch stream has been unhealthy for longer thanZ4J_SCHEDULER_ON_TIME_GRACE_SECONDS, namingwatch_unhealthywith the outage length in seconds, and answers 200 again as soon as the stream reconnects. A shorter reconnect is a blip the engine rides out on its own. Reconnect attempts back off exponentially up toZ4J_SCHEDULER_GRPC_RECONNECT_BACKOFF_MAX_SECONDS, and that penalty clears once the stream has stayed healthy for oneZ4J_SCHEDULER_RECONCILE_INTERVAL_SECONDS, so a drop after a long quiet stretch reconnects fast while a flapping stream keeps backing off./inforeports version, instance id, uptime, readiness, whether the watch stream is healthy and the number of schedules loaded./metricsis Prometheus text, behind a bearer token whenZ4J_SCHEDULER_METRICS_AUTH_TOKENis set, and absent entirely whenZ4J_SCHEDULER_METRICS_ENABLEDis false.
z4j-scheduler info reads that base URL from --url, or from
Z4J_SCHEDULER_INFO_URL, defaulting to http://localhost:7800.
z4j side enables the gRPC server with:
Z4J_SCHEDULER_GRPC_ENABLED=trueZ4J_SCHEDULER_GRPC_ALLOWED_CNS=["scheduler-prod","scheduler-staging"]Mutual TLS is required: z4j's gRPC server presents its server
cert; the scheduler presents a client cert whose CN must be in the
allow-list. The only plaintext path is a dev-only opt-in, refused unless the
environment setting on each side is exactly dev: Z4J_SCHEDULER_GRPC_INSECURE=true
on z4j (which also bypasses the CN allow-list) and
Z4J_SCHEDULER_INSECURE_GRPC=true on the scheduler.
Docker Compose - the shipped stacks
Section titled “Docker Compose - the shipped stacks”The Compose files shipped with the brain package start the scheduler next
to the brain, so a default deployment fires schedules with no further
setup. This holds for docker-compose.yml (SQLite) and
docker-compose.postgres.yml (PostgreSQL), with docker-compose.caddy.yml
layered on either. Three services take part:
scheduler-certsruns first and exits. It is the brain image runningdocker/scheduler-certs.sh(POSIX sh and openssl, nothing else) as root. It mints a private CA, a server certificate for the brain (DNS namesz4j,brainandlocalhost) and a client certificate for the scheduler, writes them into thez4j_scheduler_pkivolume, and hands them to the uid the brain and the scheduler run as. It is idempotent: a valid bundle is left alone, and a certificate is minted again only when it is missing, does not chain to the CA, does not match its key, or is within 30 days of expiry. Restart the brain and the scheduler after a renewal. The CA key stays readable by root only.z4jserves the scheduler gRPC channel on port 7701 of the private Compose network over mutual TLS, withZ4J_SCHEDULER_GRPC_TLS_CERT,Z4J_SCHEDULER_GRPC_TLS_KEYandZ4J_SCHEDULER_GRPC_TLS_CApointing at the minted files,Z4J_SCHEDULER_GRPC_ALLOWED_CNSnaming exactly one client identity, andZ4J_SCHEDULER_GRPC_REQUIRE_ALLOWLIST=true. The port is never published to the host.scheduleris the same image with its entrypoint switched toz4j-scheduler serve. It presents the client certificate (CNz4j-scheduler), starts once the brain reports healthy, and answers/health,/readyand/infoon port 7800 inside the network, where the dashboard's Schedulers page reaches it throughZ4J_SCHEDULER_INFO_URLS.
The SQLite stack runs one scheduler with
Z4J_SCHEDULER_LEADER_BACKEND=single. That scheduler binds every interface
of its container outside dev, so it logs the single-leader WARNING
described under "Outages, mismatches and shutdown" at every start; for a
one-replica stack that is expected and needs no change. The PostgreSQL stack runs two
replicas with Z4J_SCHEDULER_LEADER_BACKEND=postgres;
Z4J_SCHEDULER_LEADER_PG_DSN defaults to the stack's own database, built
from the same POSTGRES_USER, POSTGRES_PASSWORD and POSTGRES_DB values
the brain uses, so the pair elects a leader through the advisory lock and
the standby takes over when the leader stops. Add replicas with:
docker compose -f docker-compose.postgres.yml up -d --scale scheduler=3Every replica presents the same client certificate: the allow-list names a
role, not an instance, and the instance id defaults to the container
hostname. A replica beyond the second appears on the dashboard only once
its /info URL is added to Z4J_SCHEDULER_INFO_URLS on the brain.
Settings you may put in .env:
Z4J_PKI_SCHEDULER_CNchanges the minted CN and the brain's allow-list together.Z4J_SCHEDULER_METRICS_ENABLED=truetogether withZ4J_SCHEDULER_METRICS_AUTH_TOKENturns/metricson. The stacks ship it off because the scheduler binds every interface of its container and refuses to serve unauthenticated metrics outsidedev. Set both: an empty token locks the endpoint rather than opening it.Z4J_SCHEDULER_LEADER_PG_DSN, when the PostgreSQL password contains URL delimiters or the lock should live in another database.
To bring your own PKI, replace the files in the z4j_scheduler_pki volume
with your CA certificate and the two certificate and key pairs, keeping the
file names. To add a client identity signed by the generated CA, for a
scheduler that runs outside the stack, mint it from the one-shot's
container and allow-list the name on the brain:
docker compose run --rm --user 0:0 --entrypoint z4j scheduler-certs \ mint-scheduler-cert --name scheduler-ext --ca-cert /pki/ca.crt \ --ca-key /pki/ca.key --out-dir /pkiSettings
Section titled “Settings”Every setting is read from the environment with the Z4J_SCHEDULER_ prefix, or
from a .env file in the working directory.
Brain connection:
| Variable | Default | Effect |
|---|---|---|
Z4J_SCHEDULER_BRAIN_GRPC_URL |
required | Host and port of the brain's scheduler gRPC server. |
Z4J_SCHEDULER_BRAIN_REST_URL |
required | Read by check, doctor and status only; the running service never calls it. |
Z4J_SCHEDULER_TLS_CERT |
none | Client certificate PEM presented to the brain. |
Z4J_SCHEDULER_TLS_KEY |
none | Client private key PEM. |
Z4J_SCHEDULER_TLS_CA |
none | CA bundle that signed the brain's server certificate. |
Z4J_SCHEDULER_INSECURE_GRPC |
false |
Plaintext gRPC. Refused unless Z4J_SCHEDULER_ENVIRONMENT is exactly dev, and the brain must also allow it. |
Z4J_SCHEDULER_GRPC_KEEPALIVE_SECONDS |
30 (5 to 300) |
gRPC keepalive ping interval. |
Z4J_SCHEDULER_GRPC_RECONNECT_BACKOFF_MAX_SECONDS |
30 (1 to 300) |
Maximum backoff between watch-stream reconnects. |
Identity, HTTP surface and logging:
| Variable | Default | Effect |
|---|---|---|
Z4J_SCHEDULER_INSTANCE_ID |
hostname | Label shown in logs and on /info. |
Z4J_SCHEDULER_BIND_HOST |
0.0.0.0 |
Interface for the HTTP surface. A loopback bind exempts /metrics from the token requirement. |
Z4J_SCHEDULER_BIND_PORT |
7800 (1 to 65535) |
Port for the HTTP surface. |
Z4J_SCHEDULER_ENVIRONMENT |
production |
Only the exact value dev relaxes the plaintext-gRPC and metrics-token gates. Any other value, including staging, is production posture. |
Z4J_SCHEDULER_LOG_LEVEL |
INFO |
One of DEBUG, INFO, WARNING, ERROR, CRITICAL. |
Z4J_SCHEDULER_LOG_JSON |
true |
JSON log lines when true. |
Z4J_SCHEDULER_METRICS_ENABLED |
true |
Mounts /metrics; false leaves it unmounted. |
Z4J_SCHEDULER_METRICS_AUTH_TOKEN |
none | Bearer token for /metrics. Startup refuses a non-loopback bind without it outside dev. |
Leader election:
| Variable | Default | Effect |
|---|---|---|
Z4J_SCHEDULER_LEADER_BACKEND |
single |
single, postgres or postgres_per_project. |
Z4J_SCHEDULER_LEADER_PG_DSN |
none | PostgreSQL DSN for the advisory lock. Both Postgres backends refuse to start without it. |
Z4J_SCHEDULER_LEADER_NAMESPACE |
z4j-scheduler-global |
Hashed into the advisory-lock key. Two clusters sharing one database must use different namespaces. |
Z4J_SCHEDULER_LEADER_HEARTBEAT_SECONDS |
2.0 (0.5 to 60) |
Cadence of the leader's liveness probe and the standby's acquire retry. |
Tick and dispatch:
| Variable | Default | Effect |
|---|---|---|
Z4J_SCHEDULER_ON_TIME_GRACE_SECONDS |
30.0 (0 to 300) |
How late a slot may be and still fire as on-time rather than as a missed slot handed to catch-up. A brain restart takes seconds, so a slot that comes due inside this window is late, not missed. |
Z4J_SCHEDULER_RECONCILE_INTERVAL_SECONDS |
900 (60 to 86400) |
Seconds between full re-syncs of the schedule cache from the brain. Also how long the watch stream must stay healthy before a drop clears its reconnect penalty. |
Z4J_SCHEDULER_FIRE_TIMEOUT_SECONDS |
10 (1 to 3600) |
Deadline of each FireSchedule attempt, and how long shutdown waits for fires already in flight. |
Z4J_SCHEDULER_FIRE_RETRY_MAX |
3 (0 to 10) |
Retries after the first attempt, on transient gRPC errors only. |
Z4J_SCHEDULER_FIRE_RETRY_BACKOFF_SECONDS |
1.0 (0 to 60) |
Base retry delay; it doubles per attempt up to a 10 second cap, with jitter. |
Reverse trigger server, off by default. Turn it on only for a brain that calls out to the scheduler for run-now; a brain with durable schedule control fires operator triggers itself and refuses a trigger routed here:
| Variable | Default | Effect |
|---|---|---|
Z4J_SCHEDULER_TRIGGER_GRPC_ENABLED |
false |
Starts the mTLS trigger server. |
Z4J_SCHEDULER_TRIGGER_GRPC_BIND_HOST |
0.0.0.0 |
Bind interface of that server. |
Z4J_SCHEDULER_TRIGGER_GRPC_BIND_PORT |
7802 (0 to 65535) |
Bind port of that server. |
Z4J_SCHEDULER_TRIGGER_GRPC_TLS_CERT |
none | Server certificate PEM. Required when the server is enabled. |
Z4J_SCHEDULER_TRIGGER_GRPC_TLS_KEY |
none | Server private key PEM. |
Z4J_SCHEDULER_TRIGGER_GRPC_TLS_CA |
none | CA bundle that validates the brain's client certificate. |
Z4J_SCHEDULER_TRIGGER_GRPC_ALLOWED_CNS |
[] |
JSON list of accepted client CNs. Empty trusts any certificate the CA validates, with a warning. |
Z4J_SCHEDULER_TRIGGER_GRPC_REQUIRE_ALLOWLIST |
false |
Refuse to start when the allow-list is empty. |
Z4J_SCHEDULER_TRIGGER_GRPC_GRACE_SECONDS |
5.0 (0.1 to 60) |
Drain window for in-flight trigger calls at shutdown. |
Accepted for configuration compatibility and read by nothing:
| Variable | Default | Note |
|---|---|---|
Z4J_SCHEDULER_DATABASE_URL |
none | Deprecated and unread. Use Z4J_SCHEDULER_LEADER_PG_DSN. |
Z4J_SCHEDULER_PROJECTS |
* |
Deprecated and unread. Project scope comes from the certificate's project binding. |
Z4J_SCHEDULER_LEADER_POLL_INTERVAL_SECONDS |
2 |
Deprecated and unread. Use Z4J_SCHEDULER_LEADER_HEARTBEAT_SECONDS. |
Z4J_SCHEDULER_PRO_LICENSE_KEY |
none | Reserved and unread. Setting it enables nothing. |
Metrics
Section titled “Metrics”The scheduler's own /metrics endpoint exposes:
| Metric | Labels | What it says |
|---|---|---|
z4j_scheduler_is_leader |
project |
1 on the instance holding the lock. This is the failover signal. |
z4j_scheduler_schedules_loaded |
project, kind |
Schedules in the in-memory cache. Recounted after every full sync and whenever a schedule is created, deleted, or changes kind or project. |
z4j_scheduler_fires_total |
status |
Fires dispatched, by outcome. |
z4j_scheduler_fire_latency_seconds |
FireSchedule round-trip time. |
|
z4j_scheduler_fire_variance_seconds |
schedule_id, engine, project |
Difference between the due time and the dispatch. |
z4j_scheduler_tick_drift_seconds |
How late the tick loop woke. | |
z4j_scheduler_slots_discarded_total |
catch_up |
Missed slots dropped by the catch-up policy. |
z4j_scheduler_grpc_calls_total |
method, status |
Calls to the brain, by RPC method and gRPC status name (OK, UNAVAILABLE, DEADLINE_EXCEEDED, ...). A streaming call counts once, when its stream ends. |
z4j_scheduler_watch_stream_reconnects_total |
Watch-stream reconnects. | |
z4j_scheduler_watch_healthy |
project |
1 while the watch stream is connected, 0 while it is down or reconnecting. The tick engine stops dispatching and /ready turns 503 on the same signal. project is * when one stream covers every project the certificate binds. |
z4j_scheduler_engine_iterations_total |
Tick-engine iterations. | |
z4j_scheduler_engine_iteration_failures_total |
Iterations that raised. | |
z4j_scheduler_per_schedule_fires_total |
schedule_id, schedule_name, status |
Per-schedule fire counts. |
z4j_scheduler_per_schedule_fire_latency_seconds |
schedule_id, schedule_name |
Per-schedule latency. |
The brain publishes z4j_scheduler_misfires_detected_total separately; see
misfire detection.
Schedule sources
Section titled “Schedule sources”You can put schedules into z4j-scheduler's database three ways:
1. Dashboard
Section titled “1. Dashboard”The Schedules page in the dashboard (per-project) has a full CRUD UI: name, engine, kind (cron, interval, solar, clocked), expression, task name, args, kwargs, queue, catch-up policy. An audit row is written on every change. The REST surface backing the UI is the schedules API.
2. Declarative (in your app's startup hook)
Section titled “2. Declarative (in your app's startup hook)”Commit your schedules in source. Reconciler posts the dict to brain on app startup; same shape across django, flask, fastapi:
from z4j_scheduler.declarative import ScheduleSpec, reconcile
await reconcile( schedules=[ ScheduleSpec( name="hourly-cleanup", engine="celery", kind="cron", expression="0 * * * *", task_name="myapp.tasks.cleanup", ), ], project="my-app", source="declarative", brain_url="http://brain:7700", api_token=settings.Z4J_BRAIN_API_TOKEN,)Re-running reconcile with the same dict changes nothing (every row
counts as unchanged), but each call still writes one schedules.import
audit row for the batch.
3. One-shot importers (migration)
Section titled “3. One-shot importers (migration)”Import from any existing scheduler:
# Celery beat schedule -> z4jz4j-scheduler import --from celery --celery-app myapp:app \ --project myproject --brain-url https://brain.internal
# rq-scheduler -> z4jz4j-scheduler import --from rq --redis-url redis://... \ --project myproject --brain-url https://brain.internal
# APScheduler 3.x SQLAlchemy jobstore -> z4j (Redis and Mongo jobstores are not accepted)z4j-scheduler import --from apscheduler --jobstore-url postgresql://... \ --project myproject --brain-url https://brain.internal
# system crontab -> z4j. --task-prefix names the task that runs each command# line; --user-column is for /etc/crontab, which has a user field.z4j-scheduler import --from cron --crontab /etc/crontab --user-column \ --task-prefix myapp.shell.exec_command \ --project myproject --brain-url https://brain.internal
# Huey periodic tasks -> z4j (reads the @huey.periodic_task registry)z4j-scheduler import --from huey --huey-app myapp.tasks:huey \ --project myproject --brain-url https://brain.internal
# arq cron jobs -> z4j (reads WorkerSettings.cron_jobs)z4j-scheduler import --from arq --arq-settings myapp.worker:WorkerSettings \ --project myproject --brain-url https://brain.internal
# taskiq schedule labels -> z4j (reads the broker's task registry)z4j-scheduler import --from taskiq --taskiq-broker myapp.tkq:broker \ --project myproject --brain-url https://brain.internalEach importer installs its source's engine through an extra:
z4j-scheduler[celery-import], [rq-import], [apscheduler-import],
[huey-import], [arq-import] or [taskiq-import]. The huey, arq and taskiq
importers import your application module, so run them where that module and
its engine resolve. --celery-app, --huey-app, --arq-settings and
--taskiq-broker import and execute the named module, so the path is trusted
code: point them only at a module you would run.
Round-trip: every importer pairs with an exporter (--to celery, rq,
apscheduler, cron, huey, arq or taskiq; --to dramatiq prints
guidance only). z4j-scheduler export --to celery --project myproject --brain-url https://brain.internal --out beat_schedule.py renders a Python
beat_schedule file you can drop back into Celery if you ever decide to leave
z4j. --out defaults to stdout, and --source / --scheduler filter which
rows are exported.
Catch-up policy
Section titled “Catch-up policy”Per-schedule field that decides what happens to a missed slot: one that is
older than Z4J_SCHEDULER_ON_TIME_GRACE_SECONDS by the time the scheduler
reaches it, after its own outage or the brain's. A slot still inside the
grace is late, not missed, and fires under every policy: a 35 second outage
on a 10 second interval, at the default 30 second grace, fires all three
slots on recovery. A schedule whose interval is shorter than the grace can
therefore burst up to the grace divided by the interval fires when the
scheduler catches up; lower the grace if that burst is unwelcome.
skip- discard the missed slots; the next regular tick fires once.fire_one_missed- fire once on recovery, then resume normal cadence. Right for "nightly report" semantics.fire_all_missed- fire once for every missed slot. Right for "every-5-minute metric backfill" semantics. Slot computation is bounded at 10,000 missed slots, and the tick loop dispatches at most 256 fires per pass, so a long backlog is spread over several ticks rather than issued in one burst. A 12-hour outage on a per-minute schedule still means 720 fires.
Default: fire_one_missed for a schedule created from the dashboard or
through POST /schedules without a catch_up value. Imported and
declarative definitions default to skip unless they say otherwise, and
an existing schedule keeps whatever policy it has.
Overlap
Section titled “Overlap”A schedule that comes due while its previous run is still going fires anyway.
There is no per-schedule overlap control: the overlap_policy field accepts
only allow, and any other value is refused at write time rather than accepted
and quietly ignored. If a run must not overlap itself, take a lock inside the
task.
Misfire detection
Section titled “Misfire detection”A misfire is an enabled interval or cron schedule whose expected
next fire has come and gone without the scheduler firing it. The usual
cause is the scheduler process being down or partitioned - which is
exactly why detection runs brain-side, not scheduler-side: a
scheduler-side check cannot report its own death. The scheduler's tick
engine already covers the alive-but-behind case via each schedule's
catch-up policy; the misfire detector covers the case the tick engine
structurally cannot.
A periodic brain worker computes each enabled interval/cron schedule's expected next fire from its cadence (anchored on its last run, or its creation time when it has never fired) and flags any schedule whose expected fire is late past a grace window. Each detection, once per misfire episode:
- writes a
scheduler.misfire_detectedrow to the HMAC-chained audit log, naming the schedule, its expected fire, and how late it is; - fires any
schedule.misfiredautomation rule on the project; - fans out to any
schedule.misfirednotification subscription (delivery channels + in-app bell); - bumps the
z4j_scheduler_misfires_detected_total{project}Prometheus counter - a sustained non-zero rate means a scheduler is dead, partitioned, or badly behind.
When the scheduler recovers and the schedule fires, the episode ends;
the next gap is a fresh episode and alerts again. Solar and one-shot
(clocked) schedules are not misfire-detected: solar cadence depends
on a location the brain does not hold, and a one-shot has no recurring
slot to miss.
Recent misfires for one schedule are readable at
GET /api/v1/projects/{slug}/schedules/{schedule_id}/misfires
(viewer role); the project-wide history across every schedule is at
GET /api/v1/projects/{slug}/schedules/misfires - see the
schedules API. From a shell, the
z4j misfires --project <slug> CLI command
(with --json) prints the same project-wide history without the
dashboard.
Knobs (detection only runs when Z4J_SCHEDULER_GRPC_ENABLED is on):
| Variable | Default | Description |
|---|---|---|
Z4J_SCHEDULER_MISFIRE_GRACE_SECONDS |
60 |
Lateness past the expected fire before a schedule counts as misfired. The grace absorbs normal fire latency, gRPC jitter, and small clock skew. Range 5..3600. |
Z4J_SCHEDULER_MISFIRE_SWEEP_SECONDS |
60 |
Detector cadence. 0 disables misfire detection entirely. |
Fire history
Section titled “Fire history”Every fire the scheduler dispatches is recorded per-schedule: status (pending, accepted, delivered, buffered, acked_success, acked_failed, failed, or one of the terminal_* resolutions), scheduled-for vs fired-at, ack latency, and error detail, with manual "run now" fires attributed to the operator who triggered them. See Schedule fire history for the lifecycle, the API, and retention mechanics.
Replacing celery-beat - concrete steps
Section titled “Replacing celery-beat - concrete steps”# 1. Install and prepare a disabled systemd unit. Do not start it yet.pip install z4j-scheduler
# 2. Check the normalized schedule-set diff. This does not compare fire times.z4j-scheduler import --from celery --celery-app myapp:app \ --project myproject --brain-url https://brain.internal --verify
# 3. Import while celery-beat remains the only process firing schedules.z4j-scheduler import --from celery --celery-app myapp:app \ --project myproject --brain-url https://brain.internal
# 4. Verify the definitions in the dashboard. They cannot be confirmed as# firing until z4j-scheduler becomes the one live scheduler.
# 5. Cut over without ever running both schedulers. If the replacement does# not become ready, stop it and restore celery-beat immediately.sudo systemctl stop celery-beatif ! sudo systemctl start z4j-scheduler || \ ! curl --fail --retry 10 --retry-connrefused --retry-delay 1 \ http://127.0.0.1:7800/ready; then sudo systemctl stop z4j-scheduler sudo systemctl start celery-beat exit 1fi
# 6. After observing successful fires, disable celery-beat at boot. Keep its# configuration available until the rollback window closes.sudo systemctl disable celery-beat
# 7. (Optional, after the rollback window) uninstall django-celery-beat if# that separate package is no longer used. Celery workers still need Celery.pip uninstall django-celery-beatThe example assumes the standalone service has already been configured with the environment shown above and exposes its readiness endpoint on loopback port 7800. For embedded mode, prepare and test the brain configuration first, then use the same stop-old, start-new, verify-ready, or restore-old sequence.
Coexistence with celery-beat - gradual migration
Section titled “Coexistence with celery-beat - gradual migration”If you can't stop celery-beat in one shot (e.g., shared Postgres
schedule table with another team's app), keep both running and use
z4j-celerybeat as the coexistence adapter:
- celery-beat continues firing its schedules.
- z4j-celerybeat surfaces those celery-beat schedules in the z4j
dashboard - read AND write (when
django-celery-beatis installed). - z4j-scheduler can fire its OWN schedules alongside (different rows in z4j's database).
- Your dashboard shows both. Every row carries a
sourcelabel (dashboardby default; the CLI writescli), and theschedulerfield tells them apart:celery-beatfor the adapter's rows,z4j-schedulerfor z4j's own.
When you're ready to fully migrate, run the importer and disable celery-beat.
When NOT to use it
Section titled “When NOT to use it”- You only have one engine and one schedule. A single
crontabline is simpler. - You explicitly want celery-beat's exact behavior (e.g., a custom celery-beat scheduler class your team wrote). z4j-celerybeat surfaces that scheduler's existing schedules; z4j-scheduler replaces it.
- You've committed to APScheduler 4's persistence model. That model is in-process; z4j-scheduler is out-of-process. Different shape, different trade-offs.
See also
Section titled “See also”- Schedulers overview for the per-engine scheduler adapters (different surface; complementary).
- The standalone CLI:
z4j-scheduler --helplistsserve,version,check,status,info,doctor,schedules add/list/trigger/disable/edit/history/enable,import,export, andrestart(an informational stub that prints the supervisor command for your deploy shape and exits 1).