Audit webhook forwarding
The brain can optionally forward every audit-log row to an external webhook. It forwards what the audit log holds, which is successful control-plane actions plus the denials the brain records (schedule write refusals, MFA enforcement and verification denials, password-reset failures, bulk-retry refusals, source-address allowlist denials as auth.ip_denied, and an agent of an archived project as agent.auth.project_inactive; on the agent surfaces those are one row per address or agent per ten minutes). An ordinary role-based 403 on any other endpoint is never an audit row in the first place, so it can never reach the mirror; pair the forwarder with your reverse-proxy access logs if your SIEM needs those refused attempts too. The primary use cases:
- SIEM ingest: Splunk HEC, Datadog Logs, Sumo Logic, Elastic. The receiver gets the stored audit-row fields plus the row HMAC.
- A copy on a separate trust boundary: the receiver keeps a copy outside the brain's database, so an attacker who compromises the brain's database cannot reach back and rewrite what the receiver already holds. What the forwarder cannot promise is that the receiver holds everything, because the cursor it delivers from lives in the same database. Read Not a substitute for a durable sink before relying on this for evidence.
- Compliance evidence: SOC 2 / ISO 27001 auditors often want audit data in a logging stack they already control.
The brain's primary HMAC-chained audit log remains the source of truth. The forwarder delivers from a durable cursor: a receiver that is down is a delay, not a gap. Every row written after the forwarder was enabled reaches the receiver, in the order it was written, at least once.
Enabling
Section titled “Enabling”# In your env file or systemd unitZ4J_AUDIT_WEBHOOK_URL=https://siem.example.com/ingestZ4J_AUDIT_WEBHOOK_HMAC_SECRET=<32+ byte random string>Replace the example with a hostname whose DNS records resolve to public addresses. Private, loopback, link-local, and other blocked address classes are always rejected; the HTTP opt-in changes the permitted scheme, not the destination-address policy.
Restart the brain. A healthy start logs nothing for the forwarder; confirm it with GET /api/v1/admin/audit-forwarder (instance admin, session only, see Observability), the z4j_audit_forward_lag_rows gauge, or a test row at the receiver. A misconfigured URL logs:
WARNING z4j audit_forwarder: configured URL fails SSRF validation (<reason>); no row will be delivered until it is fixed, and every row waits at the cursor meanwhile. Fix Z4J_AUDIT_WEBHOOK_URL and restart.That is the URL failing the SSRF pre-flight (for example loopback, RFC1918, or plaintext HTTP without Z4J_NOTIFICATIONS_WEBHOOK_ALLOW_HTTP). There is no private-address bypass. Nothing is lost while it stays wrong: rows accumulate past the cursor and are delivered once the URL is fixed and the brain restarted.
Enabling the forwarder mirrors rows written from that moment on. The cursor is created at the audit head on the worker's first pass, so an existing brain does not replay its retained history into the receiver. To backfill, see Backfilling or skipping.
Settings
Section titled “Settings”| Variable | Default | Notes |
|---|---|---|
Z4J_AUDIT_WEBHOOK_URL |
unset | Receiver URL. Empty / unset disables the forwarder entirely. SecretStr at the Pydantic layer so a path-embedded token does not land in startup logs. |
Z4J_AUDIT_WEBHOOK_HMAC_SECRET |
unset | REQUIRED when the URL is set. At least 32 bytes. The brain refuses to start if the URL is set without an HMAC secret -- an unauthenticated mirror is worse than no mirror because downstream parsers may trust it implicitly. Mint with python -c "import secrets; print(secrets.token_urlsafe(48))". |
Z4J_AUDIT_WEBHOOK_TIMEOUT_SECONDS |
10.0 |
Per-row POST timeout, range 1.0..120.0. A slow receiver does NOT block the brain's audit write path; the forwarder is a background worker reading the audit log, not a hook on the write. |
Z4J_AUDIT_WEBHOOK_BATCH_SIZE |
100 |
Rows read past the cursor per pass, each POSTed in order. A pass that delivers a full batch runs again at once, so this bounds how long one leader-locked pass can take, not throughput. Range 1..1000. |
Z4J_AUDIT_WEBHOOK_POLL_INTERVAL_SECONDS |
5.0 |
How often the worker looks for new rows when it is caught up. Delivery latency on a quiet brain is at most this. Range 1..300. |
Z4J_AUDIT_WEBHOOK_MAX_BACKOFF_SECONDS |
300.0 |
Ceiling on the wait between attempts while the receiver keeps failing. The wait starts at one second and doubles per consecutive failure; a success resets it. Range 1..3600. |
Z4J_AUDIT_WEBHOOK_BUFFER_SIZE |
1000 |
Inert. The forwarder no longer buffers rows in memory. Still accepted so existing configurations keep loading; the value is read by nothing. |
Wire format
Section titled “Wire format”The receiver gets a POST request with the row as canonical JSON:
POST /your/path HTTP/1.1Host: siem.example.comContent-Type: application/jsonX-Z4J-Audit-Signature: sha256=<hex>X-Z4J-Audit-Timestamp: 1715515200X-Z4J-Audit-Schema: 1
{"action":"user.password_changed","api_key_id":null,"chain_generation":"...","event_id":null,"hmac_key_id":"...","hmac_version":2,"id":"...","legacy_frozen":false,"legacy_integrity_class":null,"legacy_origin":null,"metadata":{"key":"val"},"occurred_at":"2026-05-12T12:00:00.000000+00:00","outcome":"allow","prev_row_hmac":"...","project_id":null,"result":"success","row_hmac":"...","source_ip":"192.0.2.10","target_id":"user-1","target_type":"user","user_agent":"z4j-cli/1","user_id":"..."}One row per request. Fields are emitted in JSON-sorted-keys order so the signature is reproducible. The body matches the brain's internal audit row, plus the chain fields row_hmac, hmac_version, hmac_key_id, chain_generation, legacy_frozen, legacy_integrity_class and legacy_origin. To re-verify that chain in your own pipeline a receiver needs Z4J_AUDIT_CHAIN_SECRET, the dedicated audit-chain key, not Z4J_SECRET. An earlier version of this page named the wrong one, which would not have worked and would have handed an external receiver the master key that derives every agent's frame-signing key and encrypts stored TOTP secrets. Share the audit-chain key alone, and only with a receiver you would trust to verify your audit history.
Verifying the signature
Section titled “Verifying the signature”Each POST carries TWO headers:
X-Z4J-Audit-Signature: sha256=<hex>-- the HMAC digestX-Z4J-Audit-Timestamp: <unix_seconds>-- when the signature was minted
The signature is computed over the bytes <timestamp>.<body>. Folding the timestamp into the HMAC input gives replay-resistance: a captured POST replayed later still has its original signature, but the timestamp is stale so a receiver enforcing a skew window rejects it.
Python receiver (with the de-duplication at-least-once delivery requires):
import hmac, hashlib, json, time
SKEW_SECONDS = 300 # 5 minute window; tune to your fleet
# Production: replace this in-memory set with a Redis SETNX or a DB# unique-index insert. Audit row IDs are UUIDs (~10^-37 collision# probability), so a permanent dedupe table is bounded by your audit# retention window._seen_ids: set[str] = set()
def verify(body: bytes, headers: dict, secret: bytes) -> bool: sig = headers.get("X-Z4J-Audit-Signature", "") ts = headers.get("X-Z4J-Audit-Timestamp", "") if not sig or not ts: return False # Reject stale / future timestamps before the constant-time compare # so an attacker cannot use the verify call itself as a clock oracle. try: ts_int = int(ts) except ValueError: return False if abs(int(time.time()) - ts_int) > SKEW_SECONDS: return False digest_input = ts.encode("utf-8") + b"." + body expected = "sha256=" + hmac.new(secret, digest_input, hashlib.sha256).hexdigest() if not hmac.compare_digest(expected, sig): return False # De-duplicate on the row id once the signature verifies. The brain # delivers at least once: a row whose 2xx the brain did not see is # sent again with a fresh timestamp and signature, and a replayed # POST inside the skew window carries the same id. Reject repeats # here so the downstream pipeline never inserts the row twice. try: row_id = json.loads(body).get("id") except Exception: return False if not row_id or row_id in _seen_ids: return False _seen_ids.add(row_id) return TrueIf verification fails for any reason, reject the request with 401 Unauthorized. Do NOT parse the body before verifying the signature. A duplicate you have already stored is fine to answer with a 2xx as well as a 4xx: the brain only moves on from a row on a 2xx, so a receiver that answers duplicates with 401 keeps the brain retrying that row until the skew window has it give a different answer. Answering a verified duplicate with 200 OK and not storing it is the cleanest contract.
HMAC keys across brain replicas
Section titled “HMAC keys across brain replicas”Brain replicas feeding the same receiver normally share one Z4J_AUDIT_WEBHOOK_HMAC_SECRET. The signature covers <timestamp>.<body>, not a brain identity, and the signed audit-row id is the correct deduplication key. A reverse-proxy-injected source header may be useful operational metadata, but it is not cryptographically authenticated by this signature and must not be used as proof of which replica emitted a row.
Use distinct secrets only when replicas send through separately authenticated routes and the receiver can select the right verification key from trusted out-of-band routing context. The wire format itself contains no key ID or signed replica identity.
Replicas also share the cursor. The forwarder is leader-gated like the other brain workers: one replica at a time holds the advisory lock for the length of a pass, so the receiver sees one ordered stream however many brain processes are running.
SSRF protection
Section titled “SSRF protection”Every dispatch runs through the same DNS-pin protection as the generic webhook notification channel:
- Scheme must be
https://(orhttp://ifZ4J_NOTIFICATIONS_WEBHOOK_ALLOW_HTTP=true; this does not permit private destinations) - Hostname resolved to one or more IP addresses
- Each IP checked against the blocked set (loopback, RFC1918, link-local, cloud metadata, CGNAT, IPv4-mapped IPv6, 6to4, NAT64, benchmark)
- The validated IP is pinned for the TCP connect;
Hostheader + TLS SNI extension stay set to the original hostname so vhost routing and TLS certificate validation still work
A configured URL that resolves to a blocked IP is refused at dispatch time. The refusal is a failed attempt like any other: the cursor does not move, the backoff grows, and the row is retried once the DNS answer is acceptable again. It is counted under z4j_audit_forward_failures_total{reason="ssrf_or_dns"} and under z4j_swallowed_exceptions_total{module="audit_forwarder",site="ssrf_or_dns"}.
Delivery: the durable cursor
Section titled “Delivery: the durable cursor”The audit log is append-only and every row carries a strictly monotonic chain-order key, (occurred_at, id), the same order z4j audit verify walks. The forwarder keeps, in the audit_forward_state table, the key of the newest row the receiver acknowledged. Each pass:
- Takes the leader lock, so one replica forwards at a time.
- Reads up to
Z4J_AUDIT_WEBHOOK_BATCH_SIZErows strictly after the cursor, in chain order. - POSTs them one at a time, in that order. After each
2xxthe cursor moves to that row. - Stops at the first answer that is not a
2xx(a 5xx, a 4xx, a transport error, a timeout, an SSRF refusal). The cursor stays on the last acknowledged row, the failure count on the state row goes up by one, and the next attempt waitsmin(1s * 2^(failures-1), Z4J_AUDIT_WEBHOOK_MAX_BACKOFF_SECONDS). A success resets the count. - If a full batch went through and more rows are waiting, runs again at once; otherwise sleeps
Z4J_AUDIT_WEBHOOK_POLL_INTERVAL_SECONDS.
What that gives you:
- No drops. A receiver that is down for an hour, a day, or a week receives every row written in the meantime, in order, when it comes back. The only bound is audit retention: a row the retention sweep has deleted before it was acknowledged is gone, so keep
Z4J_AUDIT_RETENTION_DAYSlonger than any outage you intend to ride out. - Restart-safe. The cursor and the failure count are in the database, so a brain restart resumes where it left off and keeps backing off instead of hammering a receiver that was already failing.
- Ordered. Rows arrive in the order the brain wrote them, which is the order the chain links them.
- At least once, not exactly once. A crash between the receiver's
2xxand the cursor write re-sends that one row. A receiver that answered2xxafter the brain stopped waiting sees the row twice. The rowidis mint-once, so de-duplicate on it.
What it does not give you: the brain's audit write path never waits for the receiver, so the mirror lags by at least the poll interval on a quiet brain and by however long the receiver takes on a busy one. z4j_audit_forward_lag_rows is the number of rows the receiver has not acknowledged yet.
Backfilling or skipping
Section titled “Backfilling or skipping”The cursor is created at the audit head when the worker first runs, so rows written before the forwarder was enabled are not mirrored. There is no command for moving it, and deleting the sink's row does not replay anything: the next pass re-creates the row at the current head. To replay the retained history, stop the brain, set the cursor back with the SQL below, and start again:
-- Replay everything still retained:UPDATE audit_forward_state SET last_forwarded_occurred_at = NULL, last_forwarded_id = NULL, consecutive_failures = 0 WHERE sink_id = 'default';To skip a backlog you have decided not to deliver, set the two cursor columns to the occurred_at and id of the newest row you want to skip. Either change is a decision about what the receiver will hold, so make it with the audit trail that implies; the brain does not audit these writes itself.
Observability
Section titled “Observability”GET /api/v1/admin/audit-forwarder (instance admin, session cookie only: its tag has no API-key scope mapping, so every API key gets 403) reports the cursor and the backlog from the database, which is the brain-wide truth whichever replica holds the lock:
{ "enabled": true, "sink_id": "default", "worker": "audit_forwarder_worker", "cursor_initialised": true, "cursor_occurred_at": "2026-05-12T12:00:00.000000Z", "cursor_id": "...", "lag_rows": 0, "last_attempt_at": "2026-05-12T12:00:04.120000Z", "last_success_at": "2026-05-12T12:00:04.120000Z", "consecutive_failures": 0, "backoff_seconds_remaining": 0.0, "batch_size": 100, "poll_interval_seconds": 5.0, "max_backoff_seconds": 300.0, "updated_at": "2026-05-12T12:00:04.120000Z", "process_sent_count": 1842, "process_failed_count": 3}process_sent_count and process_failed_count are the counters of the brain process that answered the request; they are one process's view and null when that process has no forwarder constructed.
Two Prometheus series on /metrics:
z4j_audit_forward_lag_rows(gauge): rows written but not yet acknowledged, read from the database at the start of every pass. Alert on it growing for longer than your receiver's acceptable outage.z4j_audit_forward_failures_total{reason}(counter): attempts that did not get a2xx.reasonis one ofnon_2xx,post_raised(transport error or timeout),ssrf_or_dns, orcursor_conflict(the cursor moved under a pass, which only happens when a replica's lock connection died mid-pass and another replica took over; the pass stopped rather than rewind it).
The dispatch failures also increment z4j_swallowed_exceptions_total{module="audit_forwarder"} under site=non_2xx, site=post_raised and site=ssrf_or_dns, so the existing overview dashboard alert keeps firing on a failing receiver. None of these is a drop any more; they are retries waiting to happen.
Threat model
Section titled “Threat model”The forwarder is an authenticated mirror delivered at least once. Append-only retention is a property the receiver must enforce. It does not replace the brain's primary audit log; it complements it. Specifically:
- The mirror lags (the receiver is not consulted on the audit write path).
- The cursor lives in the brain's database. A role that can delete audit rows can also move the cursor past them, and the receiver cannot tell that from a quiet period.
- An attacker who compromises the brain's HMAC secret can forge rows to the mirror, so the receiver should treat the brain as one of several sources, not as a trusted oracle.
Operators wanting cryptographic non-repudiation should pair the forwarder with a receiver that re-signs every row under its own secret on receipt, so the chain extends beyond the brain's trust boundary.
Not a substitute for a durable sink
Section titled “Not a substitute for a durable sink”The threat model is explicit that a role with write access to the brain's database can delete recent audit rows and roll the authenticated chain state back with them, and that verification then reports the shortened history as clean. The remedy it points to is evidence held outside the database. This forwarder is the obvious thing to reach for, and on its own it will not carry that weight.
The reason is the cursor. Delivery is durable against outages and restarts, but the position it resumes from is a row in the same database as the audit log, written by the same role. An attacker with that role deletes the rows they want gone and moves the cursor past where they were, and the receiver's copy simply never has them. A gap in the receiver's copy is therefore still not evidence of anything on its own.
Evidence meant to survive a compromised database role needs three properties:
- Durable. Delivery is retried until it is acknowledged, and unsent rows survive a brain restart. This forwarder now has this property, for as long as the rows themselves survive in the database.
- Append-only at the sink. The receiver refuses edits and deletions in its own right, under credentials the brain never holds -- an object-store bucket with a retention lock, a WORM log service, an account whose retention policy the brain's operator cannot relax.
- Gap-detecting. The receiver can tell "nothing happened" from "something
did not arrive". The chain link itself, checked on receipt, turns a silent
gap into an alarm: every row's
prev_row_hmacmust equal therow_hmacof the row before it.
Two ways to get there, both of which can use this forwarder as one part:
- Make the sink authoritative about gaps. Have the receiver track
prev_row_hmaccontinuity and alert on a break or a stall. Rows arrive in chain order, so a break in continuity at the receiver is a row the brain's database no longer had when the forwarder reached that position. - Anchor the chain head, on a schedule, from a job that retries. Export
the current head to the durable sink, then verify against it later with
z4j audit verify --known-head. This is the smaller and more reliable mechanism, because a single head covers every row beneath it, so one successful export per interval bounds how far the log can be rolled back without detection. See HMAC audit chain.
Use the forwarder for SIEM ingest and for operational visibility, which is what it is good at. Do not let it stand in for the append-only, gap-detecting sink the threat model calls for.
Disabling
Section titled “Disabling”Unset Z4J_AUDIT_WEBHOOK_URL and restart. The forwarder is not constructed when no URL is set; no worker starts and the cursor row is left where it is, so re-enabling later resumes from it rather than from the then-current head. Delete the row from audit_forward_state if you want re-enabling to start fresh at the head.