Skip to content

Audit exports and the chain head anchor

The synchronous export on the audit page refuses result sets above its in-memory caps (50,000 rows for CSV and JSON, 25,000 for XLSX). A brain that writes more than a few hundred audit rows a day keeps a trail the download cannot return whole. Export jobs remove that limit: the dashboard or the API queues a job, the brain's export-jobs worker streams the rows page by page into an export sink, and the trail of any size lands as one object an auditor can take away.

The same sink serves a second purpose. The HMAC chain cannot, on its own, prove anything against a role that can write both audit_log and audit_chain_state (see HMAC audit chain). The documented mitigation is a chain head recorded somewhere that role cannot reach, checked later with z4j audit verify --known-head. With a sink configured, the brain writes that head on a schedule by itself.

Set Z4J_EXPORT_SINK to local or s3. The default, none, disables both exports and the head export: the API answers 409 to a job request and the worker is not started.

Terminal window
Z4J_EXPORT_SINK=local
Z4J_EXPORT_SINK_PATH=/var/lib/z4j/exports

The directory must exist and must not be a symlink. Objects are written through a temporary file in the same directory and renamed into place, so a reader never sees a partial file and a failed job leaves nothing behind. Files are created with mode 0640, and the sink refuses to write through a symlink at any level below the base, so an unprivileged account on the same host cannot redirect an export. Jobs land under audit/<project-slug>/<timestamp>-<job-id>.<format>.

A finished local-sink job can be downloaded from the dashboard or from GET /api/v1/projects/{slug}/audit/export-jobs/{job_id}/download by an auditor or an admin of the project. The job row records the object's key below the directory, not a path, and the download does not trust the row to name the file: it rebuilds the path from Z4J_EXPORT_SINK_PATH, refuses any step that is .. or otherwise not a plain key, resolves both the directory and the file and requires the one to be inside the other, follows no symlink at any level, and serves the handle it opened. A row edited to point anywhere else answers 404, so a database role that can write export_jobs cannot turn the download into a read of another file on the host. The location the API reports is still the full path, rebuilt from the directory in effect.

Terminal window
pip install "z4j[s3]"
Z4J_EXPORT_SINK=s3
Z4J_EXPORT_SINK_S3_BUCKET=z4j-evidence
Z4J_EXPORT_SINK_S3_PREFIX=z4j-exports
# Any S3 API: AWS (leave the endpoint unset), MinIO, Ceph RGW, Backblaze B2.
Z4J_EXPORT_SINK_S3_ENDPOINT_URL=https://minio.internal:9000
Z4J_EXPORT_SINK_S3_REGION=us-east-1

Credentials come from the standard AWS chain (environment variables, shared config, an instance or task role) unless both Z4J_EXPORT_SINK_S3_ACCESS_KEY_ID and Z4J_EXPORT_SINK_S3_SECRET_ACCESS_KEY are set, in which case those are used. They are secrets: they never appear in startup logs, in z4j doctor, or in any API response. Objects larger than one part go up as a multipart upload, so the whole export is never held in memory; a failed job aborts its upload. The client is imported only when the sink is first used, and a brain without the extra reports the install line rather than failing at boot.

For the head anchor, give the bucket a retention lock (S3 Object Lock, or your store's equivalent) or at least versioning, and give the brain's credentials write access only. The point of the anchor is that the database role cannot shorten it; a bucket the brain can delete from is better than nothing and worse than one it cannot.

From the dashboard, the audit page's Exports panel queues a job in the current format with the current filters and lists every job with its status, size, sink location and, for the local sink, a download button.

From the API:

Terminal window
# Queue (auditor or admin on the project; API keys need audit:read)
curl -X POST "$BRAIN/api/v1/projects/default/audit/export-jobs" \
-H "Authorization: Bearer $Z4J_API_KEY" -H "Content-Type: application/json" \
-d '{"format": "csv", "since": "2026-01-01T00:00:00Z"}'
# Poll
curl "$BRAIN/api/v1/projects/default/audit/export-jobs/$JOB_ID" \
-H "Authorization: Bearer $Z4J_API_KEY"

A job moves from queued to running and then to done or failed. While running, row_count advances after every page; when done, location names the object (a path, or an s3://bucket/key URL) and size_bytes its length; when failed, error says why. Filters, columns and bytes match the synchronous export for the same request. Every creation, completion and failure is an audit row of its own (audit.export_job.created, audit.export_job.completed, audit.export_job.failed).

The export-jobs worker runs in every brain process that has a sink and is leader-gated like the other periodic workers, so one replica drains the queue. It polls every Z4J_EXPORT_JOBS_POLL_INTERVAL_SECONDS (default 5) and claims jobs oldest first. Rows are fetched Z4J_EXPORT_JOBS_PAGE_SIZE at a time (default 2000) and encoded as they arrive; the whole result is never in memory. XLSX is spooled to a private temporary file in constant memory and then streamed to the sink, which is why its only bound is the worksheet format's row limit.

A job whose row has not moved for fifteen minutes belongs to a process that is gone (progress bumps the row on every page); the next tick marks it failed with that reason, and it can be queued again. Should the process turn out not to be gone (a replica that lost its leader lock mid-write but kept writing), the failure that was recorded stands: the late completion appends nothing to the audit trail and removes the object it wrote, locally or from the bucket, rather than leaving an orphan no row refers to. A sink that is unreachable fails the job with the error text and is reported in the log; it never stops the brain.

Terminal window
Z4J_AUDIT_HEAD_EXPORT_INTERVAL_SECONDS=3600

With a sink configured and the audit-chain key set, the worker writes the current authenticated chain head to the sink on that cadence (at least 60 seconds; 0, the default, disables it). Two objects are written each time: a stable key, audit-head/current.json, that a check can always read, and a dated copy beside it, audit-head/<timestamp>.json, so a locked or versioned bucket keeps every head ever anchored. The envelope is byte for byte what z4j audit export-head prints: six keys, canonical JSON, one trailing newline, authenticated against the configured keyring before it is written, and refused (logged, not written) when the state row does not authenticate.

Check the chain against the anchor the same way the CLI documents, reading the file the worker wrote:

Terminal window
z4j audit verify --known-head "$(cat /var/lib/z4j/exports/audit-head/current.json)"
# or, for S3
aws s3 cp s3://z4j-evidence/z4j-exports/audit-head/current.json - | \
xargs -0 z4j audit verify --known-head

CURRENT_MATCH means the chain still ends at the anchored head; VERIFIED_ANCESTOR means it has grown past it; UNPROVABLE means the log no longer contains it, which is the finding this anchor exists to produce. No audit row is written for a head export, because writing one would move the head the export just anchored.

The head export authenticates the state row; it does not walk every row first. Pair it with Z4J_AUDIT_CHAIN_VERIFY_ENABLED=true for the scheduled walk, so a chain that fails to verify is noticed within an interval rather than anchored as it stands.

Variable Default Description
Z4J_EXPORT_SINK none none, local or s3.
Z4J_EXPORT_SINK_PATH - Directory for the local sink. Required with local.
Z4J_EXPORT_SINK_S3_BUCKET - Bucket for the s3 sink. Required with s3.
Z4J_EXPORT_SINK_S3_PREFIX z4j-exports Key prefix inside the bucket.
Z4J_EXPORT_SINK_S3_ENDPOINT_URL - Endpoint for S3-compatible stores. Unset means AWS.
Z4J_EXPORT_SINK_S3_REGION - Region name. Unset defers to the AWS environment.
Z4J_EXPORT_SINK_S3_ACCESS_KEY_ID - Explicit access key. Set together with the secret, or neither.
Z4J_EXPORT_SINK_S3_SECRET_ACCESS_KEY - Explicit secret key.
Z4J_EXPORT_JOBS_POLL_INTERVAL_SECONDS 5 Queue poll cadence, 1 to 300.
Z4J_EXPORT_JOBS_PAGE_SIZE 2000 Rows fetched per page while streaming, 100 to 20000.
Z4J_AUDIT_HEAD_EXPORT_INTERVAL_SECONDS 0 Head export cadence; 0 off, otherwise 60 to 604800. Needs a sink.
  • 409 no export sink is configured: set Z4J_EXPORT_SINK and restart the brain.
  • the s3 export sink needs the aiobotocore package: install z4j[s3] in the brain's environment.
  • refusing to write through symlink: the local sink found a symlink below its base directory and stopped; replace it with a real directory.
  • A job stays queued: no replica holds the export-jobs leader lock, or the worker is not started because Z4J_EXPORT_SINK is none on the replica you are looking at; check the startup log's worker list.
  • download is served for the local sink only: the job wrote to S3; fetch the object at the job's location.