Skip to content

Audit retention and pruning

The audit log is an HMAC chain, so removing rows from it is not a plain DELETE. Every removal the brain performs is recorded in the authenticated chain state as a prune boundary, and that record is what lets z4j audit verify tell retention apart from a truncation. This page covers the three ways rows leave the table (the periodic sweep, z4j audit prune in soft mode, and an epoch cut in hard mode), the retention windows that drive them, and what verification says afterwards.

Rows are only ever removed as a contiguous prefix, oldest first. For each batch the brain verifies every row of the prefix against the chain, checks that the first row follows the current boundary and that the row after the prefix links to its last row, deletes exactly those rows, and advances the signed prune boundary in audit_chain_state to the newest row it removed, all in one transaction. The next row ever appended links to that boundary, and the verifier accepts it as the chain's anchor.

The consequence worth knowing: the prune stops at the first row that retention still keeps, whatever lies beyond it. A row the policy would remove is retained while an older row before it is kept. The chain can describe one boundary, not a set of holes.

When that happens with expired rows waiting behind the retained row, the sweep logs one WARNING per pass naming the row that holds the line: its class and age, the moment its class window keeps it until, and how many expired rows are retained behind it. A long class window is a decision to keep everything written after its oldest retained row; the warning is the size of that decision, not a fault. The dry run below prints the same row under stops at.

The periodic sweep runs on Z4J_AUDIT_RETENTION_SWEEP_INTERVAL_SECONDS, removes at most Z4J_AUDIT_RETENTION_SWEEP_BATCH_SIZE rows per statement and Z4J_AUDIT_RETENTION_SWEEP_MAX_PER_PASS per pass, and commits each batch, so an interrupted pass leaves a signed, consistent state and the next pass continues. Only one process sweeps at a time: on PostgreSQL the sweep takes an advisory lock, on SQLite the writer lock serialises it.

Z4J_AUDIT_RETENTION_DAYS (default 90) is the window for every row. Set it to your real obligation before the first sweep runs: rows past the window are deleted, not archived.

There is no ceiling. An obligation can run past ten years, so a window that long is accepted; the brain logs one WARNING at startup naming each window above 3650 days, because a value that large is more often a typo than a policy, and the table grows for the whole window.

Z4J_AUDIT_RETENTION_BY_CLASS gives a class of actions its own window. It is a JSON object mapping an action class to days; a class not listed uses the global window. The class of an action is the first dotted segment of its name: auth for auth.login, command for command.issue.requeue_dead_letter, dead_letters for dead_letters.list, audit for the chain's own markers.

Terminal window
Z4J_AUDIT_RETENTION_DAYS=90
Z4J_AUDIT_RETENTION_BY_CLASS='{"auth": 365, "command": 30}'

Validation is strict and happens when settings load: keys must look like an action class (lower-case letters, digits and underscores, no dots), values must be whole positive numbers of days. A key that names a whole action (auth.login) is rejected rather than silently matching nothing. A class window may be longer or shorter than the global one, and the prefix rule above decides what each buys:

  • Longer (auth: 365 above a global 90): every auth row is kept for a year, and so is every row written after the oldest auth row still inside that year, whatever its class. The chain holds at that row. If you need a long window for one class, expect the whole trail from that class's oldest retained row onward to stay.
  • Shorter (command: 30 under a global 90): a command row is removed once it is both past 30 days and older than every retained row before it. In a trail where classes interleave, that means the shorter window takes effect only on the rows at the very oldest end, before the oldest row any other window keeps. It never punches a hole in the middle of the chain.

The sweep and z4j audit prune apply exactly the same rule, and the dry run below shows where the prefix stops and why.

Settings are read from the environment when the brain starts. There is no write-back: nothing in the dashboard or the API changes a retention window, and changing one is a restart.

The operator-driven form of the same prune, for the cases the periodic sweep does not cover: pruning now rather than on the next tick, pruning to an explicit date, confirming what a policy change would remove before it runs, and cutting an epoch.

Terminal window
# What the configured windows would remove right now. Changes nothing.
z4j audit prune
# Remove it, under the authenticated boundary.
z4j audit prune --apply
# Prune every class before an explicit moment instead of the windows.
z4j audit prune --apply --before 2026-01-31T00:00:00Z
# Epoch cut: once the generation is fully pruned, start a fresh signed one.
z4j audit prune --hard --apply --before 2026-06-30T00:00:00Z

It requires the audit-chain key and an authenticated chain state, so it runs with the brain's settings, not with a read-only database role.

Soft mode (the default) is the authenticated prefix prune described above, driven by hand. Rows older than their class cutoff, up to the first row retention keeps, are verified, removed and recorded under the signed boundary. The chain stays in the same generation; the next row links to the boundary; verify reports the chain clean with the prune on record.

Hard mode (--hard) is an epoch cut. After the soft prune, the fully pruned generation is replaced by a fresh one: one signed audit.chain_generation_reset row becomes the new genesis and the chain state carries no prune boundary at all, so nothing in the live chain refers to the old history. It refuses, before changing anything, while any active row younger than the cutoff remains (an epoch cut never removes a row the cutoff keeps; pass --before later than the newest row you are prepared to lose) and while frozen legacy rows exist (export them first with z4j audit export-and-delete-frozen). Every head exported from the old generation reports UNPROVABLE afterwards; export a new one.

Those preconditions are checked on the preview, which runs before the command takes its leases, and a live brain can append a row in between. The reset therefore checks every active row against the same cutoffs again, inside its own transaction with the chain locks held, and refuses with exit 1 naming the row if one younger than its cutoff has appeared. The soft prune's batches stay committed under the signed boundary, the generation is not reset, and no audit.prune row is written. Run hard mode on a quiet brain (a maintenance window, or with the agents stopped) and treat the refusal as safe to rerun.

Flag Effect
(none) Dry run. Prints the mode, the cutoffs, the generation, the rows that would be removed by class, the row the prefix stops at and why, the boundary that would be recorded, and the rows that remain. Changes nothing.
--apply Execute.
--hard Epoch cut after the prune, with the preconditions above.
--before <ISO-8601> Use one explicit cutoff for every class instead of the configured windows. A timezone offset or Z is required and the value must not be in the future.

On --apply the command writes one audit.prune row about itself, in the retained chain (soft) or in the new generation after the reset marker (hard). Its metadata records the mode, the cutoff and its source (the configured windows or --before), the per-class cutoffs, the rows removed and their classes, the generation, and the boundary row's id and row_hmac; in hard mode also the new generation and the reset marker.

While it works it holds the scheduled verifier's leader lease and the periodic sweep's advisory lock, so neither can walk or prune the same rows underneath it. If the verifier is mid-walk the command refuses and says so; if the sweep is mid-pass it refuses likewise rather than reporting nothing to prune. On PostgreSQL both are advisory locks; on SQLite the single writer serialises everything, and another connection holding the writer lock past the busy timeout is the same refusal (database is locked, exit 1), not a traceback.

Each batch is its own signed transaction and the audit.prune row is written last, once every batch is through. A run interrupted in between (a killed process, a dropped connection) leaves the boundary at the last committed batch, the chain verifying clean as it stands, and no audit.prune row. Run the command again: it continues from the boundary, removes whatever is left, and records itself, with the row count of that rerun.

Code Meaning
0 Done, dry run complete, or nothing to prune.
1 Refused: no audit-chain key, chain state missing or not authenticating under the configured keys, the row count or a link in the prefix disagreeing with the signed state, a held lease or SQLite writer lock, a hard-mode precondition on the preview, or a row appended before the reset that the cutoff keeps. The message names the reason and ends with what was changed, which is nothing unless batches had already committed.
2 Settings failed to load, or --before is not an aware, past ISO-8601 timestamp.

A refusal mid-prune (REFUSING mid-prune) means batches already committed are signed and consistent and nothing after them was touched; run z4j audit verify, resolve the finding, and run the prune again.

After a soft prune z4j audit verify reports the chain clean. The pruned range is on record in the signed boundary, and --known-head tells the story exactly:

  • a head you exported at the boundary row reports PRUNE_MATCH (CURRENT_PRUNE_MATCH if nothing has been appended since);
  • a head from deeper inside the pruned range reports UNPROVABLE with no other finding: the anchor is gone, the chain is not broken. Record heads more often than your shortest window and read that result as "the anchor aged out; confirm why", not as tampering by itself.

After an epoch cut every head from the old generation reports UNPROVABLE for the same reason, and the new generation's genesis has no boundary.

Export a new head after any prune, soft or hard. UNPROVABLE is exit 1 from z4j audit verify --known-head, so a scheduled check that still anchors on a head from inside the pruned range fails from the next run on; the command's success line says so, and the scheduled head export writes a fresh anchor on its next cadence.

A deletion that did not go through the prune, a row removed by hand or by a job that bypassed the audit service, still reports as broken: the physical row count disagrees with the signed count, and the row after the gap fails its link check. z4j audit prune refuses on that chain rather than blessing the gap with a new boundary; it only ever advances the boundary over rows it verified itself.