Threat model
Actors
Section titled “Actors”| Actor | Capability | Trust |
|---|---|---|
| External attacker | Internet traffic only | Untrusted |
| Agent host (your app) | Holds an agent token; can emit any events | Partially trusted |
| Dashboard user | Has a role on one or more projects | Role-scoped trust |
| Brain operator | Full DB + env access | Fully trusted |
| Postgres admin | Direct SQL | Fully trusted (defeats both in-database boundaries: audit chain and schedule control) |
Assumptions
Section titled “Assumptions”- TLS terminates in front of z4j. No plaintext HTTP in production.
- Env vars are secret. Leaking
Z4J_SECRETcompromises frame signing, agent-token and API-key hashing, invitation, setup and password-reset tokens, the purge-queue confirm token, and the AES-GCM key that encrypts stored TOTP secrets. Treat it as compromise of the agent transport, agent/API credentials, and MFA enrollments. LeakingZ4J_SESSION_SECRETcompromises user sessions. The audit-log row chain is signed with a separate key,Z4J_AUDIT_CHAIN_SECRET, which is required outside development; keep it somewhere the database operator cannot read, since sharing it with the database would hand one party both halves. - Postgres is on a private network. Direct access is the operator's responsibility.
- Audit rows are ordered by
(occurred_at, id), whereidis a random UUID, not a sequence. The writer clamps each newoccurred_atto at least one microsecond after the authenticated head when the wall clock is not ahead. A backwards clock jump therefore cannot reorder the chain, but stored audit timestamps can move ahead of real time until the clock catches up.
In-scope threats
Section titled “In-scope threats”- Credential stuffing / brute force - defended by argon2id + rate limits + dummy-hash timing.
- Session theft - HttpOnly + Secure + SameSite=Lax cookies; session invalidation on password change.
- Prompt injection via events - z4j does not evaluate event payloads as code or instructions. Task names, exceptions, and tracebacks can be rendered into in-app and external notifications. Treat notification destinations as part of the data-exposure boundary.
- Token enumeration - agent tokens and API keys are stored only as
HMAC-SHA256 digests and resolved by digest lookup. The dummy Argon2 hash is
used only for an unknown email on the email-and-password login path; no
password or dummy hash is involved in token authentication. During a
Z4J_PREVIOUS_SECRETSrotation window, rejecting an unknown agent token performs one lookup per accepted secret and can take measurably longer than accepting a token signed under the current secret. API keys are hashed with the currentZ4J_SECRETonly: one lookup, and a key minted under a rotated-out secret stops authenticating. - Agent impersonation - tokens are per-agent. Revocation is enforced by a durable tombstone check at authenticated work boundaries; socket kicks only accelerate disconnect.
(project_id, name)uniqueness prevents duplicate-register races. - Replay attacks - WebSocket frames have monotonic seq; duplicates deduped on persist.
- SSRF via public_url -
Z4J_PUBLIC_URLis validated (no whitespace, no user:pass, http/https only). - Audit tampering by anything speaking through z4j - the HMAC chain plus the separately authenticated head and row count detect it. Tampering by a role that can write the database tables directly is out of scope, and is not detected: see below.
- Credential use from the wrong network - optional source-address allowlists per surface, below. They narrow where a leaked session, API key or agent token is usable; they are not a substitute for revoking it.
Source-address allowlists
Section titled “Source-address allowlists”Three optional lists restrict where a credential may be used from:
Z4J_DASHBOARD_IP_ALLOWLIST (the login route and every session-cookie
request), Z4J_API_IP_ALLOWLIST (every Bearer API-key request, narrowed
further by a key's own allowed_cidrs) and Z4J_AGENT_IP_ALLOWLIST (the
agent WebSocket and long-poll transports). Each is a JSON array of IPv4 or
IPv6 CIDRs, validated at startup; an empty list restricts nothing. See
settings.
Headers and proxies. The address matched is the one the brain resolved
from the socket peer and, only when that peer is inside
Z4J_TRUSTED_PROXIES, from X-Forwarded-For. A forged header from any other
source is ignored and the peer is matched instead; a malformed chain from a
trusted proxy falls back to the proxy's own address. Two consequences follow.
With no trusted proxies declared, a brain behind a reverse proxy sees every
request as coming from the proxy, so a list that omits the proxy refuses
everyone and a list that includes it admits everyone the proxy forwards: declare
the proxy first. And a trusted proxy can assert any client address, so the
lists are only as strong as the proxy's own handling of the header (strip
inbound X-Forwarded-For, append the real peer). A dual-stack socket reports
IPv4 clients as ::ffff:a.b.c.d; those match IPv4 entries.
A catch-all trusted proxy. Z4J_TRUSTED_PROXIES is validated at startup
like the three lists (a malformed entry refuses to start, entries are stored
canonical, an IPv6 zone id such as fe80::1%eth0 is refused on all four
because no resolved address ever carries one). It is not refused when it
contains 0.0.0.0/0 or ::/0, because a brain behind a service mesh may see
nothing but the sidecar as its peer and have no narrower range to write. But
with such an entry every peer is a trusted proxy, so any client can choose its
own address by sending X-Forwarded-For, and that chosen address is what the
agent connect rate limit keys on and what the three allowlists match: both are
bypassed by a header. The brain logs one WARNING naming the entry at startup.
Narrow the range to the proxy's own addresses unless every peer the brain can
see is a proxy you control.
What the lists do not protect. Loopback is never implicitly exempt, so
list 127.0.0.1/32 and ::1/128 if the brain's own host must get in. The
lists see addresses, not identities: a credential used from inside the range
is accepted as if the list did not exist, and an attacker inside the range (a
compromised host, a shared egress NAT that puts the whole office behind one
address) is not distinguished from the owner. They do not gate unauthenticated
routes (password reset, first-boot setup, health) and do not replace host
validation, TLS or rate limits. The agent list refuses a transport; it does
not revoke the token, which stays valid for the next attempt from an admitted
address. It applies to the agent WebSocket, whose hello is closed with 4403
before the bearer is read, and to both long-poll routes (the connect probe
included), which answer 403 with error: ip_denied before the bearer is
read. The dashboard list also gates the dashboard WebSocket, which closes with
4401, the code a missing session gets.
Ordering and the oracle it leaves. The dashboard and API lists are
consulted after the credential authenticates, so every refusal writes an
auth.ip_denied audit row naming the user or key presented from the wrong
place and increments z4j_auth_ip_denied_total{surface}. The 403 body is the
same whatever was presented (error is ip_denied and details names only
the surface; the address, user and key id go to the audit row, never to the
caller). The cost of attributing first is that a caller outside the list can
still tell a valid credential (403) from an invalid one (401). The login route
and the agent transports are the exceptions: the login route refuses before
the password is checked, and the agent WebSocket and both long-poll routes
refuse before the bearer is read, so none of them can be used as an oracle
from outside its list, for a token's validity or for whether the token's
project is archived. Their auth.ip_denied rows carry the address and the
path but no user or agent id, since nothing was authenticated.
- Direct write access to the database, whether from physical access,
stolen credentials, a shared reporting role, or a restored backup. It
defeats the in-database boundaries without breaking any cryptography.
Both boundaries are exposed the same way, and which engine you run decides
how much a raw client has to do. On PostgreSQL the audit-log and
schedule triggers authorise a write by reading a session setting the client
can set for itself, so table-write access is the whole requirement. On
SQLite those triggers call a connection-local function the brain registers
on its own connection, so a raw client's write errors out and fails closed
until that client registers the function itself or drops the triggers, which
schema rights allow. For
the audit log, such a role can delete recent rows and put back the
audit_chain_staterow as it stood before them; the brain signed that copy when it was current, so it still authenticates, and verification reports the shortened history as clean. For schedules on PostgreSQL, the triggers authorise a change by reading a session setting the client sets itself and by checking for a matchingschedule_change_logrow, and inserts into that log are unguarded, so the writer can supply its own evidence. Neither path involves a secret it cannot read. Treat the database credentials as equivalent to full control of z4j's integrity guarantees. The one check that survives is verifying against a head exported somewhere the database role cannot rewrite: see HMAC audit chain. - SSO / OAuth2 - not shipped. Forwarding auth through an SSO proxy in front of the brain is the supported path. (Multi-factor authentication IS shipped, and belongs in-scope: TOTP plus single-use recovery codes, trusted devices, and opt-in org-wide enforcement. See MFA.)
- Compliance certifications (SOC 2 / HIPAA / ISO 27001) - z4j is not certified. The audit log export plus Postgres backups provide raw evidence; operators own policies, controls, and external audit.
- Browser fingerprinting / anti-automation - we're not that kind of product.
What we have not fixed
Section titled “What we have not fixed”One weakness is open by our own choice rather than by oversight: the
in-database boundaries above do not hold against a role that can write their
tables. This page and the HMAC audit-chain page
are the current public record. The root SECURITY.md shipped from this source
tree records how we rate it, when it first shipped, the
mitigation available today, and what has to be demonstrably true before it
comes off the list. The standalone repository policy is sourced from
packages/z4j/SECURITY.md and carries the same limitation. A split-repository
checkout that lacks that entry predates the corrected policy; use this page or
the root policy instead.
Reporting issues
Section titled “Reporting issues”security@z4j.com. Do not file a public issue for undisclosed vulnerabilities. See disclosure.