Why a separate tier
The outbox table guarantees at-least-once event delivery: events written inside a transaction (e.g.,interaction.recorded.v1, outcome.recorded)
are durable even when the configured EventPublisher backend is down or
slow. Without a dedicated publisher tier, the BullMQ-running worker
pods own this loop alongside long-running batch jobs, and a backed-up
batch can starve the publish loop. Splitting these tiers keeps the
publish tail latency independent of batch contention.
What changed in the respond hot path
The/api/v1/respond endpoint no longer publishes events synchronously.
The interaction.recorded.v1 event is now enqueued inside the same
database transaction that writes the interaction history row.
Behavior-change note: an outbox row insertion failure now rolls back
the interaction row. This is a correctness improvement vs the prior
fail-open path — the system no longer claims outcomes whose downstream
events it can’t persist. The cost is that pathological insert failures
(JSON-too-large, constraint violation, mid-tx connection drop) surface
as 500s to the caller instead of silent drops. Operators investigating
“respond returned 500” should check outbox_events insert errors first.
Operator visibility — what surfaces when things go wrong
Recommended Prometheus alert
kaireon_outbox_pending_count gauge is registered with the platform metrics registry and refreshed by the outbox processor on every poll tick. See Metrics Reference for the full PromQL alert family.)
Who drains the outbox
Exactly one of these must be running, oroutbox_events fills up silently and
nothing is ever published:
OUTBOX_POLLER_INPROCESS is independent of WORKER_INPROCESS. Setting
WORKER_INPROCESS=0 means “a dedicated worker container consumes the job
queue”; it says nothing about the outbox. Before 2026-08-15 the two were tied
together, so a single-container deployment that opted out of in-process BullMQ
workers also silently lost its outbox drainer.
Retry schedule and dead-lettering
A failed publish is retried on an exponential backoff, and the due-time for the next attempt is persisted on the row itself inoutbox_events.nextAttemptAt.
NULL means “due now” — the state of every freshly written event and of every
event replayed out of the DLQ.
The batch claim only picks up rows that are actually due:
min(1000ms x 2^(n-1), 60000ms) for the n-th failed attempt, plus
up to 50% jitter. With the default maxRetries = 5:
So an event whose backend stays down is dead-lettered after 5 attempts spanning
roughly 15–22 seconds. The DLQ write and the status update happen in one
transaction, and the resulting
ERROR log carries dlqDepth plus an alert
level of INFO / WARNING (>10) / CRITICAL (>100).
nextAttemptAt is deliberately separate from updatedAt. The claim UPDATE
stamps updatedAt — that is what the reaper
reads to find rows orphaned by a dead worker. Anchoring the backoff on the same
column made the claim overwrite the value it was about to read, so no event with
retryCount > 0 was ever retried or dead-lettered. Fixed 2026-08-15.Outbox reaper cron — /api/v1/cron/outbox-reaper
A dedicated cron job sweeps outbox_events and resets any row stuck in
processing whose updatedAt is older than the configured staleness
threshold back to pending. This closes the failure mode where a worker
dies between claiming a row (UPDATE → processing) and either
publishing it or marking it failed — without the reaper those rows
would sit in processing forever and never re-attempted.
The cron route invokes the outbox processor’s stuck-row reaper, which
performs a single bulk SQL UPDATE driven by the configured staleness
threshold. The operation is idempotent — re-running on already-pending
rows is a no-op.
Helm wiring
helm/values.yaml. Cadence of 1–5 minutes is fine
because the operation is idempotent.
Auth
The cron route fail-closes whenCRON_SECRET is unset (route.ts:25-32).
Authenticated callers present the secret via either Authorization: Bearer <secret> or x-cron-secret. Mismatched values return 401.
Response
Configuration
Configuration knobs
OUTBOX_LIVENESS_FILE (default
/tmp/outbox-publisher.alive). The liveness probe checks this file’s
age — when the publisher loop stops touching it, the probe fails and
Kubernetes restarts the pod.
Honest known gaps
- Structured error IDs shipped on the worker tick path; not yet on every helper. The publisher’s main poll loop mints a per-tick
errorIdand threads it into the next-attempt log line so SIEM tooling can correlate retries. Other in-tier helpers (shutdown drain, reaper companion) still emit bare structured logs and are tracked as a residual for migration. SIEM correlation works for the main loop today. kaireon_outbox_pending_countgauge shipped (W10 wave). Registered with the platform metrics registry and refreshed every poll tick. The recommended Prometheus alert above can be wired today.outboxProcessedTotal+outboxEventAgefrom W8.3 still cover throughput + freshness; this gauge closes the backlog visibility gap. See Metrics Reference for full alert PromQL.