Skip to main content

Topology

The CronJob does not run business logic. It is a schedule-aware HTTP trigger — the work itself runs inside the API tier handler at /api/v1/cron/<name>. Splitting this from the API replicas means a backed-up API pool cannot prevent the cron from firing, and a slow handler does not consume an API replica’s RPS budget for the next tick (the CronJob’s concurrencyPolicy: Forbid ensures this).

Configured jobs

Seven jobs in helm/values.yaml: Disable any individual job by flipping cron.schedules.<name>.enabled to false. Disable the whole tier with cron.enabled: false (the existing in-process schedulers continue to work; you’d lose the external trigger redundancy).

In-process maintenance scheduler

The default single-container deployment does not require external CronJobs. When a cron secret (CRON_SECRET, or the legacy CRON_TOKEN alias) is set, the platform runs an in-process scheduler (src/lib/maintenance-scheduler.ts) that self-invokes /api/v1/cron/* and the legacy /api/cron/tick on fixed internal cadences (mostly mirroring the EventBridge module defaults): The flow-scheduler tick is not in this list — it runs on its own advisory-locked ticker. /api/v1/cron/campaign-scheduler fires due Run (campaign) schedules — see Campaigns → Schedule Configuration for the RRULE/cron evaluation semantics. Disable the in-process scheduler with MAINTENANCE_SCHEDULER_ENABLED=false when you bring your own scheduler (EventBridge / Kubernetes CronJob). It also no-ops (with a warning) when neither CRON_SECRET nor CRON_TOKEN is set — in that case retention purges, DSAR purge, DLQ drain, and the staging janitor will not run. When running multiple replicas, either set MAINTENANCE_SCHEDULER_ENABLED=false everywhere and use EventBridge, or accept duplicate (harmless) passes.

Auth + secret handling

CRON_SECRET is read from the existing api-secrets Secret via a secretKeyRef — it is never rendered into the values file or the container args. The cron container builds the Authorization: Bearer $CRON_SECRET header inside its shell so the secret never appears in ps output. The cron container runs as non-root, drops all capabilities, and is restricted to a bare curlimages/curl image — no node/python runtimes inside the trigger surface.

Operational checklist

  • Watch the cluster’s kube-system events for backoff-limit-exceeded errors on any cron job — repeated failures usually mean the corresponding API handler is broken or the API service DNS is wrong.
  • Set tighter activeDeadlineSeconds (default 600s) for jobs whose handlers should never run that long.
  • For staged rollouts, override cron.schedules.<name>.schedule per environment so non-prod runs less frequently.
  • Cleanup history retention: successfulJobsHistoryLimit: 3, failedJobsHistoryLimit: 5 — bump if you need more for forensics.

Operator-wired cron routes

Three /api/v1/cron/* handlers ship in the API image but are not wired into helm/values.yaml’s cron.schedules block. They exist because their cadence is policy-driven (retention windows, SIEM batch intervals, export checkpoints) rather than chart-driven, so the chart leaves the schedule choice to the operator. (/api/v1/cron/recompute-metrics, documented below, is the exception — it is wired by default in both the Helm chart and the in-process maintenance scheduler.) All three accept the same Authorization: Bearer $CRON_SECRET shape as the wired routes. Each fails closed when CRON_SECRET is unset. To run any of them, copy the wired-cron CronJob template from helm/templates/cron-jobs.yaml, swap the path: value, and pick a schedule appropriate for your retention or SIEM-batch policy. No handler-side change is needed — the only thing the chart contributes to a wired job is the schedule + the curl trigger.

/api/v1/cron/dsar-purge

Iterates every tenant, reads each tenant’s per-class retention rows (skipping any row marked legal-hold), collapses to the strictest (smallest) retentionDays across data classes, and deletes rows older than that cutoff from the decision-trace, interaction-history, and AI-attachment tables. Attachment storage blobs are deleted best-effort through the configured attachment store — when the local-filesystem file or S3 object is already gone the route logs and continues. Each tenant produces an audit_log row with the per-class purge counts. The response shape is { ok, tenantsScanned, totalBlobsAttempted, perTenant: [{ tenantId, retentionDays, decisionTraces, interactionHistory, aiAttachments, blobsDeleted, errors[] }] }.
  • Schedule: operator-supplied; no default ships in the chart.
  • Recommended cadence: once per day, off-hours. The handler scales linearly with row count so a tenant with multi-million-row purge backlogs benefits from running daily rather than weekly.
  • Pre-flight: make sure every tenant that should be purged has a retention-policy row that is not on legal hold and that has a positive retentionDays. Tenants without a config are scanned but skipped (a zero-day retention short-circuits the per-tenant loop).
  • See Retention for the per-class retention-policy schema and the legal-hold semantics.

/api/v1/cron/siem-ship

Reads the last 5 minutes of audit-log rows (up to 500) and ships them to the SIEM backend selected by SIEM_BACKEND (splunk, datadog, or elastic); the SIEM configuration is parsed and validated at startup from the SIEM environment variables listed below. When SIEM_BACKEND is unset or holds an unknown value the route returns { ok: true, skipped: "..." } and is a structured no-op. Response shape: { ok, backend, shipped, errors, windowSeconds: 300, rowsScanned }.
  • Schedule: operator-supplied; no default ships in the chart.
  • Recommended cadence: every 5 minutes — matches the hard-coded 5-minute window the handler reads. Slower ticks risk dropping the tail of the 500-row batch limit; faster ticks ship duplicate rows.
  • Required environment variables (validated at startup):
    • SIEM_BACKEND — one of splunk, datadog, elastic.
    • SIEM_ENDPOINT — destination URL for the chosen backend.
    • SIEM_API_KEY — bearer/HEC token; backend-specific.
    • SIEM_INDEX — target index (Splunk / Elastic).
    • SIEM_SOURCETYPE — Splunk source-type tag.

/api/v1/cron/export-interactions

Hive-partitioned NDJSON export of the interaction-history table. For each active tenant, reads the export checkpoint’s last-export-at timestamp, queries up to 10 000 newer interaction-history rows (the batch size is a fixed constant in the handler), writes them to exports/{tenantId}/interaction_history/year=YYYY/month=MM/day=DD/batch-{ts}.json, and advances the checkpoint. First run starts at the Unix epoch (1970-01-01) and exports all existing rows. Response shape: { ok, totalExported, tenants: [{ tenantId, tenantName, exported, filePath, error? }] }.
  • Schedule: operator-supplied; no default ships in the chart.
  • Recommended cadence: hourly. The 10 000-row batch limit means a tenant generating > 240k interactions per day needs more than one tick per hour to keep up.
  • Honest limit: the handler writes to the API pod’s local filesystem rooted at the process working directory. Production deployments must replace this with an S3 or GCS upload before flipping the schedule on; otherwise the export files are lost when the API pod restarts. A code comment at the top of the handler flags this explicitly as a production note recommending S3 uploads.
  • The route requires prisma.exportCheckpoint rows to exist or be creatable for every tenant; the first run auto-creates a checkpoint per tenant.

/api/v1/cron/recompute-metrics

Recomputes due batch-mode Behavioral Metrics. For every tenant with at least one active batch metric definition, the route recomputes each metric whose last-computed timestamp (derived from its newest MetricValue.computedAt) is older than its own batchIntervalMin — metrics that aren’t due yet are skipped on that tick, so a tenant with a mix of short- and long-interval batch metrics doesn’t redo unnecessary work every tick. Recompute also performs orphan cleanup: any MetricValue row for that metric that didn’t get refreshed on this pass (e.g. because a metric’s groupByDimensions changed) is deleted, so a Decisioning Gate never reads a stale per-dimension value. Realtime-mode metrics are unaffected by this cron — they update live on every impression/outcome event and don’t need a scheduled recompute. Auth follows the same Authorization: Bearer $CRON_SECRET shape as the other /api/v1/cron/* routes and fails closed when CRON_SECRET is unset. Response shape: { status, results: { tenantsProcessed, totalComputed, totalSkipped, perTenantErrors[] }, timestamp }.
  • Schedule: wired by defaultcron.schedules.recomputeMetrics in helm/values.yaml (*/5 * * * *) for the chart-managed CronJob path, and the in-process maintenance scheduler (maintenance-scheduler.ts, every 5 minutes) for single-service/WORKER_INPROCESS deployments. Tune the interval or set enabled: false to disable; batch metrics then only refresh on a manual “Compute Now” in the UI.
  • Recommended cadence: every 5 minutes. Because recompute is gated per-metric by batchIntervalMin, ticking more often than your fastest batch metric’s interval is harmless — each tick is a cheap due-check per metric — and keeps short-interval metrics close to on-schedule.
  • See Behavioral Metrics — Compute Modes for how batch vs. realtime metrics differ.

Legacy /api/cron/tick

The original once-per-minute tick that evaluates every active alert rule (Phase 02) and runs the report-schedule loop (Phase 03) for every tenant. Superseded by the /api/v1/cron/* tier for Kubernetes deployments — the wired-cron jobs above split the same work across narrower handlers with concurrencyPolicy: Forbid and isolated retries. This route stays in the codebase only for the EventBridge pilot path where a single AWS scheduled rule fires POST /api/cron/tick against the API service. See EventBridge Setup for that runbook. New Helm deployments should leave the legacy tick unscheduled and rely on the cron.schedules block instead. Auth on the legacy route is wider on header shape — it accepts any of x-cron-token, x-cron-secret, or Authorization: Bearer .... The v1 tier accepts Authorization: Bearer <secret> only. Both tiers now treat either CRON_TOKEN or CRON_SECRET as a valid shared secret (the v1 routes resolve CRON_SECRET || CRON_TOKEN), so a CRON_TOKEN-only environment authorizes the whole cron tier.