Database

Tables, migrations, and persistence boundaries.

Synapass keeps state in four stores, chosen so each one is the right shape for what it holds. This page covers what lives where, how migrations work, and the persistence boundaries that matter when something goes wrong.

The four stores#

StoreRoleIf it is down
PostgreSQLSystem of recordGateway reports unready; a load balancer should drain it
RedisLimits, budgets, credential cache, response cacheDegraded: limits fall back to per-process, caching stops; readiness still passes
ClickHouseTraces and analyticsAnalytics writes are dropped and counted; inference unaffected
NATS JetStreamUsage events, eval and replay jobs, audit streamAsync work is dropped and counted; inference unaffected

PostgreSQL needs transactions and foreign keys — a policy write that half-applies is worse than a failed write. Redis needs atomic increments on high-churn ephemeral data. ClickHouse is columnar and analytical; one row per request per attempt would bloat Postgres and make dashboards slow. NATS gives durable consumers, so a worker restart resumes rather than losing work.

The usage row is in Postgres and the trace is in ClickHouse on purpose: usage must be complete because it is billed, whereas traces are high-volume and analytical. They duplicate a few fields for that reason.

Postgres tables#

Core#

TableHolds
tenantsTenant identity, slug, plan, status
api_keysSHA-256 digest, prefix, scopes, expiry, routing policy binding
providersKind, base URL, status, priority, api_key_env, managed_by
provider_credentialsAES-256-GCM sealed secrets, one active per provider
provider_test_resultsConnectivity, model-listing and sample-check history
modelsContext window, max output, pricing, capabilities, priority, status
routing_policiesThe rule sets, with targets, limits, budgets and timeouts
usage_recordsPer-request tokens and estimated cost — the billable row
request_logsRequest and response metadata for the operator view
audit_eventsControl-plane mutations, with before/after values
budgetsDaily and monthly spend ceilings, plus hard per-request caps
provider_status_snapshotsHealth history
settingsRuntime settings

Intelligence#

TableHolds
policy_decisionsThe verdict per request, with the reason
task_classificationsTask label, confidence and the signals that produced it
prompt_shapesEach shaping step and its token delta
cache_entriesPrompt hash and preview, model, provider, policy, hit counts
cache_policiesPer-scope cache rules
cache_invalidationsFlush audit: scope, target, reason, actor, removed count
provider_scores, model_scoresExplainable quality scores per window
replay_jobs, evaluation_runs, evaluation_resultsOffline comparison
feedback_eventsUser-submitted scores and comments
endpointsNamed scopes with routing overrides
circuit_breaker_statePersisted breaker state across restarts
audit_overridesHealth and routing overrides

Tools and tunnels#

TableHolds
toolsThe tool registry
tool_policiesMode, bounds, allow/deny globs
tool_invocations, tool_executionsEvery model-requested call and its execution
agent_runs, agent_stepsBounded runs and their step traces
tunnel_sessionsTunnel lifecycle: target, URL, timing, stop reason

ClickHouse tables#

TableHolds
synapass.request_tracesOne row per request, with span timings
synapass.trace_attemptsOne row per provider attempt
synapass.route_decisionsChosen target, candidates, rejections
synapass.usage_eventsRaw token and cost events
synapass.usage_dailyPre-aggregated daily rollups
synapass.provider_status_snapshotsHealth over time
synapass.request_intelligenceTask, shaping, policy verdict, cache kind, scores
synapass.provider_scoresScore history
synapass.eval_resultsEvaluation and replay outcomes
synapass.cache_eventsHit, miss and bypass events

Redis keys#

PatternHolds
response:tenant:<id>:exact:<hash>Cached response bodies, tenant-namespaced
response:tenant:<id>:prefix:<hash>Prefix-tier entries
Rate-limit countersRequests and tokens per minute
Budget countersKeyed by rendered period label (d:…, w:…, m:…, total) with TTL from the label
Credential cacheResolved API keys, invalidated on revocation

Bodies live in Redis; only metadata reaches Postgres. That keeps the system of record small while leaving the dashboard able to inspect what is cached without ever serving a stored body.

NATS subjects#

SubjectCarries
Usage eventsPer-request usage for rollups
Eval and replay jobsDurable work for the workers
ar.tool.run.completedFinished run summaries, retained 30 days in SYNAPASS_EVENTS
ar.cache.invalidatedFlush summary
ar.tunnel.statusTunnel lifecycle
Audit streamControl-plane events

Migrations#

Migrations are embedded in the binary and applied in order. Each is recorded in schema_migrations with a checksum, and a mismatch on an already-applied migration is an error, not a silent re-run — a migration that changed after it was applied means the two databases are not the same shape.

MigrationAdds
0001_initCore schema
0002_phase2Intelligence tables
0003_audit_phase2Extra audit columns
0004_phase3Provider credentials and test results
0005_phase4Tool registry, policies, invocations, runs, steps
0006_phase5Cache policies and invalidation trail
0007_tunnelTunnel sessions
0008_audit_tunnelExtra audit columns
ClickHouse 0001Trace and usage tables
ClickHouse 0002Intelligence tables

Applying them#

Compose sets SYNAPASS_POSTGRES_AUTO_MIGRATE=true, so the gateway migrates at startup. Under change control, turn that off and migrate as a release step:

bash
synapass migrate

The same command migrates both Postgres and ClickHouse.

Ownership: managed_by#

Catalogue rows carry managed_by: bootstrap | api. The bootstrapper applies the configuration file only to rows it owns, so a provider, model or policy edited through the API survives a deploy. To hand a row back to the file, delete it through the API and let the seeder recreate it.

This is the mechanism that stops a dashboard edit from being silently reverted by the next docker compose up.

Persistence boundaries worth knowing#

The record pipeline is lossy on purpose. Usage rows, traces and audit events are written asynchronously and batched. A response never blocks on a write. The cost is that a crash can lose the last few seconds of telemetry, which is why synapass_async_dropped_total and synapass_async_queue_depth exist. What is synchronous — authentication, policy, budgets — is never lossy.

Audit covers the control plane, not inference traffic. A provider created through the API is audited, with before and after values. An inference request is not: its accountability need is served by usage_records, and auditing every request would dwarf the signal. An audit write failure is logged and never fails the mutation it accompanies, because refusing a completed action would leave the system matching neither the operator's intent nor the audit log.

Overrides are events, not edits. Revoking a kill switch means writing the inverse row (enabled: false), so the log keeps what was true and when.

Tool run rows are written before the first step, because agent_steps has a foreign key to agent_runs. A persistence failure does not fail the request, but it is logged, counted and flagged on the event.

Retention#

DataKept
Usage recordsUntil you prune them; they are the billing basis
Request traces in ClickHouseTTL-managed; prune by partition
Audit eventsUntil you prune them
NATS SYNAPASS_EVENTS30 days
Redis response cachecache.response_ttl, default 5m
Budget countersTTL derived from the period label

Prune ClickHouse by partition rather than row-by-row; it is built for whole-part deletes.

Common problems#

SymptomCauseFix
/ready returns 503Postgres unreachableCheck connectivity and credentials
/ready says redis: degradedRedis downLimits are per-process; caching stopped. Fix Redis, or accept degraded mode.
synapass_async_dropped_total climbingClickHouse or NATS down, or the queue is saturatedRestore the dependency; watch queue depth
migration checksum mismatchA migration file changed after it was appliedRestore the original file, or reconcile deliberately
Dashboard edits reverted on restartThe row is bootstrap-managedRe-create it through the API so it becomes api-managed
A budget never tripsCounter keyed under a different labelDo not hand-roll counter keys; use policy.PeriodKey

Related: Architecture · Observability · Providers · Back to README