Observability

Prometheus metrics, OTLP traces, JSON logs, and Grafana alerts.

Synapass emits Prometheus metrics, OpenTelemetry traces and structured JSON logs, and ships Grafana dashboards and alert rules. This page covers what each signal tells you and how to correlate them when something is wrong.

Cardinality is bounded by construction. Tenants, providers and models are labels. Request ids, user ids and prompts never are. Nothing per-request is a label — that is the single most common way an observability stack is destroyed.

Where the signals come from#

SignalEndpoint / sinkLossy?
MetricsGET /metrics on the gateway and worker listenersNever — counters are in-process
TracesOTLP to the collector in deploy/otel/collector-config.yamlBest effort
LogsJSON to stdout; Promtail ships them to LokiNever — stdout is the source of truth
Usage and tracesClickHouse, written asynchronouslyYes, by design
EventsNATS JetStreamYes, if the bus is down

The record pipeline batches and counts what it drops rather than blocking a response. synapass_async_queue_depth and synapass_async_dropped_total make that trade-off visible instead of invisible. What is synchronous — authentication, policy, budgets — is never lossy.

Metrics#

All metrics are prefixed synapass_. The workers use synapass_worker_*, so one scrape config and one dashboard cover both.

Gateway#

MetricLabelsUse it to
synapass_gateway_requests_totaltenant, provider, model, type, outcome, statusSee traffic and success rate
synapass_gateway_request_duration_secondstenant, provider, type, outcomeSee latency distribution
synapass_gateway_requests_in_flight—See saturation
synapass_gateway_request_size_bytes / response_size_bytestenant, typeSee payload growth
synapass_build_infoversion, commit, go_versionAnswer "which build is running?"

Provider#

MetricLabelsUse it to
synapass_provider_attempts_totaltenant, provider, model, outcome, attemptSee attempts per provider
synapass_provider_duration_secondstenant, provider, model, outcomeSee provider latency
synapass_provider_time_to_first_token_secondstenant, provider, modelSee streaming startup
synapass_provider_healthproviderSee breaker state (1 healthy, 0.5 degraded, 0 unhealthy)
synapass_provider_truncations_totaltenant, provider, reasonSee answers cut short by max_tokens or timeout
synapass_provider_completion_tokens_ratiotenant, providerSee headroom; a pile-up at 1.0 means the ceiling is too tight

Routing, policy and usage#

MetricUse it to
synapass_routing_decisions_totalCount routing decisions by strategy and outcome
synapass_routing_candidatesSee how many candidates survived filtering
synapass_routing_fallbacks_totalSee failover rate and which codes triggered it
synapass_policy_rate_limited_totalSee rate-limit rejections
synapass_policy_budget_blocked_totalSee budget denials
synapass_usage_tokens_totalSee prompt and completion token throughput
synapass_usage_cost_usd_totalSee estimated spend

Cache#

synapass_cache_hits_total{kind}, misses_total, bypass_total{reason}, lookup_duration_seconds, invalidations_total{scope,reason}, latency_saved_seconds{tenant,kind}, semantic_similarity. See Caching.

Intelligence#

MetricUse it to
synapass_classifier_requests_total{task}See the task mix
synapass_shaping_requests_total{step}See which shaping steps fire
synapass_guardrail_blocks_total{kind}See kill switches and caps firing
synapass_scoring_provider_scoreSee quality scores
synapass_eval_jobs_total, synapass_replay_jobs_totalSee offline work
synapass_feedback_events_totalSee user feedback volume

Tools and tunnels#

synapass_tools_runs_total{tenant,mode,status}, run_duration_seconds{mode}, invocations_total{tenant,tool,status}, persist_errors_total{tenant}. synapass_tunnel_up{target}, tunnel_sessions_total{target,outcome}, tunnel_restarts_total{target}.

Async pipeline#

synapass_async_queue_depth, dropped_total, flushed_total. A rising queue depth with drops means a dependency is slow or down.

Traces#

Spans cover classifier, shaping, cache, routing, guardrail, provider, tools and eval. Each carries the request id and trace id.

deploy/otel/collector-config.yaml fans out to a file sink by default. Point it at Tempo, Jaeger or another collector by uncommenting one exporter — the collector is the seam, so swapping backends is never a Go change.

With tracing enabled and no OTLP endpoint configured, the gateway refuses to start. Losing traces silently is worse than not starting.

Logs#

JSON to stdout. In Compose, docker compose logs -f gateway; under systemd, journalctl -u synapass -f.

FieldWhy it is there
request_id, trace_idCorrelate a log line with a trace and a request record
tenant, provider, modelFilter without reading the message
duration_msSpot slow paths
errorThe normalized message

Request bodies are not logged by default. A configurable header deny-list plus value redaction keep credentials out; the workers additionally redact values, because captured prompts would otherwise be the most likely path for personal data to escape.

Grafana#

Compose brings up Grafana on :3001 with deploy/grafana/dashboards/synapass-overview.json preloaded, and Prometheus scraping both the gateway and the workers.

The overview dashboard covers throughput, latency percentiles, success rate, token throughput, cost, provider health, fallback rate and cache hit rate.

Alert rules#

deploy/prometheus/rules/synapass.yml ships with alerts for the failures that matter:

AlertFires when
Gateway down/ready fails
High error rate5xx ratio above threshold
Provider unhealthyBreaker state stays unhealthy
Latency regressionp95 above threshold
Spend anomalyCost rate far above baseline
Telemetry lossAsync drops climbing
Truncation spiketruncations_total climbing

Load them into your own Prometheus if you are not using the Compose one.

Debugging by symptom#

Answers are stopping mid-sentence. This is the one symptom that used to be silent. Check, in order:

  1. synapass_provider_truncations_total{reason} — is it max_tokens or timeout?
  2. The response's synapass.completion block — requested_tokens, applied_tokens, budget_ms, finish_reason.
  3. The routing decision's timeout policy — a per_attempt shorter than a long generation truncates by definition.

A provider is getting no traffic. Check synapass_provider_health, then whether the policy's target list actually names it, then whether the capability filter excludes it for these requests.

Failover is not happening. Compare the error code against the policy's on_error_codes. That list is exhaustive when present.

Latency is worse than expected. Read synapass_routing_candidates — a single candidate means no choice was available, so latency is that provider's.

Telemetry seems to be missing rows. Check synapass_async_dropped_total and queue_depth before suspecting the queries.

Scripts#

bash
# Health, build identity, and where the process thinks it is
curl -s $GATEWAY/health | jq
curl -s $GATEWAY/version | jq

# Traffic and errors
curl -s $GATEWAY/metrics | grep synapass_gateway_requests_total
curl -s $GATEWAY/metrics | grep synapass_provider_health
curl -s $GATEWAY/metrics | grep synapass_provider_truncations_total

# One request, end to end
curl -s "$GATEWAY/admin/v1/requests/$REQUEST_ID" -H "Authorization: Bearer $SYNAPASS_ADMIN_KEY" | jq
curl -s "$GATEWAY/admin/v1/requests/$REQUEST_ID/explain" -H "Authorization: Bearer $SYNAPASS_ADMIN_KEY" | jq

Related: Architecture · Database · Troubleshooting · Back to README