Synapass / Changelog

What each delivery phase added, and why. For current behaviour, follow the documentation index instead — this page is history.

The project's delivery is organised in phases. Each one is independently deployable and additive; nothing in a later phase is required by an earlier one.

Phase 1 — the gateway#

The synchronous inference path.

AreaDelivered
GatewayOpenAI-compatible /v1/chat/completions (streaming and buffered), /v1/models
AdaptersOpenAI, Anthropic Messages, Ollama, vLLM, any OpenAI-compatible server
RoutingFive strategies, retry, per-attempt and total timeouts, fallback chains
HealthActive probes plus passive observation, circuit breaker per provider
PolicyMatch by model, request type, tenant, key, prompt size, capability
CostPer-request ceilings, request and token rate limits, daily and monthly budgets
AuthSHA-256-hashed API keys with scopes, revocation, caching
ObservabilityPrometheus, OTLP traces, JSON logs, ClickHouse analytics, Grafana, alert rules
ConsoleNext.js dashboard across usage, providers, policies, budgets, tenants and keys
DeploymentDocker Compose and native systemd

Phase 2 — the policy-driven control plane#

Turns the router into something that explains itself. Every request now flows through ten visible stages: authenticate → policy → classify → shape → cache → health and scores → route → execute → fall back → persist.

AreaDelivered
Policy engineTenant/key/endpoint rules, model and provider allow/deny, cost and latency ceilings, region and sensitivity constraints, batch vs interactive
ClassificationRules-first, deterministic task inference across eleven task types, mirrored in Python
Intelligent routingTask-, cost-, latency-, capability- and score-aware ordering with derived timeout, retry and shaping strategies
Prompt shapingNormalization, compression, context trimming, history summarization, guardrail injection — all visible in traces
CachingExact, prefix and semantic tiers with sensitive-request bypass, stats and invalidation
Provider scoringExplainable success/latency/cost/feedback blends with per-task breakdowns
Eval and replayAsync replay jobs over NATS, offline comparisons, golden cases, regression flags
GuardrailsKill switches, tenant overrides, hard caps, fallback blocks, forced circuits, visible deny reasons
DashboardScores, cache analytics, replay, evaluations, endpoints, audit, kill/revive, explain view

Phase 3 — management and provisioning#

Turns a view-only gateway into a manageable control plane. Providers, models, tenants, keys, policies, endpoints and overrides are created, edited, disabled and deleted through the API and the dashboard: no config-file edits, no restarts.

AreaDelivered
CRUDFull management for providers, models, tenants, keys and policies
CredentialsAES-256-GCM sealed provider secrets, metadata-only reads, audit on create and rotate
Connectivity testsThree checks — connectivity, model listing, sample completion — persisted per run
Model discoveryPopulate the registry from a provider's remote catalogue, preserving local edits
Ownershipmanaged_by so a dashboard edit survives a deploy
Key rotationSwap a secret in place; the old one stops working immediately
DashboardProvider/model/tenant/key/policy/endpoint/override management

Phase 4 — tool calling and agent runs#

Turns the gateway from a request router into an agent runtime. Models can call tools, answers can be contract-checked against a JSON Schema, and every gateway-side run leaves a durable step trace.

AreaDelivered
RegistryTool registration with an allow-listed JSON Schema subset
Execution modesClient-executed (default) and bounded gateway-side execution
SafetyExecutability derived from kind, never granted; now and echo are the only built-ins
BoundsSteps, calls, wall clock, allowed/denied globs — operator-owned, client-narrowable
Structured outputjson_object and json_schema with conformance reporting and unambiguous repair
Durabilitytool_invocations → tool_executions → agent_runs → agent_steps
DashboardTools, tool policies, agent runs with step traces

Phase 5 — the local response cache#

Makes Synapass faster and cheaper by reusing prior responses wherever it is safe, with safety as the first consideration rather than an afterthought.

AreaDelivered
TiersExact, prefix (short prompts only) and semantic over Redis
KeyingTenant, key, model, tools, generation settings, response contract, policy and endpoint — everything that can change an answer
PolicyEleven named bypass reasons; per-scope rules resolved key → endpoint → tenant → global
InvalidationScoped, audited, published on NATS, with honest blast-radius recording
ObservabilityHit/miss/bypass metrics, CacheTrace, and a full /cache dashboard page

Public landing page#

A browser reaching the gateway root now gets a small HTML page naming the real endpoints, and JSON for everything else. A gateway reached from a tunnel URL or typed into a browser should not answer with a 404 that reads as a broken deployment.

Cloudflare tunnels#

Temporary public exposure of one local service through a supervised cloudflared child process. Opt-in twice — feature flag and explicit create — and loopback targets only, so a tunnel can never become a proxy for someone else's infrastructure. See installation/cloudflare-tunnel.md.

Truncation fix#

Long answers were stopping mid-sentence for three independent reasons, all fixed:

  1. The soft latency target was reused as a hard per-attempt deadline. The default 10s target cancelled every generation longer than ten seconds. It is now a ranking signal only, and generation budgets come from the timeout policy. An explicitly configured budget is never raised behind the operator's back.
  2. The shared transport set ResponseHeaderTimeout to the first-token budget. For a buffered request a provider only sends headers once the whole answer exists, so long completions were aborted before a byte of body arrived. Removed; per-attempt deadlines come from the request context.
  3. The policy output ceiling was sent as the client's max_tokens. A client that asked for no limit was given one and nothing reported the cut. A ceiling is now only sent when the client requested one — and Anthropic's mandatory field gets a generous default rather than a small constant.

Truncation is now always reported in synapass.completion (truncated, reason, requested_tokens, applied_tokens, budget_ms), counted in synapass_provider_truncations_total, and never conflated with a completed answer. A stream that has begun is committed: the executor reports a stream-started error rather than appending a second attempt's tokens to the first.

Documentation#

This documentation set. The previous phase-by-phase notes were folded into topic-based pages so each fact is documented once and cross-linked, rather than restated per phase.


Related: Documentation index · Architecture · Back to README

Deeper context for any entry lives in the docs. Browse documentation