One endpoint in front of every provider.
Synapass is an inference gateway and control plane: a single OpenAI-compatible API that decides per request which provider serves it under policy — and records why.
Four modules, one panel.
Gateway, control plane, routing layer, and acceleration platform. Each is a module below; each links to its docs page.
- A multi-provider AI gatewayOne OpenAI-compatible endpoint in front of OpenAI, Anthropic, Ollama, vLLM, and any OpenAI-compatible API. Clients change a base URL; the gateway handles adapters, errors, retries, and provider differences.01 endpointevery provider
- An inference control planeProviders, models, policies, endpoints, tenants, keys, budgets, cache, tools, and tunnels — managed live. Shifting production traffic is an edit, not a deployment.00 restartsto re-route
- A routing and policy layerFive strategies over a registry with real context windows, capabilities, and prices — guarded by specificity-ordered policies.05 patternson panel
- Caching, tools, and analyticsExact, prefix, and semantic cache tiers; client or bounded gateway-side tool execution with step traces; throughput, latency, spend, and quality on every request.09 stepsfully traced
Every request walks the same nine steps.
Authenticate, resolve policy, enforce budgets, check cache, classify the task, shape the prompt, route, execute with failover, run the tools loop. Select a step for its explain detail.
key syn_live_… → tenant acme · scope inference
The failure modes you stop owning.
Direct integrations fail the same six ways. The gateway answers each one the same way, every time.
| Situation | Direct integrations | With Synapass |
|---|---|---|
| Provider outage | Your outage | Automatic failover to the next provider |
| New model release | SDK bump + deploy per service | Sync models, repoint an alias |
| Runaway loop | A surprising invoice | Pre-call ceiling stops it first |
| “Why this answer?” | Invoices and guessing | /explain: policy, cache, routing, cost |
| Identical repeat prompts | Paid again, slowly | Cache tiers answer in milliseconds |
| Adding a tenant | Config + code + deploy | Mint a scoped key |
Watch it work from the console.
Volume, success rate, latency, spend, provider health — and a per-request explain view for every call. Operated, not observed.
Requests · 24h
1.28M
Success
99.98%
P95
1.9s
Providers
Live bus
- openai-prodgpt-4o-mini812 ms62% bus
- anthropic-euclaude-3-5-haiku1.1 s27% bus
- local-ollamallama3.1:8b340 ms11% bus
- legacy-vllmmistral-7b—0% bus
Illustrative bus levels — the console shows your live channels after install.
Hear it run your traffic.
One command starts the full stack. Your first routed call lands in about five minutes — with the lineage to prove it.