All posts
LaunchArchitecture

One endpoint, every provider: why we built Synapass

2026-09-18 6 min read

Most inference stacks were never chosen. They accreted: a provider SDK here, a copied retry loop there, a model name hardcoded wherever it was convenient at the time. Each decision was small. Together they became an operating burden that competes with the product — one upstream hiccup becomes an outage, one runaway loop becomes an invoice, and nobody can answer which model produced last week's bad answers.

Synapass is our answer: an inference gateway and control plane that sits between your applications and your providers. It presents a single OpenAI-compatible endpoint and decides — per request, under policy — which provider serves it, what happens when that provider fails, and what it cost. Changing which model answers production traffic is a dashboard edit rather than a deployment.

Three design choices define the project. First, OpenAI compatibility is a hard constraint: any client that speaks OpenAI works by changing its base URL, and the extra synapass fields are namespaced so existing SDKs ignore them. Second, every decision is recorded — task classification, policy verdict, cache decision, routing reason — readable back per request, because a gateway that cannot explain itself cannot be operated. Third, spend is enforced before the call, not reconciled afterwards: ceilings are checked against projected cost, so a loop bug meets a limit instead of a bill.

The stack is deliberately boring technology: Go on the gateway, PostgreSQL as the system of record, Redis for limits and cache, ClickHouse for analytics, NATS as the event bus. Each store was chosen for the shape of what it holds. One command starts all of it, and re-running that command is a safe restart-and-verify — because day-two operations should be boring too.

Synapass is open source under the Apache License 2.0. If you run inference across more than one provider — or plan to — start with getting started and make your first routed call in five minutes.

Try what you just read about.

One command starts the full stack — routed calls in five minutes.