Routing strategies, explained: when to use each one
Synapass ships five routing strategies, and choosing between them is the single highest-leverage decision most teams make in the gateway. Here is how to think about each one.
priority is the default and the right starting point: an ordered list, first healthy provider wins. Use it when you have a clear preference — your best model first, with cheaper or local models as ordered fallback. It is predictable, explainable, and cheap to reason about in incidents.
weighted spreads traffic by proportion. Use it for canarying a new model (send it 5% and watch scores), for A/B comparisons across providers, or for honoring commit-based pricing across two vendors. Pair it with analytics so the split earns its keep.
lowest_cost routes each request to the cheapest eligible model that satisfies the policy's capability requirements. It is honest because the registry holds real per-model pricing — the same data the pre-call ceiling check uses. Use it for batch, background, and cost-sensitive tenants.
lowest_latency routes on observed latency, not marketing numbers: the gateway watches real response times per provider and model. Use it for interactive paths — chat, autocomplete, support drafts — where every hundred milliseconds shows up in the product.
highest_quality routes on observed scores from evals and replay. Use it where correctness dominates: code review, summarization that feeds decisions, anything with a judge. Quality data takes time to accumulate, so start new workloads on priority and graduate them.
Whichever strategy you choose, fallback chains apply the same way: ordered, bounded, with deterministic jitter — and a started stream is never appended to, only reported. The provider field in the synapass block always names who answered, not who was tried first, so failovers are visible instead of silent.