All posts
CachingGuide

Caching LLM responses without the footguns

2026-10-01 7 min read

Caching LLM responses is the highest-ROI optimization most teams never turn on, because the failure modes are scary: one tenant's answer served to another, a stale answer served after the world changed, a cache that silently degrades quality. Synapass's cache tiers are designed around those fears, and caching ships off by default so you opt in deliberately.

The three tiers answer progressively fuzzier questions. Exact serves byte-identical requests — deterministic system prompts, repeated evaluations, CI fixtures. Prefix serves requests that share a long common prefix, like a big retrieved context with a small varying question. Semantic serves paraphrases: it matches on meaning within a similarity threshold you control, with per-scope rules about what may be semantically cached at all.

Two properties make this safe. First, tenant isolation is structural: the tenant is part of the cache key and the Redis namespace, so a hit mathematically cannot cross tenants. Second, invalidation is scoped and audited: flush a scope, not the world, and the flush is recorded with who did it and what it covered.

The dashboard shows hit rates and — just as important — bypass reasons, so you can see whether misses come from policy, sensitivity flags, or genuinely novel prompts. Start with exact on deterministic workloads, measure, then widen to prefix and semantic where the bypass data justifies it.

A cache that cannot explain its misses is a cache you will eventually distrust. Ours explains both hits and misses, per request, in the same lineage block as everything else.

Try what you just read about.

One command starts the full stack — routed calls in five minutes.