A semantic cache for LLM APIs — and, more to the point, the safety gate and calibration harness that decide whether one should be turned on at all.
A normal cache is always right. A semantic cache is a classifier, and it is wrong silently: it serves a cached answer to a question that was not the same question, with a 200, in the normal latency budget, with no error anywhere.
$ python semcache.py demo
how many seats does the pro plan include?
similarity 1.000 -> SERVE
How many seats does the Starter plan include?
similarity 0.875 -> refuse (veto:entity)
a threshold-only cache at tau=0.85 WOULD have served this.
How many seats does the Pro plan not include?
similarity 0.943 -> refuse (veto:polarity)
a threshold-only cache at tau=0.85 WOULD have served this.
That second number is the thesis. Inserting the word "not" moves cosine by 5% and changes the answer by 100%. A single cosine threshold cannot be made safe, because the distance it measures is not the distance that matters.
Verified against the real API, not just against my own eval — the same three requests
through the proxy with ANTHROPIC_API_KEY set:
cold (real Claude) X-Cache=MISS sim=0.0000 reason=cold
"A semantic cache stores query results based on meaning rather than exact text..."
normalised dup X-Cache=HIT sim=1.0000 reason=exact
(same answer, no upstream call)
polarity flip X-Cache=MISS sim=0.9428 reason=veto:polarity
"A semantic cache is not a simple key-value lookup that only matches exact text..."
Claude's answer to the flipped question is genuinely different. At sim = 0.9428 a
threshold-only cache would have served the first answer to the second question — and the
user would have had no way to tell.
The harness was built to tune a semantic cache. Run against a lexical embedder, it says the cache should not be turned on:
$ python semcache.py calibrate
AUC overall 0.323
AUC adversarial only 0.158 <- inverted, not merely overlapping
lambda meaning tau* hit FH/served
1 a wrong answer costs one wasted API call 0.98 0.08 0.33
10 ... ten wasted calls 1.02 0.00 0.00
inf never serve a wrong answer 1.02 0.00 0.00
at lambda=1 tau*=0.98: true hits 8, of which NOVEL 0, bought with 4 false hits
tau* = 1.02 means "never serve anything". And NOVEL = 0: every true hit the
semantic layer found was already caught, free and risk-free, by the exact-hash layer in
front of it. The request-level replay agrees independently:
$ python semcache.py eval
served total 130 of which exact (free, zero risk) 130 semantic 0
On a lexical embedder, a correctly gated semantic cache degenerates into an exact-match
cache. That is the gate working, not the gate failing — every candidate it declined was
one a threshold-only cache would have served, and the ablation prices those at
FH/served 0.15. A pretrained embedder may well change the verdict; the seam is
embed(..., kind="openai") and it has not been run, so it is not claimed.
I am shipping the harness that produced this number rather than a cache with a cost-reduction percentage on it. The percentage would have been easy and would have been a property of a dataset I wrote myself.
If the rows were embedders this would be an embedding benchmark. The embedder is held fixed; what varies is the gate.
$ python semcache.py eval
=== gate ablation (replay of 400 requests, embedder=hashed HELD FIXED, tau=0.85) ===
gate hit FH/srv FPR(lex) FPR(sem) d(true hits) saved$
tau only (the sibling's cache) 0.39 0.15 0.22 0.24 +0 0.1363
+ numeric 0.39 0.15 0.17 0.24 +0 0.1363
+ polarity 0.38 0.04 0.06 0.24 +13 0.1330
+ entity 0.37 0.02 0.00 0.18 +13 0.1303
+ qualifier 0.37 0.02 0.00 0.18 +13 0.1303
+ order (shipped gate) 0.36 0.00 0.00 0.06 +13 0.1279
Three things to read in it:
FPR(sem) never reaches zero, and must not. Some DISTINCT pairs are provably
beyond every deterministic veto — see the permutation proof below. CI fails if that
column hits 0.00, because that means the eval broke, not that the gate improved.
d(true hits) is POSITIVE (+13). A veto does not merely delete a serve. The request
falls through to the origin and is stored, so the next identical request is caught
free and correctly by the exact layer. A pairwise eval sees only the deleted serve and
prices a veto as pure loss. It is not. This is why the project has a request-level
replay and not just a pair table.
The ablation has a price column at all. A veto ablation printed without one is an advertisement.
1. Word order. "convert 100 USD to EUR" and "convert 100 EUR to USD" have an
identical token multiset, so cosine is exactly 1.0. No threshold on any bag of words
separates them, ever. This is a unit test asserting == 1.0, not a tolerance:
def test_bag_of_words_cannot_see_word_order():
assert float(np.dot(S.embed("convert 100 USD to EUR", "hashed"),
S.embed("convert 100 EUR to USD", "hashed"))) == pytest.approx(1.0, abs=1e-12)The order veto reads raw text — strictly more information than the embedding — and that
is the entire reason vetoes exist.
2. Answers that decay. This one I planted less than I discovered, and it is the risk that survives everything:
=== the risk that survives every veto: answers that decay ===
volatility policy hit FH of which volatile saved$
allow (no policy — naive cache) 0.37 11 11 0.1323
short TTL on volatile entries 0.34 0 0 0.1193
never store them (shipped) 0.33 0 0 0.1143
11 of 11 residual false hits are volatile. "how many seats are left?" asked twice is
sim == 1.000 on an identical token multiset — so no embedder can see it, no
text-reading veto can see it, and the free exact-hash layer serves it wrong too. This
is not a semantic-cache problem; it is a cache problem that a semantic cache amplifies,
because it has more ways to reach a stale entry. The only defences are freshness and
refusal. deny ships: a TTL longer than a decay period you guessed wrong is a silent
wrong answer, and refusal is not.
Conflating these is the standard dishonesty in this space.
| name | denominator | property of |
|---|---|---|
FPR_pair |
DISTINCT pairs | the classifier — prevalence-free, transferable |
FHR_served = 1 − precision |
served hits | the product |
FHR_req |
all requests | the SLO |
Calibrate on FPR_pair, gate on FHR_served, print all three with k/n and a Wilson
interval. Every rate this project prints carries its denominator.
And attribution is measured prevalence-free, on pairs — because whether a veto fires in a replay is a fact about my traffic generator, not about the veto:
veto blocks DISTINCT costs REUSE of 24 risk / 8 value above tau
numeric 2 0
polarity 4 0
entity 8 0
qualifier 1 0
temporal 0 0 <- NOT SHIPPED (dead at this tau)
order 12 0
temporal is not shipped, and that is a measured decision: it blocks zero of the
pairs that clear tau, because changing tense changes tokens and similarity already
rejects them. A veto unreachable at the operating point is dead code, and shipping dead
code as a safety feature is exactly what this project exists to measure.
test_temporal_veto_still_unreachable_at_tau is the tripwire that says when to put it
back.
Assigning a dollar value to a wrong answer is the marketing move. This inverts it: sweep λ ("how many wasted API calls is one wrong answer worth?") and print the τ each implies. As λ → ∞ the optimal semantic cache degenerates to an exact-match cache — which is the same statement as the headline, arrived at from the cost side.
Savings are reported as an identity, never as a headline percentage:
baseline (cache disabled) $0.3658
gross savings $0.1280 (35.0%)
embedding cost -$0.000028 (paid on 100% of non-exact requests)
NET savings $0.1279 (35.0%)
break-even hit rate 0.01232%
savings ~= pi x TPR x (1 - c_embed/c_call)
TPR here 0.33 is what I measured; pi is yours. The pi in this replay is 0.70
BY CONSTRUCTION — I wrote the stream, so I chose that number.
False hits are not netted off, because pricing a wrong answer is the thing this project refuses to do.
An in-process semantic cache is a memoization dict. --shards N partitions the stream
round-robin across N independent stores — that is the uvicorn --workers N result,
deterministically, with no processes to spawn:
$ python semcache.py sweep
shards hit saved$
1 0.33 0.1143
4 0.17 0.0595
8 0.10 0.0357
The argument for Redis is shared state, not the vector index.
Do not run the lookup inline in async def. It is a CPU-bound matmul, so inline it is
the event loop's global serialization point:
$ python semcache.py bench
50,000 entries x 1536 dims
concurrency 1 lookup inline p99 thread p99
1 4.51ms 4.5ms 4.6ms
64 4.51ms 288.9ms 10.9ms
At concurrency 64 the inline p99 is worse than the API call the cache exists to avoid — a cache that makes p99 worse. At 500 entries the matmul is microseconds and inline is correct; the crossover is a number, not a preference.
No ANN index, cut with the benchmark: exact matmul is 0.27 ms at 500 entries and recall is 1.0 by construction, and 1536 floats × 50k is 300 MB — memory binds before latency. HNSW earns its keep at 10⁶.
False hits grow with cache size at fixed τ, because argmax is adversarial: more
entries means more chances for a spurious maximum. sweep --what cache-size prints it at
both gates on purpose — a flat column at the shipped gate proves nothing on its own.
A semantic cache is a first-turn cache. Prior turns are an exact-match gate (embedding only the last turn without gating history is a wrong-answer generator: "what about the second one?"), so:
turn 1 0.42 (124/294) [0.37, 0.48]
turn 2 0.03 (2/69) [0.01, 0.10]
turn 3+ 0.11 (4/37) [0.04, 0.25]
Structural, not a tuning problem.
OpenAI-compatible, so an existing client changes nothing but its base URL.
docker compose up --build
curl -sS -D- -X POST localhost:8080/v1/chat/completions \
-H 'Authorization: Bearer sk-your-key' -H 'Content-Type: application/json' \
-d '{"model":"claude-haiku-4-5","messages":[{"role":"user","content":"How many seats does the Pro plan include?"}]}'
# X-Cache: MISS | HIT | COALESCED X-Cache-Similarity: 0.9428 X-Cache-Reason: veto:polarity- Scope is a deny-list, not an allow-list. Hash the whole body minus an explicit deny-list, so a parameter the provider adds next quarter produces a new scope — cold but correct. An allow-list fails open.
- The principal is server-derived (a hash of the presented credential), never a client
header, and no credential means no cache. A shared
anonymousscope is precisely how one user's answer reaches another. The raw key never enters a scope input or a metric label. max_tokensis not in the scope key. Bumping 512 → 1024 still hits every entry that finished naturally; the length question is answered byis_servable(), which can see the stored completion length.finish_reason == "tool_calls"is never stored or replayed — that is a decision to perform a side effect against state that no longer holds.- Hits replay as SSE when
stream: true. Accepting the field and returning a JSON body is simply broken. - Single-flight: no
awaitbetween lookup and registry insert (oneawaitand two coroutines both become leader); the task is owned by the registry so a leader's disconnect cannot cancel the followers; failures are never cached. /metricsis hand-rolled — cumulative buckets,+Inf == _count,HELP/TYPEonce per family, counters ending_total, bounded label values, and deliberately nohit_ratiogauge (a ratio cannot be aggregated across instances; export the parts and divide in PromQL).test_metrics_parsesasserts every one of those rules, which is what makes hand-rolling defensible instead of lazy.
rate(semcache_requests_total{result="hit"}[5m]) / rate(semcache_requests_total[5m])
sum by (reason) (rate(semcache_refusals_total[5m]))
histogram_quantile(0.99, rate(semcache_lookup_seconds_bucket[5m]))
rate(semcache_coalesced_total[5m]) / rate(semcache_upstream_calls_total[5m])
pairs.jsonlwas committed before any veto code existed (a3178bf, thene56987b).git logis the anti-circularity evidence. I labelled 68 pairsblocked_by: semantic, predicting no deterministic rule could reach them; the vetoes reach 43 of them. I reported that as a finding and left the labels alone — editing labels after seeing the vetoes is the exact circularity the commit ordering exists to prevent, and the disagreement is information about my intuition, which was wrong in the safe direction.- Ground truth is mechanical.
answer_key = (intent, entity, qualifier, polarity, quantity, tense);REUSEiff the keys are equal. Hand labels are asserted against the derived key, which catches authoring bugs. And the label is on the answer, not the query — query-similarity labels only measure whether the embedder agrees with a human's notion of question similarity, which is circular. test_gate_never_reads_a_labelwalks the AST ofgateand every veto and fails if any of them can seelabel,family,blocked_by, orharm.- CI fails on suspicious perfection, not just on regression: a veto that blocks nothing, a blind spot that stops being blind, a permutation pair that stops scoring 1.0, fewer than 8 candidates per lookup (the replay would not be exercising argmax), or a 0.00 product metric that the prevalence-free metric contradicts.
test_stampede_negative_controlruns the same 50 concurrent requests without single-flight and asserts all 50 reach the origin — otherwise the coalescing test would also pass if the requests had merely happened to serialize.
Bugs this harness caught in my own code, all of which would have shipped green:
_argumentsadded the final content token — which in "when does the trial expire?" is the predicate, shared by the very pair the entity veto exists to catch. It missed 3/8 of its own adversarial family while firing on paraphrases.- The L1 exact layer checked the global TTL instead of the entry's, so the volatility policy was bypassed by exactly the layer volatile traffic lands on.
- The traffic generator buried volatile questions in the Zipf tail, producing one volatile request in 400 and a flattering 0.00 false-hit rate. Volatile questions are head traffic — people ask them repeatedly because the answer keeps changing.
- The decorative-veto gate fired under
--embed random, which is that ablation being asked a question it cannot answer.
semcache.py normalize, scope, embed seam, vetoes, gate, store, single-flight,
cost model, calibrate / eval / sweep / bench / demo, CLI
app.py OpenAI-compatible proxy, /metrics, SSE replay
pairs.jsonl 154 labelled pairs, 8 adversarial families, blocked_by lexical|semantic
traffic.jsonl 400 replay requests, 2 tenants, 14 scopes, multi-turn depth
test_semcache.py 39 offline tests
pip install -r requirements.txt
python -m pytest -q
python semcache.py demo && python semcache.py calibrate && python semcache.py evalRuns fully offline and deterministically — no key, no network, no clock dependence.
- RedisVL / Qdrant / any ANN — cut with the benchmark above.
prometheus_client— 60 hand-rolled lines plus a parser test carry more signal.- A TTL / eviction-policy bake-off — on a synthetic replay that measures my own generator, not the policy.
- LLM-as-judge verification of hits — tempting and wrong: it destroys the latency claim
and most of the cost claim. It pays only when
c_judge < c_call × P(false hit) × λ, which is arithmetic you can do yourself from the numbers above.
MIT