Skip to content

Repository files navigation

semantic-cache-gate

CI

A semantic cache for LLM APIs — and, more to the point, the safety gate and calibration harness that decide whether one should be turned on at all.

A normal cache is always right. A semantic cache is a classifier, and it is wrong silently: it serves a cached answer to a question that was not the same question, with a 200, in the normal latency budget, with no error anywhere.

$ python semcache.py demo

  how many seats does the pro plan include?
    similarity 1.000   ->   SERVE

  How many seats does the Starter plan include?
    similarity 0.875   ->   refuse (veto:entity)
    a threshold-only cache at tau=0.85 WOULD have served this.

  How many seats does the Pro plan not include?
    similarity 0.943   ->   refuse (veto:polarity)
    a threshold-only cache at tau=0.85 WOULD have served this.

That second number is the thesis. Inserting the word "not" moves cosine by 5% and changes the answer by 100%. A single cosine threshold cannot be made safe, because the distance it measures is not the distance that matters.

Verified against the real API, not just against my own eval — the same three requests through the proxy with ANTHROPIC_API_KEY set:

cold (real Claude)   X-Cache=MISS  sim=0.0000  reason=cold
   "A semantic cache stores query results based on meaning rather than exact text..."
normalised dup       X-Cache=HIT   sim=1.0000  reason=exact
   (same answer, no upstream call)
polarity flip        X-Cache=MISS  sim=0.9428  reason=veto:polarity
   "A semantic cache is not a simple key-value lookup that only matches exact text..."

Claude's answer to the flipped question is genuinely different. At sim = 0.9428 a threshold-only cache would have served the first answer to the second question — and the user would have had no way to tell.


The result I did not want

The harness was built to tune a semantic cache. Run against a lexical embedder, it says the cache should not be turned on:

$ python semcache.py calibrate
  AUC overall            0.323
  AUC adversarial only   0.158        <- inverted, not merely overlapping

     lambda  meaning                                  tau*    hit  FH/served
          1  a wrong answer costs one wasted API call  0.98   0.08       0.33
         10  ... ten wasted calls                      1.02   0.00       0.00
        inf  never serve a wrong answer                1.02   0.00       0.00

  at lambda=1  tau*=0.98: true hits 8, of which NOVEL 0, bought with 4 false hits

tau* = 1.02 means "never serve anything". And NOVEL = 0: every true hit the semantic layer found was already caught, free and risk-free, by the exact-hash layer in front of it. The request-level replay agrees independently:

$ python semcache.py eval
  served total 130   of which exact (free, zero risk) 130   semantic 0

On a lexical embedder, a correctly gated semantic cache degenerates into an exact-match cache. That is the gate working, not the gate failing — every candidate it declined was one a threshold-only cache would have served, and the ablation prices those at FH/served 0.15. A pretrained embedder may well change the verdict; the seam is embed(..., kind="openai") and it has not been run, so it is not claimed.

I am shipping the harness that produced this number rather than a cache with a cost-reduction percentage on it. The percentage would have been easy and would have been a property of a dataset I wrote myself.


The headline table: rows are gate configurations, not embedders

If the rows were embedders this would be an embedding benchmark. The embedder is held fixed; what varies is the gate.

$ python semcache.py eval

=== gate ablation (replay of 400 requests, embedder=hashed HELD FIXED, tau=0.85) ===
  gate                               hit  FH/srv  FPR(lex)  FPR(sem)  d(true hits)   saved$
  tau only (the sibling's cache)    0.39    0.15      0.22      0.24            +0   0.1363
  + numeric                         0.39    0.15      0.17      0.24            +0   0.1363
  + polarity                        0.38    0.04      0.06      0.24           +13   0.1330
  + entity                          0.37    0.02      0.00      0.18           +13   0.1303
  + qualifier                       0.37    0.02      0.00      0.18           +13   0.1303
  + order (shipped gate)            0.36    0.00      0.00      0.06           +13   0.1279

Three things to read in it:

FPR(sem) never reaches zero, and must not. Some DISTINCT pairs are provably beyond every deterministic veto — see the permutation proof below. CI fails if that column hits 0.00, because that means the eval broke, not that the gate improved.

d(true hits) is POSITIVE (+13). A veto does not merely delete a serve. The request falls through to the origin and is stored, so the next identical request is caught free and correctly by the exact layer. A pairwise eval sees only the deleted serve and prices a veto as pure loss. It is not. This is why the project has a request-level replay and not just a pair table.

The ablation has a price column at all. A veto ablation printed without one is an advertisement.


Two designed blind spots the eval cannot get perfect on

1. Word order. "convert 100 USD to EUR" and "convert 100 EUR to USD" have an identical token multiset, so cosine is exactly 1.0. No threshold on any bag of words separates them, ever. This is a unit test asserting == 1.0, not a tolerance:

def test_bag_of_words_cannot_see_word_order():
    assert float(np.dot(S.embed("convert 100 USD to EUR", "hashed"),
                        S.embed("convert 100 EUR to USD", "hashed"))) == pytest.approx(1.0, abs=1e-12)

The order veto reads raw text — strictly more information than the embedding — and that is the entire reason vetoes exist.

2. Answers that decay. This one I planted less than I discovered, and it is the risk that survives everything:

=== the risk that survives every veto: answers that decay ===
  volatility policy                    hit   FH  of which volatile   saved$
  allow (no policy — naive cache)     0.37   11                 11   0.1323
  short TTL on volatile entries       0.34    0                  0   0.1193
  never store them (shipped)          0.33    0                  0   0.1143

11 of 11 residual false hits are volatile. "how many seats are left?" asked twice is sim == 1.000 on an identical token multiset — so no embedder can see it, no text-reading veto can see it, and the free exact-hash layer serves it wrong too. This is not a semantic-cache problem; it is a cache problem that a semantic cache amplifies, because it has more ways to reach a stale entry. The only defences are freshness and refusal. deny ships: a TTL longer than a decay period you guessed wrong is a silent wrong answer, and refusal is not.


Three false-hit rates, never one

Conflating these is the standard dishonesty in this space.

name denominator property of
FPR_pair DISTINCT pairs the classifier — prevalence-free, transferable
FHR_served = 1 − precision served hits the product
FHR_req all requests the SLO

Calibrate on FPR_pair, gate on FHR_served, print all three with k/n and a Wilson interval. Every rate this project prints carries its denominator.

And attribution is measured prevalence-free, on pairs — because whether a veto fires in a replay is a fact about my traffic generator, not about the veto:

  veto          blocks DISTINCT  costs REUSE   of 24 risk / 8 value above tau
  numeric                     2            0
  polarity                    4            0
  entity                      8            0
  qualifier                   1            0
  temporal                    0            0   <- NOT SHIPPED (dead at this tau)
  order                      12            0

temporal is not shipped, and that is a measured decision: it blocks zero of the pairs that clear tau, because changing tense changes tokens and similarity already rejects them. A veto unreachable at the operating point is dead code, and shipping dead code as a safety feature is exactly what this project exists to measure. test_temporal_veto_still_unreachable_at_tau is the tripwire that says when to put it back.


Price nothing; publish the exchange rate

Assigning a dollar value to a wrong answer is the marketing move. This inverts it: sweep λ ("how many wasted API calls is one wrong answer worth?") and print the τ each implies. As λ → ∞ the optimal semantic cache degenerates to an exact-match cache — which is the same statement as the headline, arrived at from the cost side.

Savings are reported as an identity, never as a headline percentage:

  baseline (cache disabled)   $0.3658
  gross savings               $0.1280   (35.0%)
  embedding cost             -$0.000028   (paid on 100% of non-exact requests)
  NET savings                 $0.1279   (35.0%)
  break-even hit rate          0.01232%

  savings ~= pi x TPR x (1 - c_embed/c_call)
  TPR here 0.33 is what I measured; pi is yours. The pi in this replay is 0.70
  BY CONSTRUCTION — I wrote the stream, so I chose that number.

False hits are not netted off, because pricing a wrong answer is the thing this project refuses to do.


Systems results, each cut or kept with a number

An in-process semantic cache is a memoization dict. --shards N partitions the stream round-robin across N independent stores — that is the uvicorn --workers N result, deterministically, with no processes to spawn:

$ python semcache.py sweep
   shards    hit    saved$
        1   0.33    0.1143
        4   0.17    0.0595
        8   0.10    0.0357

The argument for Redis is shared state, not the vector index.

Do not run the lookup inline in async def. It is a CPU-bound matmul, so inline it is the event loop's global serialization point:

$ python semcache.py bench
  50,000 entries x 1536 dims
     concurrency   1 lookup   inline p99   thread p99
               1      4.51ms        4.5ms        4.6ms
              64      4.51ms      288.9ms       10.9ms

At concurrency 64 the inline p99 is worse than the API call the cache exists to avoid — a cache that makes p99 worse. At 500 entries the matmul is microseconds and inline is correct; the crossover is a number, not a preference.

No ANN index, cut with the benchmark: exact matmul is 0.27 ms at 500 entries and recall is 1.0 by construction, and 1536 floats × 50k is 300 MB — memory binds before latency. HNSW earns its keep at 10⁶.

False hits grow with cache size at fixed τ, because argmax is adversarial: more entries means more chances for a spurious maximum. sweep --what cache-size prints it at both gates on purpose — a flat column at the shipped gate proves nothing on its own.

A semantic cache is a first-turn cache. Prior turns are an exact-match gate (embedding only the last turn without gating history is a wrong-answer generator: "what about the second one?"), so:

  turn 1    0.42 (124/294) [0.37, 0.48]
  turn 2    0.03 (2/69) [0.01, 0.10]
  turn 3+   0.11 (4/37) [0.04, 0.25]

Structural, not a tuning problem.


The proxy

OpenAI-compatible, so an existing client changes nothing but its base URL.

docker compose up --build
curl -sS -D- -X POST localhost:8080/v1/chat/completions \
  -H 'Authorization: Bearer sk-your-key' -H 'Content-Type: application/json' \
  -d '{"model":"claude-haiku-4-5","messages":[{"role":"user","content":"How many seats does the Pro plan include?"}]}'
# X-Cache: MISS | HIT | COALESCED   X-Cache-Similarity: 0.9428   X-Cache-Reason: veto:polarity
  • Scope is a deny-list, not an allow-list. Hash the whole body minus an explicit deny-list, so a parameter the provider adds next quarter produces a new scope — cold but correct. An allow-list fails open.
  • The principal is server-derived (a hash of the presented credential), never a client header, and no credential means no cache. A shared anonymous scope is precisely how one user's answer reaches another. The raw key never enters a scope input or a metric label.
  • max_tokens is not in the scope key. Bumping 512 → 1024 still hits every entry that finished naturally; the length question is answered by is_servable(), which can see the stored completion length. finish_reason == "tool_calls" is never stored or replayed — that is a decision to perform a side effect against state that no longer holds.
  • Hits replay as SSE when stream: true. Accepting the field and returning a JSON body is simply broken.
  • Single-flight: no await between lookup and registry insert (one await and two coroutines both become leader); the task is owned by the registry so a leader's disconnect cannot cancel the followers; failures are never cached.
  • /metrics is hand-rolled — cumulative buckets, +Inf == _count, HELP/TYPE once per family, counters ending _total, bounded label values, and deliberately no hit_ratio gauge (a ratio cannot be aggregated across instances; export the parts and divide in PromQL). test_metrics_parses asserts every one of those rules, which is what makes hand-rolling defensible instead of lazy.
rate(semcache_requests_total{result="hit"}[5m]) / rate(semcache_requests_total[5m])
sum by (reason) (rate(semcache_refusals_total[5m]))
histogram_quantile(0.99, rate(semcache_lookup_seconds_bucket[5m]))
rate(semcache_coalesced_total[5m]) / rate(semcache_upstream_calls_total[5m])

Why you can believe the numbers

  • pairs.jsonl was committed before any veto code existed (a3178bf, then e56987b). git log is the anti-circularity evidence. I labelled 68 pairs blocked_by: semantic, predicting no deterministic rule could reach them; the vetoes reach 43 of them. I reported that as a finding and left the labels alone — editing labels after seeing the vetoes is the exact circularity the commit ordering exists to prevent, and the disagreement is information about my intuition, which was wrong in the safe direction.
  • Ground truth is mechanical. answer_key = (intent, entity, qualifier, polarity, quantity, tense); REUSE iff the keys are equal. Hand labels are asserted against the derived key, which catches authoring bugs. And the label is on the answer, not the query — query-similarity labels only measure whether the embedder agrees with a human's notion of question similarity, which is circular.
  • test_gate_never_reads_a_label walks the AST of gate and every veto and fails if any of them can see label, family, blocked_by, or harm.
  • CI fails on suspicious perfection, not just on regression: a veto that blocks nothing, a blind spot that stops being blind, a permutation pair that stops scoring 1.0, fewer than 8 candidates per lookup (the replay would not be exercising argmax), or a 0.00 product metric that the prevalence-free metric contradicts.
  • test_stampede_negative_control runs the same 50 concurrent requests without single-flight and asserts all 50 reach the origin — otherwise the coalescing test would also pass if the requests had merely happened to serialize.

Bugs this harness caught in my own code, all of which would have shipped green:

  1. _arguments added the final content token — which in "when does the trial expire?" is the predicate, shared by the very pair the entity veto exists to catch. It missed 3/8 of its own adversarial family while firing on paraphrases.
  2. The L1 exact layer checked the global TTL instead of the entry's, so the volatility policy was bypassed by exactly the layer volatile traffic lands on.
  3. The traffic generator buried volatile questions in the Zipf tail, producing one volatile request in 400 and a flattering 0.00 false-hit rate. Volatile questions are head traffic — people ask them repeatedly because the answer keeps changing.
  4. The decorative-veto gate fired under --embed random, which is that ablation being asked a question it cannot answer.

Layout

semcache.py         normalize, scope, embed seam, vetoes, gate, store, single-flight,
                    cost model, calibrate / eval / sweep / bench / demo, CLI
app.py              OpenAI-compatible proxy, /metrics, SSE replay
pairs.jsonl         154 labelled pairs, 8 adversarial families, blocked_by lexical|semantic
traffic.jsonl       400 replay requests, 2 tenants, 14 scopes, multi-turn depth
test_semcache.py    39 offline tests
pip install -r requirements.txt
python -m pytest -q
python semcache.py demo && python semcache.py calibrate && python semcache.py eval

Runs fully offline and deterministically — no key, no network, no clock dependence.

Deliberately not built

  • RedisVL / Qdrant / any ANN — cut with the benchmark above.
  • prometheus_client — 60 hand-rolled lines plus a parser test carry more signal.
  • A TTL / eviction-policy bake-off — on a synthetic replay that measures my own generator, not the policy.
  • LLM-as-judge verification of hits — tempting and wrong: it destroys the latency claim and most of the cost claim. It pays only when c_judge < c_call × P(false hit) × λ, which is arithmetic you can do yourself from the numbers above.

License

MIT

About

A semantic cache is a classifier that is wrong silently. This is the safety gate (threshold AND deterministic vetoes) and the calibration harness that measures where its safety ends — which concluded the cache should not be turned on.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages