Essays··9 min read
The Similarity Threshold You Adjust on Tuesday and Regret by Friday
Semantic caching cuts LLM inference costs by 30–60% in benchmarks and causes a tenant data leak on Monday when the namespace is global and the threshold was tuned on last week's traffic. A single cosine similarity knob controls false-positive rate, hit rate, and customer-quality risk simultaneously — a coupling the vendor guides reduce to footnotes. The production stack requires cross-encoder reranking, per-tenant isolation, and threshold recalibration tied to upstream data changes.
semantic-cachingragllm-infrastructurevector-search
Read