- Essays··11 min read
Semantic Caching Promised 73%, Finance Got a Second Bill
In January 2026, VentureBeat ran a headline about semantic caching cutting LLM bills by 73%, and by April the pitch had settled into a tighter band. Vendor materials and open-source library documentation claimed 30-70% cost reductions, latency improvements measured in multiples, and deployment described as a few hours of engineering work. The promise was …
semanticcachingpromisedfinanceRead - Essays··8 min read
Why You're Paying Twice for the Same Token
Any 2026 production agent stack without the three-layer caching pattern (engine prefix cache, API prompt cache, gateway semantic cache) is carrying a 30–60% avoidable inference bill. The pattern isn't subtle; it's just rarely implemented in the right order.
inference economicscachingfinopsllmopsRead