Problem
Saroya AI, built at TinyCheque, routes every user chat through LLM inference. Without instrumentation, costs grew unpredictably, and real routing costs through OpenRouter (a gateway that routes requests across model providers) overran the cost model whenever the router chose expensive providers.
What I did
Every LLM call records prompt tokens, completion tokens, cached-input tokens and model pricing. I added prompt caching and reply caching to maximise cache hits, and pinned OpenRouter to five providers with a fallback chain after finding that unconstrained routing picked providers whose per-token cost exceeded the budget model.
Tokens stream to mobile over server-sent events (SSE), with pacing and keepalive so connections don’t drop.
How I measured it
The cached-input ratio is computed weekly and checked in CI. If it falls below 70%, the pipeline alerts. On current tester-release traffic it is 78%.
| Control | Value |
|---|---|
| Cached-input ratio (tester traffic) | 78% |
| CI alert threshold | below 70% |
| Pinned providers | 5, with fallback |
Architecture
Limitations
The 78% ratio reflects tester-release traffic; as the user base grows and query patterns diversify it may change. The provider pin list needs periodic review as OpenRouter adds and deprecates providers. Source code is confidential.