VG
All work

LLM cost control

Every call costed, cached and routed to a pinned provider list.

Cached input, tester traffic
78%
CI alert threshold
< 70%
Pinned providers
5

Ratio = cached input tokens / total input tokens, computed weekly in CI on tester-release traffic.

Problem

Saroya AI, built at TinyCheque, routes every user chat through LLM inference. Without instrumentation, costs grew unpredictably, and real routing costs through OpenRouter (a gateway that routes requests across model providers) overran the cost model whenever the router chose expensive providers.

What I did

Every LLM call records prompt tokens, completion tokens, cached-input tokens and model pricing. I added prompt caching and reply caching to maximise cache hits, and pinned OpenRouter to five providers with a fallback chain after finding that unconstrained routing picked providers whose per-token cost exceeded the budget model.

Tokens stream to mobile over server-sent events (SSE), with pacing and keepalive so connections don’t drop.

How I measured it

The cached-input ratio is computed weekly and checked in CI. If it falls below 70%, the pipeline alerts. On current tester-release traffic it is 78%.

ControlValue
Cached-input ratio (tester traffic)78%
CI alert thresholdbelow 70%
Pinned providers5, with fallback

Architecture

Diagram scrolls sideways
Chat reqCache check5 pinned providersOpenRouterSSE outCost instrumentation

Limitations

The 78% ratio reflects tester-release traffic; as the user base grows and query patterns diversify it may change. The provider pin list needs periodic review as OpenRouter adds and deprecates providers. Source code is confidential.

Contact

Now, October 2026: looking for applied-AI and backend roles in India, the UK or Germany. Available immediately.