Problem
Saroya AI, built at TinyCheque, an AI-first venture studio, is a mobile app where creators build an AI avatar that users chat with. The original retrieval layer ran graph-database queries that took 5–8 seconds each, long enough to time out streams; many sessions ended with no answer at all.
What I did
I replaced the graph-database retrieval with Postgres: pgvector for vector similarity and full-text search on the same instance, merged by reciprocal rank fusion (RRF, which combines several ranked lists by each result’s reciprocal rank). mem0 adds conversation memory, so retrieval knows what the user has already discussed.
How I measured it
I wrote a 71-question golden set covering the creator content types and user query patterns in the tester release. The eval harness runs in CI and is gated on SLOs: no deployment passes if recall or latency regresses.
| Metric | Before (graph DB) | After (pgvector + RRF) |
|---|---|---|
| Retrieval latency | 5–8 s, timeouts | 0.6 s |
| Recall | 0.53 | 0.98 |
| MRR | not measured | 0.89 |
Architecture
A query runs two searches in parallel on one Postgres instance. RRF merges them, and mem0 supplies conversation memory before the response streams back.
Limitations
The 71-question set covers the content types in the current tester release; as creator content diversifies it will need to grow. The architecture is specific to Saroya AI and has not been tested on other retrieval workloads. Source code is confidential.