VG
All work

Retrieval rebuild

From stream timeouts to sub-second hybrid search, verified by a 71-question test set.

Retrieval latency
was 5–8 s, now 0.6 s
Recall
was 0.53, now 0.98
MRR
0.89

Measured against a 71-question golden set. CI blocks a release if either number regresses.

Problem

Saroya AI, built at TinyCheque, an AI-first venture studio, is a mobile app where creators build an AI avatar that users chat with. The original retrieval layer ran graph-database queries that took 5–8 seconds each, long enough to time out streams; many sessions ended with no answer at all.

What I did

I replaced the graph-database retrieval with Postgres: pgvector for vector similarity and full-text search on the same instance, merged by reciprocal rank fusion (RRF, which combines several ranked lists by each result’s reciprocal rank). mem0 adds conversation memory, so retrieval knows what the user has already discussed.

How I measured it

I wrote a 71-question golden set covering the creator content types and user query patterns in the tester release. The eval harness runs in CI and is gated on SLOs: no deployment passes if recall or latency regresses.

MetricBefore (graph DB)After (pgvector + RRF)
Retrieval latency5–8 s, timeouts0.6 s
Recall0.530.98
MRRnot measured0.89

Architecture

A query runs two searches in parallel on one Postgres instance. RRF merges them, and mem0 supplies conversation memory before the response streams back.

Diagram scrolls sideways
QueryVector similaritypgvectortsvector + GINFull-textRRFMemorymem0Response

Limitations

The 71-question set covers the content types in the current tester release; as creator content diversifies it will need to grow. The architecture is specific to Saroya AI and has not been tested on other retrieval workloads. Source code is confidential.

Contact

Now, October 2026: looking for applied-AI and backend roles in India, the UK or Germany. Available immediately.