VG
VG

Vansh Goenka, applied AI engineer. I rebuilt the retrieval behind a GenAI chat app: 0.6 s, down from 5–8 s.

Recall rose from 53% to 98% on a 71-question test set. I was the sole backend engineer at TinyCheque (Saroya AI), Feb–Sep 2026.

Email meDownload CV (PDF)
Kolkata+5:30London+0/+1Berlin+1/+2

Applied AI engineer in Kolkata, working in Python, TypeScript and Postgres. Open to roles in India, the UK and Germany; available now.Applied AI engineer, Kolkata. Previously sole backend engineer at TinyCheque (Saroya AI) · open to roles in India, the UK and Germany · available now.Applied AI engineer, Kolkata · available now

Four systems I shipped for one product.

Saroya AI is a GenAI app where creators build an AI avatar their audience can chat with. I was its only backend engineer; these are the four systems I’d show you first.

Retrieval rebuild

Rebuilt retrieval on Postgres hybrid search, tested against a 71-question golden set before every release.

retrieval latency
was 5–8 s, now 0.6 s
recall
was 0.53, now 0.98
How it works for Retrieval rebuild

The original graph-database retrieval was slow enough to time out streams. I replaced it with Postgres pgvector plus full-text search, merged by reciprocal rank fusion, with mem0 for conversation memory. CI blocks a release if recall or latency regresses against the golden set.

Stack: PostgreSQL, pgvector, full-text search, RRF, mem0, GitHub Actions

LLM cost control

Every LLM call is costed; prompt and reply caching keep the bill predictable.

cached input, tester traffic
78%
CI alert threshold
< 70%
How it works for LLM cost control

Each call records prompt, completion and cached-input tokens against model pricing. OpenRouter is pinned to five providers with a fallback chain, after unconstrained routing picked providers that overran the budget model. Tokens stream to mobile over SSE with pacing and keepalive.

Stack: OpenRouter, prompt and reply caching, SSE, cost instrumentation

Ingestion and deletion

Creator content arrives from X, Reddit, YouTube, LinkedIn and file uploads; a hard delete reaches every store.

stores a delete reaches
Postgres + pgvector · R2 · Redis
How it works for Ingestion and deletion

Sources are normalised into one document format, with OCR for image-heavy PDF, DOCX, XLSX and PPTX uploads. Each document is tracked through ingestion, embedding, indexing and deletion. KYC checks verify a creator before their content goes live.

Stack: OCR, Postgres, pgvector, Cloudflare R2, Redis

Agent harness

Tooling that kept coding agents spec-first and test-first while I shipped solo; I directed the agents that built it.

harness size
14 subagents · 169 tests
How it works for Agent harness

Every feature starts as a PRD or ADR, gets a spec, and only then enters implementation. Hooks refuse changes without a spec and tests. Custom MCP servers connect the agents to the codebase, docs and deploy pipeline.

Stack: Claude Code subagents, slash commands, hooks, MCP servers

What I work with.

Bold: used in production work, at Saroya AI or for a client. Regular: built and used in projects.

AI and LLM systems
In production work: RAG · hybrid retrieval (pgvector + full-text, RRF) · LLM evaluation (recall@k, MRR, golden sets) · prompt and reply caching · long-term memory (mem0) · OCR ingestion · LLM cost modelling · Also used: chunking · prompt engineering · guardrails · tool calling · multi-agent systems · MCP
Models and frameworks
In production work: OpenRouter · Gemini · Also used: OpenAI · LangChain · LangGraph · Ollama · Hugging Face · Groq · Tavily · Exa
Languages
In production work: Python · TypeScript · SQL · Also used: JavaScript · Cypher · C
Backend
In production work: Node.js · Express · REST · SSE streaming · Razorpay · Streamlit · Also used: FastAPI · WebSocket
Data
In production work: PostgreSQL · pgvector · Redis · Cloudflare R2 · Also used: Supabase · FalkorDB · ChromaDB · Apify · Kreuzberg · Tesseract
Infrastructure
In production work: Docker · GCP Cloud Run · GitHub Actions · OpenTelemetry · SigNoz · Also used: Railway · AWS (EC2, S3) · Linux
Machine learning
TensorFlow · Keras · scikit-learn · OpenCV · CNNs
Coding agents
In production work: Claude Code (subagents, slash commands, hooks, MCP) · Also used: Codex · OpenCode
Practices
In production work: test-first · spec-first (PRDs, ADRs) · DPDP-aligned data deletion · Also used: RCAs

Smaller things, in public.

  • AI-Agent-for-research

    Streamlit research assistant that writes a structured report on any topic using LangChain tools, with a hosted demo.

    Demo
  • market-research-agent

    Three-agent pipeline that turns a company or industry name into a market research report with AI use cases and datasets.

  • meta-ad-creator

    Five cooperating Gemini agents turn a product photo into a 1080 x 1080 Meta ad image. Streamlit app with Docker.

  • pizzeria-review-agent

    Local RAG question answering over restaurant reviews with ChromaDB, Ollama and Llama 3.2.

  • Evidence-Detection

    An evidence detection system using RoBERTa with Bi-LSTM layers, fine-tuned on 26k claim-evidence pairs, leveraging TensorFlow, Keras, and PyTorch.

More on GitHub

Tested before shipped.

I’m Vansh, an applied AI engineer in Kolkata. I studied Artificial Intelligence at the University of Manchester, then spent 2026 as the only backend engineer on Saroya AI at TinyCheque, an AI-first venture studio.

That meant owning everything behind the chat box: retrieval, streaming, payments, ingestion, observability and the LLM bill. Retrieval changes ran against a 71-question test set before they shipped, and CI blocked regressions.

Before that I freelanced for Chitra Goenka Crafts & Creations, a home-decor business: a Gemini image pipeline that replaced product photoshoots, and a B2B lead-finder on the Google Maps API, both behind a Streamlit dashboard their staff use daily. I work from Kolkata; the gold here is a nod to its yellow taxis.

Vansh Goenka

Where I’ve worked and studied

  1. Feb–Sep 2026TinyCheque (Saroya AI)AI EngineerSole backend engineer on a GenAI chat app, from retrieval to payments.
  2. Jan 2025–Feb 2026Chitra Goenka Crafts & CreationsFreelance AI ConsultantGemini image pipeline in place of product photoshoots; a B2B lead-finder on the Google Maps API.
  3. Sep 2021–Jul 2024University of ManchesterBSc (Hons) Artificial Intelligence, 2:1Dissertation: Simulation and Optimization of the Beta-Algorithm for Autonomous Swarm Robotics.

Practicalities

Languages
English, Hindi
Right to work
UK: needs Skilled Worker visa sponsorship
Germany: eligible for the EU Blue Card with a qualifying job offer (BSc from a UK university)

Off the clock

  • Scenic photographyWalks in the park, camera out for the view.Photos coming soon
  • CookingI love to cook.Photos coming soon
  • House musicI love house music.See the photo
  • GymI train at the gym.
  • New electronicsI go crazy for new electronics.
All of it on one page

Contact

Now, October 2026: looking for applied-AI and backend roles in India, the UK or Germany. Available immediately.