Vansh Goenka
Applied AI Engineer, Kolkata, India
vanshgoenka007@gmail.comlinkedin.com/in/vanshgoenkagithub.com/unusual9guy
Summary
Applied AI engineer in Kolkata. Builds production RAG and LLM backends; rebuilt retrieval for a GenAI app: 0.6 s, down from 5–8 s; recall 0.53 → 0.98. Open to roles in India, the UK and Germany; available now.
Experience
- Retrieval rebuild: Replaced graph-database retrieval with Postgres hybrid search (pgvector, full-text, reciprocal rank fusion): latency down from 5–8 s to 0.6 s, recall up from 0.53 to 0.98 on a 71-question golden set that gates every release in CI.
- LLM cost control: Costed every LLM call and added prompt and reply caching: 78% of input tokens served from cache on tester traffic, with a CI alert below 70%. Pinned OpenRouter to five providers with a fallback chain.
- Ingestion and deletion: Built ingestion from X, Reddit, YouTube, LinkedIn and file uploads (OCR for image-heavy documents), and a hard delete that reaches Postgres, pgvector, R2 and Redis.
- Agent harness: Shipped solo through a spec-first, test-first coding-agent harness I directed: 14 subagents, 169 tests, hooks that refuse changes without a spec, custom MCP servers.
- Gemini image pipeline in place of product photoshoots; a B2B lead-finder on the Google Maps API.
Education
- Dissertation: Simulation and Optimization of the Beta-Algorithm for Autonomous Swarm Robotics.
Skills
Each line lists what I have used in production work first, then what else I have built with.
AI and LLM systems: RAG, hybrid retrieval (pgvector + full-text, RRF), LLM evaluation (recall@k, MRR, golden sets), prompt and reply caching, long-term memory (mem0), OCR ingestion, LLM cost modelling; also chunking, prompt engineering, guardrails, tool calling, multi-agent systems, MCP.
Models and frameworks: OpenRouter, Gemini; also OpenAI, LangChain, LangGraph, Ollama, Hugging Face, Groq, Tavily, Exa.
Languages: Python, TypeScript, SQL; also JavaScript, Cypher, C.
Backend: Node.js, Express, REST, SSE streaming, Razorpay, Streamlit; also FastAPI, WebSocket.
Data: PostgreSQL, pgvector, Redis, Cloudflare R2; also Supabase, FalkorDB, ChromaDB, Apify, Kreuzberg, Tesseract.
Infrastructure: Docker, GCP Cloud Run, GitHub Actions, OpenTelemetry, SigNoz; also Railway, AWS (EC2, S3), Linux.
Machine learning: TensorFlow, Keras, scikit-learn, OpenCV, CNNs.
Coding agents: Claude Code (subagents, slash commands, hooks, MCP); also Codex, OpenCode.
Practices: test-first, spec-first (PRDs, ADRs), DPDP-aligned data deletion; also RCAs.
Languages
English, Hindi.
Selected open-source work
- AI-Agent-for-research: Streamlit research assistant that writes a structured report on any topic using LangChain tools, with a hosted demo.
- market-research-agent: Three-agent pipeline that turns a company or industry name into a market research report with AI use cases and datasets.
- meta-ad-creator: Five cooperating Gemini agents turn a product photo into a 1080 x 1080 Meta ad image. Streamlit app with Docker.
- pizzeria-review-agent: Local RAG question answering over restaurant reviews with ChromaDB, Ollama and Llama 3.2.
- Evidence-Detection: An evidence detection system using RoBERTa with Bi-LSTM layers, fine-tuned on 26k claim-evidence pairs, leveraging TensorFlow, Keras, and PyTorch.
Availability
Open to roles in India, the UK and Germany. Available now.
UK: needs Skilled Worker visa sponsorship.
Germany: eligible for the EU Blue Card with a qualifying job offer (BSc from a UK university).
Interests
Scenic photography, Cooking, House music, Gym, New electronics.