Deepak Kumar Jha

AI Demos — Senior Full-Stack Engineer · Technical Lead

Three systems, live on my own infrastructure, each with a one-click demo account against its real API — no signup, nothing mocked. This is the difference between claiming production AI experience and handing you a URL where it's running.

Privacy-First RAG

Live

Retrieval-Augmented Generation on self-hosted Llama 3 and Mistral via Ollama. Built for cases where data cannot leave the client’s infrastructure — regulated finance, government, healthcare. Upload a PDF, ask questions, get answers with citations and per-passage similarity scores.

Node.js NestJS Angular Ollama Llama 3 Qdrant (vector search)

Multi-Model Routing

Live

Routes each request between cloud LLMs (Gemini, OpenAI) and local inference based on cost, latency and data sensitivity — and shows the decision, the fallback chain and the money saved, per request and on a live dashboard.

Node.js NestJS Angular PostgreSQL Ollama Cloud LLM APIs

Agentic Workflows (MCP)

Live

A Model Context Protocol agent connected to a read-only SQL database, a live weather API and a calculator — answering multi-step questions with every tool call, result and recovery from failure streamed to the screen as it happens.

Model Context Protocol Node.js NestJS Angular PostgreSQL Ollama

A production LLM feature already shipped in a live product — see the

SmartSitting case study →

Let's Connect

Want to work together?

Open to full-time Senior Full-Stack Engineer and Technical Lead roles in Delhi NCR and Dubai/UAE. Immediate joiner.

Technical Leadership

Architecture decisions, tech strategy, and hands-on leadership for your team.

Full-Stack Engineering

A senior engineer who ships production code from day one, no ramp-up required.

Immediate Joiner

Available within 15 days of an offer.

Email Me Directly