AI · Agents · Production LLM Systems
I build production LLM systems — agents, RAG, and evaluation pipelines that hold up at scale.
AI Engineer focused on agentic workflows, agent memory architectures, and RAG systems processing 1B+ tokens/month in production.
Analyst — AI & Agent Systems at Future Standard, designing agentic memory architectures and operations-intelligence agent systems.
- 1B+
- tokens/month in production
- 1,500+
- hours/year saved
- $800K+
- ARR growth contributed
PythonJavaScriptC++TypeScriptReactNext.jsFastAPILangchainLangGraphLiteLLMPydanticPyTorchQdrantMongoDBPostgreSQLMySQLDockerKubernetesAWSGitLinearJiraJWTOAuthEmbedded CAssemblyPythonJavaScriptC++TypeScriptReactNext.jsFastAPILangchainLangGraphLiteLLMPydanticPyTorchQdrantMongoDBPostgreSQLMySQLDockerKubernetesAWSGitLinearJiraJWTOAuthEmbedded CAssembly
Experience
Where I've shipped
Selected work
Projects that show how I think
Why work with me
What I bring to a team
AI-first
Production LLM systems — agents, agentic memory, RAG, and evals built for real operational workloads.
LLM tooling
LangGraph + LiteLLM orchestration, retrieval pipelines, and evaluation at 1B+ tokens/month scale.
End-to-end ownership
From schema design to k8s deployment. I ship things that stay up.
Latest writing
Notes from the blog
2026-05-02KV Cache and PagedAttention: Why vLLM Actually WorksKV cache stores attention keys and values across tokens to avoid recomputation during autoregressive decoding. PagedAttention extends this with virtual memory paging, making vLLM's throughput possible.ml→2026-05-02The Bulkhead Pattern: Isolating Failures Before They SpreadHow the bulkhead pattern isolates failures in distributed systems — partition thread pools, connection pools, and resources so one degraded dependency cannot sink the whole service.systems→2026-05-02CRDTs: Conflict-Free Replicated Data TypesCRDTs eliminate merge conflicts by design — commutative, associative, idempotent data structures that converge to the same state regardless of operation order. G-Counters, OR-Sets, and LWW registers explained.systems→2026-05-02Multi-Paxos: From Single Decree to a Replicated LogMulti-Paxos extends single-decree Paxos into a replicated log by electing a stable leader, skipping Phase 1 for subsequent entries, and batching proposals for throughput.systems→2026-05-02Raft Log Compaction: Keeping the Log from Growing ForeverRaft logs grow forever if left unchecked. Log compaction via snapshots lets nodes discard old entries, transfer state to slow followers, and recover quickly after a restart.systems→
Contact
Let's build something reliable.
Available for full-time roles and select freelance projects. Based in Bangalore — open to remote.