Production agents
I build tool-using, conversational, RAG, and multi-agent systems that connect models to real business workflows.
Atharv Kulkarni
India
I turn ambiguous workflows into measurable agents using evaluation, observability, and disciplined software engineering.
System
Production AI loop
Build
Agents · MCP · RAG
Evaluate
Golden sets · LLM judges
Observe
Traces · Scores · Feedback
Improve
Regression gates · Release
My work sits where AI behavior meets software engineering. I build the agent, define what good looks like, trace how it behaves in production, and use that evidence to make the next release better.
93%
agent accuracy
82%
lower processing time
56%
lower token usage
600+
enterprise customers
I build tool-using, conversational, RAG, and multi-agent systems that connect models to real business workflows.
I turn product expectations into golden datasets, LLM judges, regression checks, and release decisions.
I ship the software around the model: APIs, streaming, structured outputs, provider abstractions, and cloud runtimes.
A few examples of production AI, evaluation, and engineering automation from my recent work.
Onit · Production agent and evaluation system
Rebuilt an agent that detects billing disputes, identifies disputed line items, and recommends invoice adjustments. I paired the new architecture with expert-labeled golden datasets and LLM-judge evaluation so performance could be measured before release.
Onit · Legal operations analytics agent
Built a conversational analytics agent that uses MCP data tools to answer spend and operational questions over enterprise data. Added multi-provider model support, production tracing, prompt and model comparison, and deployment on Amazon Bedrock AgentCore Runtime.
Onit · Multi-agent workflow
Built a multi-agent system that converts long requirements documents into structured user stories, automating a high-effort product workflow while keeping the output reviewable.
Onit · AI-assisted engineering
Drove the migration of QA automation across two enterprise product lines, using agentic coding for translation, refactoring, and validation while removing recurring tool licensing costs.
Independent work exploring developer infrastructure, reliable browser agents, and bounded multi-agent production workflows.
View all projectsAI developer platform
A platform for versioning prompts, routing LLM requests across providers, and observing the latency and token usage of agent and RAG applications.
Human-controlled browser agent
A desktop agent that fills browser forms while keeping review and final submission under explicit human control.
Agent-assisted publishing system
An internal system for turning a guide idea into structured content, branded PDFs, privacy-safe screenshots, product listings, and marketing assets.
Three years of hands-on GenAI work, backed by enterprise application engineering.
Agentic AI Engineer
June 2025 to present
Building and evaluating production AI agents for enterprise legal operations, with a focus on reliability, observability, and measurable releases.
Software Craftsperson, AI
October 2024 to June 2025
Developed enterprise AI products spanning multi-model chat, agents, RAG knowledge bases, deep research, streaming APIs, and document integrations.
Application Developer
November 2022 to September 2024
Built enterprise software for testing, Workday integration monitoring, and data security, while exploring GenAI through projects and prototypes.
B.E. Computer Engineering
Honors in IoT · CGPA 9.31/10
International Institute of Information Technology, Pune · 2018 to 2022
Lessons from building autonomous AI agents that can reason, plan, and execute multi-step tasks in production environments.
How publishing an IEEE paper on blockchain-based health records shaped my approach to building with large language models.