Skip to main content
Atharv Kulkarni
Atharv Kulkarni

Atharv Kulkarni

India

Agentic AI Engineer at Onit

I build reliable production AI systems.

I turn ambiguous workflows into measurable agents using evaluation, observability, and disciplined software engineering.

Open to Applied AI and Agentic AI roles in the UK and globally.

System

Production AI loop

Live
  1. 01

    Build

    Agents · MCP · RAG

  2. 02

    Evaluate

    Golden sets · LLM judges

  3. 03

    Observe

    Traces · Scores · Feedback

  4. 04

    Improve

    Regression gates · Release

Shipping at enterprise scale
What I do

The model is only part of the system.

My work sits where AI behavior meets software engineering. I build the agent, define what good looks like, trace how it behaves in production, and use that evidence to make the next release better.

93%

agent accuracy

82%

lower processing time

56%

lower token usage

600+

enterprise customers

01

Production agents

I build tool-using, conversational, RAG, and multi-agent systems that connect models to real business workflows.

PythonTypeScriptMCPRAGStrands
02

Evaluation and reliability

I turn product expectations into golden datasets, LLM judges, regression checks, and release decisions.

LLM-as-a-JudgeLangfuseGolden setsTracing
03

AI application engineering

I ship the software around the model: APIs, streaming, structured outputs, provider abstractions, and cloud runtimes.

FastAPISSELiteLLMAWS Bedrock
Selected work

Systems that moved a real metric.

A few examples of production AI, evaluation, and engineering automation from my recent work.

01

Onit · Production agent and evaluation system

Invoice Review Agent

93% accuracy82% faster56% fewer tokens

Rebuilt an agent that detects billing disputes, identifies disputed line items, and recommends invoice adjustments. I paired the new architecture with expert-labeled golden datasets and LLM-judge evaluation so performance could be measured before release.

Agentic AILLM evaluationGolden datasetsStrands
02

Onit · Legal operations analytics agent

Ask Unity

MCP data toolsProduction observabilityMulti-provider runtime

Built a conversational analytics agent that uses MCP data tools to answer spend and operational questions over enterprise data. Added multi-provider model support, production tracing, prompt and model comparison, and deployment on Amazon Bedrock AgentCore Runtime.

MCPLangfuseLiteLLMBedrock AgentCore
03

Onit · Multi-agent workflow

Requirements to user stories

~2 weeks to ~10 minutesStructured, reviewable output

Built a multi-agent system that converts long requirements documents into structured user stories, automating a high-effort product workflow while keeping the output reviewable.

Multi-agent systemsStructured outputsWorkflow automation
04

Onit · AI-assisted engineering

QA automation migration

~1 year of work in 3 weeksRecurring licensing cost removed

Drove the migration of QA automation across two enterprise product lines, using agentic coding for translation, refactoring, and validation while removing recurring tool licensing costs.

PlaywrightAgentic codingTest automation
Experience

Software foundations. AI focus.

Three years of hands-on GenAI work, backed by enterprise application engineering.

Onit

Agentic AI Engineer

June 2025 to present

Building and evaluating production AI agents for enterprise legal operations, with a focus on reliability, observability, and measurable releases.

Incubyte

Software Craftsperson, AI

October 2024 to June 2025

Developed enterprise AI products spanning multi-model chat, agents, RAG knowledge bases, deep research, streaming APIs, and document integrations.

IBM India

Application Developer

November 2022 to September 2024

Built enterprise software for testing, Workday integration monitoring, and data security, while exploring GenAI through projects and prototypes.

B.E. Computer Engineering

Honors in IoT · CGPA 9.31/10

International Institute of Information Technology, Pune · 2018 to 2022

Let's work together

Building AI that has to work in the real world?

I am looking for roles with meaningful ownership across agent engineering, evaluation, and production AI reliability. For UK opportunities, I am based in India and require Skilled Worker sponsorship to relocate.