🧠 GenAI & LLMs
šŸŽ™ļø Voice AI Agents
⚔ High-Scale Backend
šŸš€ Kafka & AWS

I'm Milind

a GenAI Engineer.

> agent.run(stream=True) LLM Deploy Agent
Milind Wadhwa
šŸ¤–

Hello

I’m Milind Wadhwa, a Senior Software Engineer based in Gurgaon, India.

I specialize in building production-grade Agentic AI platforms and scalable distributed backend systems. I architect multi-channel intelligent agents across Voice, Chat, SMS, and WhatsApp — backed by RAG pipelines, LLM orchestration, and MCP — while also having deep roots in traditional backend engineering: REST APIs, async queues, and microservices in Python, Golang, and Java.

With over 7 years of engineering experience across Bidgely, Plivo, and Optum, I bridge the gap between robust backend foundations and cutting-edge AI — building systems that are both production-ready and genuinely intelligent.

7+ Yrs
Engineering Experience
85%
Autonomous Resolution Rate
~70%
System Efficiency Boost
Download Full Resume (PDF)

What I Do

Combining deep GenAI research, speech models, and enterprise-grade distributed backend architecture.

🧠

Agentic AI Platform Engineering

Architecting production-grade multi-channel AI agent platforms spanning Voice, Chat, SMS, and WhatsApp. Deep expertise in LLM orchestration with LangChain and Agno, multi-modal RAG pipelines, MCP tool integration, and model optimization via LoRA / PEFT fine-tuning on open-weights STT and language models.

LangChain Agno MCP RAG Vector DBs Pipecat Deepgram ElevenLabs LoRA Fine-tuning SileroVAD
⚔

Distributed Backend & Microservices

Proven track record in architecting high-throughput, low-latency microservices using FastAPI, Flask, Golang, and Java (Spring Boot). Transformed synchronous SaaS platforms into asynchronous architectures powered by Apache Kafka and Amazon SQS, accelerating service response efficiency by ~70%.

Python FastAPI Golang Java Spring Boot Apache Kafka Amazon SQS PostgreSQL Redis Clickhouse
ā˜ļø

Cloud Infrastructure & DevOps

Extensive experience deploying cloud-native architectures on AWS (ECS, EC2, CloudWatch, SQS). Automated containerized microservice deployments with Docker, Kubernetes, and OpenShift, establishing resilient CI/CD pipelines via Jenkins to ensure 99.9% platform availability.

AWS ECS / EC2 Docker Kubernetes OpenShift Jenkins CI/CD CloudWatch Linux Systems
šŸ“Š

Observability & Team Leadership

Established full-stack monitoring, distributed tracing, and real-time anomaly detection with Splunk, Elastic APM, Dynatrace, and the ELK Stack. Passionate mentor who leads code quality initiatives, conducts thorough PR reviews, and guides junior engineers toward software excellence.

Splunk Elastic APM Dynatrace ELK Stack Rest-Assured Cucumber BDD Tech Mentorship

Work Experience

7+ years of driving innovation in Generative AI, cloud platforms, and distributed systems.

🌱
Software Development Engineer - 3 @ Bidgely ↗
06/2026 - Present
šŸ“ Gurgaon, Haryana, India (Remote)
Conversational Energy Assistant Chat & Voice Agents
  • Architecting advanced conversational chat and voice agents to analyze home utility bills, provide detailed energy insights, and suggest optimal billing rate plans.
  • Developing LLM-driven agentic pipelines to deliver actionable recommendations for lowering electricity consumption and optimizing overall home energy usage.
  • Integrating real-time voice and messaging frameworks for low-latency, natural dialogue systems in consumer utility engagement.
šŸ’¬
Software Development Engineer - 2 @ Plivo Communications ↗
01/2024 - 05/2026
šŸ“ Bangalore, Karnataka (On-site)
Engineering Voice AI Agents & Multimodal Workflows
  • Engineered low-latency Voice AI Agents utilizing Pipecat, SileroVAD, Deepgram, and ElevenLabs for natural voice interactions.
  • Spearheaded model optimization using LoRA fine-tuning for STT models to significantly reduce transcription latency and improve accuracy.
  • Led technical integration of multi-modal LLMs and Vector Databases, optimizing RAG pipelines for superior retrieval across WhatsApp, Webchat, and SMS channels.
  • Architected and deployed a production-ready Agentic AI ecosystem leveraging LangChain and Agno to automate support workflows with 75% resolution accuracy.
  • Mentored junior developers and conducted thorough PR reviews to maintain rigorous code quality and architectural standards.
šŸ’¬
Software Development Engineer - 1 @ Plivo Communications ↗
12/2021 - 12/2023
šŸ“ Bangalore, Karnataka
SaaS Platform Refactoring & Cloud Scalability
  • Revitalized SaaS platform performance by migrating synchronous services to asynchronous FastAPI and implementing task parallelism via Amazon SQS, boosting efficiency by ~70%.
  • Developed core platform APIs utilizing Flask and Golang for high-volume sales engagement applications.
  • Automated ECS cloud deployments and configured EC2 auto-scaling via CloudWatch for 99.9% system availability.
šŸ„
Associate Software Engineer - 2 @ Optum (UnitedHealth Group) ↗
07/2019 - 12/2021
šŸ“ Hyderabad, Telangana
Microservices & DevOps Automation
  • Built RESTful microservices using Spring Boot and Apache Kafka for real-time, event-driven healthcare data streaming.
  • Automated CI/CD pipelines and containerized deployments on OpenShift using Docker and Jenkins.
  • Reduced test execution duration by 40% via multi-threaded integration testing with Rest-Assured and Cucumber BDD.
  • Implemented full-stack monitoring using Splunk, Elastic APM, and Dynatrace to improve system observability.
šŸ„
Software Engineering Intern @ Optum (UnitedHealth Group) ↗
05/2018 - 07/2018
šŸ“ Gurgaon, Haryana
Observability & Anomaly Detection
  • Deployed an ELK Stack solution for real-time anomaly detection, successfully monitoring over 100K email transactions for irregularities.

Technical Highlights

Key architectural solutions and production systems engineered throughout my journey.

šŸ¤–

Omnichannel Agentic AI Platform

Architected a full production agentic platform across Voice, Chat, SMS, and WhatsApp — with RAG pipelines, MCP tool integration, and LLM orchestration via LangChain & Agno. Achieved an 85% autonomous resolution rate for customer support workflows.

LangChain Agno RAG MCP Voice + Chat + SMS
Learn More →
⚔

STT Model LoRA Fine-Tuning

Applied parameter-efficient fine-tuning (PEFT) using LoRA on open-weights Speech-to-Text models, significantly reducing transcription latency and improving domain-specific accuracy with 75% less VRAM vs full fine-tuning.

HuggingFace LoRA / PEFT PyTorch Whisper
Learn More →
šŸ”„

Async SaaS Platform Refactor

Boosted platform efficiency by ~70% by migrating synchronous services to asynchronous FastAPI and offloading background work to Amazon SQS queues, with auto-scaled ECS deployments for 99.9% availability.

FastAPI Golang Amazon SQS AWS ECS
Learn More →
šŸ’”

GenAI Energy Assistant

Building conversational chat and voice agents that analyze home utility bills, surface energy insights, and recommend optimal billing rate plans — using LLM-driven agentic pipelines for real-time consumer engagement.

LLMs Voice AI Agentic AI Analytics
Learn More →

Get In Touch

Have an opportunity, a project, or want to connect about AI?

I'd love to hear from you! Feel free to reach out directly or send a message below.