Architecting Ultra-Low Latency Voice AI Agents with Pipecat & SileroVAD
How we reduced end-to-end voice turnaround time from 1.8s down to sub-500ms for natural, human-like telephone conversations.
Iām Milind Wadhwa, a Senior Software Engineer based in Gurgaon, India.
I specialize in building production-grade Agentic AI platforms and scalable distributed backend systems. I architect multi-channel intelligent agents across Voice, Chat, SMS, and WhatsApp ā backed by RAG pipelines, LLM orchestration, and MCP ā while also having deep roots in traditional backend engineering: REST APIs, async queues, and microservices in Python, Golang, and Java.
With over 7 years of engineering experience across Bidgely, Plivo, and Optum, I bridge the gap between robust backend foundations and cutting-edge AI ā building systems that are both production-ready and genuinely intelligent.
Combining deep GenAI research, speech models, and enterprise-grade distributed backend architecture.
Architecting production-grade multi-channel AI agent platforms spanning Voice, Chat, SMS, and WhatsApp. Deep expertise in LLM orchestration with LangChain and Agno, multi-modal RAG pipelines, MCP tool integration, and model optimization via LoRA / PEFT fine-tuning on open-weights STT and language models.
Proven track record in architecting high-throughput, low-latency microservices using FastAPI, Flask, Golang, and Java (Spring Boot). Transformed synchronous SaaS platforms into asynchronous architectures powered by Apache Kafka and Amazon SQS, accelerating service response efficiency by ~70%.
Extensive experience deploying cloud-native architectures on AWS (ECS, EC2, CloudWatch, SQS). Automated containerized microservice deployments with Docker, Kubernetes, and OpenShift, establishing resilient CI/CD pipelines via Jenkins to ensure 99.9% platform availability.
Established full-stack monitoring, distributed tracing, and real-time anomaly detection with Splunk, Elastic APM, Dynatrace, and the ELK Stack. Passionate mentor who leads code quality initiatives, conducts thorough PR reviews, and guides junior engineers toward software excellence.
7+ years of driving innovation in Generative AI, cloud platforms, and distributed systems.
Key architectural solutions and production systems engineered throughout my journey.
Architected a full production agentic platform across Voice, Chat, SMS, and WhatsApp ā with RAG pipelines, MCP tool integration, and LLM orchestration via LangChain & Agno. Achieved an 85% autonomous resolution rate for customer support workflows.
Applied parameter-efficient fine-tuning (PEFT) using LoRA on open-weights Speech-to-Text models, significantly reducing transcription latency and improving domain-specific accuracy with 75% less VRAM vs full fine-tuning.
Boosted platform efficiency by ~70% by migrating synchronous services to asynchronous FastAPI and offloading background work to Amazon SQS queues, with auto-scaled ECS deployments for 99.9% availability.
Building conversational chat and voice agents that analyze home utility bills, surface energy insights, and recommend optimal billing rate plans ā using LLM-driven agentic pipelines for real-time consumer engagement.
I'd love to hear from you! Feel free to reach out directly or send a message below.