Microsoft
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
I build production AI systems across large language models, retrieval pipelines, and computer-vision tooling that move from notebooks to real users. Currently focused on LLM fine-tuning, RAG, and multimodal perception.
Final-year CSE student at VIT Vellore, with internships at ISRO, IIT Ropar, IIT Bombay, and AI startups. I obsess over the path from notebook to deployed system: latency, model size, retrieval quality, and the boring infra glue that makes models actually useful. I also send fixes upstream to the open-source ML stack I build on.
Founding AI Engineer
AI Engineer
ML Engineer
AI Research Intern
Machine Learning Intern
15 merged PRs across 11 projects, ranked by the organization behind each project
Microsoft
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Microsoft
Agent Package Manager
Carnegie Mellon University
End-to-End Speech Processing Toolkit
2 merged pull requests
Alibaba Qwen
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
University of Hong Kong
"Vibe-Trading: Your Personal Trading Agent"
The World's First Virtual Terminal for AI Agents
2 merged pull requests
Things I've built end to end
Open-source · LLM Security
A Python library and MCP server that detects and masks PII before it reaches LLMs. Hybrid detection engine combines regex with Microsoft Presidio and spaCy NER, using reversible tokenization and a per-request in-memory vault. Ships as a library, FastAPI middleware, CLI, and MCP server for Claude Code.
Approached by Apertu Capital (Don Sheu & Enrico) for funding.
Local LLMs · 1-bit Quantization
One-click 1-bit (IQ1_S) compressor for any local LLM on Apple Silicon. Auto-discovers Ollama and on-disk GGUF models, runs importance-matrix calibration, and quantizes via llama.cpp with Metal — shrinking Llama-3.1-8B from 16.1 GB to 2.19 GB (7.4×) at ~51 tok/s on 3.1 GB of RAM. Side-by-side playground streams live tokens/sec, RAM, and bandwidth.
Computer Vision · Identity
Face verification pipeline that compares faces across images, PDFs, and Excel documents. RetinaFace detection with quality filtering and automatic rotation handling, ArcFace embeddings via DeepFace, and cosine-similarity matching with configurable thresholds — producing detailed reports with per-document confidence scores.
Voice AI · Real-time
Real-time voice interview agent. Google-Meet-style UI with WebRTC audio, Whisper STT, GPT-4o reasoning, and ElevenLabs TTS. Streaming responses with VAD-based turn-taking for natural conversational latency.
Speech · LoRA Fine-tuning
Fine-tunes OpenAI Whisper-small for Hindi speech recognition using LoRA (PEFT) on Mozilla Common Voice 17, training under 2% of parameters to measure WER/CER gains over the baseline. Type-hinted library with YAML-driven config, runs as fast CPU/MPS smoke tests locally and full fp16 training on a Colab T4.
Research papers
Recognitions and highlights
ShieldPrompt approached by Don Sheu & Enrico
Organized 5+ technical workshops
Top 2% among 400+ teams
Showing 33 total skills
Computer Science and Engineering