- Bangalore, India
- https://www.linkedin.com/in/gauravk95/
- @gaurav_k5
Starred repositories
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per re…
This repository contains the official implementation of the research papers, "MobileCLIP" CVPR 2024 and "MobileCLIP2" TMLR August 2025
The simplest, fastest repository for training/finetuning medium-sized GPTs.
[SIGGRAPH 2025] LAM: Large Avatar Model for One-shot Animatable Gaussian Head
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
A spy satellite simulator in your browser, except the data is real. Live open source spatial intelligence on a photorealistic 3D globe.
Parallax effect made with Jetpack Compose and Sensor Manager
Open-Sora: Democratizing Efficient Video Production for All
Framework for quickly creating connected applications in Kotlin with minimal effort
An awesome list that curates the best Kotlin Multiplatform libraries, tools and more.
A simple framework for mobile system design interviews
AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
chenxwh / AniPortrait
Forked from Zejun-Yang/AniPortraitAniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
Curated coding interview preparation materials for busy software engineers
⭐️ Companies that don't have a broken hiring process
[ECCV2024] IDM-VTON : Improving Diffusion Models for Authentic Virtual Try-on in the Wild
Emote Portrait Alive: Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
Record voice notes & transcribe, summarize, and get tasks
Instant voice cloning by MIT and MyShell. Audio foundation model.
An Open Source text-to-speech system built by inverting Whisper.
The official implementation of HierSpeech++
A WebUI to create song covers with any RVC v2 trained AI voice from YouTube videos or audio files.
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translatio…
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models


