-
Qualcomm AI Research
- Seoul, Republic of Korea
- in/yangyangii
- https://scholar.google.com/citations?hl=en&user=yJjxKVQAAAAJ
- https://yangyangii.world
Lists (2)
Sort Name ascending (A-Z)
Stars
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Official implementation of "Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech"
from vibe coding to agentic engineering - practice makes claude perfect
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team at Anthropic PBC. A delectab…
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
Character-aware audio-only subtitling
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching (ICASSP '25)
Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)
Official Pytorch Implementation for "DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion" (AAAI 2024)
This repo contains the scripts, models, and required files for the Deep Noise Suppression (DNS) Challenge.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Expressive Anechoic Recordings of Speech (EARS)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
🔊 Text-Prompted Generative Audio Model
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RN…
Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"
Kandinsky 2 — multilingual text2image latent diffusion model
Tools for handling multimodal data in machine learning projects.
Keep track of big models in audio domain, including speech, singing, music etc.
AudioLDM: Generate speech, sound effects, music and beyond, with text.
Singing Voice Conversion via diffusion model
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
A playbook for systematically maximizing the performance of deep learning models.


