-
Seoul National University
- Seoul
-
05:57
(UTC +09:00) - https://medium.com/@mscheong01
- in/minsoo-cheong-606314206
Highlights
- Pro
Starred repositories
Official code for paper "Tailoring the Quantization Space for 1-Bit KV Cache Compression"
[NeurIPS 2025] NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
A comprehensive collection of process reward models.
The code implementation of ICML '26 paper "SPA-Cache: Singular Proxies for Adaptive Caching in Diffusion Language Models"
[ICLR 2026 🔥] Official pytorch implementation for "Attention Is All You Need for KV Cache in Diffusion LLMs"
[ICLR'26] Official code of paper "d2Cache: Accelerating Diffusion-based LLMs via Dual Adaptive Caching"
[NeurIPS 2025] Scaling Speculative Decoding with Lookahead Reasoning
[NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)
A collection of hardware and software projects based around the Electro-Smith Daisy Seed
📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
The official GitHub repo for the survey paper "A Survey on Diffusion Language Models".
[ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".
Embedded controller for the IK Multimedia Tonex One, Tonex Pedal, and Valeton GP5 guitar pedals
A curated list of awesome LLM agents frameworks.
A TTS model capable of generating ultra-realistic dialogue in one pass.
This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & V…
F1 Live Timing TUI for all F1 sessions with variable delay to sync to your TV. Supports replaying previously recorded sessions.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
This repository contains a collection of surveys, datasets, papers, and codes, for predictive uncertainty estimation in deep learning models.
Everything you need to build state-of-the-art foundation multimodal desktop agent, end-to-end.
Pretraining and inference code for a large-scale depth-recurrent language model




