Skip to content
View mscheong01's full-sized avatar
:shipit:
:shipit:

Highlights

  • Pro

Block or report mscheong01

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Official code for paper "Tailoring the Quantization Space for 1-Bit KV Cache Compression"

Python 3 Updated Oct 5, 2026

Preview Code for Continuum Paper

Python 108 33 Updated Aug 13, 2026

[NeurIPS 2025] NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

Cuda 3 Updated Jun 29, 2026

A comprehensive collection of process reward models.

183 6 Updated Sep 13, 2026
Python 37 7 Updated Aug 19, 2026
Python 4 2 Updated Jun 10, 2026

The code implementation of ICML '26 paper "SPA-Cache: Singular Proxies for Adaptive Caching in Diffusion Language Models"

Python 4 Updated May 23, 2026

[ICLR 2026 🔥] Official pytorch implementation for "Attention Is All You Need for KV Cache in Diffusion LLMs"

Python 43 3 Updated Jul 13, 2026

[ICLR'26] Official code of paper "d2Cache: Accelerating Diffusion-based LLMs via Dual Adaptive Caching"

Python 188 5 Updated May 14, 2026

[NeurIPS 2025] Scaling Speculative Decoding with Lookahead Reasoning

Python 69 8 Updated Oct 31, 2025

[NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)

Python 224 14 Updated Feb 11, 2026

A collection of hardware and software projects based around the Electro-Smith Daisy Seed

C++ 546 92 Updated Sep 27, 2026

📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥

2,175 103 Updated Aug 17, 2026

The official GitHub repo for the survey paper "A Survey on Diffusion Language Models".

1,228 63 Updated Oct 2, 2026
Python 38 Updated Jul 21, 2025

[ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".

Python 282 22 Updated Jul 30, 2026

Embedded controller for the IK Multimedia Tonex One, Tonex Pedal, and Valeton GP5 guitar pedals

C 416 61 Updated Oct 4, 2026

A curated list of awesome LLM agents frameworks.

Python 1,599 365 Updated Oct 4, 2026
Python 48 7 Updated Oct 16, 2025

A TTS model capable of generating ultra-realistic dialogue in one pass.

Python 19,391 1,688 Updated Nov 19, 2025

EMNLP 2025

Python 14 Updated Nov 28, 2024

This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & V…

1,315 76 Updated Mar 9, 2025

F1 Live Timing TUI for all F1 sessions with variable delay to sync to your TV. Supports replaying previously recorded sessions.

C# 907 24 Updated Jul 20, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11,959 1,986 Updated Oct 5, 2026
Python 32 3 Updated May 24, 2025

Random odd guitar pedal design in kicad

OpenSCAD 576 22 Updated Sep 19, 2025

This repository contains a collection of surveys, datasets, papers, and codes, for predictive uncertainty estimation in deep learning models.

824 80 Updated Aug 19, 2026

Everything you need to build state-of-the-art foundation multimodal desktop agent, end-to-end.

Python 44 12 Updated Sep 24, 2026

Pretraining and inference code for a large-scale depth-recurrent language model

Python 976 88 Updated Dec 29, 2025
Next