-
MiG-NJU 南京大学米格小组
- https://bradyfu.github.io/
Stars
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
A novel benchmerk specially designed to evaluate Omni-LLMs under assistant-style, real-time video chat scenarios.
EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory
OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains
[NeurIPS2026]"UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification"
[ACM MM 2026] Tango: Taming Visual Signals for Efficient Video Large Language Models
TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained GPUs.
awesome-native-multimodal-models
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
VITA-QINYU: Expressive Spoken Language Model for Role-Playing and Singing
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
[CVPR 2026 Highlight] PersonaVLM: Long-Term Personalized Multimodal LLMs
✨✨[ICML 2026] Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
A PyTorch Library for Multi-Attribution Learning in CVR prediction
🔥 A curated roadmap to the Efficient VLA landscape. We’re keeping this list live—contribute your latest work!
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.
Github repository for ACL 2025 paper: Recent Advances in Speech Language Models: A Survey.
CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
✨✨[NeurIPS 2025] VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
✨✨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
[CVPR 2026] MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension"

