Skip to content
View BradyFU's full-sized avatar
👋
👋

Organizations

@VITA-MLLM @MME-Benchmarks

Block or report BradyFU

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

Python 312 13 Updated Aug 24, 2026

A novel benchmerk specially designed to evaluate Omni-LLMs under assistant-style, real-time video chat scenarios.

HTML 32 1 Updated Aug 24, 2026

EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory

Python 58 1 Updated Jul 24, 2026

OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains

Python 64 3 Updated Jul 8, 2026
Python 9 1 Updated May 20, 2026

[NeurIPS2026]"UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification"

Python 46 5 Updated Sep 26, 2026

[ACM MM 2026] Tango: Taming Visual Signals for Efficient Video Large Language Models

Python 9 Updated Jun 12, 2026

TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained GPUs.

Python 854 79 Updated Aug 4, 2026

awesome-native-multimodal-models

Python 14 Updated Mar 6, 2026

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Python 362 3 Updated Aug 5, 2026

VITA-QINYU: Expressive Spoken Language Model for Role-Playing and Singing

Python 123 8 Updated Jul 14, 2026

VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding

Python 58 1 Updated May 1, 2026

[CVPR 2026 Highlight] PersonaVLM: Long-Term Personalized Multimodal LLMs

Jupyter Notebook 125 8 Updated Apr 16, 2026

✨✨[ICML 2026] Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

Python 156 8 Updated Mar 12, 2026

A PyTorch Library for Multi-Attribution Learning in CVR prediction

Python 9 Updated Mar 3, 2026
Python 198 5 Updated Jan 19, 2026

🔥 A curated roadmap to the Efficient VLA landscape. We’re keeping this list live—contribute your latest work!

219 7 Updated Aug 17, 2026

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

Python 522 52 Updated Mar 30, 2026

The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.

Python 165 4 Updated Oct 28, 2025

Long Video Gen Infrastructure

Python 2,654 257 Updated Sep 30, 2026

✨✨ [ICLR 2026] Think Beyond Images

Python 583 37 Updated Sep 23, 2025

Github repository for ACL 2025 paper: Recent Advances in Speech Language Models: A Survey.

225 12 Updated Aug 21, 2026

CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms

25 Updated Dec 21, 2025

✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Python 294 22 Updated May 9, 2025

✨✨[NeurIPS 2025] VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model

Python 689 62 Updated May 24, 2025

✨✨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models

Python 43 4 Updated Apr 10, 2025

[CVPR 2026] MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources

Python 217 9 Updated Sep 26, 2025

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation

Jupyter Notebook 32 1 Updated Mar 28, 2025

✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension"

Python 461 46 Updated Jun 26, 2026
Next