Skip to content

Latest commit

 

History

180 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome Video Generation Post Training

🔥🔥🔥 A continuously updated repository of research on post-training and alignment for video generation.

License: MIT Last Update GitHub stars

arXiv hf_paper

Chaoyu Li1,†,✉️, Xiaoyi Gu2,†, Yogesh Kulkarni1, Eun Woo Im1, Mohammadmahdi Honarmand3, Zeyu Wang4, Juntong Song5, Fei Du6, Xilin Jiang7, Kexin Zheng8, Tianzhi Li9, Fei Tao5, Pooyan Fazli1,✉️

1Arizona State University · 2Twitch · 3Stanford University · 4eBay · 5NewsBreak · 6Microsoft · 7Columbia University · 8University of Southern California · 9Carnegie Mellon University

†Equal contribution. ✉️Corresponding author: Chaoyu Li , Pooyan Fazli.

timeline


Important

We welcome your help in improving the repository and paper. Please feel free to submit a pull request to:

  • Add a relevant paper not yet included.
  • Suggest a more suitable category.
  • Update the information.
  • Ask for clarification about any content.

🔥 News

  • [2026.06.09] Our survey paper has been accepted to TMLR 🚀.
  • [2026.06.09] Added papers accepted to ICML 2026.
  • [2026.06.09] Added papers accepted to CVPR 2026.
  • [2026.03.23] Added papers accepted to ICLR 2026.
  • [2026.02.23] The v1 survey is now published! We have also initialized the repository.

🎯 Motivation

teaser While large-scale pretraining has significantly improved video generation models, aligning them with human intent, physical constraints, and deployment requirements remains a major challenge. As a result, post-training and alignment techniques have rapidly become a central research focus in modern video generation.

In our survey paper, we systematically review recent progress in supervised fine-tuning, self-training and distillation, preference-based optimization, and inference-time alignment methods.

This repository serves as a companion resource to the survey.
It aims to:

  • 📚 Curate recent papers on post-training and alignment for video generation
  • 📊 Collect datasets and evaluation benchmarks related to post-training and alignment

We hope this repository provides researchers and practitioners with a convenient, up-to-date reference for exploring the evolving landscape of post-training and alignment in video generation models.

📌 Citation

If you find our paper or this resource helpful, please consider cite:

@article{li2026video,
  title={Video Generation Models: A Survey of Post-Training and Alignment},
  author={Chaoyu Li and Xiaoyi Gu and Yogesh Kulkarni and Eun Woo Im and Mohammadmahdi Honarmand and Zeyu Wang and Juntong Song and Fei Du and Xilin Jiang and Kexin Zheng and Tianzhi Li and Fei Tao and Pooyan Fazli},
  journal={Transactions on Machine Learning Research},
  issn={2835-8856},
  year={2026},
  url={https://openreview.net/forum?id=YlUEWLESIu},
}

📚 Content


Datasets

Dataset Year Links
NeurIPS ChronoMagic-Pro 2024 Paper · GitHub · Website · Dataset
NeurIPS SafeSora 2024 Paper · GitHub · Website · Dataset
ICCV TIP-I2V 2025 Paper · GitHub · Website · Dataset
ICCV SynFMC 2025 Paper · GitHub · Website · Dataset
CVPR CookGen 2025 Paper · Website · Dataset
CVPR HOIGen-1M 2025 Paper · Website · Dataset
CVPR OpenHumanVid 2025 Paper · GitHub · Website
ICML PhyWorld 2025 Paper · GitHub · Website · Dataset
NeurIPS WISA-80K 2025 Paper · GitHub · Website · Dataset
NeurIPS VideoUFO 2025 Paper · GitHub · Dataset
NeurIPS TalkCuts 2025 Paper · GitHub · Website · Dataset
NeurIPS EgoVid-5M 2025 Paper · GitHub · Website · Dataset
NeurIPS OpenS2V-5M 2025 Paper · GitHub · Website · Dataset
PMLR GRADEO-Instruct 2025 Paper
arXiv PNData 2025 Paper
arXiv PairFS-4K 2025 Paper · GitHub · Website
arXiv MMVideo 2025 Paper · GitHub · Website
arXiv DAVID-X 2025 Paper
arXiv Dprim 2025 Paper
CVPR CI-VID 2026 Paper · GitHub · Dataset
arXiv GB3DV-25k 2026 Paper · Website · Dataset
arXiv MuSS 2026 Paper · GitHub

Benchmarks

Benchmark Year Links
NeurIPS FETV 2023 Paper · GitHub · Dataset
NeurIPS StoryBench 2023 Paper · GitHub
NeurIPS ChronoMagic-Bench 2024 Paper · GitHub · Website · Dataset
CVPR EvalCrafter 2024 Paper · GitHub · Website · Dataset
CVPR VBench 2024 Paper · GitHub · Website · Dataset
NeurIPS T2VSafetyBench 2024 Paper · GitHub
ICCV MTBench 2025 Paper · GitHub · Website · Dataset
ICCV FiVE 2025 Paper · GitHub · Website · Dataset
ICCV VEG-Bench 2025 Paper · GitHub · Website · Dataset
ICCV VMBench 2025 Paper · GitHub · Website · Dataset
CVPR T2V-CompBench 2025 Paper · GitHub · Website · Dataset
CVPR StoryEval 2025 Paper · GitHub · Website · Dataset
CVPR MC-Bench 2025 Paper · GitHub · Website · Dataset
CVPR Video-Bench 2025 Paper · GitHub
NeurIPS MJ-BENCH-VIDEO 2025 Paper · GitHub · Website · Dataset
NeurIPS OpenS2V-Eval 2025 Paper · GitHub · Website · Dataset
NeurIPS VideoGen-RewardBench 2025 Paper · GitHub · Website · Dataset
NeurIPS AIGC-LipSync 2025 Paper · Website · Dataset
NeurIPS WorldModelBench 2025 Paper · GitHub · Website · Dataset
ACM MM VIP-200K 2025 Paper · Website · Dataset
ACM MM RGCD 2025 Paper · GitHub · Dataset
ACM MM RecipeGen 2025 Paper · Dataset
ACM MM HVEval 2025 Paper
EMNLP Doc2Present 2025 Paper · GitHub · Dataset
ACL VidCapBench 2025 Paper · GitHub · Dataset
ICLR VideoPhy 2025 Paper · GitHub · Dataset
ICML PhyGenBench 2025 Paper · GitHub · Website · Dataset
arXiv Paper2Video 2025 Paper · GitHub · Website · Dataset
arXiv PhyWorldBench 2025 Paper · GitHub · Dataset
arXiv SeqBench 2025 Paper · GitHub · Website · Dataset
arXiv VideoVerse 2025 Paper · GitHub · Website · Dataset
arXiv MMMC 2025 Paper · GitHub · Website · Dataset
arXiv DynamicEval 2025 Paper · GitHub · Website · Dataset
arXiv V-ReasonBench 2025 Paper · GitHub · Website · Dataset
arXiv RGCD (alternate) 2025 Paper · GitHub · Dataset
arXiv VIPER 2025 Paper · GitHub · Dataset
arXiv Video Reality Test 2025 Paper · GitHub · Website · Dataset
arXiv WYD 2025 Paper · GitHub
arXiv GenVidBench 2025 Paper
arXiv DIVE 2025 Paper
arXiv UI2V-Bench 2025 Paper
arXiv PhysVidBench 2025 Paper
arXiv SurgVeo 2025 Paper
arXiv PAI-Bench 2025 Paper
arXiv MEve 2025 Paper
ICML SafeMVDrive 2026 Paper · GitHub · Website · Dataset
ICML AIGVE-60K 2026 Paper · GitHub · Dataset
ICML T2AV-Compass 2026 Paper · Dataset
ICML LoCoT2V-Bench 2026 Paper
WACV GeneVA 2026 Paper · Website · Dataset
WACV PhyEduVideo 2026 Paper · GitHub · Website
arXiv DrivingGen 2026 Paper · Website · Dataset
arXiv RISE-Video 2026 Paper · GitHub · Dataset
arXiv RBench 2026 Paper · GitHub · Website · Dataset
arXiv VC-Bench 2026 Paper · Website · Dataset
arXiv CoW-Bench 2026 Paper · GitHub · Dataset
arXiv MTAVG-Bench 2026 Paper
arXiv MSVBench 2026 Paper
arXiv SafeGen-Bench 2026 Paper
arXiv DirectorBench 2026 Paper · GitHub · Dataset
arXiv EvalVerse 2026 Paper
arXiv EntityBench 2026 Paper · GitHub · Website
arXiv MechVerse 2026 Paper · Website
arXiv WorldReasonBench 2026 Paper · GitHub · Website
arXiv WorldJen 2026 Paper · GitHub · Website · Dataset
arXiv BRITE 2026 Paper
arXiv WorldMark 2026 Paper · Website
arXiv VideoScore2 2025 Paper · GitHub · Website · Dataset

Conference Papers

Title Year Links
CVPR FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance 2026 Paper · GitHub · Website
CVPR GenHOI: Towards Object-Consistent Hand–Object Interaction with Temporally Balanced and Spatially Selective Object Injection 2026 Paper · GitHub · Website
CVPR GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling 2026 Paper
CVPR Plenoptic Video Generation 2026 Paper · GitHub · Website
CVPR SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models 2026 Paper · GitHub
CVPR EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses 2026 Paper · GitHub · Website
CVPR Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation 2026 Paper
CVPR TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models 2026 Paper · GitHub · Website
CVPR TGT: Text-Grounded Trajectories for Locally Controlled Video Generation 2026 Paper · Website
CVPR Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization 2026 Paper · GitHub · Website
CVPR Lynx: Towards High-Fidelity Personalized Video Generation 2026 Paper · GitHub · Website
CVPR ORV: 4D Occupancy-centric Robot Video Generation 2026 Paper · GitHub · Website
ICML Where Concept Erasure Should Occur: Concept-Layer Alignment in Text-to-Video Diffusion Models 2026 Paper
ICML World-R1: Reinforcing 3D Constraints for Text-to-Video Generation 2026 Paper · GitHub · Website
ICML WorldCompass: Reinforcement Learning for Long-Horizon World Models 2026 Paper · GitHub · Website
ICML Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation 2026 Paper · GitHub · Website
ICML VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation 2026 Paper · GitHub · Website
ICML TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment 2026 Paper
ICML Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model 2026 Paper
ICML RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space 2026 Paper · GitHub · Website
ICML LayerT2V: Interactive Multi-Object Trajectory Layering for Video Generation 2026 Paper · GitHub · Website
ICML Physics-Guided Motion Loss for Video Generation Model 2026 Paper
ICML SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain 2026 Paper · GitHub · Website
ICLR Neodragon: Mobile Video Generation using Diffusion Transformer 2026 Paper · Website
ICLR The Quest for Generalizable Motion Generation: Data, Model, and Evaluation 2026 Paper · GitHub
ICLR MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models 2026 Paper
ICLR Stable Video Infinity: Infinite-Length Video Generation with Error Recycling 2026 Paper
ICLR Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency 2026 Paper
ICLR MATRIX: Mask Track Alignment for Interaction-aware Video Generation 2026 Paper
ICLR Arbitrary Generative Video Interpolation 2026 Paper
ICLR Rolling Forcing: Autoregressive Long Video Diffusion in Real Time 2026 Paper · GitHub · Website
ICLR BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models 2026 Paper · GitHub · Website
ICLR MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling 2026 Paper · GitHub · Website
ICLR Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration 2026 Paper · GitHub · Website
ICLR DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing 2026 Paper
ICLR Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling 2026 Paper · GitHub · Website
ICLR Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control 2026 Paper · GitHub · Website
ICLR Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks 2026 Paper · GitHub
ICLR MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement 2026 Paper · GitHub · Website
ICLR Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model 2026 Paper · Website
ICLR UNIC: Unified In-Context Video Editing 2026 Paper · Website
ICLR EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning 2026 Paper · Website
ICLR EasyCreator: Empowering 4D Creation through Video Inpainting 2026 Paper · Website
AAAI PanFlow: Decoupled Motion Control for Panoramic Video Generation 2026 Paper · GitHub · Website
AAAI Phased One-Step Adversarial Equilibrium for Video Diffusion Models 2026 Paper · Website
AAAI Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation 2026 Paper · GitHub · Website
WACV Show Me: Unifying Instructional Image and Video Generation with Diffusion Models 2026 Paper · GitHub · Website
WACV UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models 2026 Paper · GitHub
NeurIPS Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search 2025 Paper
NeurIPS Controllable Human-centric Keyframe Interpolation with Generative Prior 2025 Paper · GitHub · Website
NeurIPS RoboScape: Physics-informed Embodied World Model 2025 Paper · GitHub
NeurIPS Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation 2025 Paper
NeurIPS Audio-Sync Video Generation with Multi-Stream Temporal Control 2025 Paper · GitHub · Website
NeurIPS Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation 2025 Paper · Website
NeurIPS DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models 2025 Paper · Website
NeurIPS EchoShot: Multi-Shot Portrait Video Generation 2025 Paper · GitHub · Website
NeurIPS Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models 2025 Paper · GitHub · Website
NeurIPS Frame In-N-Out: Unbounded Controllable Image-to-Video Generation 2025 Paper · GitHub · Website
NeurIPS GeoVideo: Introducing Geometric Regularization into Video Generation Model 2025 Paper · GitHub · Website
NeurIPS Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation 2025 Paper · Website
NeurIPS Imagine360: Immersive 360 Video Generation from Perspective Anchor 2025 Paper · GitHub · Website
NeurIPS Improving Video Generation with Human Feedback 2025 Paper · GitHub · Website
NeurIPS MoCha: Towards Movie-Grade Talking Character Generation 2025 Paper · GitHub · Website
NeurIPS PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement 2025 Paper · Website
NeurIPS Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation 2025 Paper · GitHub · Website
NeurIPS RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation 2025 Paper
NeurIPS Temporal In-Context Fine-Tuning for Versatile Control of Video Diffusion Models 2025 Paper · GitHub · Website
NeurIPS VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models 2025 Paper · GitHub · Website
NeurIPS Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image 2025 Paper · Website
NeurIPS Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance 2025 Paper · GitHub · Website
NeurIPS WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation 2025 Paper · GitHub · Website
NeurIPS WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception 2025 Paper · Website
NeurIPS Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion 2025 Paper · GitHub
EMNLP VC4VG: Optimizing Video Captions for Text-to-Video Generation 2025 Paper · GitHub
ACM MM AICL: Action In-Context Learning for Video Diffusion Model 2025 Paper
ACM MM Improving Identity Preservation in Video Generation with Multi-Branch Models 2025 Paper
ACM MM M2PE-DIFF: Music-to-Pose Encoder for Dance Video Generation Leveraging Latent Diffusion Framework. 2025 Paper
ACM MM SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation 2025 Paper
ACM MM Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation 2025 Paper · Website
ACM MM AnimeColor: Reference-based Animation Colorization with Diffusion Transformers 2025 Paper · GitHub
ICCV Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis 2025 Paper · GitHub · Website
ICCV Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models 2025 Paper · GitHub · Website
ICCV AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction 2025 Paper · Website
ICCV Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis 2025 Paper
ICCV Decouple and Track: Benchmarking and Improving Video Diffusion Transformers For Motion Transfer 2025 Paper · GitHub · Website
ICCV DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion 2025 Paper · Website
ICCV DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization 2025 Paper · Website
ICCV Dual-Expert Consistency Model for Efficient and High-Quality Video Generation 2025 Paper · GitHub · Website
ICCV DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization 2025 Paper · GitHub · Website
ICCV EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models 2025 Paper · GitHub
ICCV Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation 2025 Paper · GitHub · Website
ICCV I2VControl: Disentangled and Unified Video Motion Synthesis Control 2025 Paper · Website
ICCV LayerAnimate: Layer-level Control for Animation 2025 Paper · GitHub · Website
ICCV Learning Few-Step Diffusion Models by Trajectory Distribution Matching 2025 Paper · GitHub · Website
ICCV LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion 2025 Paper · Website
ICCV Long Context Tuning for Video Generation 2025 Paper · Website
ICCV MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization 2025 Paper · GitHub · Website
ICCV MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers 2025 Paper · GitHub
ICCV MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance 2025 Paper · GitHub · Website
ICCV Mobile Video Diffusion 2025 Paper · Website
ICCV MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent 2025 Paper · GitHub
ICCV PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation 2025 Paper · GitHub · Website
ICCV Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment 2025 Paper · GitHub · Website
ICCV Precise Action-to-Video Generation Through Visual Action Prompts 2025 Paper · Website
ICCV Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM 2025 Paper · GitHub
ICCV Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics 2025 Paper · GitHub · Website
ICCV Reangle-A-Video: 4D Video Generation as Video-to-Video Translation 2025 Paper · GitHub · Website
ICCV RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control 2025 Paper · GitHub · Website
ICCV STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution 2025 Paper · GitHub · Website
ICCV TACO: Taming Diffusion for in-the-wild Video Amodal Completion 2025 Paper · GitHub · Website
ICCV V.I.P.: Iterative Online Preference Distillation for Efficient Video Diffusion Models 2025 Paper · Website
ICCV VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation 2025 Paper · GitHub · Website
ICCV Versatile Transition Generation with Image-to-Video Diffusion 2025 Paper · Website
ICCV VideoAuteur: Towards Long Narrative Video Generation 2025 Paper · GitHub · Website
IJCAI FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model 2025 Paper
ICML FrameBridge: Improving Image-to-Video Generation with Bridge Models 2025 Paper · Website
ICML Diffusion Adversarial Post-Training for One-Step Video Generation 2025 Paper · Website
CVPR Identity-Preserving Text-to-Video Generation by Frequency Decomposition 2025 Paper · GitHub · Website
CVPR AKiRa: Augmentation Kit on Rays for Optical Video Generation 2025 Paper · GitHub · Website
CVPR AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models 2025 Paper · GitHub · Website
CVPR Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think 2025 Paper · GitHub
CVPR FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations 2025 Paper · GitHub
CVPR MotionPro: A Precise Motion Controller for Image-to-Video Generation 2025 Paper · GitHub · Website
CVPR FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance 2025 Paper · GitHub · Website
CVPR From Slow Bidirectional to Fast Autoregressive Video Diffusion Models 2025 Paper · GitHub · Website
CVPR Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise 2025 Paper · GitHub · Website
CVPR GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking 2025 Paper · GitHub · Website
CVPR High-Fidelity Relightable Monocular Portrait Animation with Lighting-Controllable Video Diffusion Model 2025 Paper · GitHub · Website
CVPR HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation 2025 Paper · GitHub · Website
CVPR InterDyn: Controllable Interactive Dynamics with Video Diffusion Models 2025 Paper · GitHub · Website
CVPR Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset 2025 Paper · Website
CVPR Mind the Time: Temporally-Controlled Multi-Event Video Generation 2025 Paper · Website
CVPR Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation 2025 Paper · GitHub
CVPR MotiF: Making Text Count in Image Animation with Motion Focal Loss 2025 Paper · Website
CVPR One-Minute Video Generation with Test-Time Training 2025 Paper · GitHub · Website
CVPR ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning 2025 Paper · Website
CVPR ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models 2025 Paper · Website
CVPR TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation 2025 Paper · GitHub · Website
CVPR TransPixeler: Advancing Text-to-Video Generation with Transparency 2025 Paper · GitHub · Website
CVPR VideoDPO: Omni-Preference Alignment for Video Diffusion Generation 2025 Paper · GitHub · Website
ICLR Boosting Camera Motion Control for Video Diffusion Transformers 2025 Paper · GitHub · Website
ICLR CameraCtrl: Enabling Camera Control for Video Diffusion Models 2025 Paper · GitHub · Website
ICLR Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model 2025 Paper · GitHub · Website
ICLR Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models 2025 Paper · GitHub
ICLR SePPO: Semi-Policy Preference Optimization for Diffusion Alignment 2025 Paper · GitHub
ICLR Trajectory attention for fine-grained video motion control 2025 Paper · GitHub · Website
ICLR VADER: Video Diffusion Alignment via Reward Gradients 2025 Paper · GitHub · Website
ICLR VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control 2025 Paper · GitHub · Website
WACV Fine-Grained Controllable Video Generation via Object Appearance and Context 2025 Paper · Website
WACV MagicStick: Controllable Video Editing via Control Handle Transformations 2025 Paper · GitHub · Website
WACV TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models 2025 Paper · GitHub
AAAI CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation 2025 Paper · GitHub · Website
AAAI CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities 2025 Paper · GitHub · Website
AAAI CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training 2025 Paper · GitHub · Website
AAAI DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation 2025 Paper · GitHub · Website
AAAI Enhancing Identity-Deformation Disentanglement in StyleGAN for One-Shot Face Video Re-Enactment 2025 Paper · GitHub · Website
AAAI Modular-Cam: Modular Dynamic Camera-view Video Generation with LLM 2025 Paper · GitHub · Website
AAAI Occlusion-Insensitive Talking Head Video Generation via Facelet Compensation 2025 Paper
AAAI PanoDiT: Panoramic Videos Generation with Diffusion Transformer 2025 Paper
AAAI TrackGo: A Flexible and Efficient Method for Controllable Video Generation 2025 Paper · Website
AAAI UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer 2025 Paper · GitHub
SIGGRAPH Asia Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy 2025 Paper · GitHub · Website
SIGGRAPH Asia FairyGen: Storied Cartoon Video from a Single Child-Drawn Character 2025 Paper · GitHub · Website
SIGGRAPH Asia CamCloneMaster: Enabling Reference-based Camera Control for Video Generation 2025 Paper · GitHub · Website
MICCAI Mission Balance: Generating Under-represented Class Samples using Video Diffusion Models 2025 Paper
MICCAI Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model 2025 Paper · GitHub
ECCV Animate Your Motion: Turning Still Images into Dynamic Videos 2024 Paper · GitHub · Website
ECCV DragVideo: Interactive Drag-style Video Editing 2024 Paper · GitHub · Website
ECCV MEVG : Multi-event Video Generation with Text-to-Video Models 2024 Paper · GitHub · Website
ECCV MoVideo: Motion-Aware Video Generation with Diffusion Models 2024 Paper · Website
ECCV SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models 2024 Paper · GitHub · Website
ECCV Stable Video Portraits 2024 Paper · Website
ECCV TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models 2024 Paper · GitHub · Website
ECCV Video Editing via Factorized Diffusion Distillation 2024 Paper · Website
ECCV VideoStudio: Generating Consistent-Content and Multi-Scene Videos 2024 Paper · GitHub · Website
COLM VideoDirectorGPT: Consistent Multi-Scene Video Generation via LLM-Guided Planning 2024 Paper · GitHub · Website
NeurIPS Exo2Ego-V: Exocentric-to-Egocentric Video Generation 2024 Paper · GitHub · GitHub
NeurIPS MotionBooth: Motion-Aware Customized Text-to-Video Generation 2024 Paper · GitHub · Website
NeurIPS SF-V: Single Forward Video Generation Model 2024 Paper · GitHub · Website
NeurIPS T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback 2024 Paper · GitHub · Website
NeurIPS Vivid-ZOO: Multi-View Video Generation with Diffusion Model 2024 Paper · GitHub · Website
EMNLP VIDEOSCORE: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 2024 Paper · GitHub · Website
EMNLP VIMI: Grounding Video Generation through Multi-modal Instruction 2024 Paper · GitHub · Website
ACM MM DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control 2024 Paper · GitHub
ACM MM DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships 2024 Paper
ICLR Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition 2024 Paper · Website
ICLR SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction 2024 Paper · GitHub · Website
ICLR VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video Generation 2024 Paper · Website
CVPR InstructVideo: Instructing Video Diffusion Models with Human Feedback 2024 Paper · GitHub · Website
CVPR SimDA: Simple Diffusion Adapter for Efficient Video Generation 2024 Paper · GitHub · Website
CVPR VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models 2024 Paper · GitHub · Website
AAAI Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free Videos 2024 Paper · GitHub · Website
NeurIPS VideoComposer: Compositional Video Synthesis with Motion Controllability 2023 Paper · GitHub · Website
ACM MM MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text 2023 Paper
ICCV DreamPose: Fashion Video Synthesis with Stable Diffusion 2023 Paper · GitHub · Website
ICCV Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models 2023 Paper · Website
ICCV Structure and Content-Guided Video Synthesis with Diffusion Models 2023 Paper · GitHub · Website
ICCV Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation 2023 Paper · GitHub · Website

arXiv Papers

Title Year Links
arXiv Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions 2026 Paper · Website
arXiv Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation 2026 Paper · Website
arXiv Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control 2026 Paper
arXiv Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment 2026 Paper
arXiv Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion 2026 Paper · Website
arXiv minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models 2026 Paper · GitHub
arXiv DeltaCam: Differential Intrinsic Camera Modeling for Video Generation 2026 Paper
arXiv Geo-Align: Video Generation Alignment via Metric Geometry Reward 2026 Paper · Website
arXiv SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models 2026 Paper · GitHub · Website
arXiv CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration 2026 Paper · Website
arXiv CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition 2026 Paper · Website
arXiv Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls 2026 Paper
arXiv PhyWorld: Physics-Faithful World Model for Video Generation 2026 Paper
arXiv GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation 2026 Paper · GitHub · Website
arXiv Video Models Can Reason with Verifiable Rewards 2026 Paper · Website
arXiv RefDecoder: Enhancing Visual Generation with Conditional Video Decoding 2026 Paper
arXiv RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO 2026 Paper · Website
arXiv DriveCtrl: Conditioned Sim-to-Real Driving Video Generation 2026 Paper
arXiv EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration 2026 Paper · Website
arXiv SEDiT: Mask-Free Video Subtitle Erasure via One-step Diffusion Transformer 2026 Paper · Website
arXiv Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation 2026 Paper · Website
arXiv CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL 2026 Paper
arXiv PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation 2026 Paper · Website
arXiv CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation 2026 Paper
arXiv CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating 2026 Paper
arXiv TIE: Time Interval Encoding for Video Generation over Events 2026 Paper · GitHub · Website
arXiv Improving Human Image Animation via Semantic Representation Alignment 2026 Paper
arXiv From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation 2026 Paper
arXiv Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models 2026 Paper · GitHub
arXiv Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers 2026 Paper
arXiv Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation 2026 Paper · Website
arXiv PhyCo: Learning Controllable Physical Priors for Generative Motion 2026 Paper · Website
arXiv AesRM: Improving Video Aesthetics with Expert-Level Feedback 2026 Paper · Website
arXiv A Systematic Post-Train Framework for Video Generation 2026 Paper
arXiv MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation 2026 Paper
arXiv Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation 2026 Paper
arXiv Not all tokens contribute equally to diffusion learning 2026 Paper
arXiv HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation 2026 Paper · Website
arXiv OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models 2026 Paper
arXiv MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling 2026 Paper · Website
arXiv Control-DINO: Feature Space Conditioning for Controllable Image-to-Video Diffusion 2026 Paper · GitHub · Website
arXiv ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation 2026 Paper · Website
arXiv Wan-R1: Verifiable-Reinforcement Learning for Video Reasoning 2026 Paper
arXiv TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets 2026 Paper · Website
arXiv LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model 2026 Paper · Website
arXiv VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward 2026 Paper · Website
arXiv DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation 2026 Paper
arXiv RefAlign: Representation Alignment for Reference-to-Video Generation 2026 Paper · GitHub
arXiv ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment 2026 Paper · GitHub
arXiv Manifold-Aware Exploration for Reinforcement Learning in Video Generation 2026 Paper · Website
arXiv PROBE: Diagnosing Residual Concept Capacity in Erased Text-to-Video Diffusion Models 2026 Paper · GitHub
arXiv Helios: Real Real-Time Long Video Generation Model 2026 Paper · GitHub · Website
arXiv 3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model 2026 Paper · GitHub · Website
arXiv AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization 2026 Paper
arXiv Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion 2026 Paper · GitHub · Website
arXiv ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation 2026 Paper
arXiv Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints 2026 Paper
arXiv DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning 2026 Paper · Website
arXiv Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards 2026 Paper · GitHub · Website
arXiv EraseAnything++: Enabling Concept Erasure in Rectified Flow Transformers Leveraging Multi-Object Optimization 2026 Paper
arXiv FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation 2026 Paper
arXiv InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions 2026 Paper
arXiv LibraGen: Playing a Balance Game in Subject-Driven Video Generation 2026 Paper · GitHub · Website
arXiv Making Video Models Adhere to User Intent with Minor Adjustments 2026 Paper · GitHub · Website
arXiv MosaicMem: Hybrid Spatial Memory for Controllable Video World Models 2026 Paper · Website
arXiv Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics 2026 Paper · GitHub · Website
arXiv SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation 2026 Paper
arXiv SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation 2026 Paper · GitHub · Website
arXiv SPIRAL: A Closed-Loop Framework for Self-Improving Action World Models via Reflective Planning Agents 2026 Paper
arXiv ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation 2026 Paper · GitHub · Website
arXiv VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment 2026 Paper · Website
arXiv ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer 2026 Paper · GitHub
arXiv AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories 2026 Paper · GitHub · Website
arXiv High-Fidelity Causal Video Diffusion Models for Real-Time Ultra-Low-Bitrate Semantic Communication 2026 Paper
arXiv EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation 2026 Paper
arXiv Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing 2026 Paper
arXiv Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing 2026 Paper · GitHub · Website
arXiv ALIVE: Animate Your World with Lifelike Audio-Video Generation 2026 Paper · GitHub · Website
arXiv PISCO: Precise Video Instance Insertion with Sparse Control 2026 Paper · GitHub · Website
arXiv TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation 2026 Paper · GitHub
arXiv Optimizing Few-Step Generation with Adaptive Matching Distillation 2026 Paper
arXiv Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving 2026 Paper · Website
arXiv LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation 2026 Paper · GitHub
arXiv Euphonium: Steering Video Flow Matching via Process Reward Gradient Guided Stochastic Dynamics 2026 Paper · GitHub
arXiv BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks 2026 Paper · Website
arXiv Unified Personalized Reward Model for Vision Generation 2026 Paper · GitHub · Website
arXiv PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards 2026 Paper
arXiv Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V 2026 Paper
arXiv Reward-Forcing: Autoregressive Video Generation with Reward Feedback 2026 Paper · GitHub · Website
arXiv Walk through Paintings: Egocentric World Models from Internet Priors 2026 Paper · GitHub · Website
arXiv Human detectors are surprisingly powerful reward models 2026 Paper
arXiv PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models 2026 Paper · Website
arXiv Inference-time Physics Alignment of Video Generative Models with Latent World Models 2026 Paper
arXiv Beyond Inpainting: Unleash 3D Understanding for Precise Camera-Controlled Video Generation 2026 Paper · Website
arXiv Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models 2026 Paper
arXiv Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models 2026 Paper
arXiv Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model 2026 Paper
arXiv DreamLoop: Controllable Cinemagraph Generation from a Single Photograph 2026 Paper · Website
arXiv Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding 2026 Paper
arXiv PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation 2025 Paper · GitHub · Website
arXiv LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation 2025 Paper · GitHub
arXiv Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion 2025 Paper · Website
arXiv Characterizing Motion Encoding in Video Diffusion Timesteps 2025 Paper
arXiv ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision 2025 Paper
arXiv DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation 2025 Paper · Website
arXiv Few-Shot-Based Modular Image-to-Video Adapter for Diffusion Models 2025 Paper · GitHub · Website
arXiv StoryMem: Multi-shot Long Video Storytelling with Memory 2025 Paper · GitHub · Website
arXiv CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization 2025 Paper · Website
arXiv Animate Any Character in Any World 2025 Paper · GitHub · Website
arXiv InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion 2025 Paper · Website
arXiv The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text 2025 Paper · GitHub · Website
arXiv Factorized Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models 2025 Paper · GitHub · Website
arXiv LongVie 2: Multimodal Controllable Ultra-Long Video World Model 2025 Paper · GitHub · Website
arXiv SneakPeek: Future-Guided Instructional Streaming Video Generation 2025 Paper
arXiv What Happens Next? Next Scene Prediction with a Unified Video Model 2025 Paper
arXiv GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation 2025 Paper · GitHub · Website
arXiv STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative 2025 Paper
arXiv SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation 2025 Paper
arXiv BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models 2025 Paper · Website
arXiv Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation 2025 Paper · Website
arXiv AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation 2025 Paper · Website
arXiv ReCamDriving: LiDAR-Free Camera-Controlled Novel Trajectory Video Generation 2025 Paper · GitHub · Website
arXiv In-Context Sync-LoRA for Portrait Video Editing 2025 Paper · GitHub · Website
arXiv Progressive Image Restoration via Text-Conditioned Video Generation 2025 Paper
arXiv Generative Video Motion Editing with 3D Point Tracks 2025 Paper · Website
arXiv SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation 2025 Paper
arXiv DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models 2025 Paper · Website
arXiv Low-Bitrate Video Compression through Semantic-Conditioned Diffusion 2025 Paper
arXiv AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement 2025 Paper · GitHub · Website
arXiv BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation 2025 Paper · GitHub · Website
arXiv WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation 2025 Paper · GitHub · Website
arXiv CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion 2025 Paper · GitHub · Website
arXiv MotionV2V: Editing Motion in a Video 2025 Paper · GitHub · Website
arXiv Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis 2025 Paper
arXiv SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation 2025 Paper · GitHub · Website
arXiv Beyond Reward Margin: Rethinking and Resolving Likelihood Displacement in Diffusion Models via Video Generation 2025 Paper
arXiv View-Consistent Diffusion Representations for 3D-Consistent Video Generation 2025 Paper · Website
arXiv Eevee: Towards Close-up High-resolution Video-based Virtual Try-on 2025 Paper · GitHub
arXiv STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution 2025 Paper · GitHub · Website
arXiv Point-to-Point: Sparse Motion Guidance for Controllable Video Editing 2025 Paper
arXiv CamC2V: Context-aware Controllable Video Generation 2025 Paper · GitHub
arXiv MultiShotMaster: A Controllable Multi-Shot Video Generation Framework 2025 Paper · GitHub · Website
arXiv LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization 2025 Paper · GitHub · Website
arXiv Taming Camera-Controlled Video Generation with Verifiable Geometry Reward 2025 Paper
arXiv IC-World: In-Context Generation for Shared World Modeling 2025 Paper · GitHub
arXiv Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Make Objects Move Slower Than in Reality 2025 Paper · Website
arXiv What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards 2025 Paper · GitHub · Website
arXiv InstanceV: Instance-Level Video Generation 2025 Paper · GitHub · Website
arXiv McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning 2025 Paper · GitHub · Website
arXiv One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer 2025 Paper · GitHub · Website
arXiv MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training 2025 Paper · Website
arXiv Video Generation Models Are Good Latent Reward Models 2025 Paper · Website
arXiv MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models 2025 Paper
arXiv Growing with the Generator: Self-paced GRPO for Video Generation 2025 Paper
arXiv Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation 2025 Paper
arXiv Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO 2025 Paper · GitHub · Website
arXiv First Frame Is the Place to Go for Video Content Customization 2025 Paper · GitHub · Website
arXiv Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation 2025 Paper · GitHub
arXiv Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models 2025 Paper
arXiv ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation 2025 Paper
arXiv LiteAttention: A Temporal Sparse Attention for Diffusion Transformers 2025 Paper · GitHub · Website
arXiv EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation 2025 Paper · Website
arXiv Video Text Preservation with Synthetic Text-Rich Videos 2025 Paper
arXiv RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation 2025 Paper · Website
arXiv PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection 2025 Paper
arXiv UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions 2025 Paper · GitHub · Website
arXiv ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation 2025 Paper · GitHub · Website
arXiv Fine-Tuning Open Video Generators for Cinematic Scene Synthesis: A Small-Data Pipeline with LoRA and Wan2.1 I2V 2025 Paper · Website
arXiv VC4VG: Optimizing Video Captions for Text-to-Video Generation 2025 Paper · GitHub · Website
arXiv CoMo: Compositional Motion Customization for Text-to-Video Generation 2025 Paper · Website
arXiv Epipolar Geometry Improves Video Generation Models 2025 Paper · GitHub · Website
arXiv In-Context Learning with Unpaired Clips for Instruction-based Video Editing 2025 Paper · GitHub
arXiv Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning 2025 Paper · GitHub · Website
arXiv Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures 2025 Paper · Website
arXiv PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning 2025 Paper
arXiv MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars 2025 Paper
arXiv Real-Time Motion-Controllable Autoregressive Video Diffusion 2025 Paper
arXiv PickStyle: Video-to-Video Style Transfer with Context-Style Adapters 2025 Paper
arXiv Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report 2025 Paper
arXiv Character Mixing for Video Generation 2025 Paper · GitHub
arXiv FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation 2025 Paper · GitHub · Website
arXiv DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder 2025 Paper · GitHub · Website
arXiv PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion 2025 Paper · Website
arXiv Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer 2025 Paper · Website
arXiv Reinforcement Learning with Inverse Rewards for World Model Post-training 2025 Paper
arXiv Echo-Path: Pathology-Conditioned Echo Video Generation 2025 Paper · GitHub
arXiv VidCLearn: A Continual Learning Approach for Text-to-Video Generation 2025 Paper
arXiv OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data 2025 Paper
arXiv PanoLora: Bridging Perspective and Panoramic Video Generation with LoRA Adaptation 2025 Paper · Website
arXiv Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis 2025 Paper · Website
arXiv Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders 2025 Paper · Website
arXiv RewardDance: Reward Scaling in Visual Generation 2025 Paper
arXiv HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning 2025 Paper · GitHub · Website
arXiv Zero-shot 3D-Aware Trajectory-Guided image-to-video generation via Test-Time Training 2025 Paper · GitHub · Website
arXiv UniVerse-1: Unified Audio-Video Generation via Stitching of Experts 2025 Paper · GitHub · Website
arXiv DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective 2025 Paper
arXiv Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation 2025 Paper
arXiv Learning Primitive Embodied World Models: Towards Scalable Robotic Learning 2025 Paper · GitHub · Website
arXiv MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation 2025 Paper · Website
arXiv ControlEchoSynth: Boosting Ejection Fraction Estimation Models via Controlled Video Diffusion 2025 Paper
arXiv SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation 2025 Paper
arXiv Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation 2025 Paper
arXiv WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception 2025 Paper
arXiv 4DNeX: Feed-Forward 4D Generative Modeling Made Easy 2025 Paper · GitHub · Website
arXiv Physical Autoregressive Model for Robotic Manipulation without Action Pretraining 2025 Paper · GitHub · Website
arXiv From Large Angles to Consistent Faces: Identity-Preserving Video Generation via Mixture of Facial Experts 2025 Paper · GitHub · Website
arXiv Animate-X++: Universal Character Image Animation with Dynamic Backgrounds 2025 Paper · GitHub · Website
arXiv SketchAnimator: Animate Sketch via Motion Customization of Text-to-Video Diffusion Models 2025 Paper
arXiv SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment 2025 Paper
arXiv DreamVE: Unified Instruction-based Image and Video Editing 2025 Paper · Website
arXiv PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation 2025 Paper
arXiv Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm 2025 Paper
arXiv LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation 2025 Paper · GitHub · Website
arXiv Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation 2025 Paper · GitHub · Website
arXiv DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework 2025 Paper · GitHub · Website
arXiv Compositional Video Synthesis by Temporal Object-Centric Learning 2025 Paper
arXiv Enhancing Scene Transition Awareness in Video Generation via Post-Training 2025 Paper
arXiv Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA 2025 Paper · Website
arXiv Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model 2025 Paper · GitHub
arXiv PUSA V1.0: Surpassing Wan-I2V with $500 Training Cost by Vectorized Timestep Adaptation 2025 Paper · GitHub · Website
arXiv Conditional Video Generation for High-Efficiency Video Compression 2025 Paper
arXiv StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation 2025 Paper · Website
arXiv Geometry-aware 4D Video Generation for Robot Manipulation 2025 Paper · GitHub · Website
arXiv Populate-A-Scene: Affordance-Aware Human Video Generation 2025 Paper · Website
arXiv TextMesh4D: High-Quality Text-to-4D Mesh Generation 2025 Paper
arXiv SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation 2025 Paper · Website
arXiv MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation 2025 Paper
arXiv Video Virtual Try-on with Conditional Diffusion Transformer Inpainter 2025 Paper
arXiv DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing 2025 Paper · GitHub · Website
arXiv Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router 2025 Paper · GitHub · Website
arXiv AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models 2025 Paper · GitHub · Website
arXiv RDPO: Real Data Preference Optimization for Physics Consistency Video Generation 2025 Paper · Website
arXiv FramePrompt: In-context Controllable Animation with Zero Structural Changes 2025 Paper · Website
arXiv DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning 2025 Paper
arXiv Causally Steered Diffusion for Automated Video Counterfactual Generation 2025 Paper · GitHub · Website
arXiv iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer 2025 Paper
arXiv PlayerOne: Egocentric World Simulator 2025 Paper · GitHub · Website
arXiv LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning 2025 Paper · GitHub · Website
arXiv DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers 2025 Paper · Website
arXiv GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning 2025 Paper
arXiv AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation 2025 Paper · Website
arXiv Enhancing Motion Dynamics of Image-to-Video Models via Adaptive Low-Pass Guidance 2025 Paper · GitHub · Website
arXiv HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation 2025 Paper
arXiv Seedance 1.0: Exploring the Boundaries of Video Generation Models 2025 Paper · Website
arXiv Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models 2025 Paper · GitHub · Website
arXiv Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models 2025 Paper · GitHub · Website
arXiv Consistent Video Editing as Flow-Driven Image-to-Video Generation 2025 Paper
arXiv From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models 2025 Paper · GitHub · Website
arXiv Restereo: Diffusion stereo video generation and restoration 2025 Paper
arXiv FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers 2025 Paper · Website
arXiv Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers 2025 Paper · GitHub
arXiv OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation 2025 Paper
arXiv Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents 2025 Paper
arXiv Playing with Transformer at 30+ FPS via Next-Frame Diffusion 2025 Paper · Website
arXiv DeepVerse: 4D Autoregressive Video Generation as a World Model 2025 Paper · GitHub · Website
arXiv Temporal In-Context Fine-Tuning for Versatile Control of Video Diffusion Models 2025 Paper · GitHub · Website
arXiv SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers 2025 Paper · Website
arXiv Video Signature: Implicit Watermarking for Video Diffusion Models 2025 Paper
arXiv Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes 2025 Paper · GitHub · Website
arXiv ATI: Any Trajectory Instruction for Controllable Video Generation 2025 Paper · GitHub · Website
arXiv FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing 2025 Paper · Website
arXiv EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance 2025 Paper · GitHub · Website
arXiv Think Before You Diffuse: Infusing Physical Rules into Video Diffusion 2025 Paper · Website
arXiv HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters 2025 Paper · GitHub · Website
arXiv Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM 2025 Paper
arXiv DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation 2025 Paper · GitHub · Website
arXiv T2VUnlearning: A Concept Erasing Method for Text-to-Video Diffusion Models 2025 Paper · GitHub
arXiv MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation 2025 Paper · GitHub · Website
arXiv ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images 2025 Paper
arXiv HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation 2025 Paper · GitHub · Website
arXiv A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation 2025 Paper
arXiv VIDSTAMP: A Temporally-Aware Watermark for Ownership and Integrity in Video Diffusion Models 2025 Paper · Website
arXiv ReVision: High-Quality, Low-Cost Video Generation with Explicit 3D Physics Modeling for Complex Motion and Interaction 2025 Paper · Website
arXiv ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance 2025 Paper
arXiv Subject-driven Video Generation via Disentangled Identity and Motion 2025 Paper · GitHub · Website
arXiv Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning 2025 Paper
arXiv UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer 2025 Paper · GitHub
arXiv Taming Consistency Distillation for Accelerated Human Image Animation 2025 Paper
arXiv Aligning Anime Video Generation with Human Feedback 2025 Paper · GitHub
arXiv I Want It That Way! Specifying Nuanced Camera Motions in Video Editing 2025 Paper
arXiv Discriminator-Free Direct Preference Optimization for Video Diffusion 2025 Paper
arXiv EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model 2025 Paper
arXiv RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism 2025 Paper · GitHub
arXiv MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators 2024 Paper · GitHub · Website
arXiv VideoFactory: Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation 2023 Paper

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.


🧑‍💻 Contributors

👏 Thanks to these contributors for this excellent work!

✉️ Contact

For questions, suggestions, or collaboration opportunities, please feel free to reach out:

✉️ Email: chaoyuli@asu.edu, pooyan@asu.edu

✨ Star History

Star History Chart

About

🔥🔥🔥 [TMLR] Video Generation Models: A Survey of Post-Training and Alignment | A continuously updated collection of papers, datasets, and benchmarks on post-training and alignment for video generation.

Resources

Stars

207 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors