Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome AI Video Awesome

A curated list of the best AI text-to-video & image-to-video models, tools, and resources.

The space moves fast — this list focuses on what's actually usable today (late 2026): hosted models, open-source models you can self-host, the tooling around them, and where to learn. Links go to official sources. Contributions welcome — see Contributing.

Legend: 🔓 open-source / self-hostable · 💲 commercial / hosted · 🆓 has a free tier

Contents

What's new

  • Seedance 2.5 — 30s single-pass generation, multi-round extension, up to 30 image / 10 video / 10 audio references, new clay-render, motion, and creative references. Listed under Hosted models.
  • GPT Image 2.5 — @Sketch, comment-based revisions, and two API models (Flare, Sunburst). Listed under Image models for keyframes & references.

Full history in CHANGELOG.md.

Hosted models

State-of-the-art commercial models you access via web app or API.

  • Seedance 2.5 💲🆓 — ByteDance's narrative-driven model, now with 30-second single-pass generation (up from 15s in 2.0) and multi-round extension that keeps characters and environments consistent. Takes up to 30 images, 10 videos, and 10 audio references per prompt, adds clay-render (pose, motion path, camera angle), motion, and creative references, and supports timestamp-level, green-screen/background-replacement, camera-perspective, and reference-based editing. Available on Dreamina (international) and via API on BytePlus ModelArk; Seedance 2.0 remains available on both.
  • Google Veo 3.1 💲 — Generates 48kHz synchronized dialogue, ambient sound, and music as part of the same diffusion pass. Usable through Google Flow and the Gemini app.
  • Runway Gen-4.5 💲🆓 — The most precise control surface for directors: motion brushes, scene consistency, and the GWM-1 world model.
  • Kling 3.0 💲🆓 — Strong native audio and lip-sync across multiple languages, with a shared audio timeline across multi-shot sequences.
  • Luma Ray3 💲🆓 — The first AI video model with native 16-bit HDR output. Part of Dream Machine.
  • MiniMax Hailuo 💲🆓 — The 2026 value pick: quality between Pika and Runway at noticeably lower pricing.
  • Pika 💲🆓 — Fast and accessible, great for short-form creative iteration and effects.
  • PixVerse 💲🆓 — Popular for stylized short-form and social video, with a generous free tier.
  • OpenAI Sora 💲 — ⚠️ Wound down: the web/app experience was discontinued April 26, 2026, and the API was scheduled to end September 24, 2026 — that date has passed. Migrate to Veo, Kling, Runway, or Seedance.

Open-source models

Models with public weights you can run yourself.

  • Wan 2.2 🔓 — Alibaba's Apache-2.0 model with a Mixture-of-Experts architecture (27B total / 14B active). The community favorite, with the largest LoRA ecosystem.
  • HunyuanVideo 🔓 — Tencent's 13B foundation model with a full open ecosystem: weights, multi-GPU inference, FP8, Diffusers and ComfyUI integrations.
  • LTX-Video 🔓 — Lightricks' DiT model built for speed — generates 1216×704 video faster than real time.
  • Mochi 1 🔓 — Genmo's 10B Asymmetric Diffusion Transformer, released under Apache 2.0.
  • CogVideoX 🔓 — Tsinghua/Zhipu's open text- and image-to-video model combining a 3D VAE with an expert Transformer.

Image models for keyframes & references

Image models used to build start frames and reference images for image-to-video and reference-driven workflows.

  • GPT Image 2.5 💲 — OpenAI's image model behind ChatGPT Images 2.5, with @Sketch input, comment annotations for revisions, more precise editing, and up to 50% lower latency than 2.0. The API ships two models: GPT-Image-2.5 Flare (fast, default) and Sunburst (premium, precise editing).

Tools & platforms

Run, chain, and deploy the models above.

  • ComfyUI 🔓 — Node-based workflow UI; the de facto home for open-source video pipelines (Wan, Hunyuan, LTX, and more). See also comfy.org.
  • Wan2GP 🔓 — "AI video for the GPU-poor" — runs Wan 2.1/2.2, Hunyuan, and LTX on low-VRAM GPUs.
  • fal.ai 💲 — Low-latency generative-media API hosting most major video models, including Seedance.
  • Replicate 💲 — Run thousands of models (many video) via a simple API, no GPU management.
  • Hugging Face 🔓🆓 — Weights, demos, and Spaces for nearly every open video model.

Prompting tools & guides

Get better, more consistent output.

  • seedance-prompt-forge 🔓 — CLI + library that turns simple inputs into clean, structured prompts for Seedance, Kling, Runway, and Veo.
  • shotlist-forge 🔓 — Expands one concept into a structured, shot-by-shot prompt sequence (built on seedance-prompt-forge).
  • Runway Help Center — Official prompting and feature guides for Gen-4.x.
  • Kling AI — In-app prompt examples and templates.

Learning & communities

Contributing

Found something missing, or a broken link? Contributions are very welcome — please read CONTRIBUTING.md and open a pull request. Keep entries factual, link to official sources, and add one concise sentence on what makes the entry worth knowing.

License

CC0

To the extent possible under law, the contributors have waived all copyright and related or neighboring rights to this work. See LICENSE.

About

A curated list of the best AI text-to-video & image-to-video models, tools, and resources (Seedance, Veo, Runway, Kling, Wan, ComfyUI, and more).

Topics

Resources

Contributing

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors