ZenStory AI

ZenStory AI Projects

Video workflow

Video Recap Skills

Turn supported video files into Chinese-narration recaps with scene understanding, scripting, voiceover, editing and subtitle assembly.

541 GitHub stars6 skills

Source on GitHub English README

Install Video Recap Skills
npx skills add zenstory-ai/video-recap-skills -y -g

What it is

Video Recap Skills turns a local video into a Chinese-narration recap inside Claude Code: five independent skills plus a thin orchestrator. It runs on local ffmpeg and one MiMo API key for speech recognition, visual understanding and voiceover, with Fish Audio as an optional TTS provider, and needs no GPU. Besides the finished video it can export an editable JianYing/CapCut draft so you can keep cutting in the editor.

What you need

  • LicenseMIT, open source.
  • HostAn agent host: Claude Code, Codex CLI, Google Antigravity, OpenCode, ZCode, OpenClaw, Reasonix, Kimi Code, DeepSeek Harness.
  • InstallOne command: npx skills add zenstory-ai/video-recap-skills -y -g

Start in 3 steps

  1. Installnpx skills add zenstory-ai/video-recap-skills -y -g
  2. Entry commandRun /video-recap in your agent host.
  3. Follow a guidePick one of the 3 guides below for your first task.

Practical guides

Writing craft

Storyboards and AI video

Video scripts and editing

How it works

  1. How do I turn a local video into a Chinese recap? Build a scene/dialogue evidence index, choose the audience promise, divide the work between pictures, original audio and narration, then voice and assemble the grounded script. The complete workflow guide explains full/cut choices, stage handoffs and expected files before optional draft export.
  2. Can I add narration when the source has none? Yes. No narration does not mean no dialogue; preserve speech evidence when people speak. Dialogue-free footage can skip ASR while still using remote VLM analysis. The original evidence-to-script examples distinguish a changed on-screen promise from a proven result, then show what can be said when only the initial attempts are available. When later footage is missing, request that evidence or propose a clearly narrower account; do not invent an ending.
  3. Which moments should keep their original sound? Let a complete original line, a revealing action sound or a meaningful wait lead. The five-beat radio example compares shortening context with omitting it, then shows captions alongside each sound: retain faithful speech captions, separate optional sound labels, and do not turn “after the move” into “the move caused the fault.” The edit preserves the wait and invents neither poverty nor a completed mix.
  4. Can I write the recap before choosing the cut? Decide the story and sound roles first. In the orchestrated cut flow, select source clips, then write narration against the actual edited output timeline. Do not copy original timestamps into that output-time narration.
  5. What if the voiceover is too long? Reduce or move the thought instead of speaking over the next protected line. Bounded speed adjustment is not unlimited; an unsafe fit can block placement rather than clipping the spoken ending.
  6. Does an audio-owner label automatically preserve the mix? No. The renderer does not parse the creative board. Actual narration timing, speech evidence and mix settings implement the plan; source audio can remain ducked beyond the written block until a safe recovery point. Listen to the rendered handoff.
  7. Can I finish the sound and captions manually? Export the existing timeline, then choose what the scene needs. The JianYing / CapCut guide compares two original rehearsal edits: move commentary after the result, or remove it for an original-audio-led cut. Move matching captions and inspect the old and new volume keyframes; moving only a voiceover does not move every dependent element.

What makes it different

  • ffmpeg handles local media processing, while configured remote providers receive the media or text required for ASR, visual understanding and speech generation; this is not an offline-only pipeline.
  • An optional CapCut/JianYing draft export supports manual finishing.
  • Automated checks locate candidate problems; real playback remains the final review boundary.

Who it is for

  • Recap and commentary channel creators
  • Editors who finish in CapCut
  • Teams turning long footage into short narrated cuts

Terms it uses

解说视频 剪映草稿 先剪后配

Sources

This page describes the source as read on 2026-09-29; star count as of 2026-09-29. Links point at that version.

Part of ZenStory AI

Six MIT-licensed projects for writers and story creators. Write web fiction with AI skills inside Claude Code or Codex, turn a novel into short-drama scripts and storyboards, a playable game or a narrated video recap — or write in the browser at the ZenStory Workbench, no agent setup needed.

ProjectFormatWhat it does
Oh Story Agent skill pack Turn a coding agent into a complete web-fiction writing workflow — chart scanning, deconstruction, drafting, de-AI editing, covers.
Drama Skills Production skill suite Develop ideas or source stories into episode scripts, visual specifications, storyboards and generation prompts, with optional production and review.
Novel to Game Adaptation skill pack A source-grounded workflow for adapting novels into playable game candidates on a chosen target runtime.
Video Recap Skills Video workflow Turn supported video files into Chinese-narration recaps with scene understanding, scripting, voiceover, editing and subtitle assembly.
Oh Story DSH DSH plugin The whole story stack inside DeepSeek Harness, with a GUI workbench.
ZenStory Workbench Web workbench Chat to create — an agent-driven novel-writing workbench that keeps files, materials, characters, chapters and versions in one place.