Video workflow
Video Recap Skills
Turn supported video files into Chinese-narration recaps with scene understanding, scripting, voiceover, editing and subtitle assembly.
541 GitHub stars6 skills
Source on GitHub English README
npx skills add zenstory-ai/video-recap-skills -y -gWhat it is
Video Recap Skills turns a local video into a Chinese-narration recap inside Claude Code: five independent skills plus a thin orchestrator. It runs on local ffmpeg and one MiMo API key for speech recognition, visual understanding and voiceover, with Fish Audio as an optional TTS provider, and needs no GPU. Besides the finished video it can export an editable JianYing/CapCut draft so you can keep cutting in the editor.
What you need
- LicenseMIT, open source.
- HostAn agent host: Claude Code, Codex CLI, Google Antigravity, OpenCode, ZCode, OpenClaw, Reasonix, Kimi Code, DeepSeek Harness.
- InstallOne command:
npx skills add zenstory-ai/video-recap-skills -y -g
Start in 3 steps
- Install
npx skills add zenstory-ai/video-recap-skills -y -g - Entry commandRun
/video-recapin your agent host. - Follow a guidePick one of the 3 guides below for your first task.
Practical guides
Writing craft
Storyboards and AI video
- How to Add a Video Frame and Logo Without Obstruction
- How to Replace an Unintended Black Video Tail
- How to Fix Clicks at Audio Edit Points Without Clipping Speech
Video scripts and editing
- How to Write a Recap Script: Give Images and Narration Jobs
- Writing Clip Titles and On-Screen Text That the Video Supports
- How to Control Editing Rhythm: Cut Waiting, Keep Action and Reaction
- How to Edit Multiple Videos into One Complete Story
- How to Research Before Writing a Video Recap
- How to Lay Out Video Subtitles and Choose Line Breaks: Lock Picture and Sound, Then Check Sample Frames
- SRT Subtitles Out of Sync with AI Narration: Offset, Drift, or One Bad Cue?
- How to Write a YouTube Video Script Viewers Can Follow
- How to Turn a Long Video into Short Clips That Stand Alone
- How to Fix Flash Frames and Stray Shots in an Edit
- How to Fix Music That Overpowers Speech
- How to Keep Approved Audio When Changing Picture
- How to Fix AI Voiceover That Runs Too Long
- How to Keep AI Edit Revisions Local
- How to Dub an English Video into Chinese
- Fix Automatic Subtitle Errors in a Video Recap
- How to Make a Video with Original Audio and No Voiceover
- How to Combine Landscape and Portrait Videos
- How to Fix Dialogue Cut Off by a Video Edit
- How to Fix Uneven Volume Between AI Voiceover Segments
- How to Fix Missing Words in AI Voiceover
- How to Fix Long Pauses Between AI Voiceover Lines
- How to Fix a Video Export with No Sound
- How to Fix Music Swells Between Voiceover Lines
- How to Fix Inconsistent AI Voiceover Voices
- How to Improve Robotic AI Voiceover Delivery
- How to Fix Repeated Lines in AI Voiceover
- How to Fix Pitch Changes After Speeding Audio
- How to Fix Video Audio Playing on One Side
- How to Fix a Blurry Video Export
- How to Fix Audio and Video Out of Sync After Editing
- Why Subtitles Disappear After Video Export and How to Fix It
- How to Fix Distorted Video Audio: Source, Mix and Export
- How to Reduce Video Export File Size Without Losing Key Content
- How to Fix Subtitles That Do Not Match the Voiceover
- How to Fix AI Voiceover That Speaks Too Fast
- When to Reveal a Twist in a Video Recap
- How to Stop AI Video Analysis From Inventing the Story
- How to Edit Video With Inaccurate ASR Timestamps
- How to Organize Video Editing Assets, Templates, and Samples
- How to Keep Subtitle Fonts Consistent Across Computers
- How to Overlay Animated Transparent Graphics on Video
- How to Prepare HDR Video for SDR Editing
- How to Fix Judder in an Exported Video
How it works
- How do I turn a local video into a Chinese recap? Build a scene/dialogue evidence index, choose the audience promise, divide the work between pictures, original audio and narration, then voice and assemble the grounded script. The complete workflow guide explains full/cut choices, stage handoffs and expected files before optional draft export.
- Can I add narration when the source has none? Yes. No narration does not mean no dialogue; preserve speech evidence when people speak. Dialogue-free footage can skip ASR while still using remote VLM analysis. The original evidence-to-script examples distinguish a changed on-screen promise from a proven result, then show what can be said when only the initial attempts are available. When later footage is missing, request that evidence or propose a clearly narrower account; do not invent an ending.
- Which moments should keep their original sound? Let a complete original line, a revealing action sound or a meaningful wait lead. The five-beat radio example compares shortening context with omitting it, then shows captions alongside each sound: retain faithful speech captions, separate optional sound labels, and do not turn “after the move” into “the move caused the fault.” The edit preserves the wait and invents neither poverty nor a completed mix.
- Can I write the recap before choosing the cut? Decide the story and sound roles first. In the orchestrated cut flow, select source clips, then write narration against the actual edited output timeline. Do not copy original timestamps into that output-time narration.
- What if the voiceover is too long? Reduce or move the thought instead of speaking over the next protected line. Bounded speed adjustment is not unlimited; an unsafe fit can block placement rather than clipping the spoken ending.
- Does an audio-owner label automatically preserve the mix? No. The renderer does not parse the creative board. Actual narration timing, speech evidence and mix settings implement the plan; source audio can remain ducked beyond the written block until a safe recovery point. Listen to the rendered handoff.
- Can I finish the sound and captions manually? Export the existing timeline, then choose what the scene needs. The JianYing / CapCut guide compares two original rehearsal edits: move commentary after the result, or remove it for an original-audio-led cut. Move matching captions and inspect the old and new volume keyframes; moving only a voiceover does not move every dependent element.
What makes it different
- ffmpeg handles local media processing, while configured remote providers receive the media or text required for ASR, visual understanding and speech generation; this is not an offline-only pipeline.
- An optional CapCut/JianYing draft export supports manual finishing.
- Automated checks locate candidate problems; real playback remains the final review boundary.
Who it is for
- Recap and commentary channel creators
- Editors who finish in CapCut
- Teams turning long footage into short narrated cuts
Terms it uses
解说视频 剪映草稿 先剪后配
Sources
This page describes the source as read on 2026-09-29; star count as of 2026-09-29. Links point at that version.
- Creative planning and stage handoffs
- Five stages plus the orchestrator
- MiMo and optional Fish Audio dependencies
- Understanding-stage prerequisites and API-key boundaries
- Playback review and advisory automated checks
- Leading sound and narration jobs
- The renderer does not parse audio-owner decisions
- Mixing, bounded fit and safe source-audio restoration
- Source and output clocks in the cut workflow
- Visual understanding can skip transcription, not remote VLM
- Video/audio/subtitle separation and authored volume placement