Article cover image

DaVinci Resolve 21.1 Gives AI Editing Agents Better Tools

Brent Schooley
Brent Schooley@chefbrent

A video editing agent needs to take what I'm trying to do and connect it to footage, make changes in the timeline, and examine the resulting video. Each of those steps needs information about the sources, precise controls, and feedback about the status of the edit.

DaVinci Resolve 21.1 improves several parts of that process. For an agent connected through MCP, the scripting updates open up practical workflows around interviews, podcasts, explainers, and other dialogue-heavy video.

MCP Server

MCP (Model Context Protocol) provides a standard way for an AI assistant to discover and call tools. Resolve’s MCP server makes its scripting capabilities accessible through a first-party MCP server that ships with Resolve.

You can enable it from the File menu:

An agent can check whether Resolve is running, launch it, search the scripting API, retrieve current documentation, and execute Python scripts against the application. The server also exposes tools for working with LUTs and DCTLs.

The combination of documentation and execution is super useful. The most important thing to me about this is that it will stay updated with each Resolve release. We no longer have to worry about third-party MCP servers or our own custom CLIs (I'll happily delete mine...) updating with each release.

Readable transcripts

One of the most important changes in the 21.1 API for agents is `MediaPoolItem.GetTranscription()`.

This exposes transcript data from the project with timed words, segments, and speaker labels when detected. That gives an agent a way to connect a request such as “make a 60-second explainer video about this product” to specific sections in the source material without requiring external transcription.

Combined with the existing timeline-assembly APIs, the agent can select passages and build a rough cut around them. Transcription has existed in Resolve before 21.1 but there was no programmatic access to it

For dialogue-heavy editing, this feature is huge.

Multicam support

The multicam additions may have an even bigger impact for podcast and interview production. It will also help for our OpenAI launch videos since we often use 2 and 3 camera setups.

Scripts can now create multicam clips, align footage, run SmartSwitch, and flatten the result. SmartSwitch exposes controls including minimum shot duration, switching delay, and wide-shot frequency.

This means I can prompt things like: “Keep the cuts relaxed and use occasional wide shots” and this is something the agent can do with much more ease than when it had to make these decisions entirely on its own.

I'll have to test this out but it should result in more repeatable workflows for multicam shoots.

Easier workflows

Audio cleanup is another area with big improvements. The new and expanded APIs cover normalization, channel mapping, native clip audio properties, and enabled states for controls such as voice isolation and dialogue leveling.

Things this helps an agent deal with include: inconsistent speaker levels, incorrect channel routing, or dialogue that needs adjustment. Having access to audio properties helps it confirm it has applied settings correctly.

Transitions, video and audio fades, clip-speed controls, and output blanking also become directly accessible through the new APIs. These operations support ordinary requests such as “add short fades, slow this B-roll clip, and dissolve into the closing shot.” Before this update, these tasks required Computer Use and sometimes these changes were difficult for an agent to apply due to tricky drag-and-drop targets.

What I would try first

I would start with a a cutdown of some sort like an interview, podcast, or a livestream. These types of videos are the best way to learn how agentic editing works in Resolve. Start by asking questions about the footage and make simple edits first before expanding into more complex edits.

This is a huge boost to agentic video workflows and I can't wait to dive in even deeper. Let me know what you're working on with this and how I can help!