Claude Code can make videos, and the honest answer is that it does it by writing code, not by imagining pictures. An AI coding agent already has the two things video production needs: it can write a script, and it can run software on your machine. Point it at a rendering framework and it will compose, narrate and render a finished MP4. What the tutorials leave out is the license, the cost, the render time, and the reason the audio never quite lines up.
Can Claude Code actually make videos?
Yes, in three different ways, and choosing the wrong one is the most expensive mistake on this page. The three paths look similar in a demo and behave nothing alike in production. One writes code you own. One buys generated footage by the clip. One rents a hosted service that does the whole job.
| Path | What you write | What comes out | Re-editable? | Marginal cost |
|---|---|---|---|---|
| Code-rendered | A prompt; the agent writes React or Python | Deterministic motion graphics, typography, charts | Yes, every frame is code | Your electricity |
| Generated clip | A prompt to a video model | Photoreal or stylised footage, a few seconds at a time | No, regenerate and hope | Per clip, billed by the model |
| Video agent | A brief to a hosted service | A finished video in the vendor's house style | Inside their editor only | Monthly subscription |
The code-rendered path is the one people mean when they say Claude Code makes videos. It is also the one worth understanding, because it is the only one where the output is a file you own outright, reproducible from source, and free of a per-video charge. A third-party Remotion skill puts the distinction bluntly: generative models "generate pixels and hallucinate", while code produces "deterministic, editable, brand-controlled motion graphics".
Code-rendered video is very good at anything with structure: explainers, data, product walkthroughs, anything that repeats with different numbers. It cannot invent footage of a real place or a real person. If you need a drone shot over a city, no amount of prompting a code framework will produce one, and that is a category limit rather than a skill gap.
How does an AI agent turn a script into a finished video?
Through six stages, and the order of the first two is the single most important thing on this page. Narration is generated before the visuals are finalized, so the pictures can be anchored to the real word-level timestamps of the real audio. Two unrelated open-source projects arrived at the same rule independently. One states it as "audio is generated before the storyboard is finalized, so visuals synchronize to word-level timestamps of the real narration". The other puts it as "generate audio first, then anchor visuals to actual timestamps rather than estimating durations upfront."
- 1Research and script. The agent gathers sources and drafts to a word budget. A rough guide used in practice: target seconds multiplied by 2.5 gives the word count, so a 15-second scene is about 34 words once a pause is subtracted.
- 2Narration. The script becomes audio, either locally or through a cloud voice. This happens second, not last, because everything downstream is timed against it.
- 3Forced alignment. A separate model reads the audio against the known transcript and returns the timestamp of every word. WhisperX and torchaudio's MMS_FA are the two common choices.
- 4Composition. The agent writes the actual scenes as code, cueing each reveal to a word timestamp rather than to a guess.
- 5Render. A headless browser or a Python renderer draws every frame, and ffmpeg encodes them into a file.
- 6Quality gates. The good pipelines check the output rather than trusting it: dead-air detection, minimum motion per scene, loudness, caption accuracy.
Stage six is what separates a pipeline from a demo. One project runs dead-air detection with ffmpeg's freezedetect on the finished video, sampled at 2 to 4 frames per second, precisely because an agent that has been told a scene is fine will report that the scene is fine.
Is Remotion free with Claude Code?
It is free for an individual and for a for-profit organization of up to three employees, including commercial use. At four people it is not, and the price jumps a long way. This is the question almost nobody answers, and it is the one that turns a free workflow into a recurring bill.
Remotion's own license permits eligible users to "use the software non-commercially or commercially for the purpose of creating videos and images". Above the free tier, its published pricing lists Remotion for Creators at $25 per seat per month and Remotion for Automators at $0.01 per render with a $100 per month minimum. Enterprise starts at $500 a month.
Headcount is not just your payroll. Remotion's license FAQ states that if you bring in "another studio, freelancers, or a consulting agency to help operate the Remotion Software on the same project, their headcount aggregates with yours for the 4-person threshold." A three-person company with two contractors on the project is a five-person team for licensing purposes.
There is one more thing worth knowing before building a business on it: Remotion's license file carries the note that "In Remotion 5.0, the license will slightly change." Check the current terms at the source rather than trusting any article, including this one, once version 5 lands.
How much does it cost to make an AI video this way?
For one person, the software cost can genuinely be zero. The floor rises in steps as soon as a fourth person joins, a cloud voice is used commercially, or a hosted service replaces the local pipeline. Those four floors, all taken from published prices, are what the chart below compares.
Read left to right, those bars are $0 for an individual or a team of three using the free Remotion license with a local voice; $6 a month if a commercially licensed cloud voice is added, which is ElevenLabs' cheapest paid tier; $18 a month for a hosted video agent, taking Synthesia's Starter plan billed yearly with its 120 minutes a year; and $100 a month as the minimum on Remotion for Automators once the team passes three people and the rendering is automated. None of these include the coding agent's own subscription, or the electricity to render.
Worth holding those numbers against the alternative. The US Bureau of Labor Statistics puts the median annual wage for a film and video editor at $75,420 as of May 2025. A pipeline that costs nothing to run does not replace an editor's judgement, but it does change which videos are worth making at all: the ones that were never going to justify a day of someone's time.
How do you add a voice-over to an AI video?
Either locally with an open-weight model, or through a cloud API with your own key. The choice is mostly a licensing decision rather than a quality one. Remotion's own agent documentation does not cover audio at all, and the most popular third-party video skill explicitly excludes voiceover generation, which is why this stage is where most people stall.
The local option that matters is Kokoro-82M, an open-weight model published under Apache 2.0 with 54 voices across 8 languages, small enough to run on a processor rather than a graphics card. One project measured it at roughly five times faster than real time on a laptop and concluded it was "not a bottleneck".
ElevenLabs is the default cloud recommendation, and its free plan explicitly does not include a commercial license. For a monetized video, the cheapest legal option there is the Starter plan at $6 a month. An Apache 2.0 model carries no such condition, which is the whole reason license-aware pipelines default to a local voice.
What goes wrong when an AI agent makes a video?
Three things, reliably, and none of them appear in the tutorials. The narration drifts out of sync with the visuals, scenes end up holding still on screen with nothing happening, and the render takes far longer than anyone expects. All three are documented by the people who build these pipelines.
Timing drift is the worst of them. Text-to-speech engines do not deliver a predictable words-per-minute rate, so a script written for 50 seconds can come back as 40 to 45 seconds of audio. The error is proportionally worse on short scenes: one toolkit's own documentation reports roughly 30% drift on a five-second scene and around 10% on a thirty-second scene. That is why the audio is generated first and the visuals are cut to it, rather than the other way around.
Dead air is the second one. One project measured 14 seconds of dead air in a generated deck and fixed it with a constant ambient motion layer, at roughly twice the render cost. An agent will happily produce a scene where nothing moves, because nothing in the code says a scene has to move.
Render time is the third. A code-rendered video is drawn one frame at a time by a headless browser: a 30-second clip at 30 frames per second means 900 screenshots. A production user filed an issue reporting that rendering animated JavaScript across a whole video took "around ~2 to 4 x the length of the input video" on an 8-core cloud machine. Remotion's own performance guidance names the expensive culprits directly, including filter: blur(), box-shadow and gradient backgrounds, and suggests replacing them with a precomputed image where possible.
Will YouTube demonetize videos made this way?
Not for being AI-made, but the policy targets the exact shape a careless pipeline produces. The stakes are worth stating plainly, because Pew Research Center found that 84% of US adults say they ever use YouTube, which makes its rules the default rules for anyone publishing video at all. YouTube's monetization policy, updated in July 2025, was renamed from "repetitious content" to "inauthentic content" and now names as not allowed "AI-generated content made with generic or unoriginal templates giving the impression of mass production without adding the creator's original, authentic insights or perspective." The governing principle it states is that "the substance of each video should be materially varied."
The most striking confirmation comes from a builder of one of these pipelines, who lists as a top risk in his own project documentation that "one template + one voice + one look is the exact signature YouTube's 2025 inauthentic-content policy targets." His mitigation is not a technical trick. It is low volume, a human in the loop, and genuine variation per video.
How do you install Claude Code skills for video?
Through the agent's own skill or plugin system, in one command. Remotion publishes official agent skills, installed with npx skills add remotion-dev/skills, and a Claude Code plugin installed with claude plugin marketplace add remotion-dev/claude-code-plugin followed by claude plugin install remotion@remotion. Restart the agent afterwards.
One caution worth repeating, because it wastes an afternoon: install commands for this stack have appeared in several different shapes across write-ups published within months of each other. Take the command from the vendor's current documentation rather than from an article. That advice applies to this page too.
The skills are not Claude Code only, despite most coverage. Remotion's documentation lists Claude Code, Codex, Kimi Code and Cursor. A genuinely cross-agent treatment is rare, which says more about how the coverage was written than about the tooling.
What is this approach not good for?
Editing footage you already shot, and anything that needs to look photographed. Code-rendered video composes; it does not film. One practitioner who published a full workflow describes editing existing footage this way as "decent but far from perfect", with rough transitions. Natural-language editing of raw footage is a different product category, not a prompt away.
The strength of a code-rendered pipeline is repetition: the same format, run again with different content, looking identical every time. For a single hero video where every second is bespoke, the iteration cost of re-rendering outweighs the benefit, and a human editor in a timeline is still the faster tool.
AI Video & Voice-Over Skill Pack
Five agent skills that carry a topic through research, script, voice-over, animation and the upload package, running on your own machine. Local Apache 2.0 voice, no API key, no subscription, no watermark. One-time purchase.
See the skill packFrequently asked questions
npx skills add remotion-dev/skills. Its Claude Code plugin installs with claude plugin marketplace add remotion-dev/claude-code-plugin then claude plugin install remotion@remotion, followed by a restart.Sources
- Pew Research Center, Americans' Social Media Use 2025 (20 November 2025, n = 5,022 US adults): 84% of US adults ever use YouTube.
- US Bureau of Labor Statistics, Occupational Outlook Handbook (page modified 27 August 2026): median annual wage for film and video editors, $75,420, May 2025.
- Remotion license and Remotion license FAQ: the three-employee free tier, the contractor aggregation clause, and the Remotion 5.0 notice.
- Remotion pricing: Creators $25 per seat per month, Automators $0.01 per render with a $100 monthly minimum, Enterprise from $500 a month.
- Remotion issue #4783: a production user's reported render time of roughly two to four times the input video length.
- Remotion performance documentation: blur, shadows and gradients named as GPU-expensive.
- Remotion Agent Skills and the Claude Code plugin: install commands and the supported agent list.
- Kokoro-82M model card: Apache 2.0 license, 54 voices, 8 languages, 82 million parameters.
- ElevenLabs pricing and commercial license policy: the free plan carries no commercial license; Starter is $6 a month.
- Synthesia pricing: Starter at $18 a month billed yearly, 120 minutes a year.
- YouTube channel monetization policies (updated 15 July 2025): the inauthentic content rule quoted above.
- nemock/video-explainer-system and digitalsamba/claude-code-video-toolkit: the audio-first rule, timing drift figures, dead-air measurement and the template-signature risk note.
- haidrrrry/claude-remotion-skill: the deterministic-versus-generative framing quoted above.
Related reading from Best Answer Hub: AI agents vs automation, which AI tools are worth paying for, and the free Best Answer Hub tools. The finished pipeline lives at AI Agent Skills.