Lately I have been trying a different way to do motion design. This text is about how to make a brand video with AI agents so that the result is not random. It is not the classic process where you prepare a storyboard, open After Effects and start assembling scenes by hand. And it is not generating video with text-to-video models either.
In my case, every frame is born as code. AI agents in Claude Code (the Claude Opus 5.5 model) propose the concept, build the individual scenes, review them, fix them and finally render the whole film right in the browser with Canvas and WebGL.
I bring the brand assets, the brief, the context and the music, and above all I decide what works, what does not and what needs to be redone.
Two quite different projects have come out of this so far: a short teaser for the new Cleevio website and a one-minute brand film for Atomika. The production principle was the same for both, only the scale was on another level.
What video as code means, and why not a text-to-video generator
Not generated. Programmed.
Every scene is a program, not a model's output. That is why the result is repeatable and can be changed with a parameter.
This is the main difference from everything written about AI video today.
Every scene is essentially a small program that knows exactly, for a given moment in time, where an object should be, how big it should be, how it should move, what the camera should do and what the final frame should look like. The animation is then drawn directly in the browser.
The advantage is that the result is the same every time and yet editable at any moment. When I want to change the timing, make the phone bigger, adjust the camera move or rework a transition between scenes, I do not have to generate the whole video again and hope the model gets it right this time. The parameters change and the scene renders again.
Further variants are just as easy to produce: another format, another timing or completely different music.

| Video as code | Text-to-video generator | |
|---|---|---|
| Control over the output | Every object, camera and moment is a parameter. Change a number and the scene recalculates. | A prompt and chance. A change means a new generation and a new result. |
| Brand consistency | Built from the real design system: fonts, colors, UI, logos. | The model imitates the brand. Fonts and UI tend to be approximate. |
| Editing a scene | Any time, parameter by parameter, without losing the rest. | In practice, only regenerating the whole scene. |
| Cost of an iteration | Seconds of rendering in the browser. | Minutes and credits for every attempt. |
| Working with real assets | Real screenshots, maps, dashboards and product photos. | Assets can only be described approximately. |
| Reproducibility | Same code, same result. A version B with different music in half an hour. | Same prompt, different result. |
Why we start with the music, not the picture
Music first, then the picture
The music is the film's timeline. The beats and the entries decide when a cut, a move or an impact lands.
With these videos I do not start with the picture. I start with the music.
First we analyze the tempo, the beats, the entries, the changes in energy and the places where the main moment should land. Only then is the picture laid onto this timeline.
Every important move gets a specific place in time. A cut, a UI change, a camera move or an object's impact is not roughly "somewhere in the music", it is tied directly to its rhythm.
Cleevio was built at 120 BPM, Atomika at 100 BPM. That is how the film's timing can be worked out precisely before all the individual scenes are finished.

Why the agents need real brand assets
Real assets instead of an invented brand
The second important part is the inputs. The agents do not just get a logo and the sentence "make a video". They get the company's real brand assets: Figma, fonts, colors, logos, product screenshots, maps, UI, presentations and other material.
Part of the preparation is research on the website and other sources. The agents need to find out what the company does, how it communicates, what products it has and what visual language it already uses.
So the final film is not made somewhere outside the brand. On the contrary, it tries to work as much as possible with what already exists and to translate the real design system into motion.

How a multi-agent workflow produces a video
Preparation → Creative → Production → QA → Final. Inside production, the loop Creator → Critic → Refinement.
1. A small studio instead of one AI assistant
The rendering itself is not what I enjoy most about this. What is more interesting is how the work is organized.
I do not use one agent that I tell "make a video" and wait for the result. The whole process is put together more like a small production studio:
Preparation → Creative → Production → QA → Final
- 01 Preparing the assets and the brief
- 02 Creative directors propose, a jury scores, the showrunner sets the direction
- 03 Creator → Critic → Refinement
- 04 QA, sound, final
Preparation covers the assets, Figma, the website, the music and references. Then several agents in the role of creative directors independently propose their own concepts. Other agents act as a jury, score the proposals and look for their weak spots. The showrunner then takes the strongest ideas and assembles the final direction and the storyboard from them.
Only then does production itself begin.
2. The loop: Creator → Critic → Refinement
Even a single scene is not made by one agent alone. I use a simple loop:
Creator → Critic → Refinement
The first agent creates the scene. The second checks it independently against the original brief and looks for problems. The third gets the scene together with the critique and works the notes in.
The scenes can be made in parallel, so instead of one long linear process, several parts of production run at the same time.
3. What the brief looks like
It all starts with plain text. For Atomika I wrote one brief: what I want, which brand assets to start from and how high to aim.
I want a video about atomika.gold: the website, the app and the company as a whole. About a minute, 16:9, with sound.
Everything is in Figma: the logo, colors, graphic style, maps, the website and the app. Go through the website and study it properly. I'm attaching the music, the original video, a LinkedIn article and the fonts Micro Grotesk and Geist Mono. Use everything that will work in the video: SVGs, maps, screens.
Work out how to explain clearly what Atomika actually is. Put the whole company of agents on it, and your best directors.
And act as if this video were going to win an award for best design. I want a wow effect.The orchestrator then broke it down into work for the individual agents. Each one gets a precise instruction for its role only, but with the same context. This is what the brief looked like for one of the five directors who independently proposed a concept for the film:
You are part of an elite motion-design studio making a 62-second, 16:9 brand film for Atomika – the open market for sodium cyanide, the reagent most of the world's gold depends on.
The film must do two things at once: win a design award, and make a first-time viewer understand what Atomika is.
Before you write anything, study the real material: the website, the Figma brand system (palette, typography, logo, 63 topographic maps, icons), the product screens, the music and the old film.
You are CREATIVE DIRECTOR A – BRAND SYSTEM. Turn Atomika's own identity into motion: contour lines drawing themselves and rising into 3D, the 45° diamond from the logo, a crosshair locking onto coordinates, sharp grotesk headlines against mono data labels. The idea: a market that never had a map finally gets surveyed.
Propose one complete concept. No tech-ad clichés – no glowing networks, no stock blockchain imagery. 14–20 shots, a hook in the first two seconds, every big reveal on a beat of the music, and a clear ending on atomika.gold.4. Sound is part of the system
The picture is of course only one part of the film. So the agents work on the sound design too.
Whooshes, clicks, impacts and other effects can be synthesized directly in code and placed on an exact frame. The music and the effects are then assembled into the final track and the mix is prepared for the final format.
This is where automation hits a clear limit. Technically you can measure volume, loudness or the exact placement of an effect, but a human still has to listen to the final sound.
Measurement tells you the mix is technically fine. It does not reliably tell you that a sound feels cheap, too aggressive or simply does not fit the picture.
How it turned out: a teaser for Cleevio and a brand film for Atomika
The same process, 18 seconds for Cleevio and 62 seconds for Atomika. What differs is the number of agents and the volume of brand assets.
1. Cleevio: a teaser for the new website
The first bigger test was a teaser for the new Cleevio website.
The core concept was FULL STOP. The new cleevio.com. I did not want to make a quick montage of website screenshots. I wanted to use elements that are part of its design.
One of the main motifs became the full stop from the website's copy, for example from the claim Built to stand out. In the film it starts to behave almost like an interactive object. It opens further layers and gradually takes us deeper into the website.
I wanted to work with the cursor in a similar way. The one from the hero video gradually comes alive as a designer's hand and starts working with real elements of the website: a MacBook, phones and project cards like Packeta or Spendee.


The showreel was split into 11 scenes and two formats were planned from the start, 16:9 and 9:16. The vertical version is not just a cropped horizontal picture. The scenes got their own compositions and adjustments so that they work on a phone too.
105 AI agents split across several workflows worked on the whole process. The first rough cut was ready in about three hours and the final 16:9 and 9:16 versions were done roughly 6.5 hours after the brief.
Later, another version was made with its own music created in Suno. The picture was retimed to the new track, some scenes were adjusted and the sound effects were mixed again under the music.
- 105
- agents
- 11
- scenes
- 120
- BPM
- 16:9 + 9:16
2. Atomika: a one-minute brand film
Atomika was a considerably bigger test.
The brief was to create a roughly one-minute 16:9 brand film that introduces the company, the product and the principle of the market, and at the same time has a strong visual "wow effect".
There was far more material at the start this time. The Figma held the logo, presentations, maps, the website, dashboards and the auction interface. On top of that came the music, part of an older video, more content and the Micro Grotesk and Geist Mono typefaces.
Out of that came the concept LEVEL. A line of equal value.
The main motif became a horizontal plane, joined by a diamond turned 45°, topographic contour lines and data cards. These elements then repeat in different forms across the whole film and hold the scenes together.
The film has 16 scenes, roughly 62 seconds and 100 BPM.

One of the main moments is a 3D landscape built from contour lines. The individual bids grow out of the surface like towers and the price plane gradually levels them to one height. A specific figure from the auction is projected into the same moment: $2,600 → $3,050/t, that is +17%.


This time the whole production system was considerably bigger. 4 agents worked on preparing the brand assets, 14 on the creative part, 72 on production and another 35 on QA.
That is 125 AI agents in total.
About 8.5 hours passed from the first brief to the finished film and website. My own active time was about 4 hours of that, mainly preparing the brand assets and the brief, checks along the way, creative decisions, reviewing the results and the final publication. In between, the individual agentic workflows ran on their own or in parallel.

- 125
- agents
- 16
- scenes
- 62 s
- 100
- BPM
- ~4 h
- of my time
What the human does in an agentic workflow
Fewer keyframes, more decisions. Taste, the brand and the final selection stay with the human.
1. What I actually do
What surprised me most about the whole experiment is how much my role shifts.
I spend less time moving keyframes by hand or making every single element, and much more time preparing the inputs, the context, the design of the process, choosing the direction and deciding what works and what has to be done again.
In a sense, my role moves from motion designer toward art director of the whole system.
AI can take over a large part of the production work, but it still needs someone who knows the brand, has a visual sense and can say: This works. This does not. And this we do again.
2. Without good inputs it does not work
The whole process can be very fast, but it is not a magic button.
Without a good Figma, the right assets, facts, music and a clear brief, a good result does not just appear.
An agentic workflow will not save a bad brief.
It will only produce it faster.
3. What changed for me
It is not really about the number 105 or 125 agents for me. And not even about the fact that a one-minute brand film can be made within a single working day.
What matters more is the change in the way of working itself.
I used to create a scene. Now I can create a system that proposes the scene, makes it, critiques it and reworks it. The individual agents have their roles, their work builds on each other's and some parts of production can run in parallel.
The human does not disappear from the process. Rather, they can focus more on the things where their decision has the greatest value: direction, idea, taste and the final selection.
Cleevio teaserThe new cleevio.comSee the result →
Atomika brand filmLevel. A line of equal value.See the result →How we did it
A six-step process that can be repeated.
If I had to sum it up for someone who wants to try this in Claude Code with Opus 5.5, the process looks like this:
- Inputs. Figma, fonts, logos, product screenshots, music, copy. Without them the agents make a generic video about nothing.
- Music as the timeline. First the analysis of the beats and the energy, only then the scenes. Every cut gets its moment.
- Concepts and a jury. Several agents propose a direction, others critique it, the showrunner assembles the storyboard.
- Production in a loop. Creator, critic, refinement. A scene goes into the film only after the third step.
- Sound. Effects synthesized in code and placed on the frame, mixed with the music. A human listens.
- Review and render. An internal jury scores, the weakest scenes are redone, the final is rendered in the browser.
On a single MacBook that means working in waves of 10 to 12 agents. A full working day for a one-minute film is a real number, not marketing.
Frequently asked questions
Answers to the questions we get most often.
Can AI agents produce production-ready motion design?
- Yes, if they get real brand assets and someone makes the decisions for them. The teaser for Cleevio and the one-minute film for Atomika were made this way and both are live. The first version is rarely the final one, though: a jury scores the scenes and the weakest are redone.
How does video production with AI agents work?
- Like a small studio. Preparing the brand assets, several creative concepts and a jury, then production scene by scene in the loop creator → critic → refinement, QA and the final render in the browser. Every scene is code, not a generator's output.
Why Claude Code instead of an AI video generator?
- Because the result is deterministic and editable. Changing the timing, the size of the phone or the music is a change of a parameter, not a new generation with an uncertain result. And the film is built from the company's real design system, not from an approximate imitation.
How many agents and how much time did the Atomika film take?
- 125 agents: 4 for preparation, 14 for creative, 72 for production and 35 for QA. About 8.5 hours passed from the brief to the finished film and website, roughly 4 hours of it human work.
What is the designer's role in an agentic workflow?
- More an art director than a keyframe operator. They prepare the inputs and the context, choose the direction, listen to the sound and decide what is good enough and what is done again. AI takes over the production work; taste and knowledge of the brand stay with the human.
What does a film like this cost and how long does it take?
- The teaser for Cleevio was done 6.5 hours after the brief, the one-minute film for Atomika in 8.5 hours. The price depends on the length, the amount of brand assets and whether you already have a design system. The fastest way to a number is a short call where we go through what you want to make.









