Work with agents
Mixture is designed and built to work visually, with a powerful web-based canvas, and programmatically, via AI and agents: hand a project to a coding agent (like Claude Code or Codex) and watch it build and edit the canvas in real time, right alongside you.
Connect your agent
Section titled “Connect your agent”Mixture connects to your agent over MCP (the Model Context Protocol) — a one-time setup, not a per-project step. On your projects page, the sidebar carries a Connect your agent card (inside a project it’s behind the agent circle, top right) — pick your client and copy what it gives you:
claude mcp add --scope user --transport http mixture https://mixture.dev/mcpcodex mcp add mixture --url https://mixture.dev/mcpAfter you approve access, Codex shows a plain “Authentication complete”
page on a 127.0.0.1 address. That’s success — close the tab.
gemini mcp add --scope user --transport http mixture https://mixture.dev/mcpAdd to ~/.cursor/mcp.json:
{ "mcpServers": { "mixture": { "url": "https://mixture.dev/mcp" } } }Add to ~/.config/opencode/opencode.json — the global config, not a
repo-local opencode.json:
{ "mcp": { "mixture": { "type": "remote", "url": "https://mixture.dev/mcp" } } }--scope user matters on the CLI clients: without it the server is registered
against the directory you ran the command in and vanishes the moment you cd
elsewhere. Cursor’s and OpenCode’s config paths above are their global ones for
the same reason.
Run the command once in your terminal. The first time your agent connects it signs in as you through Mixture’s own sign-in — no tokens to paste, no keys to manage — and from then on it can see and work on all your projects: it can list them, open an existing one by name, or create a fresh project itself. That means you can start from either end — open a canvas and ask your agent to work on it, or just say “make me a launch video” in your terminal and let the agent create the project for you. Open a project in the editor and, as soon as the agent is working in it, a green indicator appears in the top-right showing which agent is connected and what it’s doing; its edits appear on your canvas as it makes them.
What an agent can do
Section titled “What an agent can do”Start with the short version: anything you can do in the editor, it can do — because it isn’t driving a simplified API, it’s editing the same document your panels edit. There is no agent-only subset and no human-only corner.
That means the fiddly things as much as the big ones. Font size, weight and letter-spacing. Corner radius, stroke width, blend mode, opacity, masks. The easing curve on a single keyframe. Where a punch-in starts and how hard it ramps. Which filter sits on which layer, over which range. It reads the current values before it changes them, so “a bit tighter” and “make that 30% slower” are instructions it can actually carry out.
Area by area, and each one goes deep enough to have its own page:
- Layers, type and layout — shapes, text, images, video, groups, auto-layout that reflows, and the type and colour judgement to make it look composed rather than assembled.
- Motion — entrances, choreography, easing, keyframes and motion paths, including morphing one shape into another and text that animates per word or letter.
- Video and audio — cuts, speed ramps, zooms, transcripts, ducking and beat-synced edits, plus captions with word-level timing.
- Effects and 3D — filter stacks and custom WGSL it writes itself, and 3D scenes, device mockups and imported models.
- Data — layers driven by a CSV or JSON: charts, leaderboards, countdowns, audio-reactive graphics.
- Formats — one master, many cuts, reframed properly for each aspect rather than centre-cropped.
Then the mechanics that let it work at all:
- List, open, and create projects — find the right project by name, or start a new one and hand you the link.
- Read the document — see every scene, layer, and animation.
- Edit the document — add, change, and remove layers and animations, either as full edits or as typed operations.
- Upload assets — bring in images, video, and audio. Including from your own machine: the agent runs in your terminal, so “the takes are in ~/Recordings/launch” is a thing it can act on, uploading what it needs straight into the project.
- Use your library — read and place your saved components, and save new ones, at project or workspace scope. It can read each one’s exact parameters, so it fills them in rather than guessing.
- Follow your recipes — list what’s available, read one’s instructions, assets and worked example, and build the next video to that method. It can write recipes too.
- Work across workspaces — see which ones you belong to, create a project in the right one, or move one between them.
- Check its own work — render real frames from the composition and look at them, which is how it catches a layout that reads fine in the document and badly on screen.
- Roll back — list the project’s version history and restore an earlier snapshot if an edit goes wrong.
- Export the finished video — render the composition to an MP4 and hand you the file (see below).
- Show presence and post notes — appear as a collaborator and leave a short activity trail of what it changed.
Give it something to work from
Section titled “Give it something to work from”The single biggest lever on how good the result is: what the agent has to build with. Starting cold, it has to invent the footage, the type and the palette, and invented brand material looks like it.
So before the ask, or as part of it:
- Load the project (or the workspace Library) with your real material — footage, logo, typefaces, music. Anything at workspace scope is visible to every project and every agent working in one, which is what makes a series look like it came from the same studio.
- Point it at your disk for the raw takes rather than uploading them by hand first.
- Hand it the real source — the transcript, the CSV, the changelog, the script — instead of a description of them. Data can drive the graphics directly; a summary can’t.
- Keep the method once one lands well, as a recipe, so the context travels with the next ask.
And when you don’t have the shape yet, don’t make it guess: ask it to plan first. “Let’s plan this before you build.” It writes the script first — every word you’ll hear or read, beat by beat — and asks you to agree it (it will press for a real yes on the hook and the call to action, not a shrug). Then it drafts the story as a storyboard on the canvas — beat by beat, what’s on screen, what the viewer should feel, and the look it intends — and waits for you to redirect it before anything real is built. It offers this on its own for a thin brief, anything longer than about fifteen seconds, or when there’s no recipe to follow; you can ask for it any time. A wrong story is cheap to fix in a storyboard and expensive to fix in a finished composition.
Beyond the canvas: the agent brings its own toolbox
Section titled “Beyond the canvas: the agent brings its own toolbox”Your agent isn’t only a Mixture client. It’s already connected to your files, your keys and whatever other MCP servers you use — so the material a piece needs can come from anywhere it can reach: a voice-over from ElevenLabs, b-roll from Runway or Veo, a screen recording from Supercut, numbers from your analytics API, the real typeface from a brand’s own site.
Mixture needs no integration with any of them. One instruction can span three services and land in your project as ordinary assets:
“Grab my latest Supercut recording, cut the dead air, generate a voice-over, caption it from that script, and give me a 30-second vertical cut.”
What it can make itself differs by runtime — Codex has image generation, Claude Code doesn’t — so it’s told to check and say so in its first message (“I can make imagery and a voiceover here; b-roll would need to come from you”). Mixture generates images first-party regardless.
The full picture: what comes from where, and how to extend it →
Agents understand the data system
Section titled “Agents understand the data system”The documentation an agent fetches covers Mixture’s data-driven features in full — repeating layers over rows, binding properties to formulas, datasets, and transcript captions — so you can ask for data-driven work in plain language:
- “Turn this CSV into an animated leaderboard bar chart.”
- “Add captions that follow my cuts.”
- “Repeat this card over the top five rows, sorted by revenue.”
The agent authors the repeat, the bindings, and the formulas; you review the result on the canvas and adjust anything in the same panels you’d use by hand.
Built-in creative skills
Section titled “Built-in creative skills”Beyond the raw API, a connected agent loads specialised skills — creative playbooks for the jobs that come up again and again. You don’t invoke them by name; ask for the outcome and the agent picks the right one.
Some are workflows of the agent’s own:
- Subject cutouts and depth — extract the person from camera footage (using Robust Video Matting) and put them over a new background, or sandwich titles and shapes between the background and the subject for a real sense of depth. “Cut me out and put the title behind me.”
- Screen demos — the polished product-demo look for screen recordings: squircle screen framing with the right corner radius, dark backdrop and scrim, a camera bubble, punch-ins — and a cursor driven from a recorded mouse track. “Make this recording look like a launch video.”
- Website promos — the “scrolling website” video: your site full-bleed, scrolled through with rests, framed as a launch promo.
- Imagery — bespoke generated backdrops, product scenes, and textures, added straight to your asset library.
- Plan mode — on a thin brief, or anything longer than a few beats, the agent offers to plan first: it writes the script and gets you to agree the words, then storyboards the piece on the canvas as a draft composition of beat cards (what the viewer sees, hears and should feel in each beat, with mock frames for the key ones if you ask) and the look it intends, for you to redirect before anything real is built. “Let’s plan this first.”
- Design craft and self-review — a design-foundation pass that derives a consistent style ledger (palette, type pairing, spacing, easing) before composing, typography rules for pairing and hierarchy, and a frame-check pass where the agent renders real frames and critiques its own work before calling it done.
Others teach it to drive the same features you use in the editor — so you can ask in plain language and still adjust the result by hand in the same panels:
- Captions — built the data-driven way: one text layer repeated over your transcript’s words or phrases, with content and timing bound per row — so if you re-transcribe or recut, the captions follow automatically. Karaoke highlights, kinetic word reveals, and lower thirds included.
- Beat-synced motion — animations anchored to the beats Mixture detects in your audio: pulse on the beat, switch shots on every fourth, build to the drop.
- Audio mix — how the piece sounds: ducking music under narration so the bed moves instead of sitting at one compromised level, and cleaning up a recording with a high-pass and a compressor. “The music is too loud” and “I can’t hear them” are the asks it exists for.
For the full picture of what agents can do area by area, see What you can do with agents — and for walkthroughs of the common asks — text behind you, captions, voice-over, a screen-recording demo — see the Examples in the sidebar. (The agent itself fetches its own documentation over its MCP connection — there’s nothing for you to teach it.)
Finishing: review in the canvas, or let the agent export
Section titled “Finishing: review in the canvas, or let the agent export”When a build wraps up, there are two ways to get the final video — and a well-briefed agent will ask which you want (or you can settle it upfront in your first message):
- You export. The agent tells you it’s done; you review the result live on your canvas and render it yourself with the Export button — useful when you want a last look and a manual tweak before shipping.
- The agent exports. Ask for the file — “export the video and put the MP4 in this folder” — and the agent renders it headlessly using the same engine and encoder the editor’s Export button uses, so the output is pixel-identical to an in-editor export. Nothing extra to set up on your side: the same connection that lets it edit also lets it render.
Either way the export runs against the live document, so whatever you’re seeing on the canvas is exactly what lands in the file.
The Mixture skill
Section titled “The Mixture skill”Once connected, an agent fetches the Mixture skill — a single reference that teaches it the canvas API and conventions — along with topic guides and worked examples, all served over the same MCP connection. There’s nothing to install and nothing for you to teach it: the agent pulls its own documentation the first time it connects, so it knows how to read and edit safely from the first request.