Tutorial

Video-to-video: shoot the motion, generate the look.

Glass figure mid-stride overlapped by an identical twin of blue light

Text-to-video invents everything from a sentence, and motion is what it invents worst. Video-to-video splits the job: you shoot a clip that carries the motion, blocking, timing, and camera move, and the model regenerates only the appearance. The plate can be a phone clip of you crossing your own kitchen. What survives the pass is everything a prompt cannot specify: when the head turns, how long the pause lasts, where the frame puts the subject.

The short version
  • Video-to-video keeps a source clip’s structure, its motion, composition, and timing, and replaces its appearance from a prompt or a style reference.
  • Shoot a plate for anything a prompt cannot pin down: performance timing, multi-subject blocking, complex action, continuity across takes. A phone is enough; the restyle discards lens quality anyway.
  • Every tool exposes one core dial between structure and style. Low strength is faithful but timid; high strength is stylish but drifts. The midpoint is where flicker lives.
  • Watch for identity drift, background boiling, and smear on fast action, and check them in playback, not on a still.

What is video-to-video, and how does it differ?

The three generation modes differ by what you hand the model. Text-to-video gets a sentence and invents subject, motion, and look. Image-to-video gets one frame, so composition and look are pinned but every second of motion is invented. Video-to-video gets a full clip: the model reads its structure, the way subjects move through the frame over time, and renders new appearance on top of it. Luma describes its Modify Video feature as transforming footage while preserving the original motion, performance, and camera dynamics.

Open-source pipelines make the structure input explicit: ComfyUI workflows built on Wan’s VACE and Fun Control condition the generation on maps extracted from your plate, depth maps for spatial layout, edge maps for composition, pose skeletons for body motion. Closed tools like Runway’s Aleph and Kling’s video editing take the source clip whole and condition on it internally. Either way the plate is a set of instructions the model cannot ignore, which is exactly what a prompt is not.

Why shoot a plate first?

Because the things a plate carries are the things prompting fails at most reliably.

  • Timing: a prompt cannot specify a half-second hesitation before a door opens. A plate does it by containing one.
  • Blocking: two characters crossing a room, one passing behind the other, is choreography text-to-video routinely scrambles. Shot once, it is fixed geometry.
  • Complex action: pouring, catching, fighting, dancing. The failure rate of prompted action scales with complexity; the failure rate of filmed action is zero.
  • Continuity: the same plate restyled twice gives two takes with identical motion, so a look change never costs you the performance.

This is why the working method for a dialogue or action beat is to act it yourself on a phone. Steady framing, clear staging, the action readable in silhouette: those are the qualities that survive the pass. The plate’s color, grain, and resolution do not survive, which is liberating. Nobody grades a plate that is about to be regenerated.

WHAT EACH INPUT CONTRIBUTES Source plate motion · blocking · timing camera move Prompt or reference look · palette · world character design Restyle pass structure kept, look new Finished shot same take, new skin
The split that makes the workflow work: the plate decides what happens and when, the prompt decides what it looks like. Neither input can do the other’s job.

What tools restyle video in 2026?

Four paths cover the current landscape, from one-slider hosted tools to fully explicit node graphs.

PathHow it conditionsWhere it fits
Runway AlephWhole source clip as context, plus text and image references. Also handles targeted edits: relight, background swaps.Fast hosted restyle and shot surgery in one tool.
Luma Modify VideoSource clip plus a strength preset: adhere, flex, or reimagine.The clearest expression of the structure-style dial; strong on preserving performance.
Kling video editingSource clip plus image references for subject and style swaps.Restyle plus prop and wardrobe changes without leaving one ecosystem.
ComfyUI + Wan (VACE, Fun Control)Explicit control maps: depth, Canny edges, pose skeletons, chainable.Local, free, and fully inspectable; the most control and the most setup.

The hosted tools hide the conditioning; the open path shows it to you. Both obey the same physics, so skills transfer: a shot that restyles badly in ComfyUI for lack of depth separation restyles badly everywhere.

The one dial that matters

Every restyle tool, whatever it calls the control, exposes a single trade between obeying the plate and obeying the prompt. Luma names its presets honestly: adhere, flex, reimagine. In ComfyUI the same axis is control strength and denoise. At the faithful end the output tracks the plate so closely that the new style barely takes; at the inventive end the style lands fully and the structure starts to slide underneath it.

STRUCTURE VERSUS STYLE PLATE WINS (faithful, timid style) PROMPT WINS (full style, sliding structure) Adhere Flex Reimagine midpoint: flicker and boiling live here
The names change per tool; the axis does not. Commit to one end or stabilize the middle with a styled keyframe. The unstable midpoint is where the model re-decides appearance every frame.

At mid strength the model is torn: each frame settles the argument between plate and prompt slightly differently, and the disagreement reads as flicker. The standard rescue is a styled keyframe. Restyle one frame as a still image, where you have precise control and can iterate cheaply, then feed that frame as the style reference for the video pass. The model now has one consistent answer for what the new look is, instead of inventing 24 answers a second. The technique for holding a face steady through this is its own craft; the character consistency guide covers it shot by shot.

How do you keep the subject and background under control?

Full-frame restyle at high strength is where faces break first. When the shot has a hero subject, split the problem. Restyle the background plate separately, or use a tool’s masked edit to hold the subject while the world changes around them, then let a second gentle pass unify the grain. Character identity gets the same treatment as in any generative pipeline: a reference image supplied on every generation, not a description re-rolled each time. And when the plate is your own performance driving a designed character, that is its own discipline, performance transfer, with its own tools and failure modes.

Plate: phone clip, 8s, actor crosses kitchen, picks up kettle, double-take at window. Reference: styled keyframe (frame 1, restyled as a still, approved). Prompt: hand-painted animation, gouache texture, warm interior light, dusk outside the window. Strength: adhere for take 1; raise only if the style refuses to land.

That is a complete setup for a first pass. Everything else is iteration against the checklist below.

Where restyling breaks

  • Identity drift: the character’s face slides over the shot’s duration, or between shots restyled separately. Anchor with reference images and styled keyframes; cut drift-prone shots shorter.
  • Background boiling: flat, low-detail areas, walls, sky, tarmac, crawl with invented texture that never settles. The plate gives the model nothing to anchor on there. Add detail to the plate, or mask and restyle the background once as a still.
  • Motion smear: fast action turns to ghosting, worst on models generating at low native frame rates. Slow the action ten percent when you shoot the plate; restyled motion reads faster than filmed motion.
  • Flicker across cuts: two shots of the same scene restyled in separate passes land on different answers for the same wall. Restyle coverage of a scene in one batch with shared references.
  • The uncanny zone: restyles close to the source’s register, live-action into slightly stylized humans, read as wrong in a way full anime or full photoreal does not. When in doubt, push the style further from the plate, not closer.

Judge all five in playback at full size; most of them never show on a single frame. The same discipline the AI filmmaking guide applies to generation applies here: look at the pixels, in motion, before calling it done.

Questions creatives ask

What is the difference between video-to-video and image-to-video? Image-to-video animates a still: the model receives one frame and invents all the motion. Video-to-video receives a full clip and keeps its motion, blocking, timing, and camera move while regenerating the appearance. One conditions on a moment; the other conditions on a performance.

Can I film myself on a phone and restyle it into animation? Yes, and it is the workflow the tools are built around. Shoot the action on a phone with steady framing and clear staging, then run the clip through a restyle pass with a style prompt or reference image. The timing, blocking, and performance survive; the appearance is replaced. Complex action you could never prompt reliably becomes a matter of acting it once.

Why does my restyled video flicker or boil? Flicker means the model is re-deciding appearance from frame to frame, which shows up worst at mid-range style strength, where the output is torn between source and prompt. Boiling, the constant texture crawl in flat areas, appears where the source gives the model little detail to anchor on. Push strength toward one end of the dial, style one keyframe and feed it as a reference, or restyle subject and background separately.

Do I need a professional camera for the source plate? No. The restyle pass replaces appearance, so lens quality, color, and resolution of the plate matter far less than what it records: motion, timing, and composition. A phone clip with strong staging beats a cinema camera clip with weak staging. Spend the effort on blocking and performance, not on glass.

Learn this beside the people building it.

Membership is free. Masterclasses from industry leaders, hackathons where you finish something the same day, and mentor circles matched to what you want to learn. For engineers and creatives alike, across film, design, image, sound, and story.