Skip to main content
The Video node generates AI video clips from a prompt — on its own, or conditioned by the images, videos, and audio you connect. Two engines cover different jobs: Default for one coherent subject, Compose for combining several distinct elements or authoring multi-shot clips. Pick the engine by intent, not by how many inputs you have. Default keeps a single subject coherent; Compose gives you per-element control when a scene combines multiple products, characters, or views.
What you connect picks the route:
  • Nothing — text-to-video from the prompt alone.
  • Images — condition the clip as references or frames (see below). Up to 9.
  • Videos — act as references that condition the new clip’s look and motion. Up to 3. A connected video is not extended frame-by-frame — it steers the generation.
  • Audio — conditions the soundtrack and timing. Up to 3, and it needs at least one image or video alongside it.

Reference vs. frames

The image role setting decides how connected images condition the clip:
  • Reference (the default) — images are style and subject references. The clip does not start on them; instead they hold a consistent look and identity across distinct shots. This is the right choice for continuity across a sequence.
  • Frames — the clip literally starts on your first image, and optionally ends on a second. Use it to animate a specific still or to continue from a previous shot’s end frame. Connecting the same image to many shots in frames mode makes them all open identically.

Name your references in the prompt

In reference mode, a connected asset is only used when your prompt names it: @Image1, @Video1, @Audio1 — numbered by connection order. Reference every asset you want the engine to use:
  • Style transfer: “in the style of @Image1”
  • Character consistency: “@Image1 as the same subject across scenes”
  • Motion transfer: “@Image1’s subject performing @Video1’s motion”
  • Scored speech: “@Image1 speaker delivering @Audio1”

Pace multi-scene clips

One Default clip can carry several scenes. Write “Cut scene to …” between beats and “At N seconds …” to time actions — the engine honors scene cuts and timestamps, so a single generation can replace a chain of one-shot clips.
Video is the most credit-intensive node type, and longer, higher-resolution clips cost more. Failed generations are refunded automatically — see how credits work.
Once your clips are generated, wire them onward: into a Composition for editing, a Flam node as a reference, or straight into Studio for the final cut.