> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tryflowy.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Flowy is a node-based AI creative platform: you generate images, video, audio, 3D and vector on an infinite Canvas, refine on the Studio timeline, and export or publish from the same project.
> Prefer the Flowy MCP server (https://mcp.tryflowy.ai/mcp) or the REST API at https://apis.tryflowy.ai/v1 for programmatic work. Install with `flowy mcp install` from the @flowy/cli package.
> Credits are workspace-scoped. Generations reserve credits on start and only deduct on success, so failed runs refund automatically.

# Lipsync Studio

> Make any photo or clip speak a script or an uploaded voice track, with a choice of lipsync models.

Lipsync Studio syncs a character's mouth to speech. Upload a photo or a video, add a script or a voice track, pick a model, and generate. Reach for it when you need a talking-head clip without building a canvas flow.

<Frame>
  <img src="https://mintcdn.com/flamapp/F3HzOSEZ0flnKi-X/images/tools/lipsync-studio.png?fit=max&auto=format&n=F3HzOSEZ0flnKi-X&q=85&s=0f64fe3df64d42903673ae877cfd6f5f" alt="Lipsync studio" width="2880" height="1800" data-path="images/tools/lipsync-studio.png" />
</Frame>

## Where to find it

Open **/video/lipsync-studio** directly, press <kbd>⌘</kbd> <kbd>K</kbd> and search "Lipsync Studio", or find it in the **Video** area of the left sidebar.

## What you need

The panel adapts to the model you pick: some models start from a photo, others re-sync an existing clip.

| Input               | Required     | Notes                                                                                                                                |
| ------------------- | ------------ | ------------------------------------------------------------------------------------------------------------------------------------ |
| Character           | Yes          | A photo or video of the person who should speak. Whether it accepts a photo, a video, or either depends on the model.                |
| Speech text / Audio | At least one | Type a line for Flowy to speak, or upload a voice track. Which of the two is available, or required, depends on the model.           |
| Prompt              | —            | Optional scene description.                                                                                                          |
| Model               | —            | 11 models. See below. Default: **Wan 2.5 Speak Fast**.                                                                               |
| Duration            | —            | `5s` or `10s`. Default `5s`. Ignored by models that bill on the uploaded clip's or audio's own length instead. See the Models table. |

## Models

| Model                           | Provider     | Credits                    | Notes                                                                                                                                                                                                                              |
| ------------------------------- | ------------ | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Kling 2.6 Lipsync: **new**      | Kling        | 210 / s                    | Most advanced Kling lipsync model. Photo + text only. No audio upload at all. 1080p, up to 10s. Photo needs 300px+ per side (aspect 0.4–2.5), under 10 MB; script capped at 2,500 characters.                                      |
| Google Veo 3: *premium*         | Google       | 520 / s at 720p            | High-quality cinematic generation. Photo + text only. 720p or 1080p (1080p bills 2x), 4/6/8s duration, photo under 8 MB, script capped at 20,000 characters.                                                                       |
| Google Veo 3 Fast: *premium*    | Google       | 195 / s at 720p            | Faster generation, slightly lower quality. Same inputs and caps as Veo 3.                                                                                                                                                          |
| Wan 2.5 Speak                   | Wan          | 150 / s at 720p (estimate) | Next-gen generation with sound. Accepts typed text or an uploaded voice track. 480p/720p/1080p (480p bills half, 1080p bills 2x). Photo 240–8000px/side, under 25 MB; audio 3–30s, under 15 MB; script capped at 1,500 characters. |
| Wan 2.5 Speak Fast: **default** | Wan          | 75 / s (estimate)          | Faster and cheaper: renders at a fixed 480p. Same photo/audio/script limits as Wan 2.5 Speak.                                                                                                                                      |
| Kling Avatars 2.0: *premium*    | Kling        | 173 / s at Pro             | Next-gen talking avatars: requires an uploaded audio track, no typed text. Pro/Standard quality tiers; Standard bills half of Pro. Photo 300px+ per side (aspect 0.4–2.5), under 50 MB; audio 2–60s, under 5 MB.                   |
| InfiniteTalk: *premium*         | InfiniteTalk | 300 / s at 480p (estimate) | Realistic talking avatars: requires an uploaded audio track. 480p or 720p only (720p bills 2x), no 1080p tier. Up to 15s.                                                                                                          |
| Kling LipSync: *premium*        | Kling        | 21 / s (estimate)          | Fast, expressive lip sync: re-syncs an existing video to an uploaded audio track, billed on the uploaded clip's length. Source video 2–10s, 720p–1080p, under 100 MB; audio 2–60s, under 5 MB.                                     |
| Sync Lipsync 3: *premium*       | Sync         | 200 / s                    | Precise, professional lip sync: video-to-video, up to 4K, 15s. Adds three quality controls: a Temperature slider (expressiveness), Lipsync only active speaker, and Occlusion detection (keeps a covered mouth untouched).         |
| VEED Lipsync 2: **new**         | VEED         | 105 / s                    | Production-quality lip sync: video-to-video. Source clip capped at 10 minutes.                                                                                                                                                     |
| LatentSync                      | LatentSync   | 30 / s (estimate)          | Fast, open-source lip sync: video-to-video, flat-rate clip.                                                                                                                                                                        |

## What it costs

Billed per second of the final video, at the selected model's own rate (21–520 credits/second; see the table above). A model with a resolution ladder bills higher tiers at up to 2x and lower tiers at half. Flowy shows a live estimate in the panel before you run. The backend's pricing is always authoritative.

## Tips

* Match your upload to the model: Kling 2.6 Lipsync and both Veo 3 models only ever read your typed script. They have no audio input. Wan 2.5 Speak takes either text or audio. Kling Avatars 2.0 and InfiniteTalk require an uploaded audio track. Kling LipSync, Sync Lipsync 3, VEED Lipsync 2, and LatentSync re-sync an existing video's lips to an uploaded audio track instead of starting from a photo.
* Switching models changes which inputs appear: the panel adapts around your choice.

## Related

* [Voiceover](/tools/voiceover)
* [Motion Control](/tools/motion-control)
* [Change Voice](/tools/change-voice)
