STUDIO · TALKING HEADAFFOGATO STUDIOS

Give your script a face and a voice.

The Talking Head Studio turns a written script into a presenter episode. Choose a host from your Characters, a set from the catalogue or your own photo, and a voice from the library; the studio voices each paragraph, animates the host speaking it, and assembles the segments into one exported video — widescreen for YouTube or vertical for Shorts and Reels.

SCRIPT → EPISODE · UP TO 10 MIN · 16:9 OR 9:16 · TABLE READ FIRST

MADE IN TALKING HEAD STUDIO
01EXAMPLES

Any host, any set.

02WHAT YOU GET

Built for the job, not a generic prompt box.

Script in, episode out

Paragraphs become segments (up to ~45 s each) that are voiced, animated and stitched into one episode of up to ten minutes.

Your host, your set

Any saved Character presents; ten curated sets — News desk, Podcast corner, Classroom, Creator desk, Bookshelf office, Kitchen counter, City balcony, Clean studio, Midnight neon, Cozy living room — or your own set photo.

Any voice — or yours

The full voice library including custom voices, or upload your own audio clips (≤60 s each) and the host speaks in your voice.

Table read first

Voice the whole episode before spending video credits — hear the pacing, fix lines, then render.

Two engines, prices on the label

Kling AI Avatar v2 or ByteDance OmniHuman 1.5, each shown with its per-minute cost so you choose the trade-off.

Edit one line, re-render one segment

Finished audio and keyframes are cached, so a script change re-renders only the segment it touched.

03HOW IT WORKS

How to make a talking-head video from a script with AI, in three steps.

  1. 01

    Write the script and cast it

    Paste the script (or upload your own audio clips), pick a host from your Characters, a set, and a voice. Choose 16:9 or 9:16 and a framing — close-up, waist-up or wide.

  2. 02

    Table read

    Voice the episode for a few credits and listen back. Tighten lines until it reads right; nothing has been rendered yet.

  3. 03

    Generate the episode

    Pick an engine, render, and the segments are assembled and exported at 720p or 1080p to your assets — ready to post or to cut further in the Video Editor.

LENGTH UP TO 570 S (~10 MIN)FORMAT 1920×1080 · 1080×1920FRAMING CLOSE-UP · WAIST-UP · WIDESETS 10 CURATED + YOUR OWNENGINES KLING AVATAR V2 · OMNIHUMAN 1.5
04EXPLAINED

What is an AI talking-head video, and how is one made?

An AI talking-head video is a presenter speaking to camera where the presenter, the voice or both are generated. The current best approach is a pipeline: a keyframe of the host on a set, text-to-speech for each segment of the script, an avatar engine that animates the host speaking that audio with matching lips and gestures, and an assembly step that joins the segments into one video. The Talking Head Studio runs that whole pipeline from a single screen.

You bring a script and a host — a Character you created, which can be your own face. The set comes from ten curated environments or your own photo; the voice from a 58-voice library, a custom voice you designed, or your own recordings. A table read voices the entire episode for a few credits before any video is rendered, which is where most scripts get fixed. Then you choose an engine (Kling AI Avatar v2 or OmniHuman 1.5, each labelled with its per-minute cost) and render.

Episodes run up to ten minutes in widescreen or vertical, and the result is a normal video in your assets — post it, or open it in the Video Editor for captions, B-roll and a soundtrack.

05USE CASES

What people make with it.

Daily news-desk show

A recurring host on the News desk set reading the day's script.

Explainers and lessons

Classroom and Bookshelf office sets for educational series.

Vertical Shorts

9:16 episodes with waist-up framing for Reels, TikTok and Shorts.

Founder and creator updates

Your face, your own recorded voice, no camera setup.

Localised spokesperson videos

The same host, the same set, a different language per market.

06FAQ

Questions, answered.

How long can an episode be?

Up to 570 seconds — just under ten minutes — split into segments of up to about 45 seconds each. Longer pieces can be joined in the Video Editor.

Can the host be me?

Yes. Create a Character from your photo and pick it as the host; upload your own audio clips (up to 60 seconds each) if you want it in your own voice.

What is a table read?

A voice-only pass over the whole script using the chosen voice. It costs a few credits and no video credits, so you can fix pacing and wording before rendering.

Which formats are supported?

Widescreen 1920×1080 for YouTube and vertical 1080×1920 for Shorts and Reels, with close-up, waist-up or wide framing, exported at 720p or 1080p.

How is it priced?

The avatar engines bill per second of real audio (their per-minute cost is shown on the engine label), voiceover per 1,000 characters, and export by the minute. Cached segments aren't re-charged when you edit a different line.

Can I edit the finished episode?

Yes. The episode is a project in the Video Editor, so you can add captions, B-roll, titles and music, then export again.

Paste the script.
Publish the episode.

Sign in, choose a host and a set, table-read it, and render your first episode today.

NO WATERMARKS · CREDITS ROLL OVER · COMMERCIAL LICENSE INCLUDED · CANCEL ANYTIME