Your AI twin used to take five minutes of source footage. Mirage's new Avatar X drops it to ten seconds — and the twins actually look human. Here is the recipe.

Ten seconds. That is how much of yourself you need to hand over now for a talking AI twin most people can not tell isn't you. The tradeoff isn't the length of the clip — it's what you shoot in those ten seconds.
On July 28, Mirage — the New York company that owns Captions — released Avatar X, a video model that turns a 10-second clip of you into a speaking digital twin. Avatar X now powers every AI Twin and AI Avatar inside the Captions app; you enroll without leaving it. Mirage's product page pitches the tradeoff bluntly: 10 seconds to enroll, versus 15 for HeyGen and one to five minutes for Synthesia.
What actually changed under the hood is subtler and more interesting. Older twins layered lip movement onto a still frame or a stock loop — mouths moved, but eye contact drifted and the shoulders sat still. Avatar X generates the whole performance in one pass: voice, microexpressions, gaze, hand and body motion, all synced. That is why the twins finally read as people rather than puppets.

The following images were generated using Nano Banana 2. Prompt: "close-up product shot of an iPhone on a wooden desk showing the Create AI Twin enrollment screen with a 10-second countdown at 70%, warm daylight, warm neutral wood+cream palette, bright well-exposed"
Receipts: Mirage's launch post at captions.ai/blog/mirage-avatar-x, RuntimeWire's coverage at runtimewire.com/article/mirage-avatar-x-ai-twin-launch.

Worked example A — prompt: "vertical selfie phone-camera 1:1 crop of a mid-40s South Asian woman recording a 10-second enrollment clip in her bright home living room, mid-blink with a slight smirk, one hand gesturing off-camera-left, natural daylight, warm neutral+coral palette, bright well-exposed"

Worked example B — prompt: "1:1 square editorial frame of a bearded mid-30s East Asian man in a bright home podcast booth with pale grey acoustic panels and a warm brass pendant lamp, phone on a small tripod capturing him mouth mid-word with a genuine laugh, one hand gesturing to camera, warm brass and pale-grey palette, bright well-exposed"
The 10-second number is the marketing pitch. The real story is that enrollment friction has finally dropped below the friction of just recording the video yourself. Setting up a twin used to take longer than a week of just posting normally. Now it takes less time than choosing a filter. That flips the math: twins go from a "one big commercial project" tool to a "record once, use forever on the mundane stuff" tool. Ads, tutorials, localizations, follow-up posts on a thread — all of it can now come from a ten-second clip you shot in your kitchen.
The catch, and it's a real one: your enrollment clip lives inside Captions' account. If you close the account, you lose the twin. If you want to own your face — portable across whichever tools you might switch to next — you need a workflow where the reference clip is yours, and the twin engine is a service you can swap out. That is the choice worth thinking about before you enroll anywhere.
Ten seconds to a talking twin is where the industry is, this week. That is the good news. The next question is who owns the twin — the app you enrolled in, or you. If you want a talking avatar you actually control, generated from your own reference footage on a stack you can move, build your avatar with AI Avatar and keep the source material yours.
How many seconds does Mirage Avatar X actually need to build a twin?
Ten seconds of source footage, recorded directly inside the Captions app. For comparison, HeyGen asks for about fifteen seconds and Synthesia wants one to five minutes.
Does the 10-second clip need to be studio-lit?
No. Avatar X captures the environment, lighting and framing you actually record in. A well-lit phone selfie in your normal room is closer to what you will deploy than a studio clip you will never repeat.
Can I make my twin speak a different language than what I filmed?
Yes. Once the twin is trained you can change the script, voice, clothing, background and language without re-recording. That is the whole point of building a twin.
What happens if I close my Captions account?
You lose access to the twin. The enrollment reference and every generated performance are tied to the account. If ownership matters — for example, if you want a talking avatar you can port across tools — record and store your own reference clip separately.
Do I still need a real camera for the enrollment?
No. A modern phone in good light is the recommended path. The bottleneck for a good twin is what you do in the ten seconds, not the sensor.