The ChatGPT script trend needs a friend group — unless you use AI avatars. Cast each character solo, keep the deadpan, and hop TikTok's biggest August format.

Every other Reel this week is a friend group reading an AI-written scene deadpan for the camera. If you're a solo creator, the format looks locked to you — until you realize the flat delivery AI avatars get roasted for is exactly what the joke needs.
"We had ChatGPT make us a script" is TikTok's biggest August 2026 comedy trend — one prompt, one take, everyone plays it straight.
The joke requires a friend group and monotone delivery — solo creators are out unless they change the setup.
Cast each speaker as a different AI avatar and both problems disappear: you play all the parts, and the deadpan is the format's default.
Here's what's actually spreading. One person in the group asks ChatGPT for a short scene — a breakup, a job interview, a family dinner going off the rails. The model spits out four to six lines. Nobody edits, nobody rehearses more than once. They read the whole thing in a single unbroken take, staring blankly at each other while the dialogue slowly escalates into something absurd. The reveal — the on-screen text saying "we had ChatGPT make us a script" — is what makes rewatches funnier than the first viewing.
The trend has been showing up across TikTok trend roundups all August, and by Aug 23 it had crossed to Reels hard enough that both Miraflow and NextReel shipped explainers with prompt libraries. The friction is coordination — you need at least two humans on camera at the same time who won't crack. Solo creators, faceless accounts, and anyone whose "friend group" lives in a Discord server were sitting this one out.

The following images were generated using Nano Banana 2:

Prompt (Example A — Speaker A, deadpan cafe):
Portrait, 1:1. Late 20s South Asian female creator in a mustard cardigan at a cafe window seat, mid-sentence with completely deadpan expression, coffee cup and pastry on the marble table, hands loosely resting. Bright morning daylight through the window behind her, high-key soft, well-exposed. Palette: warm cream, mustard, oak, cool sage. Sharp focus, editorial photography, no text overlays.

Prompt (Example B — Speaker B, deadpan boardroom):
Portrait, 1:1. Same late 20s South Asian female creator (identity-locked, same face) in a crisp charcoal blazer over a white blouse at a modern boardroom desk, mid-word with hands lightly steepled, blank corporate deadpan expression, glass wall behind her. Bright overhead office daylight, high-key soft, well-exposed. Palette: cool white, charcoal, steel blue. Sharp focus, editorial photography, no text overlays.
Prompt ChatGPT for the scene, not the punchline. Give it a mundane premise, two to four speakers, and a hard length cap. "Write a 40-second dinner conversation between four friends where nothing important happens but the phrasing is oddly formal, roughly 12 lines total, label speakers A/B/C/D" beats "write me a funny scene." The dialogue gets weird on its own — you don't need to prompt for weird. If the first pass reads normal, ask ChatGPT to make it "slightly too specific" and regenerate once.
Cast one avatar per speaker — same face, different context. This is the flip. Instead of four different people, generate four looks of yourself with an identity-locked model like Nano Banana 2 or HeyGen Avatar V. Change the wardrobe, the setting, and the pose per speaker — but keep the face frozen. Viewers instantly clock "wait, that's the same person" before they even read the on-screen text. The two paired examples above are Speaker A and Speaker B from a real solo run.
Feed each avatar its own lines with a talking-head model. Kling 2.6 Pro, Wan 2.6, or HeyGen Avatar V will drive lip-sync from either text-to-speech or a wav upload. For this format you want the TTS to sound uninterested — most tools have a "flat" or "monotone" delivery preset, use it. If your tool defaults to expressive, dial pitch variation all the way down. The whole comedic engine of the trend runs on flat.
Stitch the clips so it feels like a single take. The look you want is a fixed camera cutting between speakers with zero flourish — no B-roll, no push-ins, no music sting. Match the framing (all mid-shots, roughly the same eye height), keep the color grade uniform across all four avatars, and skip transitions entirely. The one-take illusion sells the format; a jump-cut edit reads as production and kills the joke.
Front-load the reveal in the first frame. "we had ChatGPT make us a script" as plain white on-screen text, first frame, no fade. Both source guides stress this — the premise has to land in under a second or the payoff loses its rewatch value. Add #chatgptscript to the caption; that's where the trending audio and stitches are converging on both platforms.
When to actually use this — the trend is genuinely good for solo AI-avatar creators for one week, maybe two. It's an easy proof that your avatars can carry a scene, and the comedy engine (flat delivery + escalating absurdity) is a rare case where an AI limitation actively helps the format instead of fighting it. Don't force it if your niche is beauty, finance, or anything where deadpan reads as cold. And don't stretch it past 45 seconds — even the human-cast versions crumble past that mark.
The ChatGPT script trend was accidentally built for the exact people it seemed to lock out. Solo creators get to play every character. Faceless accounts get to have a face — several, actually. And the AI-avatar deadpan that ships as a limitation of every talking-head model on the market is the format's whole punchline. If you don't have an avatar yet, start yours here — the trend won't stay hot forever, but the workflow is yours long after it cools.
Do I need multiple different avatars, or can I use one?
You want at least two distinct-looking versions of yourself per scene — the format needs a conversation, and one talking head reads as a monologue. Same face, different wardrobe and setting is the sweet spot.
Which lip-sync model handles monotone delivery best?
HeyGen Avatar V and Wan 2.6 both expose a delivery slider — set it to "flat" or dial expression to zero. Kling 2.6 Pro is more expressive by default, which fights the format.
Do I need a cloned voice, or is text-to-speech fine?
TTS is fine for this specific format because monotone is the point — a slightly robotic voice actually helps the joke. Save your cloned voice for content where warmth matters.
How long should each finished clip be?
Thirty to forty-five seconds. Every source we checked (Miraflow, NextReel, the trend roundups) landed on the same window. Anything longer loses the one-take illusion.
Will Instagram down-rank this as AI-spun content?
The originality signal is about what you do with the trend, not whether AI is involved. A creator-driven skit with an original premise clears the bar; a template-generated clip with no premise does not. Write the ChatGPT prompt yourself — that's the original idea.