The easiest way to waste a good hook is to pour it into the wrong format.
"I stopped paying for six subscriptions after I found this" could be:
- a face reacting to a screen;
- a screen recording with commentary;
- a six-slide list;
- a founder explaining the product;
- a fake testimonial that should never have been made.
Same sentence. Five completely different levels of proof.
I run AI UGC accounts for apps, and I think about formats as reusable containers. The container decides what the viewer expects next. A gasp promises a reveal. A screen demo promises proof. A slideshow promises useful information. If the payoff does not match that promise, no amount of realism saves the post.
These are the five containers I would keep in a small app-marketing system.
Pick the job before the format
Before making anything, finish this sentence:
After watching, I want the viewer to _____.
If the answer is "understand how it works," use a demo.
If it is "feel the pain immediately," use a reaction.
If it is "save this and come back," use a slideshow.
If it is "believe somebody has lived this problem," use a real customer or founder.
If your answer is just "go viral," the brief is not done.
Format 1: the reaction hook
This is the person looking shocked, relieved, offended, curious, or caught. The face carries the interruption. Text carries the idea.
The face creates the stop. The hook text still has to earn the reaction.The format works because a reaction creates a question before the viewer has consciously read the text. "What did she just see?" buys the hook a fraction more time.
Use it for:
- surprising product discoveries;
- "how did I not know this?" hooks;
- app results that can be revealed quickly;
- stitching a human emotion to a boring screen recording;
- testing many hooks against one proven piece of motion.
Example for a flight app:
me finding out the same seat was $180 cheaper one tab over
Example for a study app:
my face when it turned 47 pages of notes into the exact quiz my professor gave us
The second example is only usable if the product genuinely does that. A good reaction cannot rescue a false claim.
Where it fails: calm, low-stakes information. Nobody gasps because a settings menu moved. If the emotion is bigger than the payoff, the format feels like bait.
The Ghostfeed production route mirrors that creative decision. I have the Ghostfeed MCP-connected agent search owned reaction templates first and the inspiration library second, judging the opening pose before the motion. If the right source is a longer owned clip, the agent returns the dashboard crop handoff; I choose the beat manually or use Smart Crop to split detected scenes into several templates.
Then I cast approved avatars into first frames and stop. Exact-motion clone is for a source performance I own and want to preserve. Prompt mode keeps the approved opening composition but directs a new action. That is usually the better route when the mechanic is useful and the original mannerism is not.
I keep a deeper swipe file of this exact mechanic in The Gasp reaction format.
Format 2: the native talking point
One person. One camera. One opinion.
A talking-head execution pairing generated video with synthetic voice. The format works only when the point sounds worth hearing.People call this a talking head, but that describes the visual and misses the job. The job is to deliver a point that sounds worth hearing from this person.
Use it for:
- warnings;
- unpopular opinions;
- personal rules;
- mistakes;
- quick explainers;
- recurring persona accounts where familiarity matters.
A weak version starts:
Here are three ways AI can improve your productivity.
There is no person inside that sentence. It could come from any account.
A stronger version:
I cancelled Notion AI after I realized I was paying it to summarize documents I never read.
Now there is a decision, a product, and an implied argument. You may disagree, which is useful.
The visual should stay simple. Eye movement, a hand entering frame, a small change in posture. The whole point is that it feels like somebody opened the camera because the thought was bothering them.
Where it fails: when the script has no opinion. A beautifully generated person reading generic advice is still generic advice.
PS: this is also the format where dead eyes, perfect skin, and wandering lip sync get punished hardest. The viewer has nothing else to inspect.
Format 3: reaction plus real screen proof
This is the format I would use most often for apps.
The opening is a face and hook. The payoff is the actual product.
Sequence:
- Reaction or talking point for one to two seconds.
- Screen recording of the product doing the thing.
- One annotation that points at the proof.
- Face or result again for the CTA.
Example for an inbox app:
"I had 18,000 unread emails when I tried this."
Then show the actual inbox, the rule being created, and the resulting state. Do not cut away right before the proof.
This hybrid is better than a fully generated demo because hands and interfaces are exactly where viewers need accuracy. Let AI create the human wrapper. Let the real screen carry the claim.
Use it for:
- app onboarding;
- before/after workflows;
- feature launches;
- comparison hooks;
- products that look boring until the result appears.
Where it fails: when the screen recording is staged or unreadable. Tiny text inside a fake phone mockup is not proof. Zoom into the one action that matters.
One useful test: mute the video and hide the caption. Can a viewer still see what changed? If not, the demo is narration wearing a screen recording as decoration.
Format 4: the saveable slideshow
A slideshow is not a cheap video. It is a different reading behavior.
01
02
03
04The viewer controls the pace. That makes slideshows good for checklists, rankings, stories, comparisons, and reference material.
Use it for:
- "save this" information;
- lists with a real ordering;
- multiple examples;
- a story that needs one beat per slide;
- testing hooks before paying to animate anything.
A six-slide structure I use:
- Hook with a specific tension.
- Second hook that makes the problem personal.
- First useful detail.
- The mistake most people make.
- The actual method.
- Payoff or CTA.
For a budgeting app:
- "the 4 subscriptions quietly charging me $91/month"
- "one of them had billed me for 19 months"
- show the first category;
- show the false saving assumption;
- show the cancellation workflow;
- "check yours before the next billing date."
One idea per slide. If a slide needs three paragraphs, you wrote an article.
Where it fails: any claim that needs motion. A screenshot of a result can support a slideshow, but six aesthetic photos cannot demonstrate a workflow.
The operational version is in How to make TikTok slideshows for apps.
Format 5: the real-person proof clip
Yes, I am including a non-AI core format in an AI UGC article.
Because pretending every job belongs to AI is how teams end up generating fake testimonials.
Use a founder, customer, or expert when the person is the proof:
- "I built this because...";
- "I used this for 30 days...";
- "Here is what changed in our account...";
- a physical product result;
- a claim that depends on credentials or lived experience.
AI can still help around the edges. Test openings before the shoot. Generate B-roll that does not imply a fake experience. Cut the real proof into multiple hook packages. Translate captions. Add reaction intros.
One useful pattern is a fast B-roll montage that shows creative range without assigning the result to a synthetic speaker:
A rapid Ghostfeed B-roll montage built around a fitness-transformation concept. Supporting footage can widen the visual vocabulary, but it cannot turn a synthetic result into proof.That footage can support a real customer's documented story. It cannot replace one.
But keep the claim attached to the person who can honestly make it.
Where it fails: over-scripting. A real customer reading your founder's paragraph word-for-word looks less trustworthy than a synthetic actor. Ask questions, capture the weird specific details, and let the customer use their own nouns.
For example, "the app saved me time" is marketing oatmeal.
"I stopped rebuilding the same Monday report in three spreadsheets" sounds like a person who had the problem.
How I choose between the five
I use two questions.
1. What is the proof object?
- Emotion: reaction.
- Opinion: native talking point.
- Product behavior: screen-demo hybrid.
- Structured information: slideshow.
- Lived result: real-person proof.
2. How expensive should this test be?
If the angle is unproven, start with the lightest format that can make the claim honestly.
A slideshow can test whether "four subscriptions charging you twice" earns saves. A reaction can test whether the surprise lands. Once the angle shows signal, build the screen demo or commission the customer story.
This is the same logic I use in AI UGC vs real creators: cheap media buys more attempts; real proof buys more belief.
A practical seven-post mix for one app
Say you market a meal-planning app.
I would not post seven talking heads. I would run:
- Reaction: "me realizing dinner was the reason my grocery bill kept winning."
- Slideshow: "5 meals from one $34 grocery basket."
- Screen demo: build the week and generate the list.
- Talking point: "meal prep advice is written for people who enjoy meal prep."
- Reaction: swap the hook, keep the motion.
- Real customer: why they stopped ordering delivery on Wednesdays.
- Slideshow: mistakes from the comments on post 2.
Now each post has a job. The comments and saves from one format feed the next instead of seven disconnected ideas going into the void.
Do not confuse formats with ideas
A gasp is a format. "Nobody told me this flight app checks nearby airports" is an idea.
A slideshow is a format. "The four subscriptions hiding inside your bank statement" is an idea.
AI makes formats cheap to reproduce. It does not make the idea good.
That is the operating rule:
Reuse the container. Rewrite the reason to care.
Ghostfeed is built around that loop. Start with a reaction or slideshow structure that already works, cast the right persona and assets, then produce enough reviewable variants to learn which hook deserves the next round.