AI Slideshow Storyboards: From Customer Insight to Consistent Images

By Kshitij (Tjay) Dhyani··9 min read
ai slideshowsai ugccreative workflowimage generationghostfeed

The expensive part of an AI slideshow is not generating seven images.

It is making seven images feel like one argument.

Most broken slideshows have the same failure pattern:

  • slide one promises one thing;
  • slide two starts a different story;
  • every image looks like a different campaign;
  • the product arrives as a random final-slide CTA;
  • the copy explains what the visual should have shown.

You do not fix that with a more cinematic image prompt. You fix it with a storyboard.

I use Claude Fable 5 as the reasoning layer and an image model such as GPT Image 2 as the renderer. The model names matter less than the handoff contract between them.

Start with one customer moment

One slideshow should resolve one tension.

Not "personal finance tips." Not "why our app is great."

Something a person can recognize:

I signed up for a free trial, forgot about it, and found the renewal after it hit.

That moment contains:

  • a person;
  • a situation;
  • an emotional turn;
  • a product-relevant problem;
  • potential proof.

Now decide what the slideshow argues:

The first defense against forgotten renewals is seeing every subscription in one place.

That is narrow enough for seven slides. It also avoids an unsupported promise that the app will save a specific amount or cancel something automatically.

Choose a narrative shape

Formats are useful because they set expectations. They are not a substitute for an argument.

Four reliable shapes:

Recognition → explanation → proof

Use when the audience already feels the problem.

Ask your agent
I forgot the trial → here is why that keeps happening → here is the screen that makes it visible

Claim → objections → evidence

Use for a contrarian position.

Ask your agent
Budgeting harder is not step one → discipline is not the only problem → invisible renewals distort the budget → visibility comes first

Mistakes → correction

Use for educational, saveable posts.

Ask your agent
three reasons your app demos look fake → mistake one → fix one → mistake two → fix two

Before → mechanism → after

Use only when the transformation is real and provable.

Do not fabricate a customer outcome because before-and-after is a familiar template.

Write a beat sheet before slide copy

The beat sheet is the argument without the prose.

SlideJobViewer question
1RecognitionIs this about me?
2EscalationWhy should I keep swiping?
3MechanismWhat is actually happening?
4ConsequenceWhy does it matter?
5Product proofCan I see the solution?
6ResolutionWhat changes for the person?
7Next actionWhat should I do now?

If two adjacent rows do the same job, combine them. If a slide has no job, delete it.

Here is one of our first-party slideshow examples:

Opening slide from a Ghostfeed slideshow exampleSecond slide continuing the Ghostfeed slideshow storyThird slide adding detail to the Ghostfeed slideshowFourth slide in the Ghostfeed slideshow sequence

The useful question is not whether each image is attractive. It is whether the sequence is legible with the sound off and the captions briefly hidden.

Put the storyboard into Ghostfeed

The handoff depends on where the storyboard came from:

  • Use From Prompt when Fable produced the audience, angle, beat sheet, and scene direction. Include the approved proof and @-mention the relevant product, collection, or avatar.
  • Use From Social when the beat sheet was extracted from a working TikTok Photo Mode post or image-only Instagram carousel. Ghostfeed keeps the source slide count and structural rhythm while replacing the wording and media.
  • Use From Scratch when every slide contract is already approved and you want to cast each frame yourself or let Claude do it through MCP.

Do not ask the image model to flatten the storyboard into final posters. Ghostfeed stores the scene and exact copy separately: backgrounds remain replaceable, text boxes remain editable, and the whole sequence remains reviewable.

Once the base is ready, use Generate variants for new hooks. The new hook controls the copy; the base controls cadence, slide count, text geometry, and visual taste. Collection-backed and avatar-backed slides rotate inside their approved pools, Pinterest slides recast, and deliberate uploads stay fixed.

Give every slide a contract

I ask the reasoning model to produce structured rows:

Ask your agent
slide: 3 job: reveal the mechanism copy: "The renewal wasn't expensive. Twelve small renewals were." visualEvidence: itemized subscription list on a phone continuity: creator: creator-reference-a location: kitchen-reference-a time: morning mustInclude: - readable item count mustAvoid: - invented bank balance - visible third-party logos - floating text generated inside the image

The renderer does not need to infer strategy. It receives the visible scene, references, composition, and exclusions.

The copy should also survive without the image, and the image should add evidence rather than serve as wallpaper.

Bad pair:

  • Copy: "I was so shocked."
  • Visual: Generic surprised face.

Better pair:

  • Copy: "I thought I had three subscriptions."
  • Visual: The creator counting a longer itemized list on the phone.

Separate scene generation from typography

Asking an image model to produce a realistic photo, preserve a character, compose a layout, and typeset exact copy in one pass creates unnecessary failure.

I usually split it:

  1. generate the clean scene;
  2. approve identity, hands, objects, product accuracy, and framing;
  3. place exact text with deterministic layout code;
  4. check the final composite.

This gives you:

  • editable copy;
  • consistent font and margins;
  • safer line breaks;
  • easier localization;
  • accessible source text;
  • fewer regenerations when one word changes.

If the image model is responsible for text, treat the result as a draft and verify every character.

Use continuity references selectively

"Use the same four references on every slide" sounds disciplined but can be wrong.

References should match the continuity requirement:

  • character reference when the same person returns;
  • location reference when the story stays in one place;
  • product reference whenever the interface or object must be correct;
  • visual-system reference for lighting, crop, grain, and color;
  • layout template at the deterministic compositing stage.

Do not force a kitchen reference into a product-only slide. Irrelevant references compete with the instruction and can create visual artifacts.

I also distinguish fixed and flexible attributes.

Fixed:

  • face identity;
  • wardrobe within one scene;
  • product;
  • room;
  • color treatment.

Flexible:

  • camera distance;
  • expression;
  • hand position;
  • prop placement;
  • negative space for copy.

The prompt becomes easier to debug because you know what drift is actually a failure.

Generate alternatives for risky slides

Not every slide deserves eight variants.

Generate more alternatives where failure is expensive:

  • the cover;
  • a recurring creator's first appearance;
  • a hand interacting with the product;
  • a product-proof slide;
  • the final CTA composition.

For connective slides, one or two may be enough.

This is a budget decision, not a superstition about a magic number of candidates.

Run three QC passes

1. Slide-level QC

  • Does the image match the contract?
  • Is the creator consistent?
  • Is the product representation accurate?
  • Is text readable in the safe area?
  • Are claims supported?

2. Sequence QC

  • Does every swipe advance the argument?
  • Does time and location continuity make sense?
  • Does the product arrive at the moment it becomes relevant?
  • Does the emotional movement feel plausible?
  • Is any slide redundant?

3. Feed-level QC

  • Does the cover read at thumbnail size?
  • Is the first line specific?
  • Does it look native to the account's established visual language?
  • Is the disclosure present where required?
  • Is the CTA honest and proportionate?

AI vision can flag mismatches. A human should approve the actual piece.

Do not invent platform folklore

You will hear precise claims about "the algorithm" rewarding swipes, audio, saves, slide two, or the first 18 hours.

Some of those may correlate with distribution. Very few justify universal causal rules.

TikTok says recommendations use signals including user interactions and content information, with weights varying by context. Read its current explanation of how TikTok recommends content instead of turning a creator anecdote into policy.

Optimize for a human sequence:

  • the first slide earns attention;
  • the second confirms the promise;
  • each next slide pays off part of the argument;
  • the final action fits what the viewer just learned.

That remains useful even when ranking systems change.

A prompt for the reasoning layer

Ask your agent
Create a seven-slide storyboard from the attached evidence table. Audience: [specific audience] Single argument: [one falsifiable claim] Required proof: [product screen, customer quote, demonstration, or source] For each slide return: - narrative job - exact overlay copy, maximum 14 words - visible scene - evidence used - continuity requirements - generation prompt - exclusions - rejection criteria Do not invent customer results, product features, quotations, or platform claims. If the evidence does not support the argument, stop and explain the gap.

The best output from that prompt may be a refusal to storyboard. That saves more money than a pretty first slide.

Ghostfeed turns the approved storyboard into an editable slideshow with explicit asset state and review. One source idea can then become several controlled executions—without losing which image, copy, and proof belong together.

Watch the storyboard become a complete, editable TikTok photo slideshow in Ghostfeed.

For format inspiration and adaptation, see TikTok Slideshow Maker for Apps. For the reasoning layer behind the workflow, read Claude Fable 5 for organic AI UGC.