AI UGC Model Routing: Use the Expensive Model Only When It Earns It

By Kshitij (Tjay) Dhyani··8 min read
ai ugcai video modelscontent operationscreative testingghostfeed

The internet version of an AI UGC stack is a shopping list:

Use model A for thinking, model B for images, model C for video, model D for voice, and suddenly two people can make 500 winners.

No.

A list of models is not a production system. "500 winners" is a contradiction. If you already know all 500 will win, you are not testing anything.

The real reason to combine models is routing: use the cheapest adequate method while uncertainty is high, then spend more only after an idea clears a meaningful gate.

Models have jobs, not status

Claude Fable 5 and Seedance 2.0 solve different problems.

Anthropic describes Fable 5 as a model for long-running knowledge work, agentic workflows, and visual analysis. That makes it useful for research synthesis, creative briefs, structured checks, and orchestration.

ByteDance says Seedance 2.0 accepts text, image, audio, and video references and supports controllable audiovisual generation. That makes it a candidate renderer when a brief needs those controls.

Neither vendor claim tells you:

  • whether a hook fits your customer;
  • whether a generated creator is right for your brand;
  • whether the result is legally usable;
  • whether the economics work at your approval rate;
  • whether another model is better for your exact format.

Treat every provider page as a capability claim to test, not an instruction to reorganize your company.

Route by production stage

I use four stages.

1. Evidence

Inputs:

  • customer language;
  • product truth;
  • previous results;
  • rights-cleared references;
  • brand and claim constraints.

Useful tools:

  • transcription;
  • retrieval;
  • a reasoning model;
  • spreadsheets or a proper database.

Output:

  • cited themes;
  • explicit uncertainty;
  • creative hypotheses.

Do not render video here. Most ideas should die as cheap text.

2. Previsualization

Inputs:

  • approved hypothesis;
  • one argument;
  • creator profile;
  • proof requirement.

Useful tools:

  • script drafts;
  • storyboards;
  • static source frames;
  • low-cost animatics;
  • human table reads.

Output:

  • an approved creative spec.

This frame is a real production artifact, not a disposable prompt result:

Approved source frame used for a Ghostfeed AI UGC reactionApproved source frame used for a Ghostfeed AI UGC reaction

If the frame does not sell the first second, better motion will not save it.

3. Candidate render

Inputs:

  • approved frame;
  • exact motion brief;
  • dialogue and audio plan;
  • duration;
  • rejection criteria.

Useful tools:

  • the least expensive video route that can represent the test faithfully;
  • deterministic editing for captions, crops, and overlays;
  • automated technical QC.

Output:

  • one or more candidates for human approval.

The operative phrase is represent the test faithfully. A cheap model is not cheap if it fails on the hand interaction central to the concept and forces eight retries.

4. Final execution

Inputs:

  • a candidate that passed creative review;
  • evidence the concept deserves more investment;
  • final delivery requirements.

Useful tools:

  • the model with the best measured approval economics for this shot class;
  • final sound, color, captions, and product compositing;
  • rights and claim review.

Output:

  • an approved asset with lineage.

Premium generation belongs here when it materially improves the outcome—not because its name looks good in a thread.

Define the routing score

For each shot class, record:

Ask your agent
expected cost per approved second = price per attempt × average attempts ÷ approved seconds

Then add review cost and latency:

Ask your agent
effective approved cost = generation cost + reviewer minutes × loaded reviewer rate + expected delay cost

A model with a higher sticker price can be cheaper if it succeeds in two attempts instead of ten.

A fast model can be more expensive if its defects are subtle and reviewers spend longer catching them.

Track at least:

  • first-pass approval rate;
  • attempts per approved asset;
  • generation latency;
  • reviewer minutes;
  • identity drift;
  • motion defects;
  • dialogue or sync defects;
  • prompt intervention count;
  • output cost;
  • downstream performance.

Do this by shot class. "Model X has a 72% approval rate" means little if simple talking heads hide its failure on product interactions.

Example routing table

Shot classFirst routePromotion conditionFinal route
Static reactionApproved still + light motionFrame and expression approvedBest tested image-to-video model
Product demoDeterministic screen captureScript and proof approvedComposite with creator footage
Dialogue sceneTable read or animaticTiming and claim review passModel with measured sync quality
Complex hand actionReference videoMotion is essential to conceptMultimodal reference-capable model
SlideshowStatic storyboardSequence and copy approvedDeterministic assembly

Notice that the product demo may not need a generative video model for the product screen at all.

Real screen capture is cheaper, more accurate, and easier to update. Generate the human reaction; preserve the interface as truth.

Promotion rules need evidence

A production artifact can be promoted because:

  • the brief passed strategic review;
  • the source frame passed identity and composition review;
  • a low-cost candidate passed human creative review;
  • an organic test produced a predefined signal;
  • a paid test produced acceptable acquisition economics;
  • a client approved the direction.

It should not be promoted because:

  • the output looks expensive;
  • someone called the model state of the art;
  • one post received unusual views;
  • the team already generated 40 variants;
  • the remaining credit balance expires.

Sunk cost is not a quality gate.

Fable 5 should orchestrate state, not invent it

A reasoning model connected to Ghostfeed MCP can move real workspace objects through the route. It searches the workspace templates first, then the Ghostfeed inspiration library only when no owned opening pose fits:

Ask your agent
brief-approved → template-library-searched → inspiration-library-searched-if-needed → source-frame-requested → source-frame-review → animation-requested → animation-qc → human-review → approved

It can:

  • assemble the next job;
  • check required fields;
  • compare output with the brief;
  • classify known defect types;
  • propose a targeted retry;
  • summarize cost and quality.

It should not silently turn rejected into approved, invent missing customer evidence, or regenerate forever.

Set:

  • maximum attempts;
  • per-asset budget;
  • timeout;
  • retryable error classes;
  • escalation owner;
  • immutable approval events.

Long-running autonomy is useful only when it has somewhere safe to stop.

The stop is not theoretical. A 30–120 second source import returns a dashboard action so a person can crop it manually or run Smart Crop. First-frame generation then returns one candidate per selected avatar and stops again for explicit approval. The agent must also ask which motion family the user intends:

  • clone for an owned or licensed source performance that should be preserved;
  • prompt for a new performance beginning from the approved source composition.
Ask your agent
You: Use Ghostfeed to build two variants from our saved desk template. Keep the composition, not the motion. Agent: I will use the template as the frame source, render two approved avatars, and stop. After frame approval I will use prompt mode, because you explicitly do not want an exact-motion clone.

That is model routing with product state, not a language model writing instructions for a disconnected renderer.

Quality control needs two layers

Mechanical QC

  • correct dimensions;
  • expected duration;
  • valid codec;
  • audio present;
  • no blank frames;
  • product screen readable;
  • captions inside safe areas.

Editorial QC

  • creator continuity;
  • believable motion;
  • useful first second;
  • honest claim;
  • correct disclosure;
  • appropriate emotion;
  • brand fit;
  • no uncanny detail a customer will notice.

The first layer can be heavily automated. The second can be assisted by vision but should remain accountable to a named reviewer.

Here are two executions from the same source-format family:

They are useful variants only if the creator change answers a real question. Otherwise they are two renders.

Buy optionality, not volume

The advantage of a routed system is not that every model runs on every job.

It is that:

  • a reasoning model can be replaced without losing production state;
  • a renderer can be benchmarked against the incumbent;
  • a price change does not break the workflow;
  • an outage pauses one route instead of erasing the job;
  • quality thresholds can differ by client and format;
  • expensive work happens after cheap uncertainty reduction.

One video. Hundreds of posts. That is a distribution capability—not permission to generate hundreds of undifferentiated files.

The good stack makes a small team more discriminating. It kills weak ideas early, spends on the few that earn it, and records enough evidence to choose a better route next month.

For the queue and persistence architecture, read How to scale AI UGC production. For turning a result into controlled branches, use the AI UGC creative variation matrix.