← Kosi Anyaegbuna

August 2026/ Generative video pipeline/ 8 min read

Building a production-grade UGC workflow for automated product reviews

Annabelle, the product this pipeline feeds — 10+ videos generated so far.

View live project

Showcase#

Three finished vertical ads the pipeline produced, the character kept consistent across them.

Finished vertical ads the pipeline produced.

UGC only works in volume. You test ten hooks to find the one that holds attention for three seconds, and each of those ten needs a person with a camera, a script, and an evening of editing. The thing the format needs most is the thing a person is slowest at.

So I built the whole run as an agent: read the product, decide the angle, write the lines, generate the character, film it, hand back something postable. A product URL goes in and a finished vertical ad comes out for $4.75 a video.[1] It has made over twenty of them.

What I never got was a cheap way to tell whether a change made it better. Every answer to that question runs back through the video model, and the video model is the entire cost.

Seven phases, two of them mine#

  1. Product URL
  2. Read the sitebrowser, screenshots, structured brief
  3. Pick the hook anglepast run ratings, plus a library of hooks that performed on social
  4. Write the scriptreads my last five review notes
  5. I approve the scriptrejecting sends it back up with my note attached
  6. approvedGenerate the anchor framethe character and the room, established once
  7. Generate the rest ×3each one references the anchorover the call ceiling — halt
  8. Animate each frame ×4eight seconds, with its own audioa clip fails twice — halt
  9. I cut them togetherby hand
  10. Finished video, rated one to five
One run. The dashed stages are mine, and the rail carries the rating back to the two phases that read it.

The two branches leaving the diagram are there because every generation bills. A model that misreads an instruction and loops costs real money, so the two generating stages count their own calls and halt the run when the count passes twelve — the expected number is eight, and anything above that means something is wrong in a way retrying won't fix.

The two places it stops for me are the only other real decision in the sequence.

The first is about money too. Everything before the script costs four cents and everything after it costs $4.70,[2] so the check sits at the four-cent mark, where I read sixty words of text before that text authorises the expensive half. It takes ten seconds and it's the cheapest place in the run to catch a bad one.

The second stop is the edit. A Veo generation caps at eight seconds and the clips come back with unpredictable tails — the line finishes at six seconds and the character sits there for the remaining two with an empty expression. Finding that cut automatically means detecting the frame where she stops talking and starts waiting. I can see it instantly, so I cut the clips together myself.

Both stops feed the next run. My script notes get appended to a file that the script phase reads before writing, and my rating goes to a second file, with the outfit and room it used, that the wardrobe phase reads before choosing. It's the cheapest possible feedback loop: the last few notes pasted into the next prompt.

The same face in every clip#

The hard technical problem is that four separately generated images have to show the same person, in the same room, wearing the same clothes. Generate them independently and you get four strangers.

The obvious approach is to chain them — feed scene one into scene two, scene two into scene three. That compounds. Every hop drifts a little further from the original face, and by the fourth clip it's someone else's sister.

What works is re-anchoring. Every scene after the first takes scene one's output as its reference image, so nothing is ever more than one hop from the original. Alongside it goes a fixed photograph of the room, passed on every single call, which holds the broad shape of the background. The clothing description is copied word for word between scenes, and the expression is the only thing I let change.

Hand raised near her mouth, lips parted, eyebrows lifted
Mid-sentence, eyes wide, looking straight down the lens
Laughing, head tilted slightly back
Three frames from one run. The expression is the only thing the prompt asks to change — holding the set even this steady was the harder half.

The character holds. Across runs the face, the hair and the outfit come back the same, clip to clip. What stays soft is everything around her. The room drifts object by object — a mug on the desk in one clip and gone in the next — and because each clip is generated alone, none of them knows what energy the last one ended on. In a still it reads as nothing. Across four clips a trained eye will catch it, and inconsistency like that is exactly what you can't afford in an ad that has to pass as a person filming herself.

The expensive models weren't the reliable ones#

I started with Opus writing the scripts. It was good — clean lines, and a second pass only now and then — and I moved off it anyway, because sixty words of dialogue didn't justify what it charged to produce them. Gemini Pro wrote the best scripts of anything I tried and wouldn't return valid JSON consistently enough to sit inside a pipeline, so runs died at phase three for reasons that had nothing to do with the writing.

What I settled on was a cheap model per phase, picked for the shape of the task, with programmatic checks on the output and my own review behind that.

Clips · Veo 3.1 Fast450¢Four generations, one per scene
Frames · Nano Banana Pro20¢Four stills, the only step that takes reference images
Script · DeepSeek V32¢The sixty words she actually says
Strategy · Gemini 3 Flash1¢Picks the hook angle and the scene structure
Site read · Gemini 2.5 Flash Lite1¢Structured extraction from the page and screenshots
Cost of one run, by phase and the model that handles it.

Here's the part I'd underestimated. I read every script myself before the expensive half runs, so a better writing model mostly buys me things my own review would have caught anyway. Paying more there changed very little about what actually got made. And the phases I kept going back to tune came to four cents of a $4.75 run.

Finding out cost $4.75 a time#

That chart is also the reason the work stopped where it did.

A prompt is tuned to a specific model, and the tuning doesn't necessarily transfer. Most of what's in my video phase exists because of how Veo behaves in particular. It renders a camera into the shot if you say "looking at the camera", so every prompt says "eye contact with the viewer" instead. It only accepts four, six or eight second durations, so clips generate at eight and get trimmed. It comes back landscape unless the ratio is specified outright, so there's a dimension check downstream. None of that is knowledge about video generation. It's knowledge about one model, and swapping the model throws it away and re-opens every prompt I'd already settled.

The models also move underneath you. The same prompt that produced a clean take last month can come back subtly worse, and there's no announcement.

So the only honest way to answer "is this better" is to generate samples and compare. One sample is $4.75, and 95% of that is the video. Comparing two models across a handful of scripts, with enough repeats to see past the variance, is a few hundred dollars before you learn anything you'd act on. The four-cent phases I could iterate on all day. The $4.50 phase — the one carrying the quality — was the one I couldn't afford to ask questions about.

What I'd fix first#

Two halves of this needed different things from me. One is infrastructure — the Telegram interface I drive it from, the orchestration holding seven phases together, authenticated upload, storage across Supabase and Cloudflare R2, and the retry and abort behaviour that keeps a half-failed run from billing twice. That's ordinary engineering and it either works or the pipeline doesn't run at all.

The other half is twelve markdown and YAML files, and that's where the videos get better or worse. It's also the half with no version control.

  • SKILL.mdorchestrator · 12 versions
  • config/
    • models.yamla model per phase · 12 versions
    • wardrobe.mdoutfits and rooms · 8 versions
    • agent-personality.mdvoice · 2 versions
  • extensions/
    • all-extensions.mdextras · 1 version
  • phases/
    • 01-site-intel.mdreads the site · 7 versions
    • 02-strategy.mdpicks the hook · 7 versions
    • 03-script.mdwrites the lines · 10 versions
    • 04-frames.mdfour stills · 11 versions
    • 05-video-gen.mdfour clips · 12 versions
    • 06-post-production.mdcaptions and delivery · 6 versions
    • 07-upload.mdupload and rate · 7 versions
The twelve files, and how many times each was rewritten across twenty-one versions. Every change meant copying the whole directory and incrementing the number.

Twenty-one of those directories exist and only one of them ever changed a single file — the rest moved two to seven at once. Combined with samples being expensive, that's the whole problem in one line: when something came back worse I could neither afford to measure it nor narrow it down to a file.

So the version history is where I'd start. Per-file history, so a phase can go back without taking the other eleven with it. A required sentence on every change, because a diff shows one model ID replacing another and can't carry the reason. And every finished video tagged with the version that produced it, which costs almost nothing and is the thing that would make the runs I'm already paying for into evidence.

I stopped at $300#

I was funding it myself. The pipeline runs and the videos are good enough to use, and the next increment of quality was priced in video generations — weeks of buying samples to find out which of my changes were real. That's a budget decision, and I am not Elon Musk unfortunately.

If I picked it up again the fix isn't a cheaper test. I already read every script before it authorises anything, so the text half was never the blind spot. What was missing is that the expensive runs left no trace — nothing recording which version produced which video, nothing tying a rating back to the files that earned it. The comparison I wanted is the same product run on two versions, side by side, and once the runs are tagged it costs nothing extra.

The arithmetic may also have moved. Video generation is cheaper than it was when I started, so the samples I couldn't justify then might be affordable now — though working out whether a cheaper model holds the quality is itself a test, and it costs money to run.

The general shape isn't specific to video. Any pipeline where the expensive stage is also the one that most needs tuning has this problem: the price of one sample sets how fast you're allowed to learn, and above some number you stop measuring and start guessing. Working that number out before starting is worth more than any individual model choice.