"Non-branded campaigns don't scale profitably."
They do. Most operators are just measuring them wrong and structuring them worse.
I've taken non-branded from 15% to 60%+ of total revenue across 20+ e-commerce brands.
The breakthrough was never better keywords. It was intent.
A searcher typing "buy organic skincare" and a searcher typing "how to reduce acne" are two completely different people.
One has a wallet open. The other is reading.
Transactional intent converts 4-6x higher than informational.
So when you bid them the same, send them to the same page, and run the same copy, Google averages the two audiences together and the math goes flat.
That's the real reason most non-brand "doesn't convert."
Here's the case that made it obvious.
A brand was spending $15k/mo on non-branded Search. Breaking even at best.
I pulled their ad groups. Transactional keywords like "buy organic skincare" sat in the same campaign as "how to reduce acne."
Same bids, same pages, same copy.
I split them. Transactional got its own campaign, its own conversion-optimized pages, its own budget.
Informational got pushed down or moved to a separate funnel.
CPA dropped 40% from the separation alone. No new budget.
Google just stopped averaging performance across two audiences that had nothing in common.
The structure I run now sits on three tiers:
• Transactional ("buy X," "X discount"): aggressive bids, pages built to convert on the spot.
• Commercial investigation ("best X," "X vs Y"): moderate bids, comparison and education pages. These people are choosing, not browsing.
• Informational ("how to X"): low bids or a separate Display funnel. Cheap reach, not a closing channel.
One messaging rule carries all three: non-brand means they don't know you.
The ad has to answer "why should I buy this product," not "why should I buy from you." Lead with their problem, not your logo.
Now the part that kills good campaigns: ROAS.
Non-brand brings in new customers.
Their first-purchase ROAS will always look worse than brand, because brand is just harvesting people who already decided.
But non-brand customers often carry 2-3x the LTV, because they're not one-and-done deal seekers who only show up for a discount code.
Judge non-brand on first-touch ROAS and you'll cut every campaign that actually grows the business.
I measure on 90-day customer value, then decide.
The scaling triggers I use:
• Transactional CPA stable for 2 weeks - increase budget 20%.
• Commercial investigation showing positive 90-day LTV - move it up to transactional budget levels.
• New winners surfacing from broad match - add them to exact match immediately.
Brand harvests demand that already exists. Non-brand is the only budget line building demand you don't have yet, and it scales the day each intent gets its own room.
Founder check, two minutes: open your non-brand campaign and read ten search terms. If "buy X" and "how to fix X" live in the same ad group, you just found the reason it "doesn't scale."
They do. Most operators are just measuring them wrong and structuring them worse.
I've taken non-branded from 15% to 60%+ of total revenue across 20+ e-commerce brands.
The breakthrough was never better keywords. It was intent.
A searcher typing "buy organic skincare" and a searcher typing "how to reduce acne" are two completely different people.
One has a wallet open. The other is reading.
Transactional intent converts 4-6x higher than informational.
So when you bid them the same, send them to the same page, and run the same copy, Google averages the two audiences together and the math goes flat.
That's the real reason most non-brand "doesn't convert."
Here's the case that made it obvious.
A brand was spending $15k/mo on non-branded Search. Breaking even at best.
I pulled their ad groups. Transactional keywords like "buy organic skincare" sat in the same campaign as "how to reduce acne."
Same bids, same pages, same copy.
I split them. Transactional got its own campaign, its own conversion-optimized pages, its own budget.
Informational got pushed down or moved to a separate funnel.
CPA dropped 40% from the separation alone. No new budget.
Google just stopped averaging performance across two audiences that had nothing in common.
The structure I run now sits on three tiers:
• Transactional ("buy X," "X discount"): aggressive bids, pages built to convert on the spot.
• Commercial investigation ("best X," "X vs Y"): moderate bids, comparison and education pages. These people are choosing, not browsing.
• Informational ("how to X"): low bids or a separate Display funnel. Cheap reach, not a closing channel.
One messaging rule carries all three: non-brand means they don't know you.
The ad has to answer "why should I buy this product," not "why should I buy from you." Lead with their problem, not your logo.
Now the part that kills good campaigns: ROAS.
Non-brand brings in new customers.
Their first-purchase ROAS will always look worse than brand, because brand is just harvesting people who already decided.
But non-brand customers often carry 2-3x the LTV, because they're not one-and-done deal seekers who only show up for a discount code.
Judge non-brand on first-touch ROAS and you'll cut every campaign that actually grows the business.
I measure on 90-day customer value, then decide.
The scaling triggers I use:
• Transactional CPA stable for 2 weeks - increase budget 20%.
• Commercial investigation showing positive 90-day LTV - move it up to transactional budget levels.
• New winners surfacing from broad match - add them to exact match immediately.
Brand harvests demand that already exists. Non-brand is the only budget line building demand you don't have yet, and it scales the day each intent gets its own room.
Founder check, two minutes: open your non-brand campaign and read ten search terms. If "buy X" and "how to fix X" live in the same ad group, you just found the reason it "doesn't scale."
❤1👍1🔥1
Everyone on X is asking "which AI video model is best."
Wrong question. It's why your ad account is bleeding.
I run YouTube ads for ecommerce brands. Last night we shipped dozens of video ads with 60+ shots, 5 hook variants each, through a full agentic AI pipeline.
For less than a stock footage subscription.
There is no best model.
There are best models per SHOT.
We route models per shot class, not per video:
1. Trust shots get the premium model.
The 5% of seconds that carry the proof: the water beading, the before/after match cut, the texture close-up.
One physics glitch there kills the whole ad, because viewers stare at those frames hunting the AI tell. Kling 3.0, $0.07/sec, no debate.
Maybe $6 of a whole production. Never save money here.
2. B-roll gets the billing arbitrage.
Some models bill per second, some bill PER VIDEO. Veo 3.1 Lite: about $0.175 for an 8-second 1080p clip.
That's $0.022/sec. Cheapest quality-per-second, because nobody does the division.
Workbench shots, driveway shots, scene-setting vignettes?
Nobody is pixel-peeping those.
Route them there.
We author b-roll beats at exactly 8 seconds to exploit the billing unit.
3. Talking heads only go to lipsync-capable models.
Most fake it. Two do it properly.
A silent mouth-flap is not a style, it's a broken file.
4. Sound costs extra. Pay it once.
Native audio is a +40% surcharge on some models. Pay it ONLY where the sound IS the appeal (the sizzle, the spray, the seal).
Everything else renders silent under one continuous voiceover.
Net effect: same perceived quality, half the spend, 5 hooks instead of 1, because hooks share the rendered body and only the opening clip changes.
A variant costs us $2-5. People pay editors $300 for what is a different first 8 seconds.
Now the part for people who don't generate video at all:
The highest-ROI step in our pipeline costs a dollar and needs zero rendering:
Before anything gets produced, the script gets torn apart by simulated buyers.
Not "does my team like it."
The skeptic burned by the last product. The spouse who controls the card. The 24-year-old who smells an ad in half a second.
Open-ended questions only:
- "retell this tomorrow, what survived?"
- "what did this assume about you that's wrong?"
- "at which second do you leave?"
Two rounds killed a line our whole team loved and rebuilt the entire opening.
Found after production, each of those is a full re-shoot. Found at script stage - a text edit.
You can do this today with any LLM and zero budget.
Most people won't.
It feels slower than creating.
It's the fastest thing we do.
Those "5 AI video prompts that go viral" threads? Recycled advice with the model name swapped. They age in days.
The questions underneath don't:
- Which shots carry the trust? Route premium.
- Which shots are wallpaper? Route cheap.
- What does the buyer retell tomorrow? That's your ad.
Stop asking which model is best. Roast the script before a single frame exists.
Wrong question. It's why your ad account is bleeding.
I run YouTube ads for ecommerce brands. Last night we shipped dozens of video ads with 60+ shots, 5 hook variants each, through a full agentic AI pipeline.
For less than a stock footage subscription.
There is no best model.
There are best models per SHOT.
We route models per shot class, not per video:
1. Trust shots get the premium model.
The 5% of seconds that carry the proof: the water beading, the before/after match cut, the texture close-up.
One physics glitch there kills the whole ad, because viewers stare at those frames hunting the AI tell. Kling 3.0, $0.07/sec, no debate.
Maybe $6 of a whole production. Never save money here.
2. B-roll gets the billing arbitrage.
Some models bill per second, some bill PER VIDEO. Veo 3.1 Lite: about $0.175 for an 8-second 1080p clip.
That's $0.022/sec. Cheapest quality-per-second, because nobody does the division.
Workbench shots, driveway shots, scene-setting vignettes?
Nobody is pixel-peeping those.
Route them there.
We author b-roll beats at exactly 8 seconds to exploit the billing unit.
3. Talking heads only go to lipsync-capable models.
Most fake it. Two do it properly.
A silent mouth-flap is not a style, it's a broken file.
4. Sound costs extra. Pay it once.
Native audio is a +40% surcharge on some models. Pay it ONLY where the sound IS the appeal (the sizzle, the spray, the seal).
Everything else renders silent under one continuous voiceover.
Net effect: same perceived quality, half the spend, 5 hooks instead of 1, because hooks share the rendered body and only the opening clip changes.
A variant costs us $2-5. People pay editors $300 for what is a different first 8 seconds.
Now the part for people who don't generate video at all:
The highest-ROI step in our pipeline costs a dollar and needs zero rendering:
Before anything gets produced, the script gets torn apart by simulated buyers.
Not "does my team like it."
The skeptic burned by the last product. The spouse who controls the card. The 24-year-old who smells an ad in half a second.
Open-ended questions only:
- "retell this tomorrow, what survived?"
- "what did this assume about you that's wrong?"
- "at which second do you leave?"
Two rounds killed a line our whole team loved and rebuilt the entire opening.
Found after production, each of those is a full re-shoot. Found at script stage - a text edit.
You can do this today with any LLM and zero budget.
Most people won't.
It feels slower than creating.
It's the fastest thing we do.
Those "5 AI video prompts that go viral" threads? Recycled advice with the model name swapped. They age in days.
The questions underneath don't:
- Which shots carry the trust? Route premium.
- Which shots are wallpaper? Route cheap.
- What does the buyer retell tomorrow? That's your ad.
Stop asking which model is best. Roast the script before a single frame exists.
Every AI video model comparison on X compares vibes. Not one of them compares invoices.
So I pulled the actual per-second math on every current model, and the "cheap" models aren't cheap, the "premium" ones aren't premium, and the best deal on the market is hiding inside a billing unit nobody bothers to divide.
The full breakdown is in the screenshot and at the end of the post. Bookmark it if you'd like.
Now the three findings that make this table worth money:
1. Per-video billing is arbitrage.
Per-second models punish long clips. Per-video models reward maxing the clip length.
If your b-roll beats are authored at exactly 8 seconds, Veo Lite undercuts everything.
Structure your shot list around the billing unit, not the other way around.
2. The audio surcharge is a tax most people pay for nothing.
Native audio adds up to 40% per second on some models.
If you're laying one continuous voiceover over the footage anyway (you should be), you're paying for audio you delete.
Render silent.
Pay the surcharge only on the beat where the sound IS the ad.
3. There is no winner in this table. That's the point.
- The proof shot goes to Kling.
- The talking head goes to Sora or Seedance.
- The wallpaper goes to Mini.
- The 8 second b-roll goes to Veo Lite.
- The revision goes to Omni.
A single production should touch four of these models.
Anyone running everything through one model is either overpaying on 60% of their seconds or under-delivering on the 5% that carry the trust.
The model wars are content for spectators.
The billing table is content for operators.
Bookmark the table.
The prices will drift, the logic won't: normalize to per-second, route by shot, never pay for audio you'll replace, and let the per-video models subsidize your b-roll.
KLING 3.0
$0.07/sec silent, $0.10/sec with audio
15 sec max. The physics king. Weak lipsync (~22%).
SEEDANCE 2.0
$0.10/sec with video input, $0.165 cold
12 sec max. Real lipsync. Multi-shot consistency is its whole pitch.
SEEDANCE 2.0 MINI
$0.03 to $0.06/sec
The cost floor. Your background shots do not need more than this.
VEO 3.1 LITE
$0.175 PER VIDEO for an 8 sec 1080p clip.
Do the division: $0.022/sec. That's the cheapest quality-per-second in the entire market, and it's invisible because it's billed per clip instead of per second. This is the single most valuable row in this table.
HAILUO 2.3
$0.15 per 6 sec clip. About $0.025/sec. Same trick, second cheapest.
WAN 2.7
$0.08 to $0.12/sec. Solid middle. No audio.
SORA 2 PRO
$0.045/sec at 720p.
The only model that holds a 25 second continuous take, and the best lipsync available. Somehow also one of the cheapest. Nobody talks about this.
GEMINI OMNI
$0.84 per generation WITH video input. It's not a generator, it's an editor. Feed it a finished clip and change one element with a sentence, instead of re-rendering the whole scene. This is a different product category and almost nobody has noticed it shipped.
So I pulled the actual per-second math on every current model, and the "cheap" models aren't cheap, the "premium" ones aren't premium, and the best deal on the market is hiding inside a billing unit nobody bothers to divide.
The full breakdown is in the screenshot and at the end of the post. Bookmark it if you'd like.
Now the three findings that make this table worth money:
1. Per-video billing is arbitrage.
Per-second models punish long clips. Per-video models reward maxing the clip length.
If your b-roll beats are authored at exactly 8 seconds, Veo Lite undercuts everything.
Structure your shot list around the billing unit, not the other way around.
2. The audio surcharge is a tax most people pay for nothing.
Native audio adds up to 40% per second on some models.
If you're laying one continuous voiceover over the footage anyway (you should be), you're paying for audio you delete.
Render silent.
Pay the surcharge only on the beat where the sound IS the ad.
3. There is no winner in this table. That's the point.
- The proof shot goes to Kling.
- The talking head goes to Sora or Seedance.
- The wallpaper goes to Mini.
- The 8 second b-roll goes to Veo Lite.
- The revision goes to Omni.
A single production should touch four of these models.
Anyone running everything through one model is either overpaying on 60% of their seconds or under-delivering on the 5% that carry the trust.
The model wars are content for spectators.
The billing table is content for operators.
Bookmark the table.
The prices will drift, the logic won't: normalize to per-second, route by shot, never pay for audio you'll replace, and let the per-video models subsidize your b-roll.
KLING 3.0
$0.07/sec silent, $0.10/sec with audio
15 sec max. The physics king. Weak lipsync (~22%).
SEEDANCE 2.0
$0.10/sec with video input, $0.165 cold
12 sec max. Real lipsync. Multi-shot consistency is its whole pitch.
SEEDANCE 2.0 MINI
$0.03 to $0.06/sec
The cost floor. Your background shots do not need more than this.
VEO 3.1 LITE
$0.175 PER VIDEO for an 8 sec 1080p clip.
Do the division: $0.022/sec. That's the cheapest quality-per-second in the entire market, and it's invisible because it's billed per clip instead of per second. This is the single most valuable row in this table.
HAILUO 2.3
$0.15 per 6 sec clip. About $0.025/sec. Same trick, second cheapest.
WAN 2.7
$0.08 to $0.12/sec. Solid middle. No audio.
SORA 2 PRO
$0.045/sec at 720p.
The only model that holds a 25 second continuous take, and the best lipsync available. Somehow also one of the cheapest. Nobody talks about this.
GEMINI OMNI
$0.84 per generation WITH video input. It's not a generator, it's an editor. Feed it a finished clip and change one element with a sentence, instead of re-rendering the whole scene. This is a different product category and almost nobody has noticed it shipped.
Grüns - a gummy brand sold for $1.2B a few months ago.
And every food founder I know took exactly the wrong note.
The wrong note: "consumables are hot, we're a consumable, we're next."
The honest note: supplements are winning because of unit economics your snack brand does not have. 70-80% gross margins, tiny dim-weight shipping, subscription behavior the customer WANTS because running out feels like a health failure.
Your hot sauce runs 45-55% gross before shipping a glass bottle. Running out of hot sauce feels like a Tuesday.
Different physics. Copying the playbook without the margins is how food brands buy growth that eats them.
So here's the split - what actually transfers and what doesn't.
Copyable:
• The hero-SKU discipline. Every big supplement exit is one product with a routine attached, not a catalog. Food brands hoard SKUs like a pantry before a storm - cut to the one thing people reorder.
• The subscription OFFER, not the subscription assumption. Coffee earns it naturally. Snacks earn it with a bundle cadence ("the monthly box"), not a checkbox at checkout.
• The 90-day scoreboard. Supplement operators live on cohort value because first orders lose money. Food first orders lose money too - most founders just haven't done the math that proves it.
Not copyable:
• The margin that forgives mistakes. North of 70% gross you can misprice CAC for a quarter and live. At half that, the same mistake is the business.
• The health-anxiety reorder loop. Nobody panic-reorders chocolate.
Which means the food version of the playbook is stricter, not looser: tighter shipping breakpoints, bundles engineered to clear the free-shipping line, and a second order you design for instead of pray for.
Before you copy anyone's playbook, run the one-line check: gross margin after fulfillment, per unit, on your best seller.
If it starts with a 4, you're playing the harder game.
And every food founder I know took exactly the wrong note.
The wrong note: "consumables are hot, we're a consumable, we're next."
The honest note: supplements are winning because of unit economics your snack brand does not have. 70-80% gross margins, tiny dim-weight shipping, subscription behavior the customer WANTS because running out feels like a health failure.
Your hot sauce runs 45-55% gross before shipping a glass bottle. Running out of hot sauce feels like a Tuesday.
Different physics. Copying the playbook without the margins is how food brands buy growth that eats them.
So here's the split - what actually transfers and what doesn't.
Copyable:
• The hero-SKU discipline. Every big supplement exit is one product with a routine attached, not a catalog. Food brands hoard SKUs like a pantry before a storm - cut to the one thing people reorder.
• The subscription OFFER, not the subscription assumption. Coffee earns it naturally. Snacks earn it with a bundle cadence ("the monthly box"), not a checkbox at checkout.
• The 90-day scoreboard. Supplement operators live on cohort value because first orders lose money. Food first orders lose money too - most founders just haven't done the math that proves it.
Not copyable:
• The margin that forgives mistakes. North of 70% gross you can misprice CAC for a quarter and live. At half that, the same mistake is the business.
• The health-anxiety reorder loop. Nobody panic-reorders chocolate.
Which means the food version of the playbook is stricter, not looser: tighter shipping breakpoints, bundles engineered to clear the free-shipping line, and a second order you design for instead of pray for.
Before you copy anyone's playbook, run the one-line check: gross margin after fulfillment, per unit, on your best seller.
If it starts with a 4, you're playing the harder game.
60 hours of agent work in a day is now real. The scarce resource is a clean 12-minute review.
The number is OpenAI's, from its heaviest users. Three of us manage $10M+ a month across 20+ brands, so I can tell you where that stat stops being impressive and starts being dangerous.
One person can't physically work 60 hours before dinner.
One operator can review 60 hours of parallel work - if the system surfaces the right decisions.
That second sentence is doing all the work, and it's the part everyone skips.
Teams are still teaching people to write prettier prompts. Workshops, internal prompt libraries, certification decks.
It genuinely annoys me, because the operators actually compounding are designing the review receipt first.
One row from ours looks roughly like this:
Job: investigate a PMax spend anomaly
Evidence: change log, query delta, inventory state
Proposed action: staged, never live
Authority: Google owner approval required
Expiry: discard if the account state changes
Delegation design is the skill now. The model keeps working long after the chat window stops being interesting, and fluency in 2026 means deciding what deserves to run - not phrasing it nicely.
The economics follow. When one operator runs research, analysis, QA, and production in parallel, headcount stops being a clean proxy for capacity.
Ours has been stuck at three on purpose.
Founder check, one email: ask the agency you pay by headcount to show you one completed agent job with its evidence, proposed action, approver, and expiry. "We check everything manually" means 60 hours of unread homework, billed to you.
Your value moves toward judgment - choosing the job, setting the boundary, killing bad output before it touches money.
I still write prompts. I just spend most of the week deciding what deserves to run.
The number is OpenAI's, from its heaviest users. Three of us manage $10M+ a month across 20+ brands, so I can tell you where that stat stops being impressive and starts being dangerous.
One person can't physically work 60 hours before dinner.
One operator can review 60 hours of parallel work - if the system surfaces the right decisions.
That second sentence is doing all the work, and it's the part everyone skips.
Teams are still teaching people to write prettier prompts. Workshops, internal prompt libraries, certification decks.
It genuinely annoys me, because the operators actually compounding are designing the review receipt first.
One row from ours looks roughly like this:
Job: investigate a PMax spend anomaly
Evidence: change log, query delta, inventory state
Proposed action: staged, never live
Authority: Google owner approval required
Expiry: discard if the account state changes
Delegation design is the skill now. The model keeps working long after the chat window stops being interesting, and fluency in 2026 means deciding what deserves to run - not phrasing it nicely.
The economics follow. When one operator runs research, analysis, QA, and production in parallel, headcount stops being a clean proxy for capacity.
Ours has been stuck at three on purpose.
Founder check, one email: ask the agency you pay by headcount to show you one completed agent job with its evidence, proposed action, approver, and expiry. "We check everything manually" means 60 hours of unread homework, billed to you.
Your value moves toward judgment - choosing the job, setting the boundary, killing bad output before it touches money.
I still write prompts. I just spend most of the week deciding what deserves to run.
YouTube Ads Manager said 0.8x ROAS. The incrementality test said 2.4x.
The platform wasn't lying on purpose. It was answering the wrong question.
Last-click attribution gives credit to the final touch before purchase.
So here's the journey it never sees.
A customer watches your YouTube ad, doesn't click, googles your brand a day later, and buys through branded Search.
Search takes 100% of the credit. YouTube gets zero.
The demand was manufactured upstream. The accounting books it downstream.
I learned this the expensive way.
Pulled a Shopify brand builder's YouTube budget because the platform said it was dead weight at 0.8x.
Two weeks later branded search volume dropped 40%, retargeting pools dried up, and direct traffic tanked.
YouTube wasn't the closer.
It was starting 60-70% of the journeys that the other channels were finishing.
Killing it didn't save money. It quietly defunded every channel downstream.
Relaunched with proper measurement. True contribution came back at 2.4x.
There are three levels of attribution, and most brands only ever see the first:
• Level 1, platform last-click, what Ads Manager reports, shows 0.8x, where the panic and the budget cuts happen
• Level 2, blended, Northbeam or Triple Whale folding in branded search lift and direct, shows 2.4x
• Level 3, incrementality, geo-holdouts and lift studies, the only number that answers what happens if you turn it off
Across the tests I've seen, YouTube drives roughly 3.4x more conversions than the platform reports.
The gap between Level 1 and Level 3 is where every bad budget decision lives.
A geo-holdout costs a few hundred a month to run.
The misallocation it catches is measured in tens of thousands.
I'd rather spend $500 to find out the truth than cut $50K of upstream demand on a number that was structurally wrong before I ever read it.
The brands making budget calls on Level 1 think they're being cautious.
They're optimizing a P&L against fiction and calling it discipline.
I check Level 2 before I touch any awareness channel's budget now, because the cut I almost made on that brand would have been the most expensive thing I did all quarter.
The platform wasn't lying on purpose. It was answering the wrong question.
Last-click attribution gives credit to the final touch before purchase.
So here's the journey it never sees.
A customer watches your YouTube ad, doesn't click, googles your brand a day later, and buys through branded Search.
Search takes 100% of the credit. YouTube gets zero.
The demand was manufactured upstream. The accounting books it downstream.
I learned this the expensive way.
Pulled a Shopify brand builder's YouTube budget because the platform said it was dead weight at 0.8x.
Two weeks later branded search volume dropped 40%, retargeting pools dried up, and direct traffic tanked.
YouTube wasn't the closer.
It was starting 60-70% of the journeys that the other channels were finishing.
Killing it didn't save money. It quietly defunded every channel downstream.
Relaunched with proper measurement. True contribution came back at 2.4x.
There are three levels of attribution, and most brands only ever see the first:
• Level 1, platform last-click, what Ads Manager reports, shows 0.8x, where the panic and the budget cuts happen
• Level 2, blended, Northbeam or Triple Whale folding in branded search lift and direct, shows 2.4x
• Level 3, incrementality, geo-holdouts and lift studies, the only number that answers what happens if you turn it off
Across the tests I've seen, YouTube drives roughly 3.4x more conversions than the platform reports.
The gap between Level 1 and Level 3 is where every bad budget decision lives.
A geo-holdout costs a few hundred a month to run.
The misallocation it catches is measured in tens of thousands.
I'd rather spend $500 to find out the truth than cut $50K of upstream demand on a number that was structurally wrong before I ever read it.
The brands making budget calls on Level 1 think they're being cautious.
They're optimizing a P&L against fiction and calling it discipline.
I check Level 2 before I touch any awareness channel's budget now, because the cut I almost made on that brand would have been the most expensive thing I did all quarter.
❤1
The same 30-second AI video ad costs $2.10 or $9.40 to render.
Same script. Same visual quality. The only difference is knowing which model renders which shot. Worth adding to your bookmarks.
We just spent a week reverse-engineering every current video model - Veo 3.1, Kling 3.0, Seedance 2.0, Sora 2 Pro, Grok 1.5, Wan 2.7, Hailuo 2.3, Gemini Omni - against their actual API pricing. Not the blog-post pricing. The real request schemas and rate cards.
Here's what nobody tells you:
1. "Which model is best?" is the wrong question.
A performance ad is not one video. It's 8-12 shots with completely different jobs:
- Proof shots (the product actually working): ~5% of your seconds, ~80% of your persuasion
- Talking head: the only place lip sync matters
- B-roll: nobody scrutinizes the workbench shot
- Filler: 1-3 second cuts between beats
Route each shot class to a different model and your cost drops 40-60% with zero quality loss where it counts.
2. Minimum billable duration is the hidden tax.
Every model has a billing floor, and they never advertise it:
- Wan 2.7 bills a true 2 seconds
- Kling 3.0 floors at 3s
- Seedance 2.0 floors at 4s
- Grok floors at 6s
- Sora on API floors at 10s
That "cheap" $0.045/sec model? A 4-second talking head bills 10 seconds. You paid $0.11/sec and never noticed.
For fast-cut ads (1-3s per shot), this one number matters more than the per-second price.
3. Per-video billing is an arbitrage.
Some tiers bill per clip, not per second. Veo 3.1 Lite: ~$0.17 for an 8-second 1080p video.
Author your b-roll AT 8 seconds, harvest 3 usable cuts from each clip, and you're paying ~$0.05 per cut. Cheaper than every per-second model on the market.
4. Lip sync is binary. Enforce it in code.
Only 2 of the 9 models can render a person speaking with usable lips. Everything else produces the uncanny half-sync that instantly reads as AI.
Our pipeline literally refuses to render a talking head on a non-lipsync model. One hard rule, zero embarrassing ads.
5. The #1 leaderboard model shouldn't render your ads.
Gemini Omni is #1 on blind arena Elo right now. We still don't generate with it:
- 3x the cost of the right model per shot
- hard-blocks prompts containing real brand names
- quality degrades after 4 sequential edits (we tested)
But as an EDITOR it's unbeatable: $0.84 to change one element in an existing clip vs $1+ to re-render the scene. Next-gen model, wrong job description.
6. Never trust the docs. Render one probe.
We ran a $0.09 test render before trusting any of this. Found two things no documentation mentions: one model returned a 10-second clip for a 6-second request, and another attaches a silent audio track that breaks naive pipelines.
One coffee's worth of API credits beats a week of confident assumptions.
---
The full routing table is in the screenshot.
Steal it, wire it into your pipeline, and stop paying premium rates for shots nobody looks at.
Same script. Same visual quality. The only difference is knowing which model renders which shot. Worth adding to your bookmarks.
We just spent a week reverse-engineering every current video model - Veo 3.1, Kling 3.0, Seedance 2.0, Sora 2 Pro, Grok 1.5, Wan 2.7, Hailuo 2.3, Gemini Omni - against their actual API pricing. Not the blog-post pricing. The real request schemas and rate cards.
Here's what nobody tells you:
1. "Which model is best?" is the wrong question.
A performance ad is not one video. It's 8-12 shots with completely different jobs:
- Proof shots (the product actually working): ~5% of your seconds, ~80% of your persuasion
- Talking head: the only place lip sync matters
- B-roll: nobody scrutinizes the workbench shot
- Filler: 1-3 second cuts between beats
Route each shot class to a different model and your cost drops 40-60% with zero quality loss where it counts.
2. Minimum billable duration is the hidden tax.
Every model has a billing floor, and they never advertise it:
- Wan 2.7 bills a true 2 seconds
- Kling 3.0 floors at 3s
- Seedance 2.0 floors at 4s
- Grok floors at 6s
- Sora on API floors at 10s
That "cheap" $0.045/sec model? A 4-second talking head bills 10 seconds. You paid $0.11/sec and never noticed.
For fast-cut ads (1-3s per shot), this one number matters more than the per-second price.
3. Per-video billing is an arbitrage.
Some tiers bill per clip, not per second. Veo 3.1 Lite: ~$0.17 for an 8-second 1080p video.
Author your b-roll AT 8 seconds, harvest 3 usable cuts from each clip, and you're paying ~$0.05 per cut. Cheaper than every per-second model on the market.
4. Lip sync is binary. Enforce it in code.
Only 2 of the 9 models can render a person speaking with usable lips. Everything else produces the uncanny half-sync that instantly reads as AI.
Our pipeline literally refuses to render a talking head on a non-lipsync model. One hard rule, zero embarrassing ads.
5. The #1 leaderboard model shouldn't render your ads.
Gemini Omni is #1 on blind arena Elo right now. We still don't generate with it:
- 3x the cost of the right model per shot
- hard-blocks prompts containing real brand names
- quality degrades after 4 sequential edits (we tested)
But as an EDITOR it's unbeatable: $0.84 to change one element in an existing clip vs $1+ to re-render the scene. Next-gen model, wrong job description.
6. Never trust the docs. Render one probe.
We ran a $0.09 test render before trusting any of this. Found two things no documentation mentions: one model returned a 10-second clip for a 6-second request, and another attaches a silent audio track that breaks naive pipelines.
One coffee's worth of API credits beats a week of confident assumptions.
---
The full routing table is in the screenshot.
Steal it, wire it into your pipeline, and stop paying premium rates for shots nobody looks at.
I ran the same cold creative on Meta Reels and YouTube Shorts. The CPM gap was 2x.
Most Shopify brands still run Shorts as a brand play.
The numbers say it prospects.
Same product, same audience signal. Here is what came back.
YouTube Shorts:
• 2.1% CTR
• $4.80 CPM
• 30+ seconds watched before the click
Meta Reels:
• 1.4% CTR
• $11.20 CPM
• 3 seconds of passive scroll
12B+ daily views on Shorts, CPMs 60-70% below Meta Reels.
The audience skews buyer, not browser.
Meta is the default cold channel for most DTC brands, so the comparison rarely gets run.
Demand Gen on YouTube reaches buyers who haven't searched for the brand, on the platform where they're already watching.
That's targeted interruption you can measure down to the purchase.
The catch, so nobody reads this as free money: Shorts burns creative faster than Meta does. The CPM edge pays for the extra production volume, or it doesn't pay at all.
If an agency runs your ads and has never shown you a Meta-vs-Shorts CPM test on your own creative, that's a one-line email to send today.
For this brand, cold traffic moved to YouTube.
Most Shopify brands still run Shorts as a brand play.
The numbers say it prospects.
Same product, same audience signal. Here is what came back.
YouTube Shorts:
• 2.1% CTR
• $4.80 CPM
• 30+ seconds watched before the click
Meta Reels:
• 1.4% CTR
• $11.20 CPM
• 3 seconds of passive scroll
12B+ daily views on Shorts, CPMs 60-70% below Meta Reels.
The audience skews buyer, not browser.
Meta is the default cold channel for most DTC brands, so the comparison rarely gets run.
Demand Gen on YouTube reaches buyers who haven't searched for the brand, on the platform where they're already watching.
That's targeted interruption you can measure down to the purchase.
The catch, so nobody reads this as free money: Shorts burns creative faster than Meta does. The CPM edge pays for the extra production volume, or it doesn't pay at all.
If an agency runs your ads and has never shown you a Meta-vs-Shorts CPM test on your own creative, that's a one-line email to send today.
For this brand, cold traffic moved to YouTube.
Everyone's selling "an AI media buyer that works while you sleep" this month. Nobody shows you what it actually finds.
Here's what ours flags in apparel accounts, ranked by how often we find it - and none of it is bid magic.
1. Dark variants. Size-level availability that never synced, so a third of the catalog reads out-of-stock to Google while it sits in the warehouse. The algorithm doesn't bid on inventory it thinks is dead. This is the single most common finding, and no human checks it weekly because it's boring.
2. Price mismatches on markdown. The site says 40% off, the feed still says full price. Google reads the click-to-page gap and disapproves the item - or its automatic updates rewrite your feed price without asking. Humans catch either in week three of the sale. An agent catches it the first morning.
3. Returns missing from conversion values. The account optimizes on revenue, the brand keeps 60-75% of it after returns, and the bidder happily scales the size-guessers. Wiring refunds back in is a one-time job that most accounts never do because nobody owns it.
The pattern across all three: the machine isn't smarter than your media buyer. It's more willing to do the unglamorous checks every single day.
One rule the "while you sleep" pitches skip: everything it drafts waits for a human signature before it spends.
The dark-variant check: open Merchant Center, filter to out-of-stock, and count how many of those sizes your warehouse says it has.
Takes ten minutes. In apparel accounts we open, the answer is rarely zero - and it's the cheapest revenue you'll recover this quarter.
Here's what ours flags in apparel accounts, ranked by how often we find it - and none of it is bid magic.
1. Dark variants. Size-level availability that never synced, so a third of the catalog reads out-of-stock to Google while it sits in the warehouse. The algorithm doesn't bid on inventory it thinks is dead. This is the single most common finding, and no human checks it weekly because it's boring.
2. Price mismatches on markdown. The site says 40% off, the feed still says full price. Google reads the click-to-page gap and disapproves the item - or its automatic updates rewrite your feed price without asking. Humans catch either in week three of the sale. An agent catches it the first morning.
3. Returns missing from conversion values. The account optimizes on revenue, the brand keeps 60-75% of it after returns, and the bidder happily scales the size-guessers. Wiring refunds back in is a one-time job that most accounts never do because nobody owns it.
The pattern across all three: the machine isn't smarter than your media buyer. It's more willing to do the unglamorous checks every single day.
One rule the "while you sleep" pitches skip: everything it drafts waits for a human signature before it spends.
The dark-variant check: open Merchant Center, filter to out-of-stock, and count how many of those sizes your warehouse says it has.
Takes ten minutes. In apparel accounts we open, the answer is rarely zero - and it's the cheapest revenue you'll recover this quarter.
Marketers keep posting the same thing: long education pages are outperforming product pages in skeptical categories.
Beauty first among them.
The discovery is real. The explanation usually stops at "it works." Here's the mechanism, because you need it to build one that survives review.
A standard PDP assumes trust and asks for money. But the beauty buyer has been burned by a decade of miracle serums - her default read of your product photo and five stars is "another one."
The education page works because it spends its first half earning the belief the PDP takes for granted.
The structure, five sections in order:
1. The problem, explained at the mechanism level. Not "dull skin" - what's actually happening and why it resists what she's tried.
2. Why the usual fixes disappoint. This is the section that buys trust, because it explains her own failed purchases back to her better than she could.
3. Your ingredient logic as evidence, not adjectives. Concentrations, the reason for each active, what the formulation deliberately leaves out.
4. The offer, arriving after belief instead of before it.
5. Objections, answered plainly - sensitivity, timeline to results, what it won't do.
Section five deserves special respect in this category: every sentence about results is an efficacy claim, and "visibly firmer in 14 days" needs a substantiation file behind it, not a copywriter's confidence.
The education format tempts you to say more. Regulators read the whole page.
Run the arithmetic on your own funnel: count how many words a first-time visitor gets between your ad and your buy button. Under 200, and you're asking a skeptic for trust you haven't built.
The limit, stated up front so you can plan for it: these pages take real work - research, drafting, a legal read - and they lose to the PDP for buyers who already know your brand.
Route cold traffic through education and warm traffic straight to product. Building one page for both audiences is how you get neither.
Beauty first among them.
The discovery is real. The explanation usually stops at "it works." Here's the mechanism, because you need it to build one that survives review.
A standard PDP assumes trust and asks for money. But the beauty buyer has been burned by a decade of miracle serums - her default read of your product photo and five stars is "another one."
The education page works because it spends its first half earning the belief the PDP takes for granted.
The structure, five sections in order:
1. The problem, explained at the mechanism level. Not "dull skin" - what's actually happening and why it resists what she's tried.
2. Why the usual fixes disappoint. This is the section that buys trust, because it explains her own failed purchases back to her better than she could.
3. Your ingredient logic as evidence, not adjectives. Concentrations, the reason for each active, what the formulation deliberately leaves out.
4. The offer, arriving after belief instead of before it.
5. Objections, answered plainly - sensitivity, timeline to results, what it won't do.
Section five deserves special respect in this category: every sentence about results is an efficacy claim, and "visibly firmer in 14 days" needs a substantiation file behind it, not a copywriter's confidence.
The education format tempts you to say more. Regulators read the whole page.
Run the arithmetic on your own funnel: count how many words a first-time visitor gets between your ad and your buy button. Under 200, and you're asking a skeptic for trust you haven't built.
The limit, stated up front so you can plan for it: these pages take real work - research, drafting, a legal read - and they lose to the PDP for buyers who already know your brand.
Route cold traffic through education and warm traffic straight to product. Building one page for both audiences is how you get neither.
Thirty-one entries from the claims blocklist our AI check runs on supplement ads.
Severities included, so you can steal the structure, not just the words. Bookmark it.
Severity 1 - blocked outright, no rewrite, no review. Disease claims, whatever verb they wear:
• cures
• treats
• mitigates (the statute's forgotten verb - it's in the same sentence as the other four)
• heals
• reverses
• prevents
• eliminates
• remedies
• diagnoses
• "fights [any disease]"
• "kills [pathogen/cell]"
• "beats depression" - a named disease puts the sentence here no matter how soft the verb
• any comparison to an Rx ("like the injections, without the needle") - an implied drug claim, and the exact pattern in the current warning-letter wave. No human review clears this one for a supplement.
Severity 2 - blocked pending rewrite. No disease named, but the sentence promises to fix a departure from normal:
• manages blood sugar
• lowers cholesterol
• reduces inflammation
• relieves anxiety (reads as diagnosis more often than not - when it does, it's severity 1)
• balances hormones
• restores testosterone
• fixes gut issues
• ends sleepless nights
• "clinically proven to [anything]" - the odd one out: you can't rewrite your way past this. Produce the studies or delete the words.
Severity 3 - routed to a human. Context decides, and pretending a regex can rule on these is how linters get people sued:
• supports / maintains / promotes (fine WITH normal-range framing, claim without it)
• customer testimonials with outcome numbers ("I lost 30 lbs" - publishing it makes it YOUR claim)
• before/after phrasing, even without images
• "without a prescription"
• "doctor-formulated" and credential adjacency
• money-back-if-it-works framings (implies a treatment outcome)
• "safe" as an absolute
• dosages positioned as protocols
• anything in second person about the reader's diagnosis
Three notes before you copy it.
This list handles claim classification - what your sentence says. Substantiation is a second lock on the same gate: a perfectly worded "supports" claim with no evidence file behind it is still an FTC problem.
The linter tells you which claims to defend, not whether you can.
The platform's gate is separate again - restricted concepts are their own list and their own fight.
And a verb list is only half a linter. "Formulated for people with type 2 diabetes" contains zero blocked verbs and names a disease - the other half is a disease-name list, and it's longer than this one.
To use it tonight: paste your five live ads and your landing page into a doc and search it against severities 1 and 2. Anything that hits severity 3, a human reads before it runs again.
Severities included, so you can steal the structure, not just the words. Bookmark it.
Severity 1 - blocked outright, no rewrite, no review. Disease claims, whatever verb they wear:
• cures
• treats
• mitigates (the statute's forgotten verb - it's in the same sentence as the other four)
• heals
• reverses
• prevents
• eliminates
• remedies
• diagnoses
• "fights [any disease]"
• "kills [pathogen/cell]"
• "beats depression" - a named disease puts the sentence here no matter how soft the verb
• any comparison to an Rx ("like the injections, without the needle") - an implied drug claim, and the exact pattern in the current warning-letter wave. No human review clears this one for a supplement.
Severity 2 - blocked pending rewrite. No disease named, but the sentence promises to fix a departure from normal:
• manages blood sugar
• lowers cholesterol
• reduces inflammation
• relieves anxiety (reads as diagnosis more often than not - when it does, it's severity 1)
• balances hormones
• restores testosterone
• fixes gut issues
• ends sleepless nights
• "clinically proven to [anything]" - the odd one out: you can't rewrite your way past this. Produce the studies or delete the words.
Severity 3 - routed to a human. Context decides, and pretending a regex can rule on these is how linters get people sued:
• supports / maintains / promotes (fine WITH normal-range framing, claim without it)
• customer testimonials with outcome numbers ("I lost 30 lbs" - publishing it makes it YOUR claim)
• before/after phrasing, even without images
• "without a prescription"
• "doctor-formulated" and credential adjacency
• money-back-if-it-works framings (implies a treatment outcome)
• "safe" as an absolute
• dosages positioned as protocols
• anything in second person about the reader's diagnosis
Three notes before you copy it.
This list handles claim classification - what your sentence says. Substantiation is a second lock on the same gate: a perfectly worded "supports" claim with no evidence file behind it is still an FTC problem.
The linter tells you which claims to defend, not whether you can.
The platform's gate is separate again - restricted concepts are their own list and their own fight.
And a verb list is only half a linter. "Formulated for people with type 2 diabetes" contains zero blocked verbs and names a disease - the other half is a disease-name list, and it's longer than this one.
To use it tonight: paste your five live ads and your landing page into a doc and search it against severities 1 and 2. Anything that hits severity 3, a human reads before it runs again.
AI got 14,430 of my dictated messages this year. My coworkers got 58.
Screenshot attached, my dictation app keeps receipts. 1,455,560 words spoken at 158 a minute. The ratio looks like bragging until you open it up, so here's what those 14,430 messages actually said, and why the voice part is what nobody copies. Bookmark this and feed it to your AI agent.
What a work order sounds like in the ad accounts we run (20+ brands, three of us):
1. Feeds: "pull last month's search terms for the top 40 SKUs and rewrite every title that doesn't contain what buyers actually type." The agent returns a rewrite table. I review diffs, not drafts.
2. Builds: "campaign skeleton for the spring launch, same structure as the one that held 2.2x, negatives preloaded from the master list."
3. Creative: "20 headline variants against the returns-eat-your-ROAS angle. Keep the claim, rotate the tension."
4. Teardowns: "every landing page this competitor runs, and what offer sits above the fold on each."
5. Anomalies: "why did Tuesday's spend spike. Check budgets first, then auction, then feed."
None of that gets typed. I call it the eye-voice split: eyes review, voice dispatches. Hands are for espresso.
The dispatch rules, since this is the part worth saving:
• Outcomes, not steps. I say "find where the money leaked last week," not ten micro-instructions. Agents are excellent at HOW and starving for WHAT.
• Ramble on purpose. I dictate the whole messy client-call context straight into the agent. Rambling is bad writing and elite context, the model compresses better than I summarize.
• Fire at 60%, steer mid-flight. "Narrower. US only. Ignore brand terms." Voice makes iteration nearly free, typing made me precious about prompts.
• One lane per agent, all lanes at once. Feed agent, creative agent, audit agent, running parallel like air traffic control, none of them waits for another to land.
• Review is the job now. Everything returns as a diff or a draft, so the human hour moved to the end of the pipe. I budget it like ad spend, because unreviewed volume is just expensive noise.
The compounding is the uncomfortable part. Speaking is 3x typing speed, and the work between prompts is free. One operator dispatching work orders ships more ad iterations before lunch than a typing team ships in a week, and volume finds winners in ads whether anyone likes it or not. The keyboard was the moat, and it drained.
Founder check: ask whoever runs your marketing what share of their day is instructions to software that does the work, versus meetings about the work. That ratio is the roadmap, and it's checkable today.
Screenshot attached, my dictation app keeps receipts. 1,455,560 words spoken at 158 a minute. The ratio looks like bragging until you open it up, so here's what those 14,430 messages actually said, and why the voice part is what nobody copies. Bookmark this and feed it to your AI agent.
What a work order sounds like in the ad accounts we run (20+ brands, three of us):
1. Feeds: "pull last month's search terms for the top 40 SKUs and rewrite every title that doesn't contain what buyers actually type." The agent returns a rewrite table. I review diffs, not drafts.
2. Builds: "campaign skeleton for the spring launch, same structure as the one that held 2.2x, negatives preloaded from the master list."
3. Creative: "20 headline variants against the returns-eat-your-ROAS angle. Keep the claim, rotate the tension."
4. Teardowns: "every landing page this competitor runs, and what offer sits above the fold on each."
5. Anomalies: "why did Tuesday's spend spike. Check budgets first, then auction, then feed."
None of that gets typed. I call it the eye-voice split: eyes review, voice dispatches. Hands are for espresso.
The dispatch rules, since this is the part worth saving:
• Outcomes, not steps. I say "find where the money leaked last week," not ten micro-instructions. Agents are excellent at HOW and starving for WHAT.
• Ramble on purpose. I dictate the whole messy client-call context straight into the agent. Rambling is bad writing and elite context, the model compresses better than I summarize.
• Fire at 60%, steer mid-flight. "Narrower. US only. Ignore brand terms." Voice makes iteration nearly free, typing made me precious about prompts.
• One lane per agent, all lanes at once. Feed agent, creative agent, audit agent, running parallel like air traffic control, none of them waits for another to land.
• Review is the job now. Everything returns as a diff or a draft, so the human hour moved to the end of the pipe. I budget it like ad spend, because unreviewed volume is just expensive noise.
The compounding is the uncomfortable part. Speaking is 3x typing speed, and the work between prompts is free. One operator dispatching work orders ships more ad iterations before lunch than a typing team ships in a week, and volume finds winners in ads whether anyone likes it or not. The keyboard was the moat, and it drained.
Founder check: ask whoever runs your marketing what share of their day is instructions to software that does the work, versus meetings about the work. That ratio is the roadmap, and it's checkable today.