Every operator scaling a venture eventually meets the same wall. Targeting is optimized. Budgets are reallocated weekly. Landing pages are tested to the second decimal. And then there is creative — the single largest driver of performance — managed almost entirely by instinct, taste, and the loudest voice in the room.
The reason is not laziness. It is that creative has been genuinely hard to measure. A targeting choice is a number. A budget is a number. A thirty-second video is a thousand decisions — a hook, a cut rhythm, a face or no face, a claim, a caption, a sound — none of which has been written down in a form a system can reason about. So creative stays a matter of opinion, and opinion does not compound.
We set out to test whether that is still true. We built an instrument to read the craft of a short-form video — not its account, not its spend, but the choices inside the frame — and we asked the question that decides whether creative can be treated as a discipline rather than a guess:
Does what makes a video work transfer across platforms, or is every platform its own black box?
This report is what we found across 150,000 videos and 2.5 million comments spanning TikTok, Instagram, and YouTube. The short version: the grammar of effective short-form is real, it is measurable, and it is largely the same grammar everywhere. The longer version is below.
The lever everyone optimizes around
In paid and organic growth alike, performance is a product of three things: who sees the work (distribution), how much you spend to reach them (budget), and whether the work itself lands (creative). The first two are instrumented to death. The third is where the leverage has quietly moved — platforms now do most of the targeting for you, which means the creative is increasingly the variable you actually control.
And yet creative is where measurement stops. Teams A/B two thumbnails and call it creative testing. They keep a folder of "winners" no one can query. They re-learn, every quarter, lessons a departing colleague already knew. The work is fast; the memory of the work is nonexistent. A company that re-learns the same lesson twice is burning the most expensive resource it has.
To change that, creative has to become data — not a vibe, but a set of named, comparable attributes you can count, correlate, and carry forward.
How we built a measurement instrument
The core of the system is a structured way to read a video's craft — a fixed, shared vocabulary of creative attributes that captures what is on screen and in the ear (who appears, the setting, the on-screen text, the role of music), how the video is built (its shot rhythm and production style), how it persuades (its hook, its narrative, its claim), and what it is trying to make you feel. Each attribute is defined precisely enough that two analysts — or two models — read the same video the same way.
We then ran every video through a deterministic pipeline (shot detection, transcription, on-screen-text extraction, audio understanding, frame embeddings) to ground the read, and a vision-language model assigned the attributes from that grounded evidence. The output is, for each video, a vector of creative attributes that is directly comparable to every other video — on any platform.
Crucially, we paired every creative vector with the video's engagement (likes, comments, plays) and acquired the corpus in a reach-normalized way, so that a big account's mediocre video and a small account's breakout are measured on the same footing.
Finding 1: Creative predicts engagement — not reach
The first result is the one that reframes everything that follows. When we ask the creative attributes to predict whether a video went viral by reach — raw views relative to the account's size — they perform at chance. They cannot. And they should not: reach at the top is dominated by distribution and luck. The algorithm decides who sees it; a perfect video shown to no one still reaches no one.
But when we ask the same attributes to predict engagement quality — whether a video lands in the top quartile of like-rate or comment-rate among its peers — the signal is strong and consistent. Across the corpus, creative attributes separate top-quartile engagement with an area-under-the-curve (AUC) in the 0.76–0.91 range, where 0.5 is a coin flip.
Creative is not how you win the lottery of reach. It is how you convert the attention you are given into a reaction. That is the part craft controls — and the part this instrument measures.
This distinction matters because it tells operators what creative is for. Stop crediting the hook for the view count. Start crediting it for the save, the comment, the share — the reactions that signal the work connected, and that platforms increasingly use to decide who sees it next.
Finding 2: The grammar is platform-portable
Here is the question we most wanted to answer. We took the model trained to read creative on one platform and asked whether the same creative→engagement relationship holds on the others. If creative were platform-specific, the signal would collapse off its home turf, and every platform would need its own playbook, its own model, its own team.
It does not collapse. The creative→engagement signal holds on all three platforms, at comparable or better strength than the baseline:
| Engagement signal | TikTok | YouTube | |
|---|---|---|---|
| Like-rate (top-quartile AUC) | 0.81 | 0.88 | 0.77 |
| Comment-rate (top-quartile AUC) | 0.76 | 0.77 | 0.85 |
| Combined like + comment | 0.80 | 0.91 | 0.79 |
Every cell is well above chance. Instagram is, if anything, the most legible to the model; YouTube comments are the most creative-predictable of all. The same vocabulary of craft that explains a TikTok explains a Reel and a Short.
The implication is direct: you do not need a different theory of creative for each platform. You need one creative DNA, read consistently, and a small amount of platform-specific tuning at the edges.
That is the difference between a capability that scales and one that fragments. A portable grammar means a winning insight discovered on one surface is an asset everywhere — the opposite of the per-platform silos most teams live in.
Finding 3: The same grammar, used the same way
A skeptic's first objection is that the model might be "generalizing" only because it is producing garbage everywhere — flat, undifferentiated tags that happen to correlate by accident. So we looked at the distributions of the creative attributes across platforms. If the grammar is real, the same attributes should discriminate in the same way on each platform, with differences that reflect genuine content mix rather than noise.
That is what we see. The structural attributes are strikingly coherent across platforms:
| Attribute (dominant value) | TikTok | YouTube | |
|---|---|---|---|
| Cast is a real person | 80% | 81% | 83% |
| No commerce call-to-action | 97% | 97% | 98% |
| Audio is music-led | 76% | 80% | 61% |
| Beat function: demonstration | 43% | 46% | 41% |
No attribute is pinned at 100% (the signature of a broken, non-discriminating read). The one real divergence is content mix: the YouTube sample skews toward health and fitness, so its claim and audio profiles shift accordingly — a fact about the corpus, not a failure of the instrument. The creative grammar is being spoken the same way on every platform; the platforms are simply saying somewhat different things with it.
Finding 4: What audiences actually do — and where it diverges
Creative attributes describe the work. To validate that the read is real — and to extract the part operators care about commercially — we also analyzed more than 2.5 million comments across the three platforms, extracting the audience's actual reaction: sentiment, whether the hook landed, expressed purchase intent, and the recurring themes, objections, and points of resonance.
The audience-response signal generalizes — the hook lands for the majority of viewers on every platform, and "delight" is the dominant observed emotion everywhere. But the commercial motion diverges sharply, and that divergence is itself a finding:
| Audience signal | TikTok | YouTube | |
|---|---|---|---|
| Positive sentiment | 0.53 | 0.61 | 0.64 |
| Hook landed | 64% | 63% | 57% |
| Expressed purchase intent | 0.09 | 0.15 | 0.18 |
| "Tag a friend" (sharing) | 3.2% | 3.8% | 0.1% |
Two platforms, two jobs. TikTok and Instagram run on social sharing — audiences tag friends to spread the work, which is how reach compounds. YouTube runs on intent — its comments carry roughly twice the expressed purchase signal and almost none of the friend-tagging. The same creative can be excellent on both, but the operator should harvest it differently: TikTok and Instagram for amplification, YouTube for capture.
The recurring resonance across all three platforms — the verbatim audience language we surface as evidence — clusters into three reusable seeds: the impulse to share ("tag a friend"), the shock of recognition ("this is so me"), and creator connection. Those are not abstractions; they are the literal hook copy and positioning truths the next brief should start from.
Finding 5: Predicted feeling versus felt feeling
The most demanding test of a creative-reading instrument is whether its prediction of intended emotional effect matches the audience's actual emotional reaction. We compared the emotion our model predicted from the craft alone against the emotion we observed in the comments. The predicted and observed emotion matched closely across platforms — far above the chance rate for a twelve-way emotional classification, and higher still when scored at the level of emotional valence rather than the exact label.
That is a real, consistent signal: the model is reading intended emotional effect from craft well enough to anticipate how an audience will feel, before a single view is bought. Emotional design stops being post-hoc rationalization and becomes a property you can set on purpose.
What this changes for operators
Put the findings together and creative stops being the last unmeasured lever:
- Creative becomes a measurable asset. Every video resolves to a comparable vector of craft attributes tied to real outcomes. "Our hook is the problem" becomes a number, not an argument.
- Insight becomes portable. Because the grammar transfers, a pattern proven on one platform is an asset on all of them — one creative DNA, not three disconnected playbooks.
- Memory compounds. The winning and losing patterns accumulate in a queryable library instead of a folder of files and a few people's memory. The work of remembering finally scales with the work of shipping.
- Emotion becomes intentional. Intended feeling can be specified up front and checked against how audiences actually react.
This is the part of the execution loop creative has always been missing: not just make the work, but read it, predict it, test it, and remember what worked — so the next brief starts ahead of where the last one ended.
The honest limits
A study earns its conclusions by being precise about what it has not shown — and about what it deliberately holds back.
- This is correlation, not a controlled experiment. The creative→engagement relationships are strong and consistent, but this public report is observational: it shows which patterns travel with engagement. Turning that into a guaranteed "change X and win" is the job of the causal, holdout-tested layer — part of the full analysis Lyberty runs for its users, not this public summary.
- Content mix is uneven. Across the platforms, verticals are not perfectly balanced; the YouTube health-and-fitness skew is the clearest example. The grammar holds across the imbalance, and per-segment precision is exactly what the full per-industry corpus sharpens.
- Comments are an outcome, never an input. We use audience reaction to validate the creative read and to build the go-to-market payload — never to tag the creative itself. Letting an outcome leak into the features would manufacture a signal that does not exist. That discipline is deliberate, and it is load-bearing.
- This is a limited public report. The figures here are faithful but deliberately partial. The full corpus, the per-industry breakdowns, the causal results, and the pre-flight prediction are reserved for Lyberty users.
None of these caveats dents the headline. They mark the line between what we publish and what we put in our users' hands.
This is the public report. The full corpus is reserved to Lyberty users.
Everything above came from 150,000 videos and 2.5 million comments — a serious corpus, and a deliberately limited public slice of the full Lyberty creative-intelligence system, which runs on more than a million videos, continuously refreshed. What we publish here is the proof that the grammar is real and portable. What we reserve for Lyberty users is the edge that comes from reading it at full scale:
- Per-industry, per-format precision. At full scale, "what works" stops being a global average and becomes specific to your category, your format, your audience — the difference between "hooks matter" and "for a DTC supplement brand on Reels, a result-first hook in the first 0.8 seconds lifts save-rate by a measurable margin."
- Causal answers, not just correlations. With enough creators contributing enough videos, the within-creator test — comparing a creator's own videos against each other to strip out the "good creators use good hooks" confound — becomes decisive, turning "this pattern travels with engagement" into "changing this will move your number."
- A living creative DNA library. A continuously-updated corpus means the grammar tracks the platforms as they shift — new formats, new sounds, new hooks — instead of freezing into last quarter's playbook.
- Pre-flight prediction. A creative scored against the full corpus can be evaluated before it ships — predicted engagement, predicted emotional response, predicted objections in the audience's own words — so spend follows evidence, not hope.
- A closed execution loop. Generate, predict, ship, measure against the ledger, and feed the result back into the DNA. Creative becomes the one part of growth that gets more certain every quarter instead of starting from zero.
That is the Execution Standard applied to the lever everyone optimizes around and no one has measured. This public report shows the grammar is real and portable. The full corpus — and the per-industry, causal, predictive analysis built on it — is what Lyberty puts in our users' hands, so that every video your team ships makes the next one better.
The figures in this report summarize Lyberty's 2026 cross-platform creative-intelligence analysis (150,000 videos across TikTok, Instagram, and YouTube; 2.5 million comments). They are observational and presented for research purposes; they are not a performance guarantee. This is a public summary — the full corpus and the analysis built on it are reserved to Lyberty users. Platform metrics and audience behavior change over time.