You spent two days generating the visuals for one video. Forty-one images, a thumbnail you were proud of, every shot approved by you personally.
Then it underperforms, and you go looking for the reason in your topic.
A creator on r/SmallYoutubers asked exactly that question — why won't this channel grow — and the first answer was not about his subject at all. Those 200% look AI generated at first glance, someone replied, and even if it was an interesting one I'd skip it and never watch. A second reply named the tells outright: some of the thumbnails kind of look like AI and a lot of people will think it's low effort and click off, maybe try simplifying the glow effects and use a different font.
Note what did not happen. That viewer never commented on the video. He scrolled. Which is why the creator was still investigating his topics.
Why can't you see your own tells?
Because you have been looking at these images for two days, and the person who decides has been looking for four tenths of a second.
That gap is the entire problem. By the time you export, you have approved every frame individually, adjusted some of them, and rerolled others. Each image has a history in your head. To you they are forty-one separate decisions.
To a stranger scrolling a feed, they are one impression. And the first one lands before your title is read — colour and glow are processed faster than text.
There is also no correction mechanism. Nobody types "your images look AI" in your comments. The people who noticed are the people who left, so the feedback you receive comes entirely from the viewers who were not bothered. You are reading a survey of the wrong sample.
So the loss is real and invisible at the same time, and it gets attributed to whatever else you can see — the topic, the title, the length. Usually the topic was fine.
What actually gives a video away?
A short list, and it splits cleanly into things inside an image and things that only exist across a set.
- The glow halo. A radial glow behind the subject, or an outer glow on thumbnail text. Currently the fastest AI signal on the platform, and the one most creators think looks professional.
- The default font. Heavy rounded sans with a stacked drop shadow. It is the default in every AI thumbnail tool, so it now reads as a signature.
- Oversaturation. Generated output ships hot. The colour is registered before the image is.
- Plastic skin. Poreless, waxy, evenly lit. Most visible at thumbnail scale on a face.
- Melted detail. Fabric patterns, crowds, background faces, teeth, jewellery. The model spends its attention on the subject and lets the rest dissolve.
- Text garble. Signs, newspapers, screens, labels. Still the most obvious tell in the set.
- Hero lighting everywhere. Every frame lit like a poster. Real footage contains flat, ugly, badly lit frames, and their absence is felt even when it is not named.
Any one of those is survivable. What makes a video read as generated is usually not on this list at all.
What are set-level tells?
The tells that no single image can reveal, because they are properties of the sequence.
Take a real example. A creator's forty-one prompts all ended with the same tail: cinematic lighting, ultra detailed, 8k, golden hour, shallow depth of field. Every one of those images is defensible on its own. Together they produce forty-one shots at the same distance from their subject, lit at the same hour of the same day, indoors and outdoors alike, every background blurred to the same degree.
Then each one got an identical six-second slow zoom in the edit.
His viewer's comment was "good story but the images all kind of look the same." That is not a taste note. It is an accurate description of a set with one lens, one light and one camera move.
- Focal sameness. A human operator moves — a wide, a mid, a tight. Yours never did.
- Grade sameness. One continuous sunset across nine minutes.
- Uniform motion. The same push on every clip, the signature of a stills pipeline. Most viewers cannot name this one but all of them feel it by minute two.
- No rest frames. Every frame maximally detailed, so the eye never rests. Density with no relief reads as generated.
- Zero real footage. One genuine archival clip re-anchors everything around it. With none, there is nothing to anchor to.
This is why per-image checking never solves it. You can approve every frame and still ship the tell, because the tell lives between the frames.
Which tells can't be fixed by regenerating?
Several, and these are the ones people burn afternoons on.
Legible text is the first. A newspaper headline, a shop sign, a screen with words on it. Some models manage short strings, most garble them, and no amount of rerolling changes that. Put the text in your editor over a plain texture — it takes ninety seconds.
Dense background faces are the second. A crowd at any real scale melts. Reframe to three people, or shoot the crowd from behind. The prompt is not the problem.
Hands at small scale and exact counts round out the list. Compose around them.
The reason this matters is time. People do not lose their evenings on faults they can fix; they lose them on faults nothing can fix, because nothing told them to stop. A tool that says "replace this shot" is more useful than one that offers a cleverer prompt.
What does the real output look like?
Here is the actual output from the sample run — a nine-minute faceless money-history video, 41 Midjourney images, one thumbnail:
TELLS 17 THUMBNAIL 6 IN-VIDEO 11 SET-LEVEL 5 TIME TELLS WHAT LANDS FIRST 0.4s thumbnail 6 the glow halo, then the airbrushed face 0-3s cold open 3 same grade, same distance, no rest frame 0-30s 4 the slow zoom repeating on every shot minute 2+ 4 melted crowd faces, garbled newspaper text
Thumbnail tells cost the click. Minute-two tells cost a viewer who already clicked. They are not worth the same.
FOCAL SAMENESS 41 of 41 shots
EVIDENCE: every prompt ends "…shallow depth of field"
WHY: a real camera operator moves. Yours never did.
FIX: 14 prompts to wides, 6 to flat deep-focus.
Delete "shallow depth of field" from 20.One line in a prompt template caused five separate set-level tells.
# MIN DROP FIX 1 4 -14 Kill the glow/bloom on the thumbnail text 2 3 -9 Swap the default rounded font 3 5 -8 Hold 12 clips still, vary the durations 4 2 -6 Desaturate the thumbnail ~15% 5 6 -10 Drop in 3 free archival clips
Two shots are routed to "replace" instead — the crowd and the newspaper headline. No reroll fixes either.
How do you run it yourself?
You paste one prompt into Claude Code and it builds the tool for you. It is a dark dashboard, pre-filled with the sample above, so it works on the first run.
It has a Settings panel for your own API key, so you can run it before every upload — which is the point, because the tells change as the models change.
Grab it below. Drop your email and the prompt is on the very next page, free. Run it on the video sitting in your editor right now, before it goes out.
Can you turn this into a side hustle?
Yes, and it is a service business with almost no setup. The tool does the production; you do the selling and the quality check.
Here is the model. Local businesses need Paste how you made your visuals — the thumbnail, the scene prompts, the motion, the footage mix. Get every tell a scroller clocks as AI, ranked by how fast they see it, plus the fix for each and the ones you can clear in twenty minutes., but they do not have the time or the skill to do it well. You do. So you run the tool, hand them a finished result, and charge for the service. Many people charge $500 a month per client for work like this.
The best part is the cost to start: a free prompt — it pays for itself on the first job. The tool does the heavy lifting in minutes, so your margin is high and you can take on more clients without more hours. To get your first client, reach out to a few local businesses you already know. Do one for free, show them the result, and ask who else needs it.
FAQ
Do I have to upload my images?
No. You paste the recipe — your thumbnail description, your scene prompts, the model, the motion treatment and the footage mix. That is deliberate: the most expensive tells are properties of the whole set, and those are visible in the recipe rather than in any one picture.
Isn't this just telling me to stop using AI?
No, and it never will. Every fix is something one person can do tonight with the tools they already named. The goal is the state a commenter described exactly: the good AI thumbnails are the ones you don't notice are AI.
How is this different from fixing a bad image prompt?
Different unit of work. Prompt fixing looks at one prompt that failed and rebuilds it. This looks at the whole video's visual set and finds the tells that only exist across it — one focal length, one grade, one camera move. Those are invisible when you check images one at a time.
Does it work if I use AI video rather than stills?
Yes. It adds the motion tells — morphing within a shot, subjects floating without weight, and the same motion arc repeating on every clip — and routes those fixes the same way.
Can I run it on every video?
That's the intent. Enter your API key once and run it as a pre-upload check. It also returns prompt-stage rules pulled from your own tells, so the next video starts without them.