Free

Making Video With AI, And Where It Actually Falls Apart

The 60 second clips in your feed took somebody 4 hours and 30 rejected generations. Here's the real workflow, including the parts nobody films.

5 min read · Last verified 2026-09-11

Every AI video demo you've seen is the good take.

They generated it 30 times. They're showing you run 30. And they're not going to mention that, because the whole business model is you believing it came out like that on the first try.

I make videos with this stuff. Here's the real version: what works, what wastes your night, and what it actually costs.

The one thing to understand first

AI does not make your video. AI makes pieces. You assemble them.

Everybody who fails at this fails the same way. They type one prompt, expect a finished video, get 4 seconds of a melting dog, and decide the technology isn't there yet.

The technology is there. The workflow is:

  1. Script — what's actually said, and in what order
  2. Voice — a real voice, or a generated one
  3. Visuals — clips, screen recordings, images, or generated footage
  4. Assembly — cuts, captions, the thing that makes it watchable

AI is genuinely great at 1 and 2. It's a coin flip on 3. It's getting useful at 4. If you skip straight to 3 you will hate this.

Step 1: the script, and why yours is boring

Your video dies in the first 2 seconds or it doesn't die at all. Everything else is downstream of that.

Don't ask for "a script about X." You'll get a LinkedIn post read aloud.

Ask like this:

Write me 10 opening lines for a short video about [thing]. Each one has to work as a standalone sentence somebody would stop scrolling for. No "in today's world," no "let me tell you," no questions that have an obvious answer. Lead with a number, a name, or something that sounds wrong until I explain it.

Then pick one and build the rest around it. The opener isn't the intro to the video. The opener is the video's entire chance.

Same rules as writing anything else with it: it goes generic when you go generic.

Step 2: voice

Three options, in order of how they land.

You, actually talking. Still the best. It's your voice, it has energy, nobody's uncanny-valley alarm goes off. Costs nothing. The only downside is you have to do it.

A cloned version of your voice. The tools do this well now from a few minutes of audio. Here's the honest part after using one heavily: a clone reads flat. It hits the words and misses the emphasis, because emphasis comes from meaning something and it doesn't mean anything. Fine for a list. Noticeably dead on anything with feeling in it.

A stock generated voice. Fastest, cheapest, most obviously AI. Works for faceless explainer content where nobody expected a person anyway.

One warning that cost me real time: some voice tools silently skip words on long inputs. No error. The file plays fine and a sentence is just gone. Feed it in chunks and listen to the whole thing before you use it.

Rest of it's yours

Drop your email and the whole thing opens up right here. No waiting on a confirmation, no download, nothing to install.

You'll also get the new ones as I make them. Leave whenever, no hard feelings.

Step 3: visuals, and the honest scorecard

This is where the money and the disappointment live.

Generated video, what it's genuinely good at:

  • Ambient motion. Water, wind, fire, light, leaves, a crowd behind somebody.
  • One person doing one simple thing. Walking away. Turning around. Reaching.
  • Camera-only motion. A push in, a drift, a POV down a street.
  • Anything where nothing has to be correct, only atmospheric.

What it still fails at, reliably:

  • Multi step physical action. Somebody doing a thing with rules, in order.
  • Two or more people coordinating. Handshakes, hugs, passing an object.
  • Hands doing anything precise.
  • Text. It'll look right in frame 1 and be gibberish by frame 40. Always check the last frame, not the first.
  • Anything where the wrong detail ruins it. A logo, a face you know, a tool being held the right way.

The rule I use: if the action has rules, don't generate it. Take a still image and pan across it instead. Same emotional beat, zero chance of a person's leg bending the wrong way. Nobody watching has ever once complained that a shot was a slow pan.

Screen recordings beat generated footage for anything you're teaching. They're free, they're real, and they're proof. If you're showing what a tool does, show the tool. Generated b-roll of "a person at a laptop" adds nothing and everybody knows what it is.

Step 4: assembly

Captions are not optional. Most people watch without sound. No captions, no video.

Cut on the beat of the audio, not on a round number. Find where the sentence actually ends and cut there. This is the single biggest difference between video that feels professional and video that feels like slides.

Pacing: if a shot sits longer than about 3 seconds without something changing, you're losing people. Every cut should be a new image, not the same clip again.

And check your aspect ratio on the actual exported file before you post it. Not a screenshot of it. A still frame can look perfect while the video plays stretched all over somebody's phone, because the stretch lives in a setting the still doesn't show. Play the finished file. Watch the whole thing.

What it actually costs

The free tier of most video generators gives you enough to learn on and not enough to publish with. Budget honestly:

  • Script and planning: free
  • Your own voice: free
  • Voice cloning: a small monthly fee, usually
  • Generated video: this is the spend. Per clip, and you'll throw most of them away
  • Editing: free tools are genuinely fine here

The trap is generating first and planning second. Every clip you generate before you've locked the script is a clip you're probably deleting. Write it, know exactly which shots you need, then spend.

The workflow I'd hand a beginner

  1. Write the script. Get the first line right before anything else.
  2. Record your own voice on your phone. Yes, really.
  3. Screen record the actual thing you're talking about.
  4. Only generate a clip where you genuinely have nothing to show.
  5. Cut it, caption it, watch the whole export before posting.

That's a publishable video with almost no spend, and it'll beat the fully generated version because it has a real person in it.

The short version

  • AI makes pieces, you assemble them
  • The opening 2 seconds is the whole game
  • Your own voice still wins
  • Generate atmosphere, never generate precise action
  • Text in generated video degrades, check the last frame
  • Screen recordings beat b-roll for anything you're teaching
  • Cut on the audio, caption everything, watch the export
  • Lock the script before you spend a cent on clips
← Back to all the free stuff