How to Make AI Videos from Text: A Beginner's Guide That Doesn't Overcomplicate It
AI video went from 'blurry nightmare' to 'genuinely usable' in about six months. Here's how to actually make something with it.
AI video is confusing because everyone calls everything the same thing
"AI video" could mean about eight completely different things right now, and nobody bothers to specify which one they're talking about. Text-to-video generation? AI avatars? AI editing of existing footage? Automated short-form content? They're all wildly different tools solving wildly different problems, and they're all lumped under "AI video" as if they're interchangeable.
They're not. So let me break this down into categories that actually make sense, explain what each is good for, and save you from buying the wrong tool for the wrong job.
Category 1: Talking head videos (for when you need a person talking to camera)
What this is: You type a script, choose a digital avatar (or create one from your own face), and the AI generates a video of that avatar speaking your words with natural lip sync, gestures, and expressions.
The tools:
synthesia is the market leader and the most polished. You pick from dozens of diverse avatars, type your script, and get a professional-looking video in minutes. The avatars look convincingly human from about two metres away. Up close? Still a touch uncanny. The eyes blink at slightly too-regular intervals, and the head movements have a quality that's hard to describe but easy to spot. Like someone who's been told to "act natural" and is trying really, really hard.
heygen is Synthesia's main competitor and in some ways it's catching up fast. HeyGen's avatar quality is roughly equivalent, but their standout feature is instant avatar cloning. Record yourself for two minutes and HeyGen creates a digital avatar that looks like you, sounds like you (including your accent), and speaks whatever script you type. The quality is genuinely impressive and also slightly terrifying.
What these are actually good for: - Corporate training videos where filming a real person is expensive or impractical - Product demos and explainer videos - Personalised sales outreach (yes, people do this, and yes, it works better than it should) - Multilingual content (record once, generate in 40+ languages)
What they're bad at: - Anything requiring emotion. The avatars can do "friendly" and "professional." They cannot do "excited," "concerned," or "funny." Trying to make an AI avatar tell a joke is painful. - Anything longer than about 5 minutes. Viewer fatigue with AI avatars is real. Short and focused works. Long and rambling doesn't.
Pricing: Synthesia starts at $22/month. HeyGen starts at about $24/month. Both offer free trials.
Category 2: Text-to-video generation (for creative and social content)
What this is: You type a text description and the AI generates a video clip from scratch. No existing footage needed. Pure AI imagination rendered as moving images.
The tools:
runway Gen-3 Alpha is currently the best text-to-video model available to the public. The output quality has improved dramatically. Consistent characters, reasonable physics, cinematic lighting. A prompt like "a woman walking through a sunlit forest, autumn leaves falling, cinematic quality" produces something genuinely beautiful. For 4-5 seconds.
And that's the key limitation. Text-to-video is impressive in short bursts. A single 5-10 second clip can look stunning. String together a 30-second sequence and the cracks show. Characters change appearance between clips. Physics stops making sense. The "AI look" (slightly too smooth, slightly too perfect, movements that are fluid in a way real movement isn't) becomes impossible to ignore.
pika is the more accessible option. Less cinematic than Runway but easier to use and cheaper. Pika 2.0 added some clever features like "scene ingredients" that let you upload reference images to guide the generation. The results are more social media than cinema, but for TikTok and Instagram content, that's fine.
What text-to-video is good for: - Short social media clips (under 10 seconds) - Concept visualisation (showing a client what you're thinking before committing to a real shoot) - Music videos and artistic projects where the "AI aesthetic" is a feature, not a bug - B-roll and background footage for longer videos
What it's bad at: - Anything longer than 10 seconds without visible artefacts - Anything requiring specific, consistent characters across multiple shots - Realistic human movement (walking still looks slightly wrong) - Text appearing in the video (garbled about 60% of the time)
Pricing: Runway starts at $15/month. Pika offers a free tier and paid plans from $10/month.
Category 3: AI-powered video editing (for existing footage)
What this is: You have real video footage and you use AI tools to edit it faster. Automatic captions, filler word removal, background noise removal, auto-cutting to remove silences, and AI-assisted colour correction.
The tools:
Descript is the standout here. It turns your video into a text transcript and lets you edit the video by editing the text. Delete a sentence from the transcript and it cuts that section from the video. It also has "filler word removal" that automatically cuts every "um," "uh," and "you know" from your footage. For podcasters and YouTubers, this single feature saves hours of editing per week.
Kapwing is more of an all-purpose online video editor with AI features bolted on. AI-generated captions (very accurate), automatic resizing for different social platforms, and a "smart cut" feature that removes dead space. Less powerful than Descript for heavy editing, but lower barrier to entry.
What AI editing is good for: - Making long-form content from raw footage without traditional editing skills - Captioning (AI captions are now accurate enough to publish without proofreading, most of the time) - Repurposing content for different platforms - Cleaning up audio in footage recorded on phone microphones
Pricing: Descript starts at $24/month. Kapwing starts at $16/month.
Category 4: Short-form content automation (for repurposing long videos)
What this is: You feed in a long video (podcast, webinar, YouTube video) and the AI identifies the most engaging clips, adds captions, resizes for vertical, and spits out 5-15 short clips ready for TikTok, Instagram Reels, or YouTube Shorts.
The tools:
Opus Clip is the best at identifying genuinely interesting moments in long videos. The AI doesn't just chop randomly; it analyses for emotional peaks, surprising statements, and complete thoughts. About 60-70% of the clips it generates are genuinely usable, which is a much better hit rate than manual clipping.
pictory does a similar thing but with more control over the output format. You can define templates, brand colours, caption styles, and aspect ratios. The clip selection AI isn't as good as Opus Clip's, but the customisation options are better.
invideo has moved into this space too, with their AI video editor that can turn a blog post or script into a short video with stock footage, text overlays, and transitions. The output is more "social media content" than "video production," but for churning out LinkedIn posts and Instagram content, it works.
What short-form automation is good for: - Podcasters who want to promote episodes on social media without spending hours clipping - YouTubers who want to create Shorts from their existing catalogue - Marketing teams who need to repurpose webinars and talks - Anyone who's been told "you need to post more on social media" and doesn't have the time
Pricing: Opus Clip starts at $19/month. Pictory starts at $19/month. InVideo offers a free tier with paid plans from $25/month.
The practical workflow (what I actually do)
Here's a real workflow I've used for creating content: Script: Write the script in a text editor, use Claude to tighten it up Talking head: If I need someone "presenting," Synthesia for professional content, or record myself and use Descript to clean it up B-roll: Runway for 3-5 second clips to fill gaps between talking head sections Edit: Descript for the main edit, auto-captions, filler word removal Repurpose: Opus Clip to extract 3-5 short clips from the finished video
Total time for a 5-minute video: about 2 hours. Two years ago this would have taken a full day with a camera setup and traditional editing software. The quality isn't broadcast-level, but for YouTube, social media, and internal company content, it's more than good enough.
The honest assessment
AI video is usable now. It was a novelty a year ago and it's a genuine tool today. But "usable" and "great" are not the same thing.
Talking head avatars still look like talking head avatars. You can spot them. Text-to-video still has that slightly dreamy, not-quite-real quality. AI editing is the most practically useful category because it works with real footage and just makes the process faster.
If you're expecting to type "make me a Super Bowl ad" and get something broadcastable, you'll be disappointed. If you're expecting to produce decent content faster and cheaper than traditional methods, you'll be genuinely impressed.
Start with the category that matches your actual need. Don't buy a text-to-video tool when what you really need is a talking head tool. Don't buy a talking head tool when what you really need is better editing software. Match the tool to the job and the results will surprise you.