GUIDES · 2026
Turning a short written story into a video usually means one of two things: paying for an AI video generator, or spending hours in a full animation program. Our free 2D Story Video Maker takes a third path β you write your story scene-by-scene, pick from a built-in library of flat-style illustrated backgrounds and characters, add narration, and export a real downloadable video. No paid AI API, no subscription, no signup β everything runs in your browser.
This is not an AI image generator. It doesn't paint a brand-new picture for every scene the way a diffusion model would β that technology needs a paid API to run, and this site's whole approach is to stay free forever. Instead, the tool ships with a hand-built library of consistent flat-style illustrations (think simple, clean, modern icon art) β 10 backgrounds and 8 "peg doll" style characters β and your job is to assemble a story from that library, the way you'd lay out a picture book. Once you see it that way, the workflow makes a lot more sense.
Open the Story Video Maker and you'll see 3 scenes already started. Each scene is one beat of your story β one setting, one moment. Type the narration text for a scene directly into its text box: this is both what gets spoken (or captioned) and what the auto-suggest feature scans to guess a matching background and characters. Add more scenes with "+ Add Scene" (up to 15), and reorder or delete them with the β² βΌ and π buttons in each scene's header.
Click "β¨ Auto-Suggest Background & Characters" on a scene and the tool scans your narration text for keywords β words like "school," "village," "forest," "beach," "night," "dog," or "grandmother" β and pre-selects a background and up to two characters that seem to match. It's a helpful starting point, not a mind-reader, so always glance at the pick and override it manually if you want something different: click any background thumbnail in the grid to select it, and use each character slot's grid plus its left/center/right position dropdown to place people (or a dog, or a cat) exactly where you want them in the frame.
This is the part worth reading carefully, because it's the one real technical limit of browser-based video tools. No mainstream browser lets a website capture the audio a synthesized (text-to-speech) voice produces into a recordable track β there's simply no reliable API for it. So the tool gives you two honest, clearly separated paths instead of pretending to combine them:
Speak each scene's narration into your mic. This is the only path that produces a downloadable video with real audio, because recorded audio (unlike synthesized speech) can be routed through the Web Audio API into a capturable stream.
Type narration text and preview it live with a computer voice. The exported video from this mode is silent, with your narration burned in as on-screen captions instead β genuinely useful, just not the same as audio.
If you want a narrated, downloadable video, use "Record My Own Voice." If you're happy with a captioned, silent video β or you just want to draft and preview your story quickly β "Type & Preview with TTS" is faster. You can also combine the two: use TTS preview mode to test pacing, then switch to Record mode and read the same lines into your mic once you're happy with the story.
Switch to the "ποΈ Record My Own Voice" tab and each scene grows a "ποΈ Record Narration" button. Click it, allow microphone access the first time your browser asks, speak your line, then click "βΉ Stop." An audio player appears right there so you can check the take, and re-record any time before you export β a scene's duration automatically follows its recorded clip's real length unless you type a manual override into the duration field.
Every scene has a "Show captions on screen" toggle, which burns the narration text onto the video as on-screen captions β turn it on for accessibility, for silent-scroll social feeds, or as the whole point if you're using TTS-preview mode's silent export. Each scene's duration field is editable too: in Record mode it defaults to your recorded clip's length, in TTS mode it's estimated from your text at roughly 150 words per minute (minimum 3 seconds) β type your own number any time to override either estimate.
Click "βΆ Preview Story" to watch the whole animated sequence play through in the canvas β backgrounds gently zoom and pan (a soft Ken-Burns effect), characters ease in with a slide-and-fade at the start of each scene and bob gently while they're on screen, and in TTS mode you'll hear each scene's narration spoken aloud in sync. Use "βΉ Stop Preview" to interrupt it early once you've seen enough.
Click "β¬ Download Video" when you're ready. This isn't instant β the tool plays through your entire story in real time while it records the canvas (and, in Record mode, mixes in your recorded audio), because that's fundamentally how browser-based capture works. A progress message tells you exactly which scene is being rendered ("Rendering scene 2 of 5β¦"), so keep the tab open and active until it finishes. When it's done, you'll get a standard .webm video file β the same format this site's Screen Recorder and Voice Recorder use β ready to save, share, or convert to MP4 with any free converter if you need that format.
Every illustration is pre-built; there's no per-image generation cost and nothing metered.
Recorded voice, story text, and the exported video are all created and kept on your own device.
Not a slideshow β a genuine .webm export you can upload, share, or convert.
Ready to try it? Head over to the 2D Story Video Maker and turn your next story idea into a shareable video in a few minutes.