Navigate the dashboard and find every feature you need for your first video
Explain the scene-based structure so you can plan your video before you build it
What Synthesia Is — And What It Is Not
Synthesia is an AI video generation platform. You type a script. You choose an avatar — a digital human presenter. You design what appears on screen alongside the avatar. You click Generate, and Synthesia renders a video — complete with lip-synced narration — without a camera, microphone, recording studio, or video editor.
The result looks like a professionally produced training video. Because it is one.
What It Is Good For
Training and onboarding videos your team will actually watch
Process walkthroughs and SOPs that used to live as text documents no one read
Compliance content that needs to be consistent, documented, and easy to update
Software demos combined with screen recordings
Content that needs regular updates — because editing is fast and cheap
What It Is Not Good For
Live, real-time content that changes daily
Moments where learners need to see a specific real person — a company founder, a patient, a key client
Highly interactive, branching simulations (use Articulate or similar for those)
💡 Pro Tip
If you are hesitating over whether to use Synthesia for something, ask yourself: "Would a recorded person be better — or just more familiar?" Most of the time, it would just be more familiar. Familiar and better are not the same thing.
Dashboard Orientation
Log in and take 60 seconds to orient yourself before you build anything. Here is what matters for your first video:
Area
What It Does
Your Priority
My Videos
Your full video library — all projects live here
High — create a folder before your first video
Templates
Pre-built video structures to study or adapt
Medium — browse one, then build your own
Media Library
Storage for logos, images, screen recordings
High — upload your logo before building
Brand Kit
Colors, fonts, logo applied automatically
High — set this up in Module 4
New Video
Creates a new project
High — you will click this in Module 3
📌 Note
Before you build your first video, create a folder in My Videos. Go to My Videos → New Folder. Name it something like "Training Videos — 2026." An organized library from day one saves significant time later.
The Scene — The Most Important Concept in Synthesia
Every Synthesia video is made up of scenes. A scene is the basic unit — like a single slide in a presentation, except it also plays narrated audio. Every scene has three parts:
Script — the words the avatar speaks. You type this.
Layout — the visual structure. The avatar's position, background, any text or images on screen.
Avatar and voice — who delivers the narration and how it sounds.
When you click Generate, Synthesia processes every scene — synthesizing the voice, animating lip movements, compositing visuals — and stitches them into a single video file.
How Long Is a Scene?
Words in Scene
Approximate Duration
Best For
50 words
~23 seconds
Intro, hook, quick transitions
75 words
~35 seconds
Most instructional content scenes
100 words
~45 seconds
Complex steps or detailed explanations
150+ words
70+ seconds
Getting long — consider splitting into two scenes
A 5-scene video where each scene is about 75 words runs approximately 3–4 minutes — a solid, complete training video for most topics.
🏁 Quick Win
Your key takeaway: Synthesia = scenes. Each scene = one script + one layout + one avatar/voice. Plan your video as a list of 5 scenes before you open the editor. This single habit separates learners who feel in control from those who feel lost.
Module 1 — Check Your Thinking
1. Which of the following is Synthesia best suited for?
2. A 5-scene training video where each scene is roughly 75 words will run approximately how long?
3. What are the three parts of every Synthesia scene?
Module 2
Write a Script That Sounds Human
The skill that determines whether your video is good or painful
Duration
~25 min
Learning Objectives
Understand why document text never sounds right when spoken by an AI avatar
Write a script in spoken English — not written English
Use the five-scene structure to plan your first video
Draft a complete 5-scene script using the provided templates
The One Rule That Changes Everything
Write for the ear, not the eye.
Written language and spoken language are completely different. When you read, your eye can scan ahead, re-read a confusing sentence, and slow down at complex parts. When you listen, you get one pass. No scroll bar. No rewind. The listener either keeps up or falls behind — and if they fall behind, they stop paying attention.
This is why copying text from a document into Synthesia never works. A paragraph that reads perfectly on paper sounds robotic, dense, and exhausting when an AI avatar delivers it at a fixed pace.
The Five Rules of Spoken English for Synthesia
Short sentences. If a sentence is longer than 20 words, split it into two.
Active voice. "The manager approves the request" — not "The request is approved by the manager."
Use contractions. "You'll learn this in step two" — not "You will learn this in step two." Contractions sound human.
No visual references. Never say "as you can see here" or "click the button below." Describe explicitly: "Click the blue Submit button in the lower right corner of the screen."
Read it aloud. Every script. Before you build. If you stumble, rewrite.
⚠️ Watch Out
The most common mistake: writing the script last. The script is the foundation — build it first, approve it before you open Synthesia. A script change after generation costs 10 minutes. A script written well before generation costs nothing extra.
The Five-Scene Structure
Your first video has five scenes. Use this structure for any training topic:
Scene
Name
Purpose
Length
1
The Hook
Name the problem or outcome. Earn the next 4 minutes. Do NOT open with "Welcome to this training..."
40–60 words
2
The Context
Give learners what they need to know before the how-to begins
60–80 words
3
The How-To
The core process — step by step, specific and sequential
75–100 words
4
The Watch-Outs
Common mistakes, exceptions, and what to do when things go sideways
50–75 words
5
The Landing
One key takeaway, next step, and where to go for help
40–60 words
🎯 Real Talk
This structure works for almost every short training video. It is not the only structure — but it is the right starting point. Once you have built five videos using it, you will naturally adapt it. Until then, follow it exactly.
Script Templates — Fill In the Brackets
🎬 Scene 1 — The Hook
"[Name the situation or problem your learner faces]. In the next [X] minutes, you'll know exactly how to [specific skill or outcome] — [brief why it matters to them personally]. Let's get into it."
EXAMPLE: "A patient calls to schedule their first appointment. You have about 90 seconds to ask the right questions and book them correctly — or create a problem that takes twice as long to fix. In this video, you'll learn the exact sequence that makes every new patient intake smooth, fast, and accurate."
🎬 Scene 2 — The Context
"Before we get into the steps, here's what you need to know. [1–3 sentences of essential background — system used, who is involved, what triggers the process, any relevant policy]. Got it? Let's walk through it."
🎬 Scene 3 — The How-To
"Here's the process. Step one: [action]. Step two: [action — and any important detail]. Step three: [action]. Step four: [action]. Once you've completed step four, [what happens next or what confirmation to expect]."
🎬 Scene 4 — The Watch-Outs
"A few things to watch out for. [Common mistake 1] — when this happens, [what to do instead]. [Common mistake 2 or exception] — in this case, [how to handle it]. If you're ever unsure, [where to go or who to ask]. Better to pause and check than to push through and create a bigger problem."
🎬 Scene 5 — The Landing
"You've got it. The most important thing to remember: [one-sentence key takeaway]. Your next step is [specific action]. If you run into something this video didn't cover, [resource or contact]. You're ready."
Controlling Pacing With Punctuation
Synthesia's voice engine reads your punctuation as performance instructions:
Punctuation
Effect on Avatar
Use For
Comma (,)
Brief natural pause
Lists, transitions, measured rhythm
Period (.)
Full stop — lets ideas land
Ending complete thoughts
Ellipsis (...)
Longer dramatic pause
Building anticipation before an important point
Paragraph break
Section-level pause
Separating distinct thoughts in one scene
✏️ Try It Now
Using the five-scene structure and the templates above, write your complete first video script right now. Target: 50–100 words per scene. Do NOT open Synthesia yet — finish the script first. When done, read every scene aloud. Fix anything that makes you stumble. Then move to Module 3.
Module 2 — Check Your Thinking
1. Why does copying text from a document into Synthesia almost never work well?
2. You are writing Scene 1 of your first video. Which opening is better and why?
3. What does Scene 4 (The Watch-Outs) contain?
Module 3
Build and Generate Your First Video
From approved script to published video — step by step
Duration
~35 min
Before You Open Synthesia — Pre-Flight Check
Confirm every item before creating a single scene:
☐
My 5-scene script is complete and approved
☐
I have read every scene aloud and fixed anything that sounds unnatural
☐
I know my video's title (format: [Topic] — [Title] — v1)
☐
I know which folder to save it in
If any of these are not checked, go back to Module 2. A video built on an unfinished script will always need to be rebuilt.
Step 1 — Create a New Video
Log in to Synthesia. Click New Video. Choose Blank — not a template. Before you do anything else, rename the video at the top of the editor. Use your naming convention: [Topic] — [Title] — v1. Save it into your folder.
💡 Pro Tip
The most common organizational mistake is building 10 videos before setting up a folder system. Set it up on Day 1. You will not regret it.
Step 2 — Configure Scene 1 (Your Configuration Scene)
Your first scene is your master template. Get these right before building more — every subsequent scene will be duplicated from this one.
Choose a layout. For instructional content, start with split-screen: avatar on one side, slide content on the other. This is the most versatile layout for training.
Enter your Scene 1 script. Paste your Hook scene (40–60 words) into the script text field. Just Scene 1 — not all five.
Choose an avatar. Filter by style: Professional, Casual, or Technical. Pick one appropriate for your topic. Upper-body avatars — showing the presenter waist up — work best for instructional content.
Choose a voice. Critical: do NOT just listen to the sample phrase. Paste your actual Scene 1 script into the voice preview field and listen. The voice that sounds great on "Welcome to today's training" may sound rushed on your actual content.
Preview Scene 1. Watch the full preview. Check: lip-sync accuracy, narration pace, and whether anything sounds off. Fix any issues before duplicating.
🏁 Quick Win
Once Scene 1 looks and sounds right, right-click it in the scene panel and select Duplicate. This gives you Scene 2 with your layout, avatar, and voice already configured. Change only the script. Repeat for all five scenes. This is the fastest way to build a consistent video.
Step 3 — Build Scenes 2 Through 5
For each remaining scene:
Duplicate the previous scene from the left panel
Replace the script with your next scene's content
Add a short on-screen text element — a headline (3–6 words), a numbered list of steps, or a key term. Not a transcript. A visual anchor.
On-screen text is not a transcript of what the avatar is saying. It is a visual anchor — short, scannable in 4 seconds, highlighting only the key point. If your slide has more text than your learner can read in 4 seconds, it has too much text.
Step 4 — Generate
Review all five scenes. Confirm each script is correct. Check on-screen text for spelling errors. Then click Generate in the top right.
Synthesia will process your video. Expect 5–15 minutes for a 5-scene video. You will receive a notification when it is ready. Do not sit and watch a progress bar — use the time to plan your next video or draft your second script.
Step 5 — Review and Fix
Watch the entire generated video before sharing it with anyone:
☐
Lip-sync looks accurate throughout all scenes
☐
Narration pace feels natural — not rushed, not sluggish
☐
On-screen text is fully visible — nothing cut off or overlapping the avatar
☐
No mispronounced words or robotic-sounding phrases
☐
Visual design is consistent across all five scenes
☐
No spelling or grammar errors in on-screen text
If you find a script error in one scene: open that scene → fix the text → click "Regenerate Scene." Only that scene re-renders (1–3 minutes). The rest of your video is untouched.
⚠️ Watch Out
Never share a first-generation video without watching it yourself first. A 2-minute personal review prevents the embarrassment of a learner catching your errors first.
Step 6 — Publish
Once your video passes the review checklist: go to your video in the library → click Share. Synthesia generates a hosted player link — a clean, professional video page anyone can watch in a browser.
For your first video, share the hosted link. That is the fastest path to done. Upload to your LMS, set up SCORM tracking, or build a formal course later. Get the video in front of learners today.
🏁 Quick Win
You just built your first Synthesia video. That is the hardest part — not because it is technically difficult, but because it is the point where most people hesitate too long. You didn't hesitate. Module 4 will make your next video look significantly more professional.
Module 3 — Check Your Thinking
1. What is the correct order of steps for building your first Synthesia video?
2. Why do you preview your actual script text when choosing a voice — not just the default sample phrase?
3. You catch a script error in Scene 3 of your generated 5-scene video. What is the most efficient fix?
Module 4
Make It Look Like You Know What You're Doing
Five visual rules and your Brand Kit — that's all you need
Duration
~20 min
Why Design Matters More Than You Think
Learners make a quality judgment within the first 5 seconds of a video. That judgment is largely visual. A video with sloppy design — mismatched fonts, cluttered slides, no visible brand — signals "someone threw this together." A video with clean, consistent design signals "this organization takes its training seriously."
You do not need to be a designer. You need to follow five rules consistently.
First: Set Up Your Brand Kit (Do This Before Anything Else)
The Brand Kit stores your organization's visual identity and applies it to any video automatically. Set it up once — it works everywhere.
Logo
Upload as a PNG with transparent background. Transparent PNG = no white rectangle around your logo on any background. Size: at least 400px wide. If you only have a JPG, get the PNG from your marketing team before proceeding.
Brand Colors
Enter as hex codes (example: #1B2A4A). Include: primary color, secondary color, and a neutral for text. If you don't know your hex codes, check your brand style guide or ask marketing. Do not guess.
Fonts
Synthesia uses Google Fonts. Choose the closest available match to your brand fonts. Document the substitution so your whole team makes the same choice every time.
To apply: open any video → video settings → enable "Use Brand Kit." Every scene updates instantly. Enable this on every video from now on.
💡 Pro Tip
Screenshot your completed Brand Kit settings and share with your team. This means everyone configures their brand correctly without hunting through settings independently.
The Five Visual Rules
Rule 1 — One Idea Per Slide
Every slide should have one dominant element. One headline. One key statistic. One numbered list. One image. If two ideas are competing for attention on the same slide, split them into two scenes. Your learner cannot absorb two competing focal points simultaneously.
Rule 2 — Less Text Than You Think You Need
The slide is not a script. If your slides echo everything the avatar says word for word, your learner doesn't know where to focus — and they disengage from both channels. A slide with three short bullets readable in 4 seconds, then attention back to the avatar. That is the right rhythm.
Rule 3 — Consistent Layout Across the Video
Choose 2–3 layouts and use them consistently. Do not switch layouts every scene — it creates visual noise without purpose. Change layout only to signal a shift in content type: standard instruction uses Layout A, knowledge check uses Layout B. That is the full system.
Rule 4 — High Contrast Text — Always
Dark text on light background, or light text on dark background. Either works. What never works: medium-gray text on a light background, light blue on white, or any text that requires squinting. Your learners are watching on everything from a 27-inch monitor to a phone screen. When in doubt, increase the contrast.
Rule 5 — Your Logo, Every Video, Same Spot
Place your logo in the same corner of every video — top left or bottom right, small enough not to compete with content. Establish this in Scene 1 and carry it through duplication. Do not move it based on what looks good per scene. Consistency beats optimization.
The Five Most Common First-Video Mistakes
Mistake
What It Looks Like
The Fix
Too much slide text
Slides echo the full script; learners don't know where to focus
Remove 60% of the text. Keep only key terms, headlines, or step numbers.
No visual hierarchy
Everything on screen is the same size and weight
Make one element dominant. Make everything else smaller and lighter.
Inconsistent layouts
Different layout every scene, no clear pattern
Pick 2–3 layouts. Use them consistently. Change only with purpose.
No brand applied
Video looks like a Synthesia demo, not your organization
Set up Brand Kit before building. Apply to every video.
Logo with white box
Logo sits on a visible white rectangle over the background
Upload as transparent PNG. Get the PNG from marketing if needed.
✏️ Try It Now
Open the video you built in Module 3. Apply your Brand Kit. Then go through each scene and apply all five visual rules: reduce text, confirm hierarchy, check logo placement, confirm layout choices. Regenerate any scenes where you also changed the script.
Module 4 — Check Your Thinking
1. What file format should your logo be for Synthesia, and why?
2. Your slides in Module 3's video echo the full script text word for word. What should you do?
Module 5
Publish, Share, and Know What Comes Next
Getting your video in front of learners — and planning your second one
Duration
~20 min
Step 1 — Add Captions Before You Publish Anything
Captions serve learners who are deaf or hard of hearing, learners watching without audio, learners for whom your video's language is a second language, and everyone whose comprehension improves with both channels active. Add them to every video.
Open your finished video in Synthesia
Find the Captions or Subtitles option in video settings or the sharing panel
Enable auto-generated captions
Review them in Synthesia's caption editor — AI is accurate but not perfect. Check proper nouns and technical terms.
Save the corrected captions
Export the SRT file from the download options — you will need this for LMS upload
📌 Note
An SRT file is a plain text file containing your caption text with timestamps. Every LMS and every major video platform accepts SRT files. Always export it alongside your MP4.
Step 2 — Choose How to Publish
Method
How It Works
Best When
Synthesia Hosted Link
Synthesia generates a shareable URL. Anyone with the link can watch in a browser.
Sharing for review, sending to a small group, embedding on a webpage
MP4 Download
Download the video file (up to 1080p). Upload anywhere.
Uploading to an LMS, sharing via drive, embedding in another platform
iFrame Embed
Copy an embed code from Synthesia. Paste into your LMS or webpage.
You want the video to stream from Synthesia inside your LMS without uploading a file
For your first video: use the Synthesia hosted link. Share it with your team. Get feedback. This is the fastest path to done. Upload to your LMS, set up SCORM tracking, and build formal course structure later. Do not let the perfect be the enemy of the published.
SCORM — The Short Version
You may have heard the word SCORM in relation to your LMS. Here is what you need to know right now:
What SCORM does: it lets course content communicate with your LMS — reporting whether a learner started, completed, and scored on content. Without SCORM, your LMS only knows a learner clicked on the content, not whether they watched or learned anything.
What Synthesia does: Synthesia exports MP4 video files and hosted player links — not SCORM packages natively. A Synthesia video by itself cannot report SCORM completion to your LMS.
Your Three Options
Option A — SCORM Wrapper: Download your Synthesia MP4 and wrap it in a SCORM shell using a free tool like iSpring Free. This turns your video into a trackable SCORM package. This is exactly how the companion Synthesia Masterclass SCORM course is built.
Option B — Separate Quiz Trigger: Host the video in Synthesia and use a short quiz on your LMS as the official completion event. The video is informational; the quiz generates the completion record.
Option C — LMS Basic Video Tracking: Many LMS platforms can track completion based on percentage watched when you upload an MP4 directly. Check your specific LMS documentation.
📌 Note
For most teams getting started, Option C (direct MP4 upload) or Option A (iSpring Free wrapper) is the right first step. Do not let SCORM complexity prevent you from publishing. Get the video published. Refine the tracking later.
Planning Your Second Video — Faster
Now that you have built one video, your second will take half the time. Your Brand Kit is configured, your avatar and voice are chosen, your folder structure exists, and you know the five-scene structure.
Open your first video. Right-click → Duplicate the entire project.
Rename the duplicate with the new video title.
Replace the script in each of the five scenes with your new content.
Update any scene-specific visuals (images, on-screen text).
Preview Scene 1. Generate. Review. Publish.
Your second video should take 60–90 minutes from script to published. Your fifth should take 45 minutes. Your tenth should take 30. The system gets faster every time you use it.
🎯 Real Talk
The single most important thing you can do after this course: build your second video this week. Not next week. This week. The gap between Video 1 and Video 2 is where momentum dies. Close it fast.
Your Production Checklist — Use This Every Time
☐
Script complete, read aloud, approved
☐
Video named: [Topic] — [Title] — v1
☐
Video saved in correct folder
☐
Brand Kit applied
☐
Avatar and voice previewed with actual script text in Scene 1
☐
All five scenes labeled: Hook, Context, How-To, Watch-Outs, Landing
☐
Generated video watched in full before publishing
☐
Captions enabled and reviewed
☐
SRT caption file exported
☐
Video published and link or file shared with intended audience
What Comes Next
You have completed the Synthesia Quick Start. You have built and published your first video, applied your brand, and have a checklist and production system you can use going forward.
Here is your next sequence:
This week: Build your second video using the duplicate-and-replace method.
Next week: Set up three scene templates in your Synthesia library — intro, content, knowledge check — so you stop configuring from scratch.
When ready: Take the Synthesia Masterclass to go deeper. Advanced scripting, visual design principles, translation, SCORM, team workflows, analytics, and a full Capstone Project.
Module 5 — Final Check
1. Synthesia natively exports which of the following?
2. What is the fastest way to build your second Synthesia video?
3. Why should you add captions to every Synthesia video?
🎬
Quick Start Complete!
—
You have completed all 5 modules of the Synthesia Quick Start.
Your completion and score have been reported to your LMS.
Your next moves:
1. Build your second video this week
2. Set up three scene templates in your library
3. Take the Synthesia Masterclass when you're ready to go deeper