Quick answer: what makes a good training video? Our recommendation is to build each training video around one clear learning goal and remove anything that doesn’t help achieve it. Break complicated material into meaningful sections, and use visual cues to show learners exactly where to focus. Use video when seeing movement, sequence, or a process unfold actually helps explain the subject. Then pair the video with questions, practice, or other activities that make learners use what they’ve learned.
Those recommendations line up with a large body of multimedia learning research. In a 2022 overview of multimedia design research published in Review of Educational Research, Noetel and colleagues combined 29 previous systematic reviews covering 1,189 studies and 78,177 participants. They found significant positive effects on learning for 11 multimedia design principles, including signaling, segmentation, coherence, and appropriate use of animation.
Start every training video with a clear learning objective
Before you write the script, decide what someone should know or be able to do after watching.
A video explaining a new expense policy has a different job from a video teaching an employee how to submit an expense report. The first may only need to communicate a few rules. The second probably needs to show the actual workflow.
We recommend using that objective to decide what belongs in the video, what doesn’t, and whether video is even the right format. If the learner only needs a short reference list, a document may work better. If the learner needs to see a process happen, video has much more to offer.
Don’t make learners just watch
One of the most directly relevant studies for L&D comes from researchers affiliated with MIT working with Accenture employees.
In the peer-reviewed 2018 workplace learning study published in PLOS ONE, 99 employees completed a two-day experiment using an eight-minute, 20-second DevOps training video already used by Accenture. Researchers randomly assigned employees to normal video viewing, unstructured discussion, instructor-guided structured discussion, or questions inserted throughout the video.
When researchers tested memory 20 to 35 hours later, employees who answered interpolated questions remembered 26% more of the material than employees who simply watched the video. Structured discussion improved memory by 25%. Unstructured discussion didn’t significantly improve memory.
This study is especially useful because it tested actual employees with existing corporate training rather than creating an artificial classroom exercise. Accenture funded the research and helped recruit volunteers, but the authors reported that the company had no role in study design, data collection, analysis, the decision to publish, or manuscript preparation. The researchers also made their data publicly available.
A much larger 2025 meta-analysis of active learning in video combined 54 studies reported across 46 articles. Ninety-three percent of the studies focused primarily on adults. Compared with passive video viewing, active-learning strategies improved retention, with an effect size of g = 0.33, comprehension at g = 0.28, and transfer at g = 0.43. Embedded questions were the most common active-learning strategy studied.
Here’s how we’d translate those findings into an L&D lesson:
| What to add | What you’re asking the learner to do | Example |
|---|---|---|
| Interpolated questions | Retrieve information while learning | Pause after explaining a policy and ask which rule applies to a scenario |
| Practice | Perform the process being taught | Have an employee complete the software workflow after watching it |
| Scenario questions | Apply knowledge to a new situation | Present a customer interaction and ask what the employee should do next |
| Structured discussion | Work through defined ideas with other learners | Give a team two specific questions to discuss after the video |
The first and fourth approaches map directly to the Accenture experiment. The other examples are our recommendations for bringing active practice and transfer into workplace training.
How long should a training video be?
Research doesn’t support one ideal length for every training video.
The popular six-minute rule comes largely from Guo, Kim, and Rubin’s 2014 ACM study of MOOC video engagement. The researchers analyzed 6.9 million video-viewing sessions across four edX courses and found that shorter videos generally generated more engagement. But they measured engagement mainly through viewing time and whether learners attempted post-video assessment problems. The study didn’t test whether six minutes produced better learning than seven, ten, or fifteen minutes.
Research on segmentation gives us a more useful principle. Rey and colleagues’ 2019 meta-analysis in Educational Psychology Review combined 56 investigations and 88 pairwise comparisons. Learners performed better on retention and transfer measures when multimedia instruction was divided into meaningful, coherent segments instead of presented continuously. Segmentation also reduced overall cognitive load, although learners spent more time learning.
Our recommendation is to focus on structure rather than a magic number. A coherent 12-minute software walkthrough can make more sense than chopping the same workflow into three arbitrary four-minute videos.
Show learners exactly where to look
If you’re explaining one button in a complicated interface, don’t make the learner search the entire screen for it.
Schneider and colleagues examined this idea in a 2018 meta-analysis of signaling in multimedia learning published in Educational Research Review. Signaling means adding cues that direct attention toward relevant information. The analysis included 103 studies and 12,201 participants spanning 46 years of research.
Signaling improved retention, with an effect size of g = 0.53, and transfer at g = 0.33. The analysis also found significantly lower cognitive load when signaling techniques were used.
In a training video, that can mean highlighting a menu, zooming into a relevant part of the screen, adding an arrow, or visually emphasizing an important phrase.
There is also research specifically on software training. A 2023 study of video-based software tutorials had 114 undergraduate students learn software through different tutorial designs. The researchers found that signaling and the way practice was arranged each made unique contributions to task performance and self-efficacy.
That’s a useful principle for software training: the learner should be thinking about why they’re taking an action, not trying to figure out where your cursor went.
Use video when learners need to see something happen
Video is most defensible when movement, sequence, timing, or changes over time are part of the information being taught.
Höffler and Leutner’s peer-reviewed meta-analysis in Learning and Instruction compared dynamic and static instructional visuals across 26 studies and 76 pairwise comparisons. Dynamic visuals produced a moderate overall learning advantage, d = 0.37. The effect was larger when the material was highly realistic or video-based, d = 0.76, and when learners were acquiring procedural-motor knowledge, d = 1.06.
There’s an important limitation: only six of the 26 studies actually used video clips. Most tested computer animations against static pictures. So this isn’t evidence that every training topic should become a video. It supports the narrower idea that dynamic visuals are especially useful when the dynamic process itself carries information.
That’s a natural fit for software workflows, SOPs, equipment demonstrations, and other sequential procedures. Visla’s Screen Step Recorder records clicks, typing, screens, and other actions, then turns distinct steps into editable scenes with on-screen text and visual guidance. For conventional walkthroughs, AI Screen Recording can capture the workflow, split the recording into scenes, clean up the presentation, and open it in the Scene-Based Editor.
Remove visuals that don’t help teach the lesson
Interesting additions can hurt learning when they compete with the information learners actually need.
A 2026 meta-analysis in Educational Psychology Review examined 177 effect sizes from 50 studies on “seductive details,” meaning interesting but instructionally irrelevant material. The researchers found a small negative effect on overall learning, g = -0.16, with negative effects on recall, comprehension, and transfer. Their analysis found that increased extraneous cognitive load was the primary mechanism explaining the effect.
The study wasn’t specifically testing corporate B-roll, so we shouldn’t pretend it was. Our takeaway is to apply the same coherence principle when editing training video. A diagram that explains the system earns its place. An arrow showing the correct menu earns its place. Relevant B-roll can earn its place, too.
Generic footage added purely because the screen hasn’t changed recently deserves more skepticism.
Use AI to cut production work, then spend that time on teaching
We don’t recommend AI avatars because we think the avatar itself automatically improves learning.
A randomized 2026 study of AI-avatar instruction compared two ways of teaching introductory Adobe Illustrator to 81 first-year university students. One group received static text-and-image materials. The other watched an AI-avatar video covering the same learning objectives and worked example. The text-and-image group performed better on the immediate knowledge test, while the researchers found no significant difference in practical-task performance or the seven-day delayed test.
That experiment doesn’t prove that AI avatars reduce learning. The formats also differed in pacing, navigation, signaling, verbal presentation, and other ways, so the researchers explicitly warned against attributing the result to avatar presence alone.
The clearer case for AI in L&D is production.
Traditional presenter-led video can involve finding a presenter or actor, scheduling a shoot, setting up a location, arranging cameras and audio, recording multiple takes, and editing everything afterward. Updating one section later may require some of that work again.
Visla’s AI Video Agent can take an SOP, script, webpage, PDF, presentation, or other source material and turn it into an editable video draft. AI Director Mode can plan and generate custom multi-scene videos while maintaining recurring characters, objects, products, environments, and visual style. For presenter-led material, Advanced Avatar can provide a presenter with lip sync, expressions, gestures, and body language without requiring someone to record every version of the lesson.
Our recommendation is to use that saved production time deliberately. Spend less time coordinating shoots, retakes, routine edits, and updates. Spend more time designing questions, building practice, checking instructional accuracy, gathering learner feedback, and teaching.
Make the video serve the lesson
A polished training video can’t rescue a poorly designed lesson. Make video production easier, then put the time you save into helping people actually learn.
Spend less time producing training videos and more time designing the lesson. Create your next training video with Visla.
FAQ
For prerecorded training videos with meaningful audio, captions are a strong default because they make the content accessible to employees who are Deaf or hard of hearing, and WCAG 2.2 includes prerecorded captions as a Level A success criterion. Captions can also help people watching without sound and employees who are less fluent in the video’s spoken language. A 2025 meta-analysis of 49 experiments found a medium positive effect of captioned viewing on second-language vocabulary learning, although that research doesn’t prove captions improve every type of workplace learning. For L&D teams, the safest practical approach is to caption training by default and review automatically generated captions for accuracy, speaker identification, and important non-speech audio.
Don’t stop at views, completion rates, or a post-video satisfaction survey; the more important question is whether employees can use what they learned on the job. A 2026 systematic scoping review of 45 workplace-learning studies found that training-transfer evaluation remains fragmented, with organizations and researchers using standalone instruments, combined tools, and measures tailored to specific workplaces. A separate 2025 systematic review of workplace e-learning found that self-reports were by far the most common transfer measure, while objective measures appeared in only seven of the 31 studies it reviewed. When possible, pair a knowledge check with a job-relevant performance measure, such as whether an employee can complete the trained workflow correctly or whether errors fall after training.
No; research doesn’t show that simply putting a human instructor on screen consistently improves learning. A 2023 meta-analysis covering 27 years of instructor-presence research found a benefit for retention but not transfer, while instructor presence also improved some social and motivational measures and reduced attention to relevant visual information. Three 2024 field experiments conducted inside actual university courses likewise found no general improvement in learning outcomes from showing the instructor, although some learners reported greater social presence or well-being. For L&D teams, that means a presenter should have a specific job to do, such as modeling a conversation, establishing a personal connection, or demonstrating a physical behavior, rather than appearing simply because training videos are expected to have a talking head.
May Horiuchi
May is a Content Specialist and AI Expert for Visla. She is an in-house expert on anything Visla and loves testing out different AI tools to figure out which ones are actually helpful and useful for content creators, businesses, and organizations.

