Published November 8, 2025 in Tutorials · Updated September 28, 2026
How to Make Software Tutorial Videos Customers Finish
Founder & Content Creator at Vibrantsnap · 15 min read
A customer writes to support: how do I export last month's inspections as a PDF? Your customer success lead replies with a link to a 12-minute recording of the whole reports module, made in one take on an account called "Test Company 2". The answer is in there, around minute seven. The customer never gets that far and opens a second ticket.
The problem with that video isn't production value. It answers the wrong question, starts on the wrong screen and hides the one click that matters. This guide is for the people who record tutorials at small B2B software companies: founders, product managers and customer success leads, often narrating in a second language. It covers how to plan, record, edit, publish and maintain a software tutorial your customers watch to the end, with the research behind each practice where research exists. The method works with any tool.

The short version
A software tutorial video is a short screen recording that teaches an existing customer one task in your product, narrated step by step. To make one that customers finish:
- Scope one task per video and title it in the customer's words. Aim for one to three minutes.
- Start on the screen your customer starts on, with their role and their kind of account.
- Record in a demo account with realistic data, never with real customer records.
- Plan the steps, then talk. A script is a plan, not something to read aloud.
- Name every button exactly as it is labeled, while you click it.
- Pause after each click so the viewer sees what changed.
- Cut everything that isn't the task, zoom in on each click and add captions.
- Publish the video where the question comes up, with the steps written below it.
- Measure views per video and tickets on the topic.
- Track which video shows which screen, and re-record when a screen changes.
The rest of this guide explains each step and why it works.
Why software tutorials need their own rules
A tutorial is not a demo. A demo shows a prospect what the product can do for them, and we cover it in how to create SaaS demo videos that convert. A tutorial is for someone who already bought, often in the middle of a task, with your product open in the next tab. They came with one question, and they leave as soon as it's answered.
One of the largest studies of how people watch instructional video supports that picture. Philip Guo, Juho Kim and Rob Rubin analyzed 6.9 million video-watching sessions across four edX courses (Guo et al., 2014). Median engagement time never went past 6 minutes, whatever the length of the video, and viewers often stopped before the halfway point of videos longer than 9 minutes. Tutorials behaved differently from lectures: viewers watched 2 to 3 minutes of each tutorial on average, regardless of its length, and re-watched tutorials more often. The authors' advice for tutorials was to support re-watching and skimming.
Those viewers were students in online courses, and the study was observational, not a controlled experiment. But a customer who opens your help center with one question is in the same position as those tutorial viewers: looking for the part that answers it. Your job is to make that part the whole video.
The second body of evidence is Richard Mayer's research on multimedia learning: controlled experiments on how people learn from narration combined with visuals. His 2023 summary reports, for each design principle, how many experimental tests supported it and the median effect size, based on his 2021 meta-analysis. Several of those principles map directly onto the steps below.
If you're recording for your own staff rather than for customers, with an LMS and completion tracking, our guide on how to make training videos covers that case.
Step 1: Scope one task per video
Scope each video to one thing a customer wants done: "Schedule a recurring inspection", not "The inspections module". Use the verb and the nouns a customer would type into your help center search, because that's how they'll look for it.
One task per video pays off three times:
- Customers find the answer by title, and the video ends when their question does.
- People learn better in segments. Mayer's segmenting principle, that people learn more deeply from a lesson broken into learner-paced parts than from one continuous presentation, held in 7 of 7 tests (median effect size 0.67). Shorter videos are one way to segment, as Cynthia Brame notes in her review of educational video research.
- Updates stay cheap. When a screen changes, one three-minute video goes stale instead of a twenty-minute tour.
If a task needs more than about three minutes, it's usually two tasks: "Create an inspection template" and "Assign a template to a site". Number them and link each one to the next.
Step 2: Start where the user starts
When you sit down to record, you already have the right page open, you're logged in as an admin, and your account is full of data. Your customer is on the dashboard, with a standard role and maybe an empty account. Recording from where you are, not from where they are, is a common reason customers can't follow a tutorial.
Decide three things before you press record:
- The first screen. Start where the customer is when the question comes up, usually the home screen, and show the path to the task. What's obvious after two years in the product isn't obvious in week one, so don't skip the navigation.
- The role. If the task needs admin rights, say so in the first sentence. If menus change by role, record with the role most customers have.
- The account state. New customers see empty states. An onboarding tutorial recorded in a full account shows lists and buttons they don't have yet. Record those in a fresh account.
Then open with the outcome and the prerequisite in one or two sentences: "In two minutes you'll export last month's inspections as a PDF. You need the Manager role." Skip the greeting and the product introduction. The customer already owns the product.
Step 3: Set up realistic demo data
Viewers read everything on screen. A customer list full of "Test Test" and "asdf" makes the product look unfinished, and it makes the steps harder to follow because nothing on screen means anything. Real customer data is worse: once the video is shared, so is their data.
- Build one demo account and reuse it for every tutorial, so the same sites, people and records appear across your library.
- Use records from your customers' world: believable sites, work orders, dates and amounts, spelled correctly.
- Follow one record through the video. If the task changes a record, show the same record before and after.
- Clean the frame. Record a single browser tab when you can, turn notifications off, close unrelated tabs, and raise the browser zoom a step so small labels stay readable.
- Do a 30-second test recording to check the microphone, the recorded area and the cursor.

Step 4: Write the script as a plan, not a teleprompter
Reading a script word for word makes most people sound flat, and it's harder still in a second language. Guo's team recommended extemporaneous speaking, and found that videos with a more personal feel could be more engaging than high-fidelity studio recordings. What you need is a plan: the steps in order, the exact button names and what each step changes on screen. Then you talk.
TASK Export last month's inspections as a PDF
WHO Site managers (Manager role needed)
START Dashboard
OUTCOME In two minutes you'll have last month's inspections in one PDF.
STEP 1 Click Reports in the left menu → the report list opens
STEP 2 Set Date to Last month → only last month's inspections remain
STEP 3 Click Export, then PDF → the download starts
RESULT Open the PDF: one page per inspection, photos included
NEXT Next video: send this report automatically every month.
Three rules for the words:
- Talk to one person. Say "you" and "your", not "the user". In Mayer's research, conversational wording ("your lungs" rather than "the lungs") improved learning in 13 of 15 tests (median effect size 1.00).
- Describe each step as action, then result. What you click, then what changes. Viewers check their own screen against yours after every step.
- Leave out everything that isn't the task: the feature's history, the other export formats, the shortcut only power users need. Mayer's coherence principle, that people learn more deeply when extraneous material is left out, held in 18 of 19 tests (median effect size 0.86).
The free demo script generator has a customer tutorial template: list your steps and it returns a timed plan of what to say and what to show.
Step 5: Record: name it, click it, pause
- Say each button's name exactly as it's labeled. If the button says Export, don't say "download it". If the menu path is Reports, then Scheduled, say that. Customers scan their own screen for the words they just heard, and exact labels keep captions and translations matching the interface.
- Say it while you click it. Mayer's temporal contiguity principle, that people learn better when words and the matching visuals arrive at the same time, held in 8 of 8 tests (median effect size 1.31). Pointing the cursor at a control as you name it is also a cue, and cues that highlight the essential material (the signaling principle) helped in 26 of 28 tests.
- Pause after each click. Hold still for a beat so the page loads and the viewer sees the result. Don't narrate over a spinner. This isn't about speaking slowly: Guo's team found viewers engaged more with instructors who spoke fairly fast and with enthusiasm, and advised against slowing down on purpose. The pause is for the screen, not for your voice.
- Move the cursor with purpose. No circling, no selecting text while you think. Viewers' eyes follow the cursor.
- Don't restart on a stumble. Pause, say the sentence again and carry on. Guo's team gave instructors the same advice: speak naturally and remove pauses and filler words in editing.
Narrating in a second language
At small software companies, the person narrating tutorials often speaks English as a second language, for customers who may not be native speakers either. The usual reaction is to write out every sentence and read it, which brings back the stiffness Step 4 warns about. These help more:
- Keep interface words exactly as on screen, in the interface's language, even when the sentence around them is in another language. That's what the viewer matches against.
- Use short sentences, one step each. They're easier to say, to caption and to translate.
- Narrate in the language you think in, if your tool allows it. Some tools transcribe your narration and rewrite it into another language, so a founder can explain in French and publish in English. Check that product names and interface labels keep their spelling.
- Always add captions. Morton Ann Gernsbacher's review of more than 100 studies found that captions improve comprehension of, attention to and memory for video, and that they particularly help people watching in a non-native language. Mayer's research finds that on-screen text repeating the narration can distract in fast lessons, but that printed words can help when they are technical terms or in the viewer's second language, or when the viewer controls the pace. A help center tutorial usually meets all three conditions.
- Don't rule out a synthetic voice. The classic research favored human voices over machine voices (6 of 7 tests in Mayer's summary). Newer work questions how far that holds with modern speech engines: in a 2017 randomized trial, Scotty Craig and Noah Schroeder found that a modern text-to-speech voice, used by an on-screen virtual human, was rated as credible and as helpful for learning as a recorded human voice, and outperformed an older engine.
This is the workflow Vibrantsnap is built around. You narrate in your own words, in the language you're comfortable in. The AI transcribes your narration, rewrites it into a clean script and reads it in one of 30 studio voices, in any of 10 languages: English, French, Spanish, Portuguese, German, Italian, Japanese, Chinese, Korean and Hindi. You can record in one language and publish in another, and interface labels and product names keep their spelling. You can also keep your own voice, and switch at any time.
Step 6: Edit for a viewer with the product open
Four edits do most of the work:
- Cut everything that isn't the task. Loading screens, wrong clicks, "so, let's see". If it doesn't move the task forward, it goes.
- Zoom in on each click. In an embedded player, a full SaaS interface shrinks labels until they're hard to read on a laptop and unreadable on a phone. A close-up on the control being clicked, then back out, tells the viewer where to look.
- Add captions, then read them once. Tutorials get watched with the sound off: in open offices, during meetings, on job sites. Automatic captions can still misspell product names, so check them.
- Leave out music, animated intros and effects between steps. For someone following along, they're extraneous material.
If your editor supports on-screen text, a few words per step ("Step 2: Filter by date") work as signposts for people skimming back to one step, which is what Guo's team suggested for tutorials. Keep them short.
In Vibrantsnap, these edits happen once you stop recording. The rewrite drops hesitations and false starts, and the video is retimed to match the new narration. Auto-zoom gives every click a close-up, using the cursor data the Chrome extension captures in web pages. On desktop software or an imported video there is no cursor data, so you place the zooms yourself. Captions are generated word by word and stay editable.

Step 7: Publish where the question comes up
A tutorial nobody finds does nothing. Put each video where the customer is when they need it, and give each placement its own measure of success.
| Placement | When the customer sees it | Length | What to measure |
|---|---|---|---|
| Help center article | Searching for an answer mid-task | 1 to 3 min, one task | Views; tickets on that topic |
| Support reply or macro | Right after asking your team | The exact task they asked about | Repeat tickets on the same question |
| Onboarding email | First days after signup | One first task per email | Share of new accounts that complete the task |
| In your product | On the screen where the task happens | Under 2 min | Use of the feature after publishing |
| Release notes | When a feature ships or changes | 30 to 90 s | Adoption of the feature |
A few details make the difference:
- Write the steps below the video. People skimming and people who can't play sound get the answer, and search engines and AI assistants can read text in a way they can't read a video.
- Turn autoplay off in help center embeds, and title the video with the task.
- End with the next task, linked: "Next: send this report automatically every month."
- In your product, link or embed the video from the help widget, empty state or tooltip your product already has.
For the program around these videos, see our guide to customer education videos.
Step 8: Measure what the tutorial changed
- Views per video show which tasks customers actually need help with. A video with no views is either hard to find or not needed. Check the placement before you re-record anything.
- Tickets on the covered topic. Tag tickets by topic and compare the same number of weeks before and after publishing. If customers still write in after receiving the video link, the video probably skips a step.
- Help center searches with no good result are your list of the next videos to record.
- Second-by-second retention helps at scale. If you need it, use a video host that reports it. Many small teams learn more from the three signals above.
In Vibrantsnap, Pro and higher plans show the views on each share link, day by day. It doesn't report completion, drop-off points or who watched.

Keep tutorials current when the UI changes
SaaS interfaces change often, and a tutorial showing an old screen teaches it with full confidence. Libraries that stay accurate follow a routine:
- Keep an inventory: one row per video with the task, the screens it shows, every place it's published, the script and the date someone last checked it.
- Add one question to your release checklist: which tutorials show a screen this release changes?
- Re-record when the screen changes. With one task per video, that's a three-minute recording, not a project.
- Re-voice when only the words change. A renamed plan, a step described badly or a new language doesn't need new footage if the screen is still right. Rewrite the script and regenerate the voice.
- Update every placement in the inventory, not only the help center.
In Vibrantsnap, each chapter of a video has its own editable script, so fixing one step's narration doesn't touch the rest, and a new recording is processed in about two minutes.
Mistakes that make customers stop watching
- Opening with a greeting, your name and a product introduction.
- Covering three tasks in one video.
- Starting on a screen the customer never sees.
- Hunting for the button on camera.
- Saying "click here" instead of the label.
- Skipping the "obvious" step, like where the menu is.
- Test data, or worse, real customer data.
- Ending without a next step.
Where Vibrantsnap fits
Vibrantsnap is a product demo and tutorial video maker for software teams who record their own videos. The workflow in short:
- Record a browser tab, a window or your whole screen with the Chrome extension, narrating as you click, or import a video you already have. Recording length is never capped.
- Let the AI edit. Your narration is transcribed, rewritten into a clean script and read in a studio voice, and the video is retimed to match. Clicks get a close-up and captions are generated word by word.
- Adjust. Edit the script chapter by chapter, change the voice or the language, move a zoom, and pick a background and a format: 16:9, 4:5, 1:1 or 9:16.
- Share. Publish a share link with a call-to-action button, embed the video in your help center, or export an MP4 in 1080p or, from Pro, 4K.
Recording, editing and previewing with every AI feature is free; the free share link carries a watermark and a Vibrantsnap badge, and exporting needs a paid plan. Pro is $49/month, or $40/month billed yearly, with 25 min of export a month, link analytics, 4K and no watermark. See pricing for Scale and Business.
Frequently asked questions
How long should a software tutorial video be? Long enough for one task, which usually means one to three minutes. In an edX study of 6.9 million viewing sessions, viewers watched tutorial videos for 2 to 3 minutes on average whatever their length, and median engagement never went past 6 minutes for videos of any length. If a task needs longer, split it into numbered videos.
Do I need to show my face in a software tutorial? Not for a software tutorial. The research is mixed: in the edX study, cutting between the instructor's face and slides made videos more engaging than slides alone, while Richard Mayer's summary finds that a static image of the speaker does not reliably improve learning. For a tutorial, the practical point decides it: the interface is the lesson, and a camera bubble covers part of it. Put the effort into a clear voice and a close-up on each click instead.
Should I write a script for a tutorial video? Write a plan, not a text to read aloud: the outcome, any prerequisite, then each step as the button you click and what changes on screen. Then narrate in your own words. Read-aloud scripts tend to sound stiff, and conversational wording that talks to the viewer as "you" helps people learn.
How do I narrate a tutorial if English is not my first language? Narrate in the language you are comfortable in and clean up the words afterwards, or record in your own language and publish a rewritten script in your customer's language with an AI voice. Either way, say button names exactly as they appear on screen and add captions, which research shows help people watching in a non-native language.
How do I keep tutorial videos up to date when the UI changes? Record one task per video, keep the script with the video, and add a question to your release checklist: which tutorials show a screen this release changes? Re-record a video when its screen changes. If only the words are wrong, rewrite the script and regenerate the voice instead of recording again.
Where should I publish software tutorial videos? Where the question comes up: embedded in the help center article for the task, linked in support replies, in onboarding emails, next to the feature in your product and in release notes. Then track views per video and the number of tickets on each covered topic before and after publishing.
What equipment do I need to record a software tutorial? A computer, a browser and a decent microphone in a quiet room. A USB or headset microphone is enough, and soft furnishings cut echo. You don't need a camera or lights for a screen tutorial. In the edX study, videos with a more personal feel could be more engaging than high-fidelity studio recordings.
Sources
- Philip J. Guo, Juho Kim and Rob Rubin, "How video production affects student engagement: an empirical study of MOOC videos", Proceedings of the first ACM Conference on Learning @ Scale, 2014 (PDF, ACM Digital Library)
- Richard E. Mayer, "Research-based principles for designing multimedia instruction", in C. E. Overson et al. (Eds.), In Their Own Words, Division 2 of the American Psychological Association, 2023 (PDF)
- Cynthia J. Brame, "Effective educational videos: principles and guidelines for maximizing student learning from video content", CBE—Life Sciences Education, 2016 (CBE—LSE)
- Morton Ann Gernsbacher, "Video captions benefit everyone", Policy Insights from the Behavioral and Brain Sciences, 2015 (PubMed Central)
- Scotty D. Craig and Noah L. Schroeder, "Reconsidering the voice effect when learning from a virtual human", Computers & Education, 2017 (ScienceDirect)
Comparison