Create an SOP from a screen recording
Turn a screen recording into a step-by-step guide with screenshots, then download it as PDF or Word. Runs on your device; nothing is uploaded.
Loading tool…
- Screenshots are taken where the picture changes and where the speaker moves to the next step; what is said becomes the instructions.
- Using the speech in the video downloads an open speech-recognition model once (about 81 MB), kept by your browser for next time.
- Works best in Chrome, Edge or Safari on a computer. Keep the tab open while it works.
- Everything runs on this device; nothing is uploaded.
Record yourself doing a task once, and this tool turns the recording into a written standard operating procedure: a title, a short introduction, numbered steps each with a screenshot, warnings and tips, and a closing line. It works on the video file in your browser. The frames are compared to find the moments where the screen changes, the speech is transcribed on your device with Whisper (the same open speech-recognition model as the video to transcript tool), and the two are combined: a step starts where the screen changes or where you say "first", "next", "then" or "after that" (or "أولاً" and "بعد كده" in Arabic).
The instructions are what you said, cleaned of filler words such as "um", "you know" and "يعني". Sentences with "be careful", "don't" or "انتبه" become warnings, and sentences with "tip" or "you can also" become tips. This is a set of simple rules, not a writing AI: it cannot understand what is on the screen and it does not invent steps, so read the draft and fix it. Everything is editable before you download: rename steps, rewrite the text, move steps up or down, delete them, pick a better screenshot from a strip of frames or from the video player, or add a step at the moment the player is showing.
The finished guide downloads as a PDF with a cover page, numbered steps with their screenshots and page numbers, laid out right to left when the guide is in Arabic, or as a Word document you can keep editing. You can also start from the screen recorder: record, then click "Create SOP from this recording" and the video (and its transcript, if you made one) opens here without being uploaded. Videos must be in a format your browser can play (MP4 or WebM from the screen recorder always work); the speech model downloads once, about 44 to 81 MB depending on the quality you choose.
How to use it
- Record the task with the screen recorder (speak through the steps as you go), or drop an existing screen recording here.
- Keep "Use what is said in the video" on, choose the spoken language and the Fast or More accurate speech model, then click Create steps.
- Wait while the tool looks for screen changes, transcribes the speech and takes a screenshot for each step.
- Edit the title, introduction, steps, warnings and tips; reorder or remove steps, choose another screenshot, or add a step at the player position.
- Download the guide as PDF or Word.
Frequently asked questions
Is my recording uploaded?
No. The video is played, sampled and transcribed inside your browser tab, and the PDF and Word files are made on your device. The only download is the speech model (once, from this site), which your browser keeps for next time.
Does it use AI to write the steps?
It uses Whisper, an open speech-recognition model, to turn your voice into text on your device. The steps themselves are put together by simple rules: screen changes and phrases such as "next" start a step, filler words are removed and cue words mark warnings and tips. Nothing is rewritten or invented, so the quality depends on how clearly you explained the task while recording.
What if I recorded without speaking?
Turn off "Use what is said in the video". Steps are then cut where the screen changes, each with its screenshot, and you type the instructions. If the screen barely changes, the video is split into evenly spaced steps you can adjust.
Can I change the screenshot of a step?
Yes. "Choose another screenshot" shows six frames from that step, and "Use the player’s frame" takes the frame currently shown in the video player, so you can pause exactly where you want.
Does it work in Arabic?
Yes. Choose Arabic as the spoken language (or let it detect it). Arabic cue words such as "أولاً", "بعد كده", "انتبه" and "نصيحة" are recognised, and the PDF and Word files are laid out right to left. Use the Arabic page for Arabic labels in the guide.
How long can the recording be?
Up to 60 minutes on a computer and 20 minutes on a phone, the same limits as the transcription tools. Short, focused recordings of one task give the clearest guides.
Related tools
- Record your screen onlineRecord your screen, a window or a tab with system audio, microphone and an optional webcam bubble.
- Video to text (transcript)Turn speech in a video into a timestamped transcript on your device, then export TXT, Word, SRT or VTT. Nothing is uploaded.
- Extract frames from a videoSave a video frame as a full-resolution PNG or JPG, or extract many frames at once as a ZIP.
- Summarise a PDF, Word or text documentSummarise a PDF, Word or text file on your device: key points, action items, tables to Excel, details found and passage search with page numbers.
- Convert Word to PDFConvert a .docx into a clean PDF with headings, lists, tables and images, in your browser.
- Fix my fileDrop any file to find out why a site, portal or app rejects it (size, format, location data, damaged PDF, video codec) and fix it in one click.