Is this the right transcription workflow?
This page is for creating timed subtitle text from spoken media, with SRT export available as a Pro format after transcription and review.
Typical sources for this workflow include: podcast audiograms, lessons, interviews, presentations, social video audio, and other recordings that need timed captions.
Decision boundary: Use MP3, M4A, WAV, or MP4 to Text when a plain editable transcript is the primary result. SRT and VTT exports require Pro.
When this page is the right fit
Use this workflow when the source and intended result match before you begin. That choice keeps the upload, review, and export steps aligned with a reviewed SRT subtitle file.
- Source fit: podcast audiograms, lessons, interviews, presentations, social video audio, and other recordings that need timed captions.
- Preparation priority: Use the final media cut whenever possible; changing the edit after subtitle generation can make every later cue drift out of sync.
- Review priority: Check every subtitle against the audio and inspect cue timing, reading order, line breaks, names, numbers, and gaps after edits.
What to check before transcription
Play a short section before uploading. Confirm that it contains audible speech, opens correctly, and matches the recording you intend to transcribe.
- File check: confirm that an audio or video recording is complete and not an empty, damaged, or incorrectly named file.
- Speech check: listen for clipping, very low volume, overlapping speakers, loud music, or persistent background noise.
- Context check: keep a short spelling list for names, organizations, places, abbreviations, and specialist vocabulary.
- Task-specific check: Use the final media cut whenever possible; changing the edit after subtitle generation can make every later cue drift out of sync.
Three Steps to an Editable Transcript
This is batch transcription. A microphone recording is captured first; after you stop and review it, you submit it for transcription. Text does not appear live while you speak.
Record or Upload
Record a voice note, or choose a supported audio or video file from your device.
Submit for Processing
After recording or upload, send the file to the transcription service and follow its progress.
Review and Export
Review and edit the transcript, then copy it or export it in the format you need.
How to review and export the result
Treat the first transcript as an editable draft, not as a certified record. Check every subtitle against the audio and inspect cue timing, reading order, line breaks, names, numbers, and gaps after edits.
Listen again wherever wording changes a decision, quotation, instruction, amount, date, address, or legal or medical meaning. Correct the transcript before sharing, publishing, or using it as a source document.
The free plan provides TXT and JSON exports. SRT, VTT, PDF, and DOCX are Pro export formats. Copying and editing the transcript before export lets you correct the content once and then choose the appropriate output.
| Intended result | Available route | What to verify |
|---|---|---|
| Editable text | Copy, TXT, or JSON; TXT and JSON are available on the free plan | Paragraph breaks, names, numbers, terminology, and omitted words |
| Timed subtitles | SRT or VTT export on Pro | Text, cue timing, reading order, and line breaks against the media |
| Formatted document | PDF or DOCX export on Pro | Transcript accuracy and formatting before creating the final file |
Accuracy, privacy, and practical limits
Results vary with language, recording quality, background noise, and speaker clarity. A human review is necessary whenever the wording matters.
Clear speech and a clean recording usually require fewer corrections than distant speech, heavy compression, echo, music, or simultaneous speakers. The service does not promise a fixed accuracy percentage for every file.
Processing requires a secure upload to the transcription service. Review the site privacy information and your own confidentiality obligations before submitting private, regulated, or third-party material.
- Page-specific risk: a media edit that changes after transcription, overlapping dialogue, music over speech, long unbroken cues, or timing that was not reviewed.
- You remain responsible for checking the transcript and for having permission to process and use the recording.
Choose this page or an adjacent workflow
Choose by the source you already have and the output you need. The closest-sounding page is not always the most useful route.
| Your situation | Best route | Why |
|---|---|---|
| You have an audio or video recording and need a reviewed SRT subtitle file | Use this page | This page is for creating timed subtitle text from spoken media, with SRT export available as a Pro format after transcription and review. |
| Your source or final output falls outside this page’s scope | a format-specific plain-text transcription page | Use MP3, M4A, WAV, or MP4 to Text when a plain editable transcript is the primary result. SRT and VTT exports require Pro. |
| You have another supported audio or video format | Use the general speech-to-text workspace | It accepts the complete supported format set without forcing a format-specific route. |
What the transcription workspace looks like
The interface separates source selection from processing. You first upload a supported file or switch to the recorder, then submit the prepared recording.
The image below is an authentic capture of the current workspace. Use it to identify the source controls and the processing action before working with your own recording.

Completion checklist before you use the transcript
A transcript is ready only after the source, wording, and intended output have been checked together.
- Confirm that the transcript belongs to the correct recording and that the recording was processed from beginning to end.
- Correct Check every subtitle against the audio and inspect cue timing, reading order, line breaks, names, numbers, and gaps after edits.
- Recheck the parts most exposed to this risk: a media edit that changes after transcription, overlapping dialogue, music over speech, long unbroken cues, or timing that was not reviewed
- Open the exported file and verify that its text, timing, or document formatting matches your intended use.