How to transcribe audio to text
Start with the recording you already have. Upload it, check the language and price, then review the words before you use them.
1. Save the recording as a file
TranscribeCat accepts MP3, WAV, M4A, MP4, FLAC, OGG, WebM and Opus files, up to 500 MB and ten hours long. A video file works too: the transcript comes from its audio. You do not need to extract the soundtrack first.
Use the original recording when possible. Converting an MP3 to WAV does not restore detail already lost during compression. If you recorded on a phone, save or share the voice memo to Files, then upload the M4A file. See the specific guides for MP3 and M4A voice recordings.
2. Upload and check the price
Drop the file on the homepage to see its price. Sign in to continue to payment. Choose the language spoken in the recording, or keep automatic detection. The language list explains the available choices and limitations.
The rate is $2.00 per hour with a $2.00 minimum per file. A 20-minute file costs $2.00; a 90-minute file costs $3.00. Several short files each have their own minimum. The pricing page shows the calculation before you start.
Payment first authorizes a card hold. It is captured when the transcript is ready, or cancelled if processing fails. Processing time varies, and an email tells you when the transcript is ready.
3. Listen while you review
Open the transcript alongside the audio. Rename speaker labels so the conversation is easier to follow, and correct the text where needed. Save your edits before downloading.
- Check names, dates, prices and technical terms against the recording.
- Listen again wherever two people speak at once.
- Verify an exact quotation in the audio before publishing it.
- Do not invent words to fill a passage you cannot hear.
The English and Japanese sample transcripts show unedited output. They are examples you can inspect, not an accuracy guarantee for your recording.
4. Download the format you need
- TXT: plain text for notes, search or pasting into another tool.
- DOCX: a Word document for editing and sharing.
- SRT or VTT: timestamped text for subtitles. Review cue text and timing in the player where you will use it.
- PDF: a document for reading when the transcript’s writing system is supported by the PDF font.
TXT, DOCX and PDF include saved text edits. SRT and VTT keep speaker renames but use the original transcript wording. Correct subtitle wording, timing and line breaks in a subtitle editor before publishing.
Transcription writes down the spoken language. It does not translate the recording, summarize it, or guarantee publication-ready captions. For a longer interview, use the interview workflow; for a final pass, use the review checklist.