Skip to content
Trustample

What is audio transcription?

The plain definition, the two styles nobody explains, and the words it's constantly confused with.

By Anubhav Jain, Trustample founder · Updated July 2026 · Claims verified against the live product

Free: 60 min per 30 days, files up to 30 min / 50 MB · Paid plans: 10 to 100 hrs, files up to 2 GB. Big video? Upload just its audio track — same transcript, much faster upload.

Encrypted in transit · audio & video auto-delete in 30 days · never used to train AI

Optional settings

Set this for the most accurate result on quiet or mixed audio.

Get a second copy translated into another language (Pro).

Uncommon words the AI might misspell — list them so they come out right.

Audio transcription is the process of converting the spoken words in a recording into written text. The recording can be audio or video; the output is a document you can read, search, quote and edit. That's the whole idea — everything else is detail about who does it, how faithfully, and what you get at the end.

It's done one of two ways. A human listens and types, taking roughly 4–6 hours per hour of audio and charging $1–2 per audio minute — still the standard where certified accuracy is required. Or AI transcription does it automatically in about 2–5 minutes per hour of audio, typically 95–99% accurate on clear speech in a major language, with the remaining errors fixed by a quick human pass. The AI route is what most people mean by transcription today, and why the job went from a paid service to something you do yourself before lunch.

One distinction worth knowing up front: transcripts and captions aren't the same thing. A transcript is readable text. Captions are that same text cut into short timed lines that appear on screen, saved as an SRT or VTT file. Modern tools generate word-level timestamps during transcription, so one recording can produce both — but only if the tool exports subtitle formats, which is worth checking before you commit to one.

Frequently Asked Questions (FAQs)

What's the difference between verbatim and clean-read transcription?

Verbatim keeps everything — every "um", false start, stutter and repetition — which research, legal work and linguistic analysis often require. Clean read (also called intelligent verbatim) removes filler and tidies false starts into readable sentences, which is what you want for meeting notes, articles and show notes. AI tools generally produce something close to clean read by default.

Transcription vs translation vs transliteration — what's what?

Three jobs people constantly conflate. Transcription: speech to text in the same language. Translation: text from one language into another. Transliteration: writing one script in another (Japanese kana into romaji, say) without changing the language. If you have a French recording and want English text, that's transcription followed by translation — in that order, because only text can be translated.

What is audio transcription used for?

Meeting minutes and searchable records; research interviews for coding and analysis; journalism, where accurate quotes matter; subtitles and captions for video; lecture notes; podcast show notes and repurposed content; and accessibility, which is a legal requirement for a lot of published video. The common thread: text is searchable, quotable and skimmable, and audio is none of those things.

Is audio transcription accurate?

AI transcription typically reaches 95–99% on clear audio in a major language and drops to 80–90% on noisy recordings, heavy crosstalk, strong dialect or specialist jargon. Accuracy depends on your recording more than on the tool — which is why any honest service publishes ranges rather than a single number, and offers a free tier so you can test on your own audio.

Can I transcribe audio myself for free?

Yes. Trustample's free plan transcribes 60 minutes a month with no credit card, using the same engine and synced editor as the paid plans, with TXT export. Doing it manually is also free and costs you 4–6 hours per audio hour — worth it only when the audio is so poor that AI genuinely can't cope.

What formats does a transcript come in?

Plain text (TXT) for pasting anywhere; DOCX for editing and sharing; PDF for distribution; and SRT or VTT when you need subtitles rather than a document. A good tool exports all of them from one transcript, so you're not re-doing the work per format.

An honest note: Transcription reproduces words, not meaning. It won't reliably capture tone, sarcasm or who was joking, and it can't tell you what a garbled passage really said — it will guess plausibly and quietly. That's why every transcript worth relying on gets a pass against the audio, and why an editor synced to the recording matters more than a vendor's accuracy claim.

Related pages

Explore everything in Answers or browse all transcription tools.

Chat with us