Skip to content
Trustample

Convert audio to text

A free-to-start audio transcriber for any format, any language: one accurate, editable transcript.

Free: 60 min per 30 days, files up to 30 min / 50 MB · Paid plans: 10 to 100 hrs, files up to 2 GB. Big video? Upload just its audio track: same transcript, much faster upload.

Encrypted in transit · audio & video auto-delete in 30 days · never used to train AI

Optional settings

Naming the language beats leaving this on auto. If the recording has more than one, pick the one spoken most.

Get a second copy translated into another language (Pro).

Uncommon words the AI might misspell. List them so they come out right.

Last updated 6 August 2026

Trustample is a general-purpose audio transcriber: MP3, WAV, M4A, AAC, FLAC, OGG and more, converted to text in multiple languages. If it plays, it almost certainly transcribes.

You get more than raw text: every word carries its timestamp, so the transcript stays synced to a built-in audio player. Click a sentence to hear it; click a word to fix it; export in five formats. That editor is what separates a real audio-to-text converter from a dictation toy: transcription does the typing, and the synced editor turns the correction pass into a few minutes of clicking rather than an hour of re-listening. Whether that editor is worth choosing over a cheaper converter is the question the comparison with Transkriptor sets out to answer.

A note on names, because the category is muddled: "speech to text" usually refers to live dictation (you talk, it types), while transcribing audio to text means converting a recording that already exists. The recording route wins on everything except immediacy: full-context punctuation, long files, multiple speakers, many languages, and an editor with the audio attached. If your words are already recorded, this is the tool; if you want to type by talking, your device's dictation is free and built in.

Does the audio format affect accuracy? We ran all five

People convert files before uploading because they assume a lossless format will transcribe better. In August 2026 we tested that directly: one spoken recording, encoded as MP3, WAV, M4A, FLAC and OGG, with every version sent through the identical transcription request.

All five returned the same transcript, word for word. WAV, being uncompressed, was 2,381 KB. The MP3 of the same audio was 433 KB, roughly five and a half times smaller, and produced text that was byte-identical to the WAV. FLAC, the lossless compressed format people reach for when they want quality, landed in between at 936 KB with, again, exactly the same words.

So the format is not the variable. Converting a file before uploading it cannot improve a transcript, and it can only lose information if the conversion is lossy, which is the sort of thing our M4A test actually caught. Upload whatever you have, and if you get a choice, pick the smaller file purely because it uploads faster.

What does change the result is the recording. Microphone distance, background noise, people talking over each other and how clearly they speak account for nearly all of the variation you will see. The same test made this visible in miniature: one name in the script came back slightly wrong in every single format, because the difficulty lived in the word rather than the file.

Getting the best transcript from the audio you already have

Since the recording is what matters, the highest-value habits happen before and after transcription rather than during it. Before uploading, set the spoken language explicitly instead of relying on auto-detect, which is the single biggest avoidable source of error, and type any recurring names, company names or specialist terms into the vocabulary box, a field available on every plan including free.

Afterwards, read the names first. Speech recognition is most reliable on ordinary conversational words and least reliable on proper nouns, so corrections cluster there rather than being spread evenly through the text. If a name recurs, fix it once and search for the other spellings.

For anything longer than a few minutes, the synced editor is the part that saves real time: clicking any line replays that exact moment, so verifying a quote or a figure is a two-second job rather than a hunt through the audio. That is the difference between a transcript you trust and one you have to listen through again.

How it works

  1. 1Drop in any audio file; format detection is automatic.
  2. 2Pick the spoken language (or leave auto-detect) and let the AI work.
  3. 3Edit inline with the synced player, then export TXT, DOCX, PDF, SRT or VTT.

Frequently Asked Questions (FAQs)

Which audio formats are supported?

All common ones: MP3, WAV, M4A/AAC, FLAC, OGG/Opus, WMA and more, plus video containers like MP4 and MOV, from which speech is read directly.

Is there really a free plan?

Yes. It resets every month with 60 minutes of transcription, takes files up to 30 minutes each, exports plain text, and never asks for a credit card. Basic ($12/mo) adds a 10-hour monthly pool, every export format and speaker labels; Pro ($19/mo) adds 20 hours, AI chat and translation. Prices include taxes.

How is this different from live dictation?

Dictation transcribes as you speak and can't go back. Trustample transcribes finished recordings, so it can use full context (punctuation, long files, multiple speakers, noisy audio) and gives you an editor afterwards.

Can I search my old transcripts?

Yes. Your dashboard lists every transcript with instant title search, so a recording from months ago is a few keystrokes away.

Can it transcribe a recording with several speakers?

Yes. On every paid plan, segments arrive labeled Speaker 1, Speaker 2 and so on, renameable once for the whole transcript. Meetings, interviews and panel recordings are the everyday cases this exists for.

How long does transcription take?

Short clips finish in seconds and an hour of audio takes a few minutes, most of which is the upload rather than the transcription. Jobs run in the background with automatic retries, so you can send several files at once and let the dashboard collect the results.

Worth knowing: Accuracy follows the recording, not the format. Our five-format test used one clean, close-mic sample, which is a best case and is why every version matched: it shows the container is not the variable, not that every recording transcribes perfectly. Heavy crosstalk or a far-away microphone can push any engine, ours included, into territory where the editing pass takes real time. Test with your own audio on the free plan before committing to a big batch.

Related tools

See all converters on the transcription tools page.

Browse all transcription tools.

Chat with us