Every voice, word for word.
How accurate is the transcription?
One file, one price. No subscription.
- No account
- First 2 minutes free
- Private: deleted after 24 hours unless you order
Payment canceled. Your preview is still here.
Free preview: the first 2 minutes
·
Your full transcript
$8.90One file, one price. No subscription.
- Verbatim and edited version
- 6 files: Word, TXT and Markdown
- Every speaker labeled, names you confirm
- Summary and key points
- Money back within 7 days, no questions
Secure payment by Stripe: card, Apple Pay, Google Pay
Cards are only charged once your transcript has passed our check.
Sent. The link works for 24 hours.
Dettalo uses ElevenLabs Scribe v2 for speech recognition. Here are the independent benchmark numbers, what they mean, and why the first two minutes of your own recording tell you more than any table.
- Verbatim + edited version
- 6 files: Word, TXT, Markdown
- 90+ languages
- Every speaker labeled
- Money back within 7 days
Two versions of every transcript
Verbatim Every word as spoken, fillers included.
So, um, how did you, how did you start the bakery?
Well, I I started in 2019 with, uh, one oven and my friend Marco Bell Lucci. The first year was really hard.
Edited Corrected, laid out, with headings and a summary. Nothing added.
How it started
So, how did you start the bakery?
Well, I started in 2019 with one oven and my friend Marco Bellucci. The first year was really hard.
The names you type in fix the spelling: “Marco Bell Lucci” becomes “Marco Bellucci”.
What word error rate means
Word error rate (WER) counts the words a transcript gets wrong (replaced, missing or added) and divides them by the number of words actually spoken. A WER of 3% means about 3 errors in every 100 words. Lower is better. Two numbers are only comparable when they come from the same test.
Spanish, Italian, French and German
The Open ASR Leaderboard paper (v4, 30 March 2026, by Hugging Face and others) compares models in a multilingual track on read speech (FLEURS, MLS, CoVoST-2). Word error rate, lower is better:
| Model | Spanish | Italian | French | German |
|---|---|---|---|---|
| ElevenLabs Scribe v2 (used by Dettalo) | 2.33% | 2.58% | 3.28% | 2.27% |
| AssemblyAI Universal-3 Pro (March 2026 version) | 2.34% | 3.94% | 3.74% | 2.34% |
| Cohere Transcribe (open model) | 2.81% | 3.44% | 4.05% | 3.84% |
| Speechmatics Enhanced | 2.78% | 5.58% | 5.04% | 2.84% |
| Voxtral Small 24B (open model) | 3.04% | 3.91% | 4.13% | 3.01% |
| Whisper large-v3 (open model) | 3.65% | 4.69% | 6.36% | 4.26% |
Scribe v2 has the lowest rate in all four languages; in Spanish, AssemblyAI is 0.01 points behind. AssemblyAI is also our fallback when ElevenLabs is unavailable.
English
The Artificial Analysis speech-to-text leaderboard is mostly English, with about 8 hours of conversations, parliament speeches and earnings calls. Word error rates: MAI-Transcribe-2 2.0%, ElevenLabs Scribe v2 2.2%, Gemini 3.5 Transcribe 2.6%, Voxtral Small 2.8%, AssemblyAI Universal-3 Pro 3.1%, OpenAI GPT Transcribe 3.3%, Deepgram Nova-3 5.2%. Here Scribe v2 is second.
ElevenLabs puts English, Spanish, French, Italian, German, Dutch, Portuguese and Japanese in its best tier (“Excellent”, WER of 5% or less), and Mandarin Chinese and Cantonese in the next (“High accuracy”, 5–10%), according to its documentation.
Why real recordings score worse
Read speech is easier than real recordings. Noise, distance from the microphone, people talking over each other and strong accents raise the error rate. A benchmark tells you which engine is good. It can't tell you how your file will turn out.
Check your own file
Upload it and read the free preview of the first 2 minutes, speakers labeled, in both versions. Look at names, numbers, technical terms and passages where people talk at once. If the detected language is wrong, pick the right one and the preview runs again.
To get fewer errors: record close to the speakers in a quiet room, let one person talk at a time, tell us how many people speak, and type up to 60 names and terms to get them spelled right.
If you need a person to check every word, Rev offers human transcription at $1.99 per minute, “99%+ accurate, delivered in 12 hours or less” (as of October 2026).
How it works
Upload
Drop an audio or video file, up to 3 hours. No account needed.
Read the free preview
In about 2 minutes you read the first 2 minutes, speakers labeled, verbatim and edited.
Pay once
Enter your email and pay for this file only. No subscription.
Download
We email you a private link. Confirm the speaker names and download 6 files.
What a transcript costs
| Price | How you pay | |
|---|---|---|
| Human transcription (Rev) | $1.99 per minute: about $119 for one hour | Per minute of audio |
| Subscription apps (TurboScribe, Otter) | $16.99–20 a month, or $8.33–10 a month billed yearly | Monthly or yearly plan |
| Dettalo | $8.90 per file, up to 3 hours | Once per file, no subscription |
Questions and answers
Which transcription engine is the most accurate?
It depends on the language and the test. On the mostly English Artificial Analysis leaderboard, MAI-Transcribe-2 has the lowest word error rate (2.0%), followed by ElevenLabs Scribe v2 (2.2%). In the Open ASR Leaderboard's multilingual track, Scribe v2 had the lowest word error rate in Spanish, Italian, French and German.
Will my transcript be perfect?
No. No speech recognition is perfect: names, numbers, crosstalk, noise and strong accents all cause errors. Read the free preview of the first 2 minutes of your own file before you pay.
What is a good word error rate?
For reference, ElevenLabs calls 5% or less “Excellent” and 5–10% “High accuracy”. At 5%, about one word in twenty is wrong. Your own recording decides where you land.
Why is my transcript worse than the benchmark?
Read speech is easier than real recordings. Noise, distance from the microphone, people talking over each other and strong accents all raise the error rate.
Does the edited version fix recognition errors?
It corrects clear recognition errors and names, and lists every change at the end. It is written from the verbatim text, not from the audio, so it can only fix what the context makes clear. Check names, numbers and quotes yourself.
What if I'm not happy with the accuracy?
Within 7 days of your order you get a full refund with one click on your order page, no reason needed. Your files are deleted at once. See Refunds.