
Quick answer: “99% accurate” is meaningful only if the vendor tells you what is being counted, against what reference, under which audio conditions, and whether speaker labels, punctuation, names, numbers, and formatting are included. In automatic speech recognition, a common metric is Word Error Rate (WER): substitutions + deletions + insertions divided by the number of reference words. Under a simplified accuracy = 1 - WER interpretation, 1% WER corresponds to 99% word accuracy - but commercial services may use different definitions, and WER alone does not tell you whether the errors were harmless fillers or critical medical, legal, or numeric terms.
“99% accuracy” is one of the most repeated claims in transcription marketing. It sounds precise. Often, the measurement behind it is not stated.
VerbalScripts’ own guidance emphasizes that responsible providers should not promise one percentage for every recording because source quality, speaker count, terminology, and review conditions change. See How Accurate Is Professional Transcription?.
NIST uses Word Error Rate as a standard measure in ASR evaluation. The basic formula is:
WER = (Substitutions + Deletions + Insertions) / Number of words in the reference transcript
Where:
• Substitution: the system outputs the wrong word.
• Deletion: a reference word is missing.
• Insertion: the system adds a word that was not spoken.
Reference:
The witness did not enter the building.
System output:
The witness did enter a building.
Potential errors include deletion of “not” and substitution of “the” with “a.” The sentence is still fluent, but its meaning has changed dramatically.
That is why accuracy evaluation should consider error consequence, not just count.
Imagine a 9,000-word transcript. If a service truly has a 1% word error rate on that file, a simple approximation is about 90 word errors.
Those 90 errors could be mostly low-impact fillers - or they could include:
• a party’s name;
• a medication;
• the word “not”;
• a dollar amount;
• a parcel number;
• the identity of a speaker;
• a quoted phrase used in research.
The percentage alone cannot tell you which.
Also note that WER can exceed 100% in pathological cases because insertions add to the numerator. So “accuracy = 1 - WER” is a convenient intuition, not a universal quality definition for every scoring situation.
Depending on the evaluation setup, WER may normalize or exclude punctuation, capitalization, formatting, or other elements. Even when word recognition is excellent, a transcript can still be unusable because of:
“Attorney” and “Witness” can be reversed while every word is correctly recognized.
A long stream of correct words can become difficult to interpret if punctuation is poor. In legal or medical text, punctuation can also affect meaning.
A surname may count as one word error, but correcting 50 participant names across a research dataset consumes significant time.
A clinical note can contain the right words in the wrong section. A hearing transcript can lose Q/A structure. A focus group can have unstable participant labels.
A transcript can be word-accurate but have timestamps shifted enough to make source retrieval frustrating.
Human services often advertise 99% or 99%+ accuracy, but buyers should ask the same questions:
• Is the number measured or aspirational?
• Is it an average across all audio or only clear audio?
• Who creates the reference transcript?
• Are names and speaker labels counted?
• Are inaudible segments excluded?
• Is accuracy measured before or after customer corrections?
• Is there an independent QA sample?
Human review improves the ability to resolve context and uncertainty, but humans are not error-free. A mature vendor should be able to explain the conditions under which accuracy degrades.
Use a scorecard with more than one metric.
Word Error Rate: Overall word recognition errors
Critical Entity Error Rate: Errors in names, dates, amounts, drugs, case IDs, etc.
Speaker Attribution Accuracy: Whether words are assigned to the right person
Omission Rate: Whether meaningful phrases/segments are dropped
Timestamp Accuracy: Ability to return to source
Formatting Compliance: Whether client template/style is followed
Uncertainty Quality: Whether inaudible speech is marked honestly
Customer Cleanup Time: Real operational burden after delivery
For many professional buyers, cleanup time is the most revealing commercial metric.
A 99.2% transcript with one wrong medication dose may be less acceptable than a 98.8% transcript whose errors are filler words in a podcast.
Create a “critical terms” list for the use case.
• names;
• dates;
• exhibit numbers;
• amounts;
• statute/case citations;
• negation;
• speaker identity.
• patient/provider names;
• medications/doses;
• anatomy/laterality;
• measurements;
• negation;
• diagnoses.
• participant ID;
• quoted phrases;
• domain terms;
• speaker label;
• code-switching;
• timestamps.
• motion wording;
• vote result;
• case/parcel numbers;
• conditions;
• names.
An accuracy score can be distorted if the reference excludes hard sections. Real customers do not get to exclude the hard section from their case.
Ask the vendor how it handles:
• fully inaudible passages;
• uncertain single words;
• overlap;
• low-volume speech;
• names not verifiable from audio.
A transparent marker such as [inaudible 00:18:42] is not automatically an “accuracy failure.” It can be the most accurate representation of the source.
This is central to poor-quality audio transcription.
1. Choose 30-60 minutes of representative audio.
2. Include both easy and hard sections.
3. Create or commission a carefully verified reference transcript.
4. Normalize the same way for WER.
5. Score speaker labels separately.
6. Score critical terms separately.
7. Track cleanup time.
8. Test formatting and deadlines.
9. Review how each vendor marks uncertainty.
10. Repeat on more than one file.
Do not compare one vendor’s “99%” claim to another vendor’s “96%” claim unless they use the same test.
• How do you define accuracy?
• Is the figure measured using WER or another metric?
• What audio conditions are included?
• Does it include speaker labels?
• How are punctuation and formatting scored?
• Are inaudible passages excluded?
• Is the figure for AI output, human-reviewed output, or both?
• What happens when the source audio is poor?
• Can I test my own sample before a high-volume order?
A trustworthy answer may be less impressive than a marketing slogan - and far more useful.
Human review is strongest when the reviewer has access to:
• original audio;
• domain glossary;
• case/study context;
• speaker roster;
• reference documents;
• a second-review process;
• explicit authority to mark uncertainty.
The goal is not to produce text that looks correct. It is to produce text that has been checked against the source.
For AI-specific failure patterns, see Can You Trust AI Transcription for Accents, Crosstalk, and Noisy Audio?.
Under a simple word-accuracy interpretation, roughly yes. But the exact relationship depends on the metric, normalization, and treatment of insertions. Ask how the vendor defines the number.
No. WER is a useful word-recognition metric, but it may not capture speaker attribution, punctuation, formatting, timestamps, or the consequence of specific errors.
Yes. Because insertions are included, a system can produce more total errors than reference words in extreme cases.
No responsible vendor should guarantee the same accuracy on every recording. If speech is masked or missing, neither humans nor AI can reconstruct it reliably.
Use word-level scoring plus critical-term, speaker, number/negation, formatting, and cleanup-time metrics relevant to the document’s use.
If accuracy matters enough to influence a case, study, patient record, board decision, or client report, test the vendor with representative audio. Request a VerbalScripts sample assessment or quote and define the critical terms you want checked.
• NIST, Open Speech Analytic Technologies 2020 Evaluation Plan (WER definition)
Quick answer: “99% accurate” is meaningful only if the vendor tells you what is being counted, against what reference, under which audio conditions, and whether speaker labels, punctuation, names, numbers, and formatting are included. In automatic speech recognition, a common metric is Word Error Rate (WER): substitutions + deletions + insertions divided by the number of reference words. Under a simplified accuracy = 1 - WER interpretation, 1% WER corresponds to 99% word accuracy - but commercial services may use different definitions, and WER alone does not tell you whether the errors were harmless fillers or critical medical, legal, or numeric terms.
“99% accuracy” is one of the most repeated claims in transcription marketing. It sounds precise. Often, the measurement behind it is not stated.
VerbalScripts’ own guidance emphasizes that responsible providers should not promise one percentage for every recording because source quality, speaker count, terminology, and review conditions change. See How Accurate Is Professional Transcription?.
NIST uses Word Error Rate as a standard measure in ASR evaluation. The basic formula is:
WER = (Substitutions + Deletions + Insertions) / Number of words in the reference transcript
Where:
• Substitution: the system outputs the wrong word.
• Deletion: a reference word is missing.
• Insertion: the system adds a word that was not spoken.
Reference:
The witness did not enter the building.
System output:
The witness did enter a building.
Potential errors include deletion of “not” and substitution of “the” with “a.” The sentence is still fluent, but its meaning has changed dramatically.
That is why accuracy evaluation should consider error consequence, not just count.
Imagine a 9,000-word transcript. If a service truly has a 1% word error rate on that file, a simple approximation is about 90 word errors.
Those 90 errors could be mostly low-impact fillers - or they could include:
• a party’s name;
• a medication;
• the word “not”;
• a dollar amount;
• a parcel number;
• the identity of a speaker;
• a quoted phrase used in research.
The percentage alone cannot tell you which.
Also note that WER can exceed 100% in pathological cases because insertions add to the numerator. So “accuracy = 1 - WER” is a convenient intuition, not a universal quality definition for every scoring situation.
Depending on the evaluation setup, WER may normalize or exclude punctuation, capitalization, formatting, or other elements. Even when word recognition is excellent, a transcript can still be unusable because of:
“Attorney” and “Witness” can be reversed while every word is correctly recognized.
A long stream of correct words can become difficult to interpret if punctuation is poor. In legal or medical text, punctuation can also affect meaning.
A surname may count as one word error, but correcting 50 participant names across a research dataset consumes significant time.
A clinical note can contain the right words in the wrong section. A hearing transcript can lose Q/A structure. A focus group can have unstable participant labels.
A transcript can be word-accurate but have timestamps shifted enough to make source retrieval frustrating.
Human services often advertise 99% or 99%+ accuracy, but buyers should ask the same questions:
• Is the number measured or aspirational?
• Is it an average across all audio or only clear audio?
• Who creates the reference transcript?
• Are names and speaker labels counted?
• Are inaudible segments excluded?
• Is accuracy measured before or after customer corrections?
• Is there an independent QA sample?
Human review improves the ability to resolve context and uncertainty, but humans are not error-free. A mature vendor should be able to explain the conditions under which accuracy degrades.
Use a scorecard with more than one metric.
Word Error Rate: Overall word recognition errors
Critical Entity Error Rate: Errors in names, dates, amounts, drugs, case IDs, etc.
Speaker Attribution Accuracy: Whether words are assigned to the right person
Omission Rate: Whether meaningful phrases/segments are dropped
Timestamp Accuracy: Ability to return to source
Formatting Compliance: Whether client template/style is followed
Uncertainty Quality: Whether inaudible speech is marked honestly
Customer Cleanup Time: Real operational burden after delivery
For many professional buyers, cleanup time is the most revealing commercial metric.
A 99.2% transcript with one wrong medication dose may be less acceptable than a 98.8% transcript whose errors are filler words in a podcast.
Create a “critical terms” list for the use case.
• names;
• dates;
• exhibit numbers;
• amounts;
• statute/case citations;
• negation;
• speaker identity.
• patient/provider names;
• medications/doses;
• anatomy/laterality;
• measurements;
• negation;
• diagnoses.
• participant ID;
• quoted phrases;
• domain terms;
• speaker label;
• code-switching;
• timestamps.
• motion wording;
• vote result;
• case/parcel numbers;
• conditions;
• names.
An accuracy score can be distorted if the reference excludes hard sections. Real customers do not get to exclude the hard section from their case.
Ask the vendor how it handles:
• fully inaudible passages;
• uncertain single words;
• overlap;
• low-volume speech;
• names not verifiable from audio.
A transparent marker such as [inaudible 00:18:42] is not automatically an “accuracy failure.” It can be the most accurate representation of the source.
This is central to poor-quality audio transcription.
1. Choose 30-60 minutes of representative audio.
2. Include both easy and hard sections.
3. Create or commission a carefully verified reference transcript.
4. Normalize the same way for WER.
5. Score speaker labels separately.
6. Score critical terms separately.
7. Track cleanup time.
8. Test formatting and deadlines.
9. Review how each vendor marks uncertainty.
10. Repeat on more than one file.
Do not compare one vendor’s “99%” claim to another vendor’s “96%” claim unless they use the same test.
• How do you define accuracy?
• Is the figure measured using WER or another metric?
• What audio conditions are included?
• Does it include speaker labels?
• How are punctuation and formatting scored?
• Are inaudible passages excluded?
• Is the figure for AI output, human-reviewed output, or both?
• What happens when the source audio is poor?
• Can I test my own sample before a high-volume order?
A trustworthy answer may be less impressive than a marketing slogan - and far more useful.
Human review is strongest when the reviewer has access to:
• original audio;
• domain glossary;
• case/study context;
• speaker roster;
• reference documents;
• a second-review process;
• explicit authority to mark uncertainty.
The goal is not to produce text that looks correct. It is to produce text that has been checked against the source.
For AI-specific failure patterns, see Can You Trust AI Transcription for Accents, Crosstalk, and Noisy Audio?.
Under a simple word-accuracy interpretation, roughly yes. But the exact relationship depends on the metric, normalization, and treatment of insertions. Ask how the vendor defines the number.
No. WER is a useful word-recognition metric, but it may not capture speaker attribution, punctuation, formatting, timestamps, or the consequence of specific errors.
Yes. Because insertions are included, a system can produce more total errors than reference words in extreme cases.
No responsible vendor should guarantee the same accuracy on every recording. If speech is masked or missing, neither humans nor AI can reconstruct it reliably.
Use word-level scoring plus critical-term, speaker, number/negation, formatting, and cleanup-time metrics relevant to the document’s use.
If accuracy matters enough to influence a case, study, patient record, board decision, or client report, test the vendor with representative audio. Request a VerbalScripts sample assessment or quote and define the critical terms you want checked.
• NIST, Open Speech Analytic Technologies 2020 Evaluation Plan (WER definition)
Get latest updates for our Articles & Blogs. We post fresh content every week.
Sign up for our monthly newsletter