What Does “99% Transcription Accuracy” Actually Mean? Word Error Rate vs. Human Review
Aug 8, 2026

What Does “99% Transcription Accuracy” Actually Mean? Word Error Rate vs. Human Review

by Verbalscripts2 minute read

Quick answer: “99% accurate” is meaningful only if the vendor tells you what is being counted, against what reference, under which audio conditions, and whether speaker labels, punctuation, names, numbers, and formatting are included. In automatic speech recognition, a common metric is Word Error Rate (WER): substitutions + deletions + insertions divided by the number of reference words. Under a simplified accuracy = 1 - WER interpretation, 1% WER corresponds to 99% word accuracy - but commercial services may use different definitions, and WER alone does not tell you whether the errors were harmless fillers or critical medical, legal, or numeric terms.

“99% accuracy” is one of the most repeated claims in transcription marketing. It sounds precise. Often, the measurement behind it is not stated.

VerbalScripts’ own guidance emphasizes that responsible providers should not promise one percentage for every recording because source quality, speaker count, terminology, and review conditions change. See How Accurate Is Professional Transcription?.

What is Word Error Rate?

NIST uses Word Error Rate as a standard measure in ASR evaluation. The basic formula is:

WER = (Substitutions + Deletions + Insertions) / Number of words in the reference transcript

Where:

Substitution: the system outputs the wrong word.

Deletion: a reference word is missing.

Insertion: the system adds a word that was not spoken.

Simple example

Reference:

The witness did not enter the building.

System output:

The witness did enter a building.

Potential errors include deletion of “not” and substitution of “the” with “a.” The sentence is still fluent, but its meaning has changed dramatically.

That is why accuracy evaluation should consider error consequence, not just count.

Why “99%” can still mean a lot of corrections

Imagine a 9,000-word transcript. If a service truly has a 1% word error rate on that file, a simple approximation is about 90 word errors.

Those 90 errors could be mostly low-impact fillers - or they could include:

a party’s name;

a medication;

the word “not”;

a dollar amount;

a parcel number;

the identity of a speaker;

a quoted phrase used in research.

The percentage alone cannot tell you which.

Also note that WER can exceed 100% in pathological cases because insertions add to the numerator. So “accuracy = 1 - WER” is a convenient intuition, not a universal quality definition for every scoring situation.

WER usually ignores several things buyers care about

Depending on the evaluation setup, WER may normalize or exclude punctuation, capitalization, formatting, or other elements. Even when word recognition is excellent, a transcript can still be unusable because of:

Speaker attribution errors

“Attorney” and “Witness” can be reversed while every word is correctly recognized.

Punctuation and sentence boundary errors

A long stream of correct words can become difficult to interpret if punctuation is poor. In legal or medical text, punctuation can also affect meaning.

Proper-noun errors

A surname may count as one word error, but correcting 50 participant names across a research dataset consumes significant time.

Formatting errors

A clinical note can contain the right words in the wrong section. A hearing transcript can lose Q/A structure. A focus group can have unstable participant labels.

Timestamps

A transcript can be word-accurate but have timestamps shifted enough to make source retrieval frustrating.

Human transcription “accuracy” has the same measurement problem

Human services often advertise 99% or 99%+ accuracy, but buyers should ask the same questions:

Is the number measured or aspirational?

Is it an average across all audio or only clear audio?

Who creates the reference transcript?

Are names and speaker labels counted?

Are inaudible segments excluded?

Is accuracy measured before or after customer corrections?

Is there an independent QA sample?

Human review improves the ability to resolve context and uncertainty, but humans are not error-free. A mature vendor should be able to explain the conditions under which accuracy degrades.

A better way to evaluate transcription quality

Use a scorecard with more than one metric.

Word Error Rate: Overall word recognition errors

Critical Entity Error Rate: Errors in names, dates, amounts, drugs, case IDs, etc.

Speaker Attribution Accuracy: Whether words are assigned to the right person

Omission Rate: Whether meaningful phrases/segments are dropped

Timestamp Accuracy: Ability to return to source

Formatting Compliance: Whether client template/style is followed

Uncertainty Quality: Whether inaudible speech is marked honestly

Customer Cleanup Time: Real operational burden after delivery

For many professional buyers, cleanup time is the most revealing commercial metric.

Critical errors should be weighted differently

A 99.2% transcript with one wrong medication dose may be less acceptable than a 98.8% transcript whose errors are filler words in a podcast.

Create a “critical terms” list for the use case.

Legal

names;

dates;

exhibit numbers;

amounts;

statute/case citations;

negation;

speaker identity.

Medical

patient/provider names;

medications/doses;

anatomy/laterality;

measurements;

negation;

diagnoses.

Research

participant ID;

quoted phrases;

domain terms;

speaker label;

code-switching;

timestamps.

Government

motion wording;

vote result;

case/parcel numbers;

conditions;

names.

What should happen with inaudible speech?

An accuracy score can be distorted if the reference excludes hard sections. Real customers do not get to exclude the hard section from their case.

Ask the vendor how it handles:

fully inaudible passages;

uncertain single words;

overlap;

low-volume speech;

names not verifiable from audio.

A transparent marker such as [inaudible 00:18:42] is not automatically an “accuracy failure.” It can be the most accurate representation of the source.

This is central to poor-quality audio transcription.

How to compare two vendors fairly

1. Choose 30-60 minutes of representative audio.

2. Include both easy and hard sections.

3. Create or commission a carefully verified reference transcript.

4. Normalize the same way for WER.

5. Score speaker labels separately.

6. Score critical terms separately.

7. Track cleanup time.

8. Test formatting and deadlines.

9. Review how each vendor marks uncertainty.

10. Repeat on more than one file.

Do not compare one vendor’s “99%” claim to another vendor’s “96%” claim unless they use the same test.

Questions to ask any transcription provider claiming 99% accuracy

How do you define accuracy?

Is the figure measured using WER or another metric?

What audio conditions are included?

Does it include speaker labels?

How are punctuation and formatting scored?

Are inaudible passages excluded?

Is the figure for AI output, human-reviewed output, or both?

What happens when the source audio is poor?

Can I test my own sample before a high-volume order?

A trustworthy answer may be less impressive than a marketing slogan - and far more useful.

Where human review changes the outcome

Human review is strongest when the reviewer has access to:

original audio;

domain glossary;

case/study context;

speaker roster;

reference documents;

a second-review process;

explicit authority to mark uncertainty.

The goal is not to produce text that looks correct. It is to produce text that has been checked against the source.

For AI-specific failure patterns, see Can You Trust AI Transcription for Accents, Crosstalk, and Noisy Audio?.

Frequently asked questions

Does 99% accuracy mean one mistake per 100 words?

Under a simple word-accuracy interpretation, roughly yes. But the exact relationship depends on the metric, normalization, and treatment of insertions. Ask how the vendor defines the number.

Is WER the same as transcription quality?

No. WER is a useful word-recognition metric, but it may not capture speaker attribution, punctuation, formatting, timestamps, or the consequence of specific errors.

Can WER be over 100%?

Yes. Because insertions are included, a system can produce more total errors than reference words in extreme cases.

Is 99% human transcription possible on poor audio?

No responsible vendor should guarantee the same accuracy on every recording. If speech is masked or missing, neither humans nor AI can reconstruct it reliably.

What metric should legal or medical buyers use?

Use word-level scoring plus critical-term, speaker, number/negation, formatting, and cleanup-time metrics relevant to the document’s use.

Ask for evidence, not a percentage

If accuracy matters enough to influence a case, study, patient record, board decision, or client report, test the vendor with representative audio. Request a VerbalScripts sample assessment or quote and define the critical terms you want checked.

Authoritative references

NIST, Open Speech Analytic Technologies 2020 Evaluation Plan (WER definition)

NIST, OpenASR20 Challenge Evaluation Plan

Subscribe to our newsletter.

Get latest updates for our Articles & Blogs. We post fresh content every week.

Weekly articles
Stay updated with our weekly articles covering various topics.
No spam
We respect your inbox. No spam, just valuable content.