
Dom DiFurio // accessiBe
1 / 2Tools making the most significant improvements in word errors
As shown in the chart above tracking word error rates, Google and IBM's audio transcription AI performed worse in 2023 than in 2022. Google Standard had a 28% error rate while Google Video had a 14% error rate—an increase of 2 percentage points and 0.7 percentage points since 2022, respectively. Meanwhile, IBM Watson had a 25% error rate, a 1.5 percentage point error increase since 2022.
However, other platforms saw slight improvements with lower error rates over the same period, including Rev AI (-3.4 percentage points), Microsoft (-0.91 percentage points), and Speechmatics (-1.11 percentage points). They all had a 10% or less error rate for 2023.
Other transcription tools tracked include Assembly AI (8% error rate), the multilingual OpenAI Whisper: Large (8% error rate), and the English-only OpenAI Whisper: Tiny (15% error rate), all of which didn't have 2022 data to make a year-over-year comparison.
Word error rates describe the number of times a transcription engine might interpret the wrong word in an uploaded audio file. Generally, the models 3Play Media analyzed made some progress from 2022 to 2023 in reducing the number of word errors they make, but not in all cases.
Another factor that affects AI's reliability for transcription is its frequency of punctuation errors, which affect readability. In this realm, OpenAI's model is most accurate, but it is still only 85% reliable.
Open AI offers multiple speech recognition models of varying complexity and power. The "tiny" version only performs English language transcription, whereas its large model is multilingual. The company launched these for the first time in 2022, and 3Play Media only studied them for the 2023 year. AssemblyAI's speech recognition tool was also more recently released and has no comparable prior-year data. Developers trained its 2023 speech recognition software on 1 million hours of audio—it is also capable of English language transcription.
In a testament to how quickly advancements in the space are moving, Assembly released a successor just this year based on 12 times the training data, and which it advertises as multilingual and "hallucinates" 30% less often than OpenAI's competing service.
AI hallucination refers to large language models' tendency to invent misstatements in their output. In a chatbot, this might look like a confidently stated fact that isn't true. These remaining shortcomings and the current state of AI development make using it for accessibility reasons difficult.
That's a major drawback for companies that take accessibility and the surrounding laws and requirements seriously, giving them pause before entirely handing the reins to AI for transcribing audio.









