Deep Learning with PolyAI

Is word error rate just a vanity metric?

Team PolyAI

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 28:42

Send us Fan Mail

Voice AI vendors love to quote word error rate, usually somewhere around 2 to 3%, as proof their system understands customers. Oliver Shoulson, Agent Design & Engineering Lead at PolyAI, thinks it's one of the most misleading numbers in the business. He and Nikola Mrkšić go through a month of analysis across more than 100,000 real production call turns to explain why. 

Turns out not all words carry the same weight, most transcription errors are harmless, and task success barely moves even when the system mishears, because today's models recover the way a person would. 

Listen to the full episode, and see how PolyAI builds dialog agents that get this right at https://poly.ai?utm_source=youtube&utm_medium=podcast&utm_campaign=podcast&utm_content=podcast

  • Follow PolyAI on LinkedIn
  • Watch this and other episodes of the Deep Learning pod on YouTube
SPEAKER_00

It's clear to me now more than ever that like not all words are created equal, right? In context, you know, missing a yes or no could be everything or it could be nothing. Now that we're have we're starting to deal with audio models that like take prompting in and situational context in addition to the audio itself. Like just going on like raw, how many words did we get right? Like is a pretty bad proxy for all the stuff that we actually care about.

SPEAKER_01

Hello everyone and welcome to another episode of Deep Learning with PolyAI. Today I've got Oliver back on the show with me. Good to see you again. Great to be here. Yeah. Oliver is one of our elite Asian designers and a real kind of like force behind a lot of work, whether it's, you know, pioneering Open Claw or just, you know, bringing the good spirit of linguistics in. And today, after we spoke about some very interesting data that he showed me, I thought it would be really interesting to just kind of like have a whole episode that really looked into the magical notion of a word error rate, right? So we build voice agents, and how well these systems work feeds downstream from whether you're able to recognize what someone has said or not, right? And I think most companies, especially those that focus predominantly on that component, really just talk about improving the word error rate and getting to a better and better one in their public benchmarks. And it's kind of treated as a holy grail by a lot of people buying software, deciding what component to use, deciding whether it's good or not. And we'll talk a lot about whether it's good or not. I think like one anecdote from my PhD supervisor was really good where I think it was the early days of deep learning, and we were all we just thought it would all be done in two years. I think that's the prevalent attitude that's continued ever since. And I remember Steve telling me, like, whoever thinks that speech recognition will be solved in two years has worked on the problem for less than two years. And, you know, I've now spent more than two years working on the problem. I feel like it was very but Oliver, like maybe over to you to talk a bit about like the word error rates and the findings that you've had over the past kind of few months.

SPEAKER_00

Yeah, totally. I so I mean, I have much less experience like working on this like word error rate in an academic context than you, but you know, what I can tell you is that But more than two years. More than two, uh yeah, I guess so. So I first of all, I know that ASR is not solved, and second of all, I know now that word error rate is a pretty terrible proxy for conversation quality. Like if you're using that as a way to evaluate both how well your ASR is performing in context and just as a way to understand like how well are my customers being understood by the virtual agent, like it is just utterly confounded with noise. And you know, the thing I've been saying as I've been going through a lot of this data is like it's clear to me now more than ever that like not all words are created equal, right? Like in context, you know, missing a yes or no could be everything or it could be nothing. And depending on the turn length, and that could be 50% word error rate or it could be 1% word error rate. And when now that we have these models that are able to reason around ASR mistranscriptions and reason in context, and especially now that we're have we're starting to deal with audio models that like take prompting in and situational context in addition to the audio itself, like just going on like raw how many words did we get right, like is is a pretty bad proxy for all the stuff that we actually care about.

SPEAKER_01

100%. I think that I remember, you know, and data sets evolve and change and gradually become more difficult. But you know, we went from kind of like talking about a 5%, you know, state-of-the-art word error rate to now, you know, recently I was speaking with a CIO of a company and he's like, and you were like a 2% like everyone else. And I'm like, no one is a two percent for anything that really objectively matters in a complex conversation. And I think kind of like you know, my my favorite kind of like peasant logic thing to break this is like to anyone who has Alexa or Google Assistant, it's like, does it work 98% of the time? And it's like, I think it's really far from that. But it somehow managed to permeate the consciousness of people where it's like, oh yeah, it's all but done, voice is solved, and I don't know. I guess we'll dig into your data and take a look at why it's not, but yeah.

SPEAKER_00

Totally. Okay, well, so I'm happy to pull up just some basic stats of like what we what we've done here. You know, I ran this analysis over the course of about a m a month on on some of our really high volume production calls, analyzed over 100,000 individual user turns. What we used kind of as the gold standard for word error rate here was a consensus between two different speech recognition vendors, uh, um, or I should say audio, audio LLM vendors, um, which were able to take the entire conversation and all of the context and that definitely perform in a vacuum with the best word error rate compared to any ASR. They're of course a bit slower and they don't produce like continuous transcription, so they're a bit harder to use in production. Um, but they they perform really well. And so when when we retro retroactively do analysis like this, we can use that kind of consensus between these two vendors as like a really good, confident gold standard to compare against. And so this was over the course of around 25,000 calls. And what we did was when we found turns that were that differed in production from the gold standard that the consensus produced, we had an LLM judge basically classify those mistranscriptions, you might call them, as either benign. So these are like little formatting errors, kind of filler words, stuff that doesn't ultimately change the meaning. Semantic, which is changes that alter the meaning of the user input, but that could be reasoned around in context or that you could ask clarification about. And then critical. And these are specifically things like what we would call entity flips. So things that change the value of a slot that the user is providing in such a way that obfuscates or changes the value in a critical way. So this is something like if you're collecting a phone number and you mishear a digit, that's a critical error because you need to get that 100% right. And it's not necessarily obvious in context that you misheard something. The the agent has no reason to believe it misheard something if it heard, you know, for some reason, three instead of eight. But so that would be something critical. And I can pull up for you just like a little to show you kind of the proportions that we're dealing with of these errors of the of the missed turns. So up to 42%, so so around 28% of mistranscriptions were deemed benign. Up to potentially 42, some of them weren't judged for a variety of reasons. We have a relatively small proportion of just of these kind of meaning-changing semantic errors. So this is when like content words get flipped, but in context, they can often be reasoned around, or it's clear the what the clarification question should be. And then about half of the errors we saw are these entity flipping errors. And what this tells me actually is that the the proportion that are benign is a huge amount of obfuscating and confounding noise, basically, in the word error rate signal. I'll stop there for a second to like hear, like, I'm sure you have thoughts about this, but like that's my first takeaway from this.

SPEAKER_01

No, no, I mean it's like I think it's like the empty calories and this half where, you know, like anything people need to understand it's not like a word is missed. Like, yeah, not all words are equal, but in particular, you know, articles, like whether you pick up and or the or whether someone might have said an or not, and whether that's in or not is completely material. It's like semantically empty. I think, you know, beyond that, like plural or singular might matter for the meaning, but again, like an L-based system is probably gonna make it through about just fine, right? Similarly, you know, like a not omitted can be really, really devastating and creating like an antonym and opposite meaning of other things. So between all that, like it it can be really hard to like decipher like you know which ones really matter. And it's very easy to have even an you know something with a much higher word error rate performing much worse. And it's plagued evals on dialogue systems forever in that you improve something, you measure on this like intermediate task, you see a higher number, you roll it into production, and god forbid you made changes downstream as well. At that point, you're completely lost for like where did you make progress? Where did you degrade? And it's it's tough, right? And the more we integrate these components, the the harder it gets still, right?

SPEAKER_00

Yeah, and it's not even like we can identify like particular classes of words that are semantically like like in, you know, we have this basic like distinction in linguistics between content words and function words, where you have like your words that sort of mark syntactic information, like which are your function words, so things like articles and and modal verbs and stuff like that. And then you have your content words, which actually sort of introduce the mental content into the sentence. But it's not even like you can necessarily just classify sort of an inventory of content and function words and know which ones are matter and which ones aren't. Like it's obviously all contextual. So if the question is, you know, are you calling to check the status of an order and the user just says yes and you mishear the yes, like that's a disastrous mishearing. Like you you missed the whole content of what they said. But if they say, Yes, I am, and you miss the yes, but you hear the I am, that's basically completely benign. Like the like you'll proceed and and interpret that exactly the same as if they had just said yes. So it's like Yeah.

SPEAKER_01

And I mean, you know, rabbit hole, but like in different languages, it also like behaves very, very differently, right? I mean, like the morphology of Slavic languages in the cases, for instance, means that word error rates tend to be quite a bit higher. For all intents and purposes and understanding, it doesn't really matter. In fact, there are entire groups of say, I don't know, Serbian population that get cases wrong or use fewer than like what language does. And you know, like it doesn't really matter. Like, not for this.

SPEAKER_00

That's so interesting because like the stem is it gets like the stem right, but it like it's yeah, it's like inflected wrong, and so that's counted as a word error. Yeah. Oh, that's so interesting.

SPEAKER_01

And you know, I think there's almost like you would then look at character level error rates, but I feel like it would probably have the exact same empty calories, and maybe more actually, than the word error rate.

SPEAKER_00

Yeah. You know, one of the things that we were looking to sort of correlate or not correlate word error rate with is our proprietary evaluation, which is what we call our polyscore, which gets run on all of our production calls. And so I thought I'd take a second to just sort of talk a little bit about polyscore and what it consists of. So actually, maybe I feel like you're maybe better suited to speak to like a little of the history of polyscore and sort of what how that came about and how we're I'm happy to.

SPEAKER_01

I mean, like look, I mean, I think that stuff like predict predictive, like MPS scores and C sets and just like a measure of like, is it a good call or not? Has been something that's always needed. And you know, whether you trust the vendor and showing you whether this is good or not is one piece. The other piece where it's just objectively very useful, 21 launching a system is you want to see the great calls and you want to see the terrible calls, right? So I think that, you know, as a measure of that, it's something that's been honed for years and you know, continues to be honed. It really composes it's composed of a measure of the quality of the conversation. Like, you know, is the fluidity is the UX is it a natural human-sounding conversation? Are we interrupting each other? Or if they're interrupting, are we stopping? And very, very many things that have to do with everything from you know the quality of like the audio piece to kind of just like the fluidity of the back and forth. And then maybe most importantly, like task success. Like, did the AI agent actually do what the agent is supposed to do in that particular uh conversation? And you know, sometimes the caller is not happy with that. If they want a refund and you're not able to give them a refund, you know, like the score may end up being high because the system behaved and did the right thing, but yeah, it's pretty nuanced. And, you know, again, kind of like the LLM judges from the other one, it's it's not perfect, but it's pretty good.

SPEAKER_00

Yeah. And it's really nice that it it returns this kind of granular, like constituent broken down score because we actually get to see some like interesting ways in which the number of meaning-changing errors that occur over the course of the call affect conversation quality. So obviously, like the need for clarification and repeating of the same questions, but actually don't affect task success nearly as much as you might expect, which is kind of interesting. So, you know, the customer tends to pay in friction, but in terms of like the actual effect it has on getting where they're trying to go or even call outcomes like escalations and handoffs, you're actually it's kind of amazing how effective the model is at like uh retaining that containment, even at the expense of a little bit of friction of re-asking those questions. Um, and just to sort of hammer in something I was saying before about this empty calories, I like that metaphor you're using. You know, benign only word error rate, like as you might expect, varies widely and like predicts nothing on the polyscore side. Like we can we can account for maybe up to 0.5% of variance in polyscore via the benign word error rate, which first of all goes to kind of validate the the LLM judge that was determining the severity of errors, but also tells you exactly what you'd expect, which is that like word error rate is confounded by all of this noise. And if you isolate the benign stuff, like it's just it's literally just noise.

SPEAKER_01

Yeah. Yeah. I mean, like the the zero correlation is almost like it looks it almost looks rigged, Oliver.

SPEAKER_00

Yeah, the task success zero correlation, it does, it does almost look rigged. I was kind of amazed by that. And I you can actually see that even more. It's amazing how how the task success stays super flat, even as we see the compounding effects of meaning changing errors. And this is sort of the flip side to word error rate, which is where we see these actually, I think, really valuable correlations, which is how many semantic or critical errors happen over the course of the call. And you see this really clear monotonic relationship between the number of these meaning-changing errors and conversation quality on the polyscore, for instance, which is this left graph over here, or handoff rate, which you can see in this right graph, which crawls up with each meaning-changing error that happens over the course of the call. That little dotted line you see on the top of that left graph over there is the task success. So again, like amazingly flat the entire time. And really, it's just the user is paying in additional turns and clarification. But I like my sort of rose-colored glasses take on that is like that's an amazing vindication of the capacity of LLMs to recover from these mishearings. Because even if it requires a question to be asked more than once or a follow-up question to be asked, the user is still getting the task accomplished that they want to. They're just paying in a little bit of friction. And I think like that's something that, you know, back in the deterministic days, that like was just a complete non-starter. You could not ask targeted follow-ups, you could not repeat questions in a way to actually glean the clarification that you wanted.

SPEAKER_01

No, totally. I think like, you know, maybe just to give like some contextual examples, but you know, if you're being asked for, I don't know, your British car license plates, and you said LO23 EWN, and then like it repeated it and you know, got the last character wrong. And then you would be like, no, no, no. It's like and for November, right? At that point, you're like, okay, like that last one, yeah, good. Or like repeat the last three system repeats, it's that, not that poof, right? And that's like very much how humans would talk to each other, and it's like that's exactly it, and it's just really you're right, like how much, and again, this has to be implemented, right? It doesn't always work by default, right? It can still be frustrating, and it is frustrating, we see it from the coal quality score. But I think like it's interesting in that graph how you have like the steepness of the coal quality falling is higher, right? Because I guess the quality is decreasing, but the functional capability of the system is there as long as the human stays in the coal, which again is something that I guess we would have to kind of like derive from the data, right?

SPEAKER_00

Yeah, though, I mean, I I guess another thing that I was sort of pleasantly surprised by is to see how while we do have absolutely that slope downward over those first two errors, like it's not as steep as you might think it would be. Like you might think that people, when they get one critical or semantic error and it's clear that they've been misunderstood, they're like, screw this, let me talk to a human. And so, and on on neither round, on neither graph, do you see this incredibly steep either degradation in call quality, meaning that like we're not able to recover from those errors, or an extremely steep increase in handoffs, which you might expect to see, which also points to this capacity for the model to retain engagement, even after, you know, polyscore, which detracts for repeated or clarified questions, like is claiming that the call quality has degraded. And so only until up until that like third semantic or critical error, which is where we see that real cliff in conversation quality, you know, the first couple are not free, but like not as disastrous as you might expect them to be.

SPEAKER_01

Yeah. So so wait, like the three plus on the handoffs going to 68. What is that one about?

SPEAKER_00

That one's, you know, our Wilson band there, our 95% confidence interval is quite wide. So I don't want to make any claims about that. You can see those little whiskers, but I think we just don't have enough data there. Because, and and again, this is also like I was happy to see our cohort of three plus critical or semantic errors in in in our entire database of over 100,000 turns is in the like tens of calls. So, you know, we have very few calls have three plus meaning-changing turns. And the other thing to emphasize here, again, is like, you know, I wouldn't be I wouldn't be very scientific if I didn't like whip out the correlation is not causation, because other thing like you need to keep in mind is that things that affect conversation quality and handoff rate also affect things like word error rate and like critical error. So like if like the user is in a very noisy environment, like that's this external factor that's going to affect all of this stuff. And so you can't necessarily say, oh, the ASR was bad, so it led to this compounding effect of errors. You can say, you know, the user is driving 70 miles an hour down the highway blasting music, and it's not going to lead to a good experience. Yeah, yeah.

SPEAKER_01

And I mean, also like it's it's a very complex if, right? Because like to those conversations that get to have, I guess like the caveat, just knowing what I know about us, would be like we don't tend to proactively not hand off often if we get three like really bad turns, right? So in places where they get to accumulate, especially if there's very few of them, those are going to be those very, very long calls. And I think just anecdotally, you know, you kind of like the longer you stay in gameplay, maybe if you start the game with three lives, you know, like it refills like for every minute that you've stayed in the conversation, the sun cost of rage quitting and demanding to be handed off. You don't know if the human can seamlessly continue the conversation. We have many clients where the flow has been implemented to kind of like persist and fill out to CRM partially, and it's great when it does, but for the most part, and industry-wide, it's not usually implemented well. So you're kind of used to just like you get out of that process, and then it's like cool. So what was your name? What was your like social security number? And you're like, oh my god, right? Like that's that that's that C set is zero, right? So I think like the yeah, like if you if you manage to last that long, then people are pretty committed. So it makes sense that the handoff would would fall, because you know, you're about to achieve something, you will persist even if your last name is Merkshic and you really have to, you know, somehow convey that to an automated thing. What was the word error rate in that aggregate analysis?

SPEAKER_00

Um, our headline word error rate for English was 6.4. Yeah. And for Spanish, which we had a smaller cohort of, was 8.8. Yeah. Which gives us a total of 6.7 weighted.

SPEAKER_01

Yeah. I mean, you know, when you look at like how people benchmark against, you know, like data sets, they're talking about 2%, 3%. And I think you know, it completely amidst the fact that like there's a different low pass filter on the phone, and would avoid that, you know, you might get several of them hitting the data and then just like lossy packets and stuff. So I really almost think that we should start producing a data set of those. And there are some that are. I remember like evaluating. We we used to collaborate with uh a bunch of kind of like car makers for their voice agents back back at Cambridge, and those in car word error rates would literally go around Cambridge in a car and record samples there. It would be between 20 and 40 percent. Now that's like car music, everything, and like an era of different models.

SPEAKER_00

But that's where people are calling customer service from. So, you know. Sure. It's very, very it's not a yeah.

SPEAKER_01

Anything else that you kind of like think in this one is worth kind of maybe sharing with the audience?

SPEAKER_00

Yeah, I mean, I guess the the last thing I wanted to share, which again I think is kind of intuitive if you think about it, is the difference between like in terms of handoffs versus quality, whether like the different effects that semantic errors have versus critical errors. And, you know, what's kind of interesting is that like we punish the semantic errors on quality a lot more, but they don't move the handoff needle very much. And that makes a lot of sense when you consider that like semantic errors are easy to sort of disambiguate around. And so if we're punishing them on the quality score because we're re-asking a question or we're asking a clarifying question, but they're not leading, but they have basically no effect on handoff rate. Whereas interestingly, you know, the critical errors have half of the quality degradation that semantic errors do. And that's because critical errors are getting escalated quickly. Which is what you'd want to see. Like when the model like misses an entity and is and fails to look up an account or fails to retrieve an account or validate the user or something, and something has gone really wrong to like obstruct the path of the flow. Like we want to escalate to an agent quickly and in a way that does not degrade the quality. So you can see this like minus 0.27 effect on quality that critical errors have versus the minus 0.55 that semantic errors have. And again, that has a lot to do with how Polyscore scores quality and how it punishes clarification and disambiguation. But like again, I was happy to see that there's not this huge quality degradation around critical errors, which shows that we are escalating appropriately at the right times.

SPEAKER_01

Yeah, yeah, yeah, yeah. This is a bit of a chicken and egg. Like, I'm not sure if we can claim good science, but it's it's really interesting. I mean, it's also like the the whole like you know, confirming stuff and like, you know, how you attribute that to quality if you, you know, just kind of like the best call is what's your name? What's your date of birth? Ideally, I'm an AI oracle, I get everything right, and we don't go through the awkward like S-H-O-U-L-S-O-N, right? Whatever. I mean, that's just like really not there, not there yet in terms of good design. But if you correct that, you know, you're doing it with a human just the same as you are with a machine.

SPEAKER_00

Right. I mean, that's why that's why I kept trying to emphasize that like a lot of this, not an artifact, but a result of the way that Polyscore works, which is that it tends to punish that kind of like sort of the opposite of what you're saying, which is where we can't, we don't just like asked an answer to every single question. So, you know, we should interpret conversation quality through that lens and understand that a lot of these cases where we're seeing reduced quality, they're still navigating the conversation in a very natural and human-like way. And so that's why it's helpful to know sort of what that score consists of.

SPEAKER_01

Yeah. Okay, so maybe like if we break out a bit from just, you know, like the analysis itself. You know, if you could have like one wish of speech recognition, like one thing that it could do better to make dialogue systems a whole lot better, like what is the thing that is like most annoying in practice?

SPEAKER_00

I mean, I think that one thing that we encounter a ton in practice is the difference between in performance, between models that perform really well on short turns and models that perform well on longer turns. And because that's so unpredictable in the conversation, like you don't know necessarily. Like you can you can sort of predict if you're asking a yes or no question, you can imagine that the utterance is the response is probably gonna be short, but you don't know. And it's frustrating to have to configure different models on every single turn based on whether you're expecting a shorter or longer response, whether you're expecting a yes-no versus like kind of numeric or more value-driven response. So that's super annoying in practice. And then I think also just like the reason that these audio models perform so well and we can treat them as the gold standard probably has a lot to do with the way that they are being prompted in addition to just returning sort of blind transcriptions. Um, like they get to see the conversation context, they get text context, and they evaluate and they understand the audio tokens in the context of all of that text. And so I think like that's really the direction that this has to go, probably.

SPEAKER_01

100%. I mean, maybe like to take a quick digression that is very like topical. Have you seen the new like GPT Live?

SPEAKER_00

No. Oh, oh, I saw I saw their like little promo video for it. Yes.

SPEAKER_01

Yeah. I find it really interesting because you know, I think that talking about short terms, the quick yeses and no's. I think you know, it's not as much our world and that we're really focused on like real conversations that drive like business and and and customer needs. But they have this one where like let's say that you're talking and I'm going, yes, yes, yes, no. It's clear that someone has just like fixated on the problem and really wanted to show that like full duplex, like I'm speaking while listening, and kind of like inputs are coming in and out, which we as humans do, and when we do, we just tend to be annoying, right? I think I was curious about how you felt about that piece because I think like you know, like Raven Omni and our you know internal voracious appetite for it, amazement at like everything that it can do, both in terms of understanding using all that context to just you know fix problems that we never could fix before is one thing. But how do you feel about the full duplex piece? Do you think it's even good as a feature for a conversational system?

SPEAKER_00

Yeah, I do. Like, I think I can't count the number of times I've seen like a delay in a barge in taking effect such that we only hear the tail end of what someone said and then totally misinterpret it. Like I think, I think just in terms of at least like responsiveness to barge in and interruption, it's necessary to have like maybe that's not full duplex, like we just need to know whether they're speaking or not, not necessarily what they're saying. But yeah, yeah, I I don't know. I would want to think about it more, but I feel I have to feel like it's necessary because it's such a key part, as you said, of like the way humans navigate conversations.

SPEAKER_01

Yeah, I mean, I feel like you know, they've done one UX thing, you know, a grave, grave, grave, grave crime there, which is because it's built on Chat GPT, which is used to, you know, sociopathically uh inducing more more more interaction from you, right? It is built and designed for you to interrupt it rather than to have a natural flow of back and forth conversation, right? Because of that, I kind of feel like they've taken for granted that that should be the interaction mode between two parties, right? That I just speak and that you will interrupt me, and that's why we have the whole fascination with bargin and stuff. Although bargeing matters, right? If I'm asking a question, you go, yep, and I don't hear you, like that's obviously bad, right? But I feel like now it's almost like turning into a separate, you know, like academic discipline of just like, and look, like this model is like taking inputs, producing outputs. Like, we as humans, if we're talking, and if you try to like slot something in, like I don't really, for the most part, one want you to do it, right? Like, sure, if it's a critical piece of information, you should interrupt me. But I think like for the most part, we see people react very weirdly when they get interrupted by AI because they just feel like they're not being heard, right?

SPEAKER_00

Yeah, yeah, no, and I I I see what you're saying about how it feels like it feels like they're excusing themselves for the way that the model is just going to monologue at you and like not actually design information delivery in a way that like makes it makes you not want to interrupt the model as much. And just saying basically, like, we're not concerned with that problem, like we'll just let you interrupt. Yeah, and as a conversation designer, that does hurt. Fair.

SPEAKER_01

Well, on that note, I think we I think we're out of time for today. But Oliver, always a pleasure. And always a pleasure. Please, you know, like, share, subscribe, and we'll see you on the next one.