Accent and Dialect Bias
ASR Performs Significantly Worse on Non-Standard Accents
66 patterns in this category
Voice AI agents fail in production because four largely independent layers — audio capture quality, speech recognition, dialog/conversation management, and voice synthesis — each have their own failure modes, and a breakdown in any single layer degrades the whole call even when the other three layers work perfectly. A caller can be recognized accurately and responded to with a perfectly-worded reply, and the interaction still fails if the agent talks over the caller, mispronounces its own brand name, or loses the recording to a mid-call disconnect. Speech-and-audio failures are distinct from generic LLM failures because they carry hard real-time constraints (sub-second turn-taking, audible dead air) that a text-based chatbot never has to solve.
| Goal | Covers | Patterns |
|---|---|---|
| Audio Handling | Signal-path degradation, acoustic interference, and call-session lifecycle problems upstream of ASR | 6 |
| Speech Recognition | ASR mishearing — accents, names, homophones, numbers, confidence handling — on a signal that has already reached the recognizer | 8 |
| Conversation Flow | Turn-taking mechanics, persona integrity, data capture, and business-logic compliance across a multi-turn dialog | 44 |
| Voice Synthesis | TTS pronunciation, prosody, emotional tone, and voice-persona consistency on the output side | 8 |
Total: 66 patterns
The four goals map onto the voice-agent pipeline in order: Audio Handling governs everything before a spoken word becomes a signal worth transcribing — network quality, device variance, noise, echo, and disconnection. Speech Recognition governs turning that signal into text — the accuracy of the transcript itself. Conversation Flow governs what the agent does with that transcript — when to respond, what to say, what data to extract, and which business rule applies, all while managing the sub-second timing of a live call. Voice Synthesis governs turning the agent’s chosen response back into audio the caller can understand and trust. A failure early in the pipeline (bad audio) can masquerade as a failure later in it (wrong intent classification), so root-causing a production issue means checking upstream before assuming the layer where the symptom appeared is where the bug lives. To localize an incident by symptom: garbled or dropped audio → Audio Handling; wrong words in an otherwise-clean transcript → Speech Recognition; the agent says something odd, mistimed, or non-compliant given a correct transcript → Conversation Flow; the agent’s own voice sounds wrong, robotic, or inconsistent → Voice Synthesis.
Conversation-flow sits at the intersection of two large problem spaces that don’t overlap elsewhere: real-time audio turn-taking (a problem unique to voice, with no text-chatbot equivalent) and multi-turn business-logic dialog management (a problem any structured conversational agent faces, voice or text). See Conversation Flow for the full breakdown into sub-clusters.
Check Audio Handling first — background noise, packet loss, or echo frequently produces exactly the symptom of “misheard me” and is the more common root cause than the ASR model itself. Only move to Speech Recognition once audio quality is confirmed clean, since accent bias, homophones, and domain vocabulary gaps are the ASR-layer explanations for the same complaint.
Rarely on its own. Across all four goals, the documented pattern is that a stronger model narrows the failure rate but doesn’t eliminate it — the reliable fixes are architectural: audio pre-processing paired with confidence-gated confirmation, custom pronunciation lexicons paired with SSML overrides, and dialog-state validation gates paired with explicit business-logic checks.
Audio-handling failures are signal-quality problems that occur before ASR ever runs — packet loss, noise, echo, device variance. Speech-recognition failures are what happens when ASR runs on a signal (clean or not) and still mishears specific content classes — names, accents, homophones, numbers. A noisy recording that gets mistranscribed is an audio-handling problem with a speech-recognition symptom; a clean recording that still mishears a caller’s name is a pure speech-recognition problem.
Voice Synthesis — mispronunciation, flat or mismatched emotional tone, audio artifacts (clicks/pops), and voice-persona drift across a session are all synthesis-side failures distinct from anything happening on the recognition or dialog-management side of the pipeline.
ASR Performs Significantly Worse on Non-Standard Accents
Agent Claims It Will Perform Actions Beyond Its Capabilities
Agent Mishandles Questions About Being AI/Bot
TTS Produces Clicks, Pops, Distortion, or Other Audio Artifacts
Poor Audio Quality from Network or Device Issues
Agent's Acknowledgment Cues ("uh-huh", "right") Occur at Wrong Moments
ASR Accuracy Degrades Significantly in Noisy Environments
User Cannot Interrupt Agent's Speech
TTS Mispronounces Brand Names, Domain Terms, or Key Phrases
Agent Fails to Handle Mid-Call Disconnections Gracefully
Performance Varies Significantly Across Devices
Agent's Filler Words and Speech Disfluencies Don't Match Persona or Context
ASR Fails to Recognize Industry-Specific Terms
Agent's Own Audio Interferes with User Speech Detection
Agent Overuses Laughter, Exclamation Marks, and Emotional Cues
Voice Emotion Doesn't Match Message Content or Context
Agent Can't Reliably Detect When User Has Finished Speaking
ASR Misinterprets or Fails to Filter Conversational Fillers
Call Ends Awkwardly Without Natural Closure
Agent Pressures Hesitant Callers Instead of Providing Low-Friction Options
ASR Selects Wrong Word Among Sound-Alike Options
Agent Can Be Manipulated Into Adopting Different Personas or Revealing Prompt
Agent Waits Until All Fields Collected Before Saving Data
Similar Intents Misclassified Due to Overlapping Definitions
Agent Reveals Backend Operations, Routing Logic, or Internal Details
Agent Doesn't Properly Handle User Corrections and Interruptions
Agent Cannot Communicate Due to Unsupported or Incomprehensible Language
ASR Confidence Scores Not Used Effectively
Agent Delivers Long Feature Lists or Information Without Pausing for Engagement
Agent Asks for Multiple Data Fields in a Single Turn
Agent Can't Distinguish Between Multiple Speakers
Agent Loses Context Across Conversation Turns
Agent Fails to Match or Maintain Caller's Language Choice
ASR Fails to Correctly Transcribe Personal and Business Names
Long "Never Say X" Lists Inadvertently Prime the Model to Output Banned Content
ASR Incorrectly Transcribes Numbers, Dates, and Times
Agent Opening Doesn't Account for Caller's Greeting State
Call Outcomes Incorrectly Classified, Affecting Follow-up Actions
Agent Fails to Detect or Respond to Conversation-Stopping Signals
Agent Ends Call When Caller Pauses, Interrupts, or Shows Confusion
Oversized System Prompts Cause Dead Air Due to Time-to-First-Token Delays
TTS Mispronounces Words, Names, or Domain Terms
Speech Rhythm, Stress, and Intonation Don't Match Content
Agent Skips Required Steps or Asks Questions Out of Sequence
Agent Fails to React to Caller's Personal Comments or Emotional Cues
Agent Continues Planned Flow Instead of Adapting to Caller's Response
Agent Takes Too Long to Respond
Agent Answers Questions or Provides Information Outside Approved Knowledge
Agent Deviates from Required Phrases, Boundaries, or Conversation Structure
Agent Incorrectly Interprets Pauses and Silence
Critical Information Extracted Incorrectly or Inconsistently
Agent Goes Silent During Tool Execution Without Acknowledgment
Numbers, Dates, and Formatted Text Not Converted to Spoken Form
Speech Markup Not Processed Correctly
Real-Time Transcription Changes Mid-Utterance
System Prompts Designed for Text Chatbots Fail in Voice Conversations
Vague or Poor Tool Descriptions Cause Wrong Tool Calls or Bad Parameters
Agent and User Speak Over Each Other
Agent Makes Promises Outside Allowed Scope (No Spam, Follow-up Guarantees, Delivery Promises)
Agent Sounds Robotic, Scripted, or Inauthentically Enthusiastic
Agent Requests Personal Information Beyond What's Needed
Agent Uses Caller-Provided Information Without Verification
Agent Accepts Ambiguous or Passive Consent as Full Permission
Agent Produces Long Responses Despite Explicit Length Constraints
Voice Characteristics Change Unexpectedly During Conversation
Agent Fails to Detect or Handle When Call Reaches Unintended Recipient