Healthcare

38 patterns in this category

Healthcare AI-agent failures cluster around nine clinical domains — diagnosis, treatment, medication safety, documentation, compliance — and within each domain, failures are not hallucinations or knowledge gaps but rather architectural blind spots: checks scoped too narrowly, information dropped at agent handoffs, data treated as ground truth without verification, or a model’s learned pattern applied with inappropriate confidence to an individual case. Across all nine goals, the pattern is the same: an agent can reason correctly within its scoped input, but the scoped input is narrower than the clinical reality it is supposed to represent.

Key Takeaways

  • 38 patterns documented across 9 clinical-domain goals reveal a recurring failure taxonomy orthogonal to model capability: scope gaps (checks exclude relevant data), handoff drops (findings identified upstream fail to reach downstream agents), retrieval mismatches (similarity search or fuzzy matching substitutes wrong-but-plausible data), and individualization gaps (generic defaults applied where patient- or context-specific overrides are needed).
  • Multi-agent handoff failures account for 5 of 38 patterns — a recurring architectural problem where information established by one agent is lost or deprioritized when handed to the next.
  • Demographic bias and atypical-presentation blindness in diagnosis recur because training data reflects historical healthcare biases (female underrepresentation in cardiac studies, elderly underrepresentation in acute-disease research), and models reproduce those biases without explicit debiasing intervention.
  • Embedding-retrieval mismatches appear across three goals (adverse-drug-interaction, lab-result-interpretation, medication-reconciliation) and trace to the same root cause: similarity search over text (drug names, assay names) is used where exact-identifier matching is required.

Healthcare Goals

GoalCoversPatterns
Adverse Drug InteractionPairwise and multi-way interaction gaps, supplement exclusion, organ-function dosing, causal attribution7
Clinical DocumentationUnverified allergy fields, upcoding/downcoding from inflated documentation2
Compliance & LiabilityDe-identification quasi-identifier risk, informed-consent source fidelity, consent-scope handoff drops3
Diagnosis SafetyAnchoring, demographic/presentation-type bias, history truncation, rare-disease misses, data-interpretation gaps10
Lab Result InterpretationCritical-value alert routing, reference-range retrieval mismatches, patient-identity verification3
Medication ReconciliationDischarge reconciliation gaps, LASA drug substitution, interaction-flag handoff drops3
Mental-Health TriageIndirect risk-language blindness, risk-disclosure handoff drops2
Telehealth TriageUrgency-flag handoff drops, missing-vitals default-to-normal2
Treatment PlanningCare-goal drift, comorbidity neglect, guideline-conflict blindness, specialist-contraindication handoff drops, outdated guidelines, pediatric-dosing errors6

Total: 38 patterns

How the Goals Relate

The nine healthcare goals are mostly parallel concerns rather than a strict pipeline, because a healthcare agent can fail at any domain independently of whether others succeeded. Diagnosis Safety failures in the assessment step do not determine whether Medication Reconciliation or Treatment Planning will succeed, though a missed diagnosis can certainly propagate into wrong recommendations downstream. To localize an incident by symptom: an agent identifies the wrong diagnosis entirely → Diagnosis Safety; a drug-drug or condition-drug interaction is missed → Adverse Drug Interaction; a medication from home is silently omitted from discharge → Medication Reconciliation; a specialist-noted contraindication to a procedure approach is not incorporated into the final plan → Treatment Planning; a critical lab value is not immediately notified → Lab Result Interpretation; vital signs are missing and the acuity score defaults to normal → Telehealth Triage; a patient-specific care goal set in prior visit is silently overridden → Treatment Planning; risk factors disclosed during intake are not propagated to scheduling → Mental-Health Triage or Telehealth Triage; documentation inflates or misrepresents what was discussed → Clinical Documentation or Compliance & Liability.

Frequently Asked Questions

How do healthcare AI-agent failures concentrate on scope gaps rather than model capability?

Models are often capable of reasoning correctly within a well-posed question; the failures documented here are not about reasoning but about what input reaches the reasoning step. A pairwise interaction checker works correctly; the problem is it was built to check pairwise interactions and never sees condition-based contraindications. A similarity-search lookup works correctly; the problem is it retrieves a name-similar but clinically distinct drug. An agent is asked to generate a note without source verification; it generates plausible-sounding boilerplate that the source never supported. Scope-gap failures are architecture and design questions, not capability questions.

How do you reduce multi-agent handoff failures?

Require explicit, structured fields in handoff payloads for findings that downstream agents need to act on (risk flags, contraindications, narrowed consent scopes). Implement automated reconciliation checks comparing upstream reasoning against downstream schema fields before the handoff is considered complete. Require downstream agents to explicitly acknowledge or resolve any flag present in the handoff payload before issuing their own conclusion. Consider replacing agent-local conversational summaries with a single shared case record that both agents read from and write to.

How does demographic bias recur across diagnosis and treatment?

Training data reflects historical healthcare biases: women are underrepresented in cardiovascular-disease research, so models learn “female + chest pain” as lower-risk; elderly patients are historically underrepresented in acute-disease studies, so models learn age-driven discounting of acuity. Without explicit debiasing — stratified training data, fairness constraints, demographic-stratified outcome auditing — models reproduce the training-data biases. Mitigate by ensuring representative training data, auditing outcomes by demographic stratum, and applying demographic-aware differential expansion or symptom-threshold adjustment at inference time.

Can a single model update fix multiple patterns across different goals?

Only partially. Some fixes are architectural (add a field to a handoff schema, split a check into two stages); some are about data and training (representative training data, calibrating to indirect risk language); some are about governance (updating guidelines quarterly, maintaining a LASA-pair list). Across all 38 patterns, the recurring theme is that verification, grounding, and structured propagation matter more than model capability.

How do you audit for healthcare AI-agent failures in production?

Implement stratified outcome monitoring across all nine domains: diagnostic accuracy stratified by demographic group and presentation type; medication error rates stratified by polypharmacy tier; critical-value notification latency; medication-reconciliation omission rates; multi-agent handoff completeness checks comparing upstream findings against downstream schema; care-plan goal continuity; guideline freshness. Establish clear alert thresholds for each domain so systematic failures are caught before they propagate to many patients.

  • Document Processing — covers OCR and multimodal failures that upstream feed into clinical note extraction and EHR data quality
  • Knowledge Retrieval — covers guideline retrieval, knowledge freshness, and citation accuracy failures that treatment-planning and diagnosis agents depend on

Age Bias in Symptom Interpretation

Frequency: Very Common
Category:

Model Misinterprets Symptoms Differently Based on Patient Age; Younger Patients Undertreated, Older Patients Overtreated

Anchoring Bias on First Diagnosis

Frequency: Very Common
Category:

Agent Fixates on the First Plausible Diagnosis Suggested Early in the Conversation and Discounts Later Contradictory Evidence

Atypical Presentation Blindness

Frequency: Common
Category:

Model Trained Predominantly on "Textbook" Symptom Presentations Misses Atypical Presentations Common in Women, Elderly, and Diverse Populations

Contraindication Omission

Frequency: Common
Category:

Agent Recommends or Approves a Medication Without Checking Patient-Specific Contraindications Beyond Drug-Drug Interactions

Embedding Retrieval Matches Look-Alike/Sound-Alike Drug Name

Frequency: Occasional
Category:

A Medication-Reconciliation Agent Matching a Free-Text or Handwritten Medication Entry Against a Structured Formulary Database Uses Similarity-Based Lookup That Resolves the Entry to a Lexically Similar but Pharmacologically Different Look-Alike/Sound-Alike (LASA) Drug, and the Reconciled Medication List Carries the Wrong Drug Forward Into the Patient's Active List

Embedding Retrieval Matches Similarly Named Lab Panel With Different Reference Range

Frequency: Occasional
Category:

An Agent Interpreting a Lab Result That Looks Up the Applicable Reference Range Via Semantic Search Over a Reference-Range Knowledge Base, Rather Than an Exact Assay-Code Match, Retrieves the Range for a Differently Named but Textually Similar Test -- Such as Confusing "Vitamin D, 25-Hydroxy" With "Vitamin D, 1,25-Dihydroxy" -- and Flags or Clears the Result Against the Wrong Range

Embedding Retrieval Matches Structurally Similar, Different-Class Drug for Interaction Check

Frequency: Occasional
Category:

An Agent Checking a Medication List for Drug-Drug Interactions, Using Semantic Similarity Search Over an Interaction Knowledge Base to Find the Relevant Interaction Profile for a Given Drug, Retrieves the Profile for a Structurally or Name-Similar but Pharmacologically Distinct Drug, and Clears or Flags the Combination Based on the Wrong Drug's Interaction Data

Empty Allergy-Query Result Documented as Confirmed No-Known-Allergies

Frequency: Occasional
Category:

A Clinical-Summary or After-Visit-Note Agent Queries a Structured EHR Field for the Patient's Allergy History, the Query Returns Zero Records (Because the Field Was Never Populated, Not Because a Clinician Affirmatively Confirmed the Patient Has No Allergies), and the Agent's Note-Generation Step Renders This as "No Known Drug Allergies" -- an Affirmative Clinical Statement the Underlying Data Never Supported

Guideline Conflict Resolution Failure

Frequency: Common
Category:

Agent Cannot Reconcile Conflicting Recommendations From Different Clinical Guideline Bodies and Defaults to an Arbitrary or Most-Recently-Retrieved Source

Imaging Report Discrepancy Blindness

Frequency: Common
Category:

Agent Summarizes the Radiology Impression Line Without Reconciling It Against the Ordering Clinician's Stated Question or Prior Comparison Imaging

Informed-Consent Documentation Gap

Frequency: Occasional
Category:

Agent Drafts or Summarizes Clinical Documentation Implying Informed Consent Was Obtained and Discussed in Detail That Was Not Actually Covered in the Encounter

Multi-Agent Handoff Drops Disclosed Risk Factor Between Intake and Scheduling Agent

Frequency: Occasional
Category:

A Chat-Based Mental-Health Intake Agent That Elicits and Records a Significant Risk Disclosure During Conversation Captures That Finding Only in Its Own Conversational Reasoning or a Free-Text Summary, and When the Case Is Handed Off to a Downstream Scheduling/Routing Agent That Acts on a Structured Acuity Field to Determine Appointment Urgency, the Disclosed Risk Factor Never Crosses the Handoff Boundary, So the Case Is Scheduled at a Routine Priority as if the Disclosure Had Never Occurred

Multi-Agent Handoff Drops Flagged Interaction Between Reconciliation and Pharmacy-Review Agent

Frequency: Occasional
Category:

A Medication-Reconciliation Agent That Identifies a Specific Drug-Drug Interaction Risk Between a Newly Prescribed Medication and a Continuing Home Medication Records That Finding Only Within Its Own Free-Text Reasoning or Conversational Summary, and When the Reconciled Medication List Is Handed Off to a Downstream Pharmacy-Review Agent That Consumes Only the Structured Medication List, the Interaction Flag Never Crosses the Handoff Boundary, So the Pharmacy-Review Agent Approves the List as if No Interaction Risk Had Been Identified

Multi-Agent Handoff Drops Narrowed Consent Scope Between Intake and Billing Agent

Frequency: Occasional
Category:

An Intake Agent That Records a Patient's Narrowed Consent -- For Example, Consent to Treatment but Explicit Refusal of Consent to Share Records With a Specific Third-Party Payer or Research Registry -- Captures That Restriction Only as a Note Within Its Own Free-Text Reasoning or Conversation Summary, and a Downstream Billing or Records-Release Agent That Acts on a Structured Patient-Status Field Never Receives the Restriction, Proceeding as if Full Consent Were Granted

Multi-Agent Handoff Drops Specialist-Noted Contraindication Before Care-Plan Finalization

Frequency: Occasional
Category:

A Specialist-Consult Agent That Identifies, in Its Own Consult-Note Reasoning, a Contraindication to a Specific Treatment Approach Hands That Finding Off to a Primary Treatment-Planning Agent Through a Structured Consult Summary That Has No Field for Contraindications, So the Treatment-Planning Agent Finalizes a Care Plan Including the Approach the Specialist Had Ruled Out

Patient-Identity Mismatch in Tool-Retrieved Lab Payload Accepted Without Verification

Frequency: Rare
Category:

An Agent Calls a Structured EHR/FHIR Tool to Retrieve a Patient's Latest Lab Results, and Because the Underlying Record System Has a Duplicate or Merged Medical-Record-Number Entry, the Returned Payload Belongs to a Different Patient; the Agent Treats the Tool's Structured Response as Ground Truth and Interprets the Wrong Patient's Values Without Cross-Checking the Payload's Own Patient Identifiers Against the Request