Age Bias in Symptom Interpretation
Model Misinterprets Symptoms Differently Based on Patient Age; Younger Patients Undertreated, Older Patients Overtreated
38 patterns in this category
Healthcare AI-agent failures cluster around nine clinical domains — diagnosis, treatment, medication safety, documentation, compliance — and within each domain, failures are not hallucinations or knowledge gaps but rather architectural blind spots: checks scoped too narrowly, information dropped at agent handoffs, data treated as ground truth without verification, or a model’s learned pattern applied with inappropriate confidence to an individual case. Across all nine goals, the pattern is the same: an agent can reason correctly within its scoped input, but the scoped input is narrower than the clinical reality it is supposed to represent.
| Goal | Covers | Patterns |
|---|---|---|
| Adverse Drug Interaction | Pairwise and multi-way interaction gaps, supplement exclusion, organ-function dosing, causal attribution | 7 |
| Clinical Documentation | Unverified allergy fields, upcoding/downcoding from inflated documentation | 2 |
| Compliance & Liability | De-identification quasi-identifier risk, informed-consent source fidelity, consent-scope handoff drops | 3 |
| Diagnosis Safety | Anchoring, demographic/presentation-type bias, history truncation, rare-disease misses, data-interpretation gaps | 10 |
| Lab Result Interpretation | Critical-value alert routing, reference-range retrieval mismatches, patient-identity verification | 3 |
| Medication Reconciliation | Discharge reconciliation gaps, LASA drug substitution, interaction-flag handoff drops | 3 |
| Mental-Health Triage | Indirect risk-language blindness, risk-disclosure handoff drops | 2 |
| Telehealth Triage | Urgency-flag handoff drops, missing-vitals default-to-normal | 2 |
| Treatment Planning | Care-goal drift, comorbidity neglect, guideline-conflict blindness, specialist-contraindication handoff drops, outdated guidelines, pediatric-dosing errors | 6 |
Total: 38 patterns
The nine healthcare goals are mostly parallel concerns rather than a strict pipeline, because a healthcare agent can fail at any domain independently of whether others succeeded. Diagnosis Safety failures in the assessment step do not determine whether Medication Reconciliation or Treatment Planning will succeed, though a missed diagnosis can certainly propagate into wrong recommendations downstream. To localize an incident by symptom: an agent identifies the wrong diagnosis entirely → Diagnosis Safety; a drug-drug or condition-drug interaction is missed → Adverse Drug Interaction; a medication from home is silently omitted from discharge → Medication Reconciliation; a specialist-noted contraindication to a procedure approach is not incorporated into the final plan → Treatment Planning; a critical lab value is not immediately notified → Lab Result Interpretation; vital signs are missing and the acuity score defaults to normal → Telehealth Triage; a patient-specific care goal set in prior visit is silently overridden → Treatment Planning; risk factors disclosed during intake are not propagated to scheduling → Mental-Health Triage or Telehealth Triage; documentation inflates or misrepresents what was discussed → Clinical Documentation or Compliance & Liability.
Models are often capable of reasoning correctly within a well-posed question; the failures documented here are not about reasoning but about what input reaches the reasoning step. A pairwise interaction checker works correctly; the problem is it was built to check pairwise interactions and never sees condition-based contraindications. A similarity-search lookup works correctly; the problem is it retrieves a name-similar but clinically distinct drug. An agent is asked to generate a note without source verification; it generates plausible-sounding boilerplate that the source never supported. Scope-gap failures are architecture and design questions, not capability questions.
Require explicit, structured fields in handoff payloads for findings that downstream agents need to act on (risk flags, contraindications, narrowed consent scopes). Implement automated reconciliation checks comparing upstream reasoning against downstream schema fields before the handoff is considered complete. Require downstream agents to explicitly acknowledge or resolve any flag present in the handoff payload before issuing their own conclusion. Consider replacing agent-local conversational summaries with a single shared case record that both agents read from and write to.
Training data reflects historical healthcare biases: women are underrepresented in cardiovascular-disease research, so models learn “female + chest pain” as lower-risk; elderly patients are historically underrepresented in acute-disease studies, so models learn age-driven discounting of acuity. Without explicit debiasing — stratified training data, fairness constraints, demographic-stratified outcome auditing — models reproduce the training-data biases. Mitigate by ensuring representative training data, auditing outcomes by demographic stratum, and applying demographic-aware differential expansion or symptom-threshold adjustment at inference time.
Only partially. Some fixes are architectural (add a field to a handoff schema, split a check into two stages); some are about data and training (representative training data, calibrating to indirect risk language); some are about governance (updating guidelines quarterly, maintaining a LASA-pair list). Across all 38 patterns, the recurring theme is that verification, grounding, and structured propagation matter more than model capability.
Implement stratified outcome monitoring across all nine domains: diagnostic accuracy stratified by demographic group and presentation type; medication error rates stratified by polypharmacy tier; critical-value notification latency; medication-reconciliation omission rates; multi-agent handoff completeness checks comparing upstream findings against downstream schema; care-plan goal continuity; guideline freshness. Establish clear alert thresholds for each domain so systematic failures are caught before they propagate to many patients.
Model Misinterprets Symptoms Differently Based on Patient Age; Younger Patients Undertreated, Older Patients Overtreated
Agent Fixates on the First Plausible Diagnosis Suggested Early in the Conversation and Discounts Later Contradictory Evidence
Model Trained Predominantly on "Textbook" Symptom Presentations Misses Atypical Presentations Common in Women, Elderly, and Diverse Populations
Agent Regenerates the Care Plan at Each Visit From the Current Problem List Alone, Losing Continuity With Previously Agreed Patient Goals
Model Recommends Treatment Optimal for Primary Condition But Dangerous for Comorbidities Patient Has
Agent Anchors on a Prior Visit's Diagnosis and Discounts New Evidence Contradicting It
Agent Recommends or Approves a Medication Without Checking Patient-Specific Contraindications Beyond Drug-Drug Interactions
Agent Summarizes a Critical Lab Result Within a Routine Note Instead of Triggering Immediate Notification
Model Biased Toward/Against Certain Demographics (Race, Gender, Age); Different Accuracy for Different Groups
Agent Generates Discharge Medication List Without Reconciling Against Pre-Admission Home Medications
Agent-Generated Clinical Note Language Inflates or Deflates Billing Code Level Relative to Actual Care Delivered
Agent Recommends Standard Adult Dosing Without Adjusting for Impaired Renal or Hepatic Clearance
Model Recommends Drug Combination Without Flagging Known Dangerous Interactions
A Medication-Reconciliation Agent Matching a Free-Text or Handwritten Medication Entry Against a Structured Formulary Database Uses Similarity-Based Lookup That Resolves the Entry to a Lexically Similar but Pharmacologically Different Look-Alike/Sound-Alike (LASA) Drug, and the Reconciled Medication List Carries the Wrong Drug Forward Into the Patient's Active List
An Agent Interpreting a Lab Result That Looks Up the Applicable Reference Range Via Semantic Search Over a Reference-Range Knowledge Base, Rather Than an Exact Assay-Code Match, Retrieves the Range for a Differently Named but Textually Similar Test -- Such as Confusing "Vitamin D, 25-Hydroxy" With "Vitamin D, 1,25-Dihydroxy" -- and Flags or Clears the Result Against the Wrong Range
An Agent Checking a Medication List for Drug-Drug Interactions, Using Semantic Similarity Search Over an Interaction Knowledge Base to Find the Relevant Interaction Profile for a Given Drug, Retrieves the Profile for a Structurally or Name-Similar but Pharmacologically Distinct Drug, and Clears or Flags the Combination Based on the Wrong Drug's Interaction Data
A Clinical-Summary or After-Visit-Note Agent Queries a Structured EHR Field for the Patient's Allergy History, the Query Returns Zero Records (Because the Field Was Never Populated, Not Because a Clinician Affirmatively Confirmed the Patient Has No Allergies), and the Agent's Note-Generation Step Renders This as "No Known Drug Allergies" -- an Affirmative Clinical Statement the Underlying Data Never Supported
Agent Cannot Reconcile Conflicting Recommendations From Different Clinical Guideline Bodies and Defaults to an Arbitrary or Most-Recently-Retrieved Source
Agent Checks Prescription-Drug Interactions but Ignores Patient-Reported Supplements and Herbal Products
Agent Produces a "De-Identified" Summary, Research Extract, or Support Ticket That Still Contains Re-Identifiable Information
Agent Summarizes the Radiology Impression Line Without Reconciling It Against the Ordering Clinician's Stated Question or Prior Comparison Imaging
Agent Drafts or Summarizes Clinical Documentation Implying Informed Consent Was Obtained and Discussed in Detail That Was Not Actually Covered in the Encounter
Agent Applies Generic Adult Reference Ranges to Lab Values Without Adjusting for Age, Sex, Pregnancy, or Assay-Specific Ranges
A Chat-Based Mental-Health Intake Agent That Elicits and Records a Significant Risk Disclosure During Conversation Captures That Finding Only in Its Own Conversational Reasoning or a Free-Text Summary, and When the Case Is Handed Off to a Downstream Scheduling/Routing Agent That Acts on a Structured Acuity Field to Determine Appointment Urgency, the Disclosed Risk Factor Never Crosses the Handoff Boundary, So the Case Is Scheduled at a Routine Priority as if the Disclosure Had Never Occurred
A Telehealth Triage Bot That Reasons, in Its Own Free-Text Assessment, That a Patient's Symptoms Warrant Urgent Same-Hour Clinician Contact Hands the Case to an On-Call Clinician Queue Through a Structured Ticket That Carries Only a Generic Priority Field, So the Urgency the Bot Actually Concluded Never Reaches the Queue's Sort Order
A Medication-Reconciliation Agent That Identifies a Specific Drug-Drug Interaction Risk Between a Newly Prescribed Medication and a Continuing Home Medication Records That Finding Only Within Its Own Free-Text Reasoning or Conversational Summary, and When the Reconciled Medication List Is Handed Off to a Downstream Pharmacy-Review Agent That Consumes Only the Structured Medication List, the Interaction Flag Never Crosses the Handoff Boundary, So the Pharmacy-Review Agent Approves the List as if No Interaction Risk Had Been Identified
An Intake Agent That Records a Patient's Narrowed Consent -- For Example, Consent to Treatment but Explicit Refusal of Consent to Share Records With a Specific Third-Party Payer or Research Registry -- Captures That Restriction Only as a Note Within Its Own Free-Text Reasoning or Conversation Summary, and a Downstream Billing or Records-Release Agent That Acts on a Structured Patient-Status Field Never Receives the Restriction, Proceeding as if Full Consent Were Granted
A Specialist-Consult Agent That Identifies, in Its Own Consult-Note Reasoning, a Contraindication to a Specific Treatment Approach Hands That Finding Off to a Primary Treatment-Planning Agent Through a Structured Consult Summary That Has No Field for Contraindications, So the Treatment-Planning Agent Finalizes a Care Plan Including the Approach the Specialist Had Ruled Out
Model Uses Medical Guidelines That Have Been Superseded by Newer Research; Recommends Treatment No Longer Best-Practice
Medical Diagnosis Model Trained on Limited Patient History; Misses Patterns Evident Only in Full Longitudinal Record
An Agent Calls a Structured EHR/FHIR Tool to Retrieve a Patient's Latest Lab Results, and Because the Underlying Record System Has a Duplicate or Merged Medical-Record-Number Entry, the Returned Payload Belongs to a Different Patient; the Agent Treats the Tool's Structured Response as Ground Truth and Interprets the Wrong Patient's Values Without Cross-Checking the Payload's Own Patient Identifiers Against the Request
Agent Extrapolates Adult Weight-Based or Fixed Dosing Formulas to Pediatric Patients, Producing Unsafe Doses
Patient on 5+ Medications; Pairwise Interaction Checking Misses Three-Way or Four-Way Drug Interactions
Model Trained on Common Diseases; Misses or Misdiagnoses Rare Conditions
Agent Triages a Telehealth Encounter as Lower Acuity Because Objective Vitals Are Simply Unavailable, Not Because They Are Normal
An Agent Reviewing a Patient's Chart for a Suspected Adverse Drug Reaction Attributes a New Symptom to Whichever Medication Was Most Recently Started, Based on Temporal Proximity Alone, Instead of Applying a Structured Causality Assessment, Producing a Confident-Sounding Attribution That May Implicate the Wrong Drug or Miss the Actual Cause
Agent Triages a Message Containing Indirect Self-Harm Indicators as Low Acuity Because No Explicit Keyword Was Present
Model Anchors on Initial Symptom; Misses True Diagnosis Because Anchored to Wrong Hypothesis