Building trustworthy AI in medicine: The role of explainability and cognitive load

The effectiveness of AI-powered CDSSs in clinical practice hinges on clinicians' trust and their willingness to rely on AI recommendations. AI models often operate as “black boxes,” producing results without clear reasoning, which can lead to skepticism among medical professionals. The explainability of AI systems - or the ability to provide clear, interpretable justifications for their outputs - has emerged as a critical factor in fostering trust and acceptance among clinicians.

Building trustworthy AI in medicine: The role of explainability and cognitive load
Representative Image. Credit: ChatGPT

Artificial Intelligence (AI) has become a transformative force in healthcare, particularly in assisting medical professionals with diagnostic accuracy, workflow efficiency, and decision-making. Among various applications, AI-driven Clinical Decision Support Systems (CDSSs) play a crucial role in breast cancer detection, aiding radiologists and oncologists by analyzing medical images, identifying abnormalities, and providing diagnostic recommendations. However, despite the potential benefits, the successful integration of AI into clinical workflows depends on clinicians' trust and acceptance of these systems.

A recent study titled "Explainability and AI Confidence in Clinical Decision Support Systems: Effects on Trust, Diagnostic Performance, and Cognitive Load in Breast Cancer Care", authored by Olya Rezaeian (Stevens Institute of Technology), Alparslan Emrah Bayrak (Lehigh University), and Onur Asan (Stevens Institute of Technology), investigates how varying levels of explainability and AI confidence scores influence clinicians' trust, cognitive load, and diagnostic performance. Published in January 2025, the study examines how healthcare professionals interact with AI-based CDSSs and provides insights into designing AI systems that are transparent, reliable, and seamlessly integrated into medical decision-making processes.

The importance of explainability and AI confidence in clinical AI systems

The effectiveness of AI-powered CDSSs in clinical practice hinges on clinicians' trust and their willingness to rely on AI recommendations. AI models often operate as "black boxes," producing results without clear reasoning, which can lead to skepticism among medical professionals. The explainability of AI systems - or the ability to provide clear, interpretable justifications for their outputs - has emerged as a critical factor in fostering trust and acceptance among clinicians. Additionally, AI confidence scores, which indicate the system's certainty in its predictions, influence how much weight clinicians give to AI-generated recommendations.

The study was designed to evaluate how different levels of AI explainability and confidence impact clinicians' diagnostic decision-making. A controlled web-based experiment was conducted, where 28 healthcare professionals, including radiologists and oncologists, interacted with an AI-powered CDSS that provided varying levels of explanation and confidence in breast cancer diagnoses. The research sought to answer three fundamental questions: (1) How does explainability impact clinicians' cognitive load? (2) How do AI confidence scores affect trust and diagnostic performance? (3) How do demographic factors shape clinicians' interaction with AI-driven CDSSs?

Evaluating clinicians' interaction with AI-based CDSS

The study employed an interrupted time series design, where clinicians evaluated ultrasound images of breast tissues both independently and with varying levels of AI assistance. Participants completed diagnostic tasks under five conditions, progressing through different levels of explainability and confidence cues provided by the AI system.

The Baseline Condition (Independent Diagnosis) required clinicians to analyze breast ultrasound images and classify them as healthy, benign, or malignant without any AI assistance. In Intervention I (Classification Only), the AI provided a diagnostic classification but no additional explanations. In Intervention II (Probability Distribution), the AI displayed confidence scores for each diagnosis, giving participants insight into the system's certainty levels. Intervention III (Tumor Localization) introduced visual tumor localization, helping clinicians understand which regions of the image contributed to the AI's decision. Finally, Intervention IV (Enhanced Localization + Confidence) combined localized tumor detection with confidence indicators, marking areas where the AI was highly or moderately confident.

Throughout the study, participants' trust in AI, cognitive load, stress levels, and diagnostic accuracy were measured through surveys and system interactions. This methodology allowed for a comprehensive analysis of how explainability and confidence levels influence clinician behavior and decision-making.

Explainability, AI confidence, and clinician trust

AI Explainability Affects Cognitive Load and Stress Levels

The study found that while increased explainability generally reduced mental demand, it also led to higher stress levels, particularly in the tumor localization with probability estimates condition. Clinicians experienced more cognitive strain when explanations were too detailed or misaligned with their expectations.

Simpler explanations, such as those providing classification-only recommendations or probability distributions, had little impact on cognitive load. However, when the AI system introduced localized tumor detection with probability estimates, stress levels significantly increased. This suggests that while visual explanations can enhance understanding, they may also introduce cognitive overload if not designed carefully. The findings highlight the need for clear and concise AI explanations that provide meaningful information without overwhelming clinicians.

AI Confidence Scores Influence Trust, Agreement, and Diagnostic Performance

The study revealed that AI confidence scores play a significant role in shaping clinician trust and reliance on AI recommendations. Low-confidence AI outputs decreased trust and agreement, causing clinicians to spend more time on diagnoses and rely more on their own judgment. In contrast, high-confidence AI outputs increased trust but also led to overreliance, resulting in a slight decline in diagnostic accuracy.

When AI predictions were accompanied by low confidence scores, clinicians exhibited more cautious behavior, questioning AI suggestions and extending their diagnosis duration. Conversely, high confidence scores increased clinicians' reliance on AI, sometimes leading to automation bias, where clinicians deferred to AI recommendations without critically assessing their accuracy. These findings suggest that AI confidence indicators must be carefully designed to encourage appropriate levels of human oversight and decision-making.

Demographic Factors Shape Perceptions of AI in Healthcare

The study also examined how age, gender, and professional role influenced clinicians' trust, stress levels, and perceptions of AI usability. Older clinicians were more likely to trust AI and view it as useful compared to younger clinicians, who reported higher stress levels and cognitive demand when interacting with AI-based CDSSs. Gender differences also emerged, with female clinicians perceiving AI-driven tasks as more complex and cognitively demanding than their male counterparts.

Professional role significantly influenced AI adoption, with radiologists generally perceiving AI as a valuable diagnostic tool, while oncologists found AI-based decision-making more challenging and stressful. These findings underscore the importance of designing AI systems that accommodate different levels of expertise, experience, and cognitive preferences across clinician groups.

Implications for AI-driven clinical decision support systems

The study provides several important insights for AI developers, policymakers, and healthcare institutions seeking to integrate AI-driven CDSSs into clinical workflows. Explainability features should be designed to enhance understanding without introducing cognitive overload, ensuring that AI-generated explanations are clear, relevant, and aligned with clinicians' mental models.

AI confidence scores must be calibrated to prevent overreliance, ensuring that clinicians remain actively engaged in diagnostic decision-making. Additionally, demographic differences must be considered when designing AI-based CDSSs, as different clinician groups have varying levels of trust, cognitive load tolerance, and AI adoption readiness.

  • FIRST PUBLISHED IN:
  • Devdiscourse
Give Feedback

Use this form for editorial or site feedback. We usually reply within 2 to 3 working days.

By submitting, you agree that we may use your email address to respond.