By: Akanksha
While medical artificial intelligence (AI) tools are increasingly being used in medical imaging, as a “second pair of eyes” to flag possible abnormalities, doctors should exercise caution when relying on them.
A team of researchers at the International Institute of Information Technology - Hyderabad (IIIT-H) audited four vision-language models and found that the visual heatmaps generated by these tools to indicate disease in chest X-rays do not always match the areas radiologists consider clinically relevant. The results revealed that an AI system may seem to highlight the correct part of an X-ray without pinpointing the disease the way a human would.
Consequently, the findings underscore the need for clinicians to independently assess AI-generated findings rather than relying on them entirely for a final diagnosis.
Led by Prof. Parameswari Krishnamurthy at IIIT-H’s Language Technologies Research Centre (LTRC), the research team set out to answer a question that extends beyond whether an AI model arrives at the correct diagnosis: does the visual highlight truly match where a doctor sees the disease?
“We essentially wanted to examine whether the heatmaps created by vision-language models actually correspond to where radiologists, who look at the image, would say the disease lies,” explained Dr. Syed Faizan, principal investigator of the study titled, “How Well Do Chest X-Ray VLM Attention Overlays Match Radiologist Boxes? A Cross-Model Audit and Radiologist Reader Study,” to The Hindu.
To evaluate this, the researchers benchmarked four leading vision-language models — MAIRA-2, MedGamma-4B, LLaVA-Med-1.5 — against thousands of publicly available chest X-rays. To establish a ground-truth human benchmark, two radiologists reviewed the datasets with marked abnormality boxes and evaluated anonymized overlays, allowing the team to directly compare machine-generated heatmaps against clinical assessments.
“An AI model may appear to highlight the correct part of an image, but that does not necessarily mean it has identified the disease in the same way a radiologist would. The model may first arrive at a diagnosis and then use that diagnosis to determine where to place its heatmap. In other words, it may be working backwards,” said Dr. Syed Faizan to The Hindu.
To test this possibility, the team removed the diagnostic information and examined how the models localized abnormalities. Their performance dropped, suggesting that some of what looked like image-based reasoning could actually be influenced by an anatomical expectation learned from the diagnosis, rather than the model independently identifying the affected region from the image.
The human assessment revealed a notable difference from the automated audit : while MAIRA-2 was ranked higher for overlap with radiologist-identified regions, the two radiologists rated MedGemma higher, researchers said this difference may arise from how radiologists interpret visual information, highlighting the importance of assessing AI tools from a clinical perspective.
“This suggests that a model paying attention to a particular area — which may be the exact spot where the disorder lies — is not exactly helpful to a radiologist. A radiologist might want to look at not only the area with the disease, but also a broader surrounding area to know the extent of disease spread. MedGemma is doing that; it provides a broader area,” Dr. Faizan explained to The Hindu. A broader heatmap could therefore be more clinically useful than one that tightly overlaps with the suspected lesion.
Also Read: How Conversational AI Is Improving Patient Communication in Modern Healthcare
The findings have been accepted at the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2026, and are to be presented at the iMIMIC Satellite Event in Strasbourg.
The researchers said their work addressed a gap in the evaluation of medical AI systems, as earlier studies had largely focused on prediction and diagnostic accuracy rather than on whether AI-highlighted regions matched those identified by radiologists.
“Studies have always been conducted to test whether medical AI models are correct in their predictions or diagnosis. No prior study has ever tested whether heatmaps actually match the radiologists’ ‘bounding boxes,” said Dr. Faizan.
The LTRC is also investigating another potential weakness in vision-language models: their sensitivity to the wording of medical questions. In collaboration with a radiologist from CMC Vellore, the researchers developed different categories of ways in which a medical question could be rephrased and examined how vision-language models responded, with the study accepted at EMNLP.
Doctors may phrase the same clinical query differently — for example, asking whether a chest X-ray shows pneumonia or whether pneumonia can be ruled out – and may use technical terminology or more familiar expressions.
This study on paraphrase robustness explored how changes in phrasing affected the responses generated by the models. The larger question the researchers are seeking to answer is whether these systems genuinely understand the clinical query or depend excessively on the precise wording used.
“Our lab’s efforts are focused on how we use natural language processing (NLP) for healthcare. The goal is to leverage large language model (LLMs) and vision language model (VLMs) to help clinicians such that their time spent is cumbersome, mundane tasks such as documentation or report writing or patient-doctor communication is better utilized in some other place,” remarked Prof. Parameswari Krishnamurthy, adding that there is a specialized course being offered by their lab titled, “NLP for healthcare.”
The aim is to reduce the time doctors spend on routine administrative work and enable them to devote more attention to clinical decision making. Ultimately, the researchers emphasized that the objective of medical AI research should be to support clinicians, not replace them, leaving human judgment central to healthcare.
Also Read: New AI Tool Detects Widely Underdiagnosed Heart Condition
Reference:
Faizan, Syed, Mahesh Vasamsetti, Shivam Dilipbhai Kotak, Madhavi Latha Gundamaraju, and Parameswari Krishnamurthy. “How Well Do Chest X-Ray VLM Attention Overlays Match Radiologist Boxes? A Cross-Model Audit and Radiologist Reader Study.” In Medical Image Computing and Computer Assisted Intervention—MICCAI 2026 Workshops and Challenges. Springer Nature Switzerland, 2026. https://papers.miccai.org/miccai-2026-sat/iMIMIC_010.html
(Rh/APC)