From Foundational to Multimodal Models for Medical Imaging
Tutorial covering foundational models, multimodal models, and emerging approaches for medical imaging and healthcare. Jointly presented with Yunsoo Kim and Ismail Bin Ayed.
Keynotes, plenary lectures, invited talks, tutorials, and other presentations spanning artificial intelligence, medical imaging, healthcare AI, multimodal systems, and trustworthy AI.
Selected presentations and accompanying slides are collected here as a record of research ideas, technical directions, and perspectives shared with academic, industry, and healthcare communities.
The next frontier of high-stakes AI is not only generating answers, but ensuring their correctness. This talk explores the transition from generation-centric AI toward verification-centric systems capable of checking, explaining, and improving AI outputs demonstrated in the domain of radiology AI on chest X-ray image interpretation.
Tutorial covering foundational models, multimodal models, and emerging approaches for medical imaging and healthcare. Jointly presented with Yunsoo Kim and Ismail Bin Ayed.
Generative AI is rapidly transforming radiology, but in high-stakes clinical settings, generating a plausible report is not enough, we need to know whether it is correct. This talk introduces Verification AI, a new class of domain-specific models designed to detect, measure, correct, and ultimately prevent errors in AI-generated clinical reports. Drawing upon recent work of students and collaborators appearing in conferences such as CVPR, MICCAI, and NeurIPS, I will highlight approaches for fine-grained, location-grounded verification and automated correction of chest X-ray reports, including fact-checking discriminative models and retrieval-augmented LLM judges. Together these advances point to a fundamental shift in clinical AI from generating text to ensuring its correctness, thus creating a critical layer of trust and safety for the next generation of healthcare AI.
Understanding the human brain—arguably the most complex biological system—demands integrative frameworks that bridge molecular, structural, functional, and clinical levels of observation. This in turn requires integrating diverse data types including genomic, neuroimaging, and clinical data, which has traditionally been addressed using early, mid, or late fusion methods. Through long-term academic–industrial collaborations, Dr Syeda-Mahmood’s group has been developing novel neural architectures tailored for multimodal fusion. They have applied these architectures to model fusion across diverse diseases ranging from cancer and cardiac diseases to neurodegenerative disorders such as Alzheimer’s disease. Dr Syeda-Mahmood will present a range of approaches, from statistical fusion methods such as sparse canonical correlation analysis to generalized neural fusion models based on multiplexed and multi-layer graph neural networks. She will highlight the translational impact of these methods through a recent study that identified 23 genes associated with cardiac morphology, illustrating the broader potential of multimodal fusion to drive biologically meaningful discovery in brain research and precision neuroscience.
Vision-language models (VLM) bring image and textual representations close together in a joint embedding space, which is useful for tagging and retrieval from content stores. However such associations are not very stable in that a synonymous textual query does not retrieve the same set of images or with a high degree of overlap. This is due to the absence of linkages between semantically related concepts in vision-language models. In contrast, the episodic memory store in the brain has linkages to the semantic conceptual memory subsystem which helps in both the formation and recall of memories. In this paper, we exploit this paradigm to link a VLM to a semantic memory thereby producing a new semantic vision-language model called SemCLIP. Specifically, we develop a semantic memory model for the language of object-naming nouns reflecting their semantic similarity. We then link a vision language model to the semantic memory model through a semantic alignment transform. This leads to a richer and more stable understanding of the concepts by bringing synonymous visual concepts and their associated images closer. Both the semantic memory model and the alignment transform can be learned from word knowledge sources thus avoiding large-scale retraining of VLMs from real-world image-text pairs. The resulting model is shown to outperform existing embedding models for semantic similarity and downstream tasks of retrieval on multiple datasets.
Neuroscience has been a building block for artificial intelligence since this field began, and as recognized by the recent Nobel prize in Physics for the work on Hopfield networks. Building computational models of different brain systems can give us new insight into the development of next generation computer systems. At IBM Research, we are pioneering one such ambitious project to develop a computational model of human declarative memory system in the brain. This is intended to serve both as a virtual memory prosthetics for memory-impaired patients and to guide the design of next-generation computer storage systems. In this talk I will describe our emerging work on new neural architectures for semantic and episodic memories as well as representation and encoding methods in the tri-synaptic circuit of the hippocampus. Specifically, I will present our latest work on developing a foundational model for knowledge to emulate human semantic and episodic memories using cross-linked neural embeddings for visual and language concepts. I will also describe a new neural model of storage and retrieval called cross-modal Hopfield encoding networks which make the Hopfield networks more practical for use in computer systems. I will discuss how these new neural architectures are beginning to influence IBM’s hardware accelerated AI/ML infrastructure solutions including the recently announced IBM content-aware storage system jointly with NVIDIA that convert passive storage devices to active content-understanding systems.
An overview of foundational and multimodal models and their applications to medical imaging.
18+ keynote lectures across AI, computer vision, medical imaging, and healthcare.
6 plenary lectures addressing emerging directions in AI and healthcare technology.
25+ invited talks at universities, research institutions, conferences, and industry organizations worldwide.
Tutorials and advanced training programs covering foundation models, multimodal AI, and medical imaging.
Moving AI from generation toward verification, factuality, reliability, and trustworthy deployment.
Multimodal foundation models, retrieval, memory, and intelligent systems.
Clinical AI, medical imaging, decision support, precision AI, and healthcare transformation.
Translating research advances into enterprise platforms, products, and organizational impact.
I welcome invitations to speak on AI strategy, scientific direction, healthcare AI, multimodal foundation models, verification AI, and the translation of research into real-world systems.