Artificial intelligence in healthcare is entering a new phase. Earlier generations of medical AI were generally designed to perform one narrowly defined task. A model might detect diabetic retinopathy from a retinal photograph, identify a suspicious lesion in a medical scan, predict a patient’s risk of a particular disease, or classify a laboratory result. These systems demonstrated that machine learning could become highly capable within carefully defined clinical problems, but their intelligence remained largely specialised.
Foundation models are changing that architecture. Instead of training an artificial intelligence system for only one task, foundation models are developed through large-scale pretraining and can subsequently be adapted to multiple applications. When these models learn from different forms of medical information simultaneously, including clinical text, medical images, laboratory measurements, physiological signals, pathology slides, genomic information and other patient data, they begin to resemble a more general computational layer for medicine.
This shift is significant because human clinicians do not diagnose disease by looking at one isolated dataset. A physician may read a patient’s history, examine laboratory results, interpret an image, consider symptoms, review previous treatments and observe how the patient’s condition has changed over time. Medicine is inherently multimodal.
The emerging generation of healthcare foundation models is attempting to move AI closer to this reality. Recent research describes multimodal foundation models as systems capable of integrating different types of clinical information and transferring their learned representations across tasks and settings. At the same time, researchers caution that strong benchmark performance does not automatically translate into safe or reliable clinical use.
The central question is therefore not simply what happens when AI learns from more medical data. The more important question is what happens when AI learns to connect different kinds of medical information and begins to reason across them.
From Task-Specific AI to Foundation Models
Traditional medical AI typically begins with a clearly defined task. Researchers collect a dataset, label the relevant examples and train a model to produce a particular prediction. This approach can be highly effective, but the resulting model often has limited flexibility.
A foundation model follows a different philosophy. It is trained on large and diverse datasets to develop representations that can later be adapted to different tasks. The same underlying model can potentially support classification, prediction, summarisation, information extraction, image interpretation or clinical decision-support applications.
In healthcare, this flexibility is particularly attractive because medical information is fragmented across many formats. A patient’s clinical story exists partly in structured databases and partly in narrative notes. Imaging contains visual information, while laboratory results provide numerical measurements. Electrocardiograms contain physiological signals, and pathology slides contain microscopic visual patterns.
A foundation model that can learn relationships across these different forms of information has the potential to become considerably more versatile than a collection of isolated algorithms.
A 2026 comprehensive review of healthcare foundation models describes scalability, multimodal processing and adaptability as important characteristics of the emerging technology while also identifying privacy, bias, interpretability and sustainability as major barriers to clinical translation.
What Makes Medical Data Multimodal?
The human body generates information in many different forms. Medical imaging provides visual evidence of anatomy and disease. Pathology provides microscopic evidence at the cellular level. Clinical notes capture symptoms, observations and reasoning. Laboratory tests describe biochemical processes. Genomic information provides insight into biological variation. Wearable devices can generate continuous physiological measurements.
Each modality provides only part of the overall picture.
A chest scan may reveal an abnormality, but the clinical significance of that abnormality can depend on the patient’s age, symptoms, previous imaging and medical history. A laboratory value may appear abnormal in isolation but have a different meaning when considered alongside medications, previous measurements and other tests.
Multimodal AI attempts to integrate these sources rather than treating them as completely independent.
This is one reason healthcare is becoming an important testing ground for foundation models. Medicine provides enormous quantities of heterogeneous data, but it also demands careful integration because clinical decisions often depend on relationships between different forms of evidence.
When AI Learns From Images and Text Together
Medical imaging provides one of the clearest examples of the potential and limitations of multimodal foundation models.
A radiological image contains visual information, while the accompanying clinical description provides contextual information. A foundation model can potentially process both at the same time.
However, recent research has revealed an important complication. A June 2026 Nature Communications study evaluated eight multimodal foundation models across 1,090 multimodal medical cases and found that image predictions were strongly influenced by text. The models generally performed better when informative text was available, but misleading clinical descriptions could substantially reduce performance. In one experiment, introducing misleading text reduced the accuracy of one evaluated model from 84% to 28%.
This finding is extremely important for understanding medical multimodal AI. Combining modalities does not automatically mean that the system understands each modality independently and then integrates them perfectly.
A model may instead rely heavily on whichever information source is most statistically useful to it.
The lesson is that multimodal intelligence must be evaluated carefully. A system that appears to “see” a medical image may sometimes be relying more heavily on the accompanying text than clinicians realise.
Learning Across the Electronic Health Record
Electronic health records represent another major opportunity for foundation models.
An EHR contains structured information such as laboratory values, medications, diagnoses and vital signs, but it also contains unstructured information in clinical notes. A patient’s medical history may therefore exist as a mixture of tables, codes, measurements and narrative descriptions.
Multimodal generative models are increasingly being developed to work across these formats. A 2026 study in npj Health Systems described a multimodal generative model designed to integrate structured and unstructured EHR information, reflecting the broader movement toward systems capable of processing the different forms in which clinical information is recorded.
The potential application is substantial. Instead of asking a model to examine a single laboratory result or clinical note, healthcare systems could potentially use foundation models to interpret a patient’s broader longitudinal record.
This could allow AI to identify relationships between previous diagnoses, treatment history, laboratory trends and clinical observations that might otherwise remain distributed across different parts of the record.
The computational challenge is enormous, but so is the potential value.
The Rise of Multimodal Clinical Reasoning
The ultimate ambition is not simply to make AI capable of processing multiple inputs. It is to make those inputs useful for clinical reasoning.
A doctor does not simply collect information. The clinician determines which information matters, identifies contradictions, considers alternative explanations and decides what additional evidence may be necessary.
Foundation models are increasingly being evaluated for similar capabilities.
In July 2026, researchers reported a multimodal reasoning foundation model called QoQ-Med3 that was designed to transfer representations across clinical modalities, tasks and heterogeneous datasets. The model was evaluated across previously unseen modalities and held-out clinical datasets, illustrating the effort to move medical foundation models beyond narrow task-specific applications.
Such research represents an important direction because healthcare environments are inherently variable. A model trained in one dataset must ideally remain useful when it encounters different hospitals, devices, patient populations and clinical contexts.
However, generalisation remains one of the central challenges of medical AI.
Foundation Models Could Connect Previously Separate Data
One of the most exciting possibilities is that foundation models could become a common computational layer connecting different medical disciplines.
Radiology, pathology, cardiology, genomics and clinical medicine have historically developed specialised data systems and analytical methods. A foundation model could potentially learn representations that connect information across these domains.
Cancer research provides a strong example. Cancer diagnosis and treatment frequently involve imaging, pathology, molecular information, clinical history and longitudinal treatment data. A 2026 Nature Communications perspective describes multimodality, reasoning, expert subnetworks, agents and external tools as emerging components of cancer foundation models across screening, diagnosis, treatment, prognosis and follow-up.
This could eventually change how cancer AI is designed. Instead of using separate systems for image analysis, pathology interpretation and clinical summarisation, healthcare organisations may increasingly experiment with coordinated foundation-model architectures.
The goal would not necessarily be one enormous model that does everything. It could instead be a system in which different specialised components communicate through a shared computational framework.
Learning From Physiological Signals
Multimodal foundation models are also expanding beyond text and medical images.
Physiological signals such as electrocardiograms and photoplethysmograms provide information about cardiovascular function. These signals can be collected in hospitals, clinics and increasingly through wearable or portable devices.
A 2026 Nature Machine Intelligence study described a multimodal foundation model trained using data from 1.7 million individuals for cardiac health assessment across different scenarios and devices. The research illustrates how foundation-model approaches can potentially learn representations that transfer across datasets and measurement environments.
This direction is particularly important because health does not occur only during clinical appointments. Patients generate physiological information continuously.
If foundation models can learn from clinical measurements alongside other forms of health information, the computational representation of a patient could eventually become more longitudinal and dynamic.
From Diagnosis to Prediction
The potential of foundation models extends beyond recognising disease that is already visible.
Researchers are increasingly investigating whether these systems can identify patterns associated with future disease risk and progression.
A March 2026 study of RETFound Plus, a time- and person-sensitive retinal foundation model, used more than 1.3 million fundus photographs from more than 300,000 participants across multiple visits. The model incorporated temporal information and improved risk prediction for several systemic and ocular diseases compared with an earlier retinal foundation model.
This represents a broader transformation in medical AI. A medical image does not necessarily need to be treated as a static snapshot. With longitudinal information, it can become part of a model of how health changes over time.
Foundation models could therefore contribute to a shift from disease detection toward disease forecasting.
The Promise of Personalised Medicine
Multimodal foundation models may also support personalised medicine because individual patients rarely fit neatly into population-level categories.
Two patients can have the same diagnosis but respond differently to treatment because of differences in genetics, physiology, previous treatments, lifestyle and other factors.
A model capable of integrating multiple dimensions of patient information could potentially identify these differences more effectively than a system focused on a single variable.
This does not mean that foundation models will automatically solve personalised medicine. Personalisation requires high-quality data, reliable validation and careful interpretation. Nevertheless, the architecture is well suited to the problem because it can potentially integrate the multiple factors that contribute to individual variation.
The broader vision is a healthcare system in which AI does not simply ask whether a disease is present but also considers how that disease behaves in a particular individual.
Foundation Models and the Computational Patient
The rise of multimodal foundation models connects directly with the emerging idea of the computational patient.
A computational patient is a dynamic representation of an individual’s health constructed from longitudinal information. Foundation models could provide the underlying intelligence needed to connect different sources of information within that representation.
Clinical notes could describe symptoms and observations. Imaging could represent anatomy. Laboratory results could capture biochemical changes. Wearable signals could provide continuous physiological information. Genomic information could describe biological variation.
A multimodal foundation model could potentially learn relationships across these data streams and create a more comprehensive representation of the patient’s state.
This does not mean that a perfect digital twin of every human being is imminent. Current systems remain limited in their understanding of causality, generalisation and biological complexity. But foundation models may provide some of the computational infrastructure required for increasingly sophisticated patient-level modelling.
The Problem of Data Bias
The power of foundation models comes partly from the scale of their training data. This also creates one of their greatest risks.
Medical datasets are not neutral reflections of humanity. They can contain differences in healthcare access, demographic representation, geographic distribution, clinical documentation and disease prevalence.
If a foundation model learns primarily from data representing certain populations or institutions, its performance may decline when applied elsewhere.
The problem becomes more complicated with multimodal data because each modality may have its own biases. Imaging devices can differ, clinical documentation practices vary and laboratory systems are not always standardised.
A model can therefore appear highly capable during development while performing differently in real-world environments.
The 2026 Nature Biomedical Engineering perspective on biomedical imaging foundation models specifically identifies data scarcity, domain heterogeneity, limited interpretability and insufficient prospective validation as barriers between benchmark performance and clinical value.
Hallucinations and Incorrect Clinical Reasoning
Generative foundation models introduce another challenge: they can produce plausible but incorrect information.
A model may generate a medically convincing explanation that is not supported by the underlying evidence. This is particularly concerning in healthcare because confidence and fluency can be mistaken for accuracy.
Multimodal systems add another layer of complexity. A model may misinterpret an image, over-rely on accompanying text or combine contradictory signals incorrectly.
Recent research therefore increasingly emphasises evaluation of reasoning behaviour rather than simply measuring final-answer accuracy.
Clinical AI needs to be evaluated under realistic conditions, including incomplete information, misleading information, unusual cases and conflicting modalities.
The ability to produce a fluent medical explanation is not equivalent to possessing reliable clinical reasoning.
Interpretability and Clinical Trust
Clinicians need to understand enough about an AI recommendation to determine whether it deserves consideration.
A foundation model that simply produces a diagnosis without useful supporting evidence may be difficult to integrate safely into clinical practice.
Interpretability research is therefore becoming increasingly important. A 2026 Nature Communications review of interpretable graph-based approaches to multimodal biomedical data integration emphasised the importance of explainability for clinical translation and examined methods such as saliency, sensitivity analysis, SHAP and graph masking.
The ideal future system may not expose every internal computational process. Instead, it should provide clinically meaningful evidence about why a conclusion was reached, which data influenced it and where uncertainty remains.
Privacy Becomes More Complex With Multimodal Data
A foundation model trained on medical data creates substantial privacy considerations.
Medical information is sensitive even when stored in a single format. Combining imaging, clinical notes, genomic information, physiological signals and longitudinal records creates an even richer representation of an individual.
The potential consequences of data leakage therefore become more significant.
Healthcare organisations and AI developers must consider secure data storage, controlled access, appropriate consent, governance and mechanisms for monitoring how models are trained and deployed.
Recent reviews of healthcare foundation models consistently identify privacy and security as major barriers to adoption. Research on large language model safety in healthcare has also highlighted vulnerabilities involving both models themselves and their integration into clinical workflows.
The Risk of Treating the Foundation Model as a Medical Oracle
The greatest conceptual mistake may be to assume that a larger model automatically becomes a better doctor.
Foundation models can process enormous amounts of information and demonstrate impressive performance, but medical reasoning is not simply a matter of pattern recognition.
Clinical decisions involve uncertainty, causality, patient preferences, ethical considerations and responsibility.
A foundation model may recognise that two findings frequently occur together without understanding whether one causes the other. It may generate a plausible explanation without possessing a biological model of the underlying disease.
This is why recent research increasingly argues for human-guided clinical AI rather than fully autonomous medical systems. The 2026 perspective on biomedical imaging foundation models explicitly argues that the immediate role of foundation models should be to augment rather than replace clinical expertise.
From One Giant Model to Coordinated Medical AI
The future may not be dominated by a single universal healthcare model.
Medicine contains enormous domain-specific complexity. A pathology model may need different representations from a cardiology system. A radiology model may require different forms of validation from a genomic model.
An alternative future involves coordinated systems in which foundation models provide general capabilities while specialised clinical models, tools and human experts handle particular tasks.
This architecture could combine the flexibility of foundation models with the precision of domain-specific systems.
Recent research on multimodal AI agents in healthcare reflects this direction. A 2026 scoping review found applications across clinical decision support, documentation, health management and medical education, while examining how foundation models, tools and agentic workflows can be combined.
The result could be less like a single artificial doctor and more like an interconnected computational healthcare team.
What Clinical Validation Must Prove
Before multimodal foundation models can become routine clinical infrastructure, they must demonstrate more than impressive benchmark scores.
Researchers need to establish whether models work across different populations, hospitals, devices and clinical environments. They also need to determine whether AI actually improves patient outcomes rather than merely improving an isolated prediction metric.
Prospective evaluation is especially important.
A model can achieve excellent performance on historical data while producing limited benefit when introduced into a real hospital. Clinicians may interact with it differently than expected, workflows may change, or the model may encounter data that differs substantially from its training distribution.
Clinical value must therefore be demonstrated in the environment in which the system will actually operate.
The Future of Multimodal Medical Intelligence
The long-term significance of healthcare foundation models may lie in their ability to connect information that medicine has historically treated separately.
A patient does not experience their radiology image, laboratory results, genome and clinical history as separate datasets. These are different representations of the same biological system.
Multimodal AI attempts to build computational systems that recognise this interconnects.
If successful, future clinical platforms could analyse a patient’s history, interpret new medical images, monitor physiological signals, summarise relevant evidence and identify patterns that deserve professional attention.
Such systems could support earlier diagnosis, personalised treatment, clinical research, drug development and continuous health monitoring.
But the future will depend on disciplined development rather than technological enthusiasm. Better models will not automatically solve problems of bias, privacy, interpretability, causality or accountability.
Conclusion
Foundation models represent a significant evolution in healthcare AI because they move the field beyond narrowly defined algorithms toward flexible systems capable of learning from and connecting multiple forms of medical information.
Text, medical images, pathology, laboratory results, physiological signals, genomic information and longitudinal health records can potentially become part of a common computational framework. Recent 2026 research demonstrates both the promise of this approach and its limitations. Multimodal models can transfer knowledge across tasks and datasets, support clinical reasoning and generate powerful representations, but they can also over-rely on one modality, inherit bias, struggle with generalisation and produce errors that are difficult to interpret.
The most important transformation may therefore not be the creation of an artificial doctor that knows everything. It may be the creation of a computational system that can bring together the different pieces of information that clinicians already use, identify relationships across them and present those relationships in a clinically useful form.
In that sense, multimodal foundation models could become a new layer of healthcare infrastructure.
The future of medical AI may not be defined by machines replacing doctors, but by machines becoming capable of understanding more of the patient’s story at once. When AI can learn from images, words, numbers, signals and longitudinal history together, healthcare moves one step closer to a model in which intelligence is not attached to a single dataset, but distributed across the complete picture of human health.