Why Medicine Needs More Than Just AI
With “ChatGPT for Clinicians,” OpenAI has introduced a specialized version of its language model that was developed explicitly for everyday medical practice. It has garnered significant international attention, and the hopes placed in it are well-founded. However, the role of a single large language model (LLM) is systematically overestimated. For the promises in healthcare to truly be fulfilled, we need many AI models working together.
Amid all the excitement surrounding “ChatGPT for Clinicians,” two points are often overlooked:
First, the application is primarily built around general medical information such as guidelines. However, researching this information is not the actual time-consuming task in daily clinical practice. The far greater effort lies in reviewing individual patient records and gaining an overview of the patient’s medical history. This is where the majority of the roughly three hours that doctors spend on documentation each day goes.
Second, “ChatGPT for Clinicians” is essentially based on a single model architecture. A single LLM—as powerful as it may be—cannot fully meet the multifaceted demands of medicine. Anyone who wants to create sustainable added value must think beyond the use of isolated models.
From General Guidelines to Patient-Specific Data
As long as AI relies exclusively on publicly available sources, its clinical utility remains limited. Only when AI systems gain access to individual patient data does the picture change fundamentally. In Europe, it is important to note that the use of U.S. models for patient-specific data involves significant data protection and compliance hurdles. Solutions are needed that fully comply with German data protection regulations.
There is another challenge as well: About 90 percent of patient data in hospitals is unstructured—in the form of PDF documents or free-text entries. To make it usable for AI, it must first be structured. This requires two steps: structuring in the background and presenting it in an understandable format in the foreground. Both must be carried out with the utmost precision—because in medicine, errors can have serious consequences.
Not the best solo LLM—the best AI team
Not every LLM is equally good at reliably standardizing, extracting, and presenting unstructured content in an understandable way. Some models excel at text generation, while others excel at information extraction or classification. It is therefore crucial to combine different AI models in such a way that their strengths complement each other and their weaknesses are offset. This could look like this, for example
- One model extracts medical entities from free-form text.
- Another model maps the information to standardized codes (e.g., ICD or SNOMED).
- A third model uses this information to generate clear, conversational text for the physician.
This type of “model orchestration” can also solve the problem of hallucinations: Multiple LLMs work together, cross-check each other, and evaluate the quality of the results. When supplemented by rule-based systems and evidence-based databases, the risk of misinterpretations can be reduced to nearly zero.
From Theory to Practice: Averbis Medical Summary
Such solutions are closer to real-world practice than most people realize. At the Bosch Health Campus in Stuttgart, Averbis has tested one such application in an AI real-world lab—a controlled environment that enables the development, training, testing, and validation of innovative AI systems in consultation with regulatory authorities. Routine operation in the first departments of the Robert Bosch Hospital is now imminent.
Averbis’ “Medical Summary” covers the entire process: from the structured presentation of the patient’s medical history to the drafting of physician reports. The application extracts relevant information from all patient data available at the hospital, summarizes it clearly, and makes it interactive. Physicians can ask specific questions—for example, about allergies, medications, or the date of the most recent imaging.
The results speak for themselves: Using the Medical Summary, doctors can access relevant medical documentation five times faster and spend up to 50 percent less time drafting medical reports. For a facility like the Robert Bosch Hospital, this equates to approximately 100,000 physician work hours—which are now available for patient care.
The success at the Bosch Health Campus has convinced other institutions: The Medical Summary is now also being put into practical use at the Rheine Clinic. And the hospital information system (HIS) provider Meierhofer has integrated the application into its hospital information system.
Conclusion: Orchestration Instead of a Monolith
The future of AI in healthcare does not lie in the dominance of a single system. The difference will be made by platforms that strategically combine different models—to make the most of their strengths and compensate for their weaknesses. This is the only way to strike a balance between efficiency and reliability. And this is the only way AI will live up to the hopes that many in the medical field currently place in the technology.