In a new study published in Radiology, our team successfully leveraged large language models (LLMs) to automatically assess the completeness of clinical histories accompanying imaging orders.
This research spans two of the AIDE Lab’s research foci: algorithm development, and quality assessment, in which we strive to use the latest machine learning techniques to not only maintain high-quality AI algorithms but also enable robust quality assessments.
The Problem: Incomplete Clinical Histories
Radiologists depend on concise, relevant, and complete clinical histories to accurately interpret imaging studies and provide a more useful radiology report. However, the clinical histories accompanying imaging orders are often incomplete, inconsistent, or missing key details. Previous efforts to improve the quality (completeness) of clinical histories have required extensive manual processes that are tedious and expensive. An automated approach could both speed up such assessments and make dedicate quality improvement efforts easier to adopt.
The Methodology
The AIDE Lab leveraged existing LLMs for this task. We adapted both smaller open-source and larger closed-source LLMs, including Mistral-7B, LLaMA-7B, and ChatGPT-4-Turbo. Model training and adaptation included in-context learning (all models) and fine-tuning through QLoRA (open-source models only) using a set of deidentified clinical histories accompanying imaging orders from the adult and pediatric emergency department of Stanford Medical Center.
To assess “completeness”, the models were trained to extract five clinical history elements: “past medical history,” “what,” “when,” “where,” and “clinical concern”.
Key Results and Impact
Why This Matters
Automating the assessment of clinical histories helps address many of the previous challenges associated with dedicated quality improvement efforts to improve the completeness of clinical histories. Our method provides a way to reliably compare and assess the quality of clinical histories at scale. Up until recently, there was no way of utilizing information from unstructured data sources like clinical histories to learn about communication quality. But with the advent of LLMs, this can be done automatically, and reliably at scale.
Along these lines, the study results also highlight the overall potential of AI tools for quality assessments of text, which may be more broadly generalizable and have wider applications across quality assurance and improvement in radiology.
Our next steps include exploring the possibility of not only using the tool to measure and monitor quality of information shared within Stanford Medicine but also as an educational tool for clinical trainees to help demonstrate “complete” vs “incomplete” clinical histories. Additionally, this model could serve as an assessment and monitoring tool for AI-assisted clinical summary tools as they become available to assist in providing clinical histories to radiologists.
To help facilitate similar quality improvement projects, our model and code are fully open-sourced: https://github.com/stanfordaide/clinical-history-eval
Evaluate your own clinical histories through this interactive demo: https://huggingface.co/spaces/stanfordaide/clinic-hist-eval-demo-hf
Read the full paper here: https://pubs.rsna.org/eprint/N43E56EYFWGHSBHPZXIW/full
This work was previously featured in the RSNA 2024 Daily Bulletin: https://dailybulletin.rsna.org/en/2024/tue/tue07