
In everyday clinical work, patient care is recorded in many places, including progress notes, imaging reports, medication records, and discharge summaries. Because this information is often incomplete or written in different ways, it can be difficult to know whether care followed recommended clinical guidelines. Medical process conformance checking is a method that compares what happened to what should normally happen according to professional guidance.
A recent study by Leonardi and colleagues, titled “Orchestrating Large Language Models to Support Medical Process Conformance Checking,” examined how large language models, such as ChatGPT, could support this process. Large language models are computer systems that can understand and produce human language. In this study, they were used to read clinical documents and guidelines, identify important information, and compare patient care with recommended practice (Leonardi et al., 2026).
The researchers used several language models rather than asking one system to complete the entire task. Each model had a specific role. One helped identify clinical events, another examined guideline recommendations, and another converted these recommendations into computer-readable rules. This step-by-step approach can make the process easier to check and correct.
The study focused on stroke care. The system examined discharge letters and created a simple timeline of each patient’s care. It then compared these timelines with selected recommendations for stroke treatment. The results suggested that most patient records followed the expected care pathway. This shows that language models may help hospitals review large numbers of clinical records more efficiently.
However, the results must be interpreted carefully. A medical record may not include everything that happened during treatment. A missing action in the documentation does not always mean that the action was not performed. In addition, patients may require different treatments because of their age, medical history, risks, or personal circumstances.
For therapists and other healthcare professionals, this technology should be viewed as a support tool, not as a replacement for clinical judgment. It may help identify repeated delays, missing documentation, or differences between departments. These findings can then guide team discussions, quality improvement, and better communication between professionals.
One important advantage of this approach is that it can make the reasoning process more visible. Users may be able to see the original clinical text, the guideline recommendation, and the rule used to evaluate care. This is safer than relying on a single unexplained answer from an artificial intelligence system.
There are also important ethical concerns. Patient information must be protected, and healthcare organizations must carefully control how records are used. Language models can misunderstand information or produce biased results, especially when documentation is incomplete. Clinicians and institutions must remain responsible for decisions influenced by these systems.
In conclusion, large language models may help connect clinical documentation with medical guidelines and support safer, more organized healthcare. Their role should remain supportive and transparent. Future research should examine how accurate these systems are in different hospitals and specialties, and whether they can genuinely improve clinical practice without reducing care to simple numbers.
References
Leonardi, G., Montani, S., Striani, M., Canessa, A., & Ferrandi, D. (2026). Orchestrating large language models to support medical process conformance checking. Neuroscience Informatics, 6(4), Article 100294. https://doi.org/10.1016/j.neuri.2026.100294
