Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Computational Semantics for Intelligent Digital Health Applications

  • Akhila Naz Kuppassery Abdulnazar

Studienabschlussarbeit: Dissertation

Abstract

Introduction. The introduction of electronic health records (EHRs) has revolutionized healthcare by digitizing patient information and improving its accessibility and reuse. However, interoperability remains a significant challenge due to variations in terminologies, data formats, and system architecture across institutions, hindering seamless data exchange. This lack of standardization limits the potential of intelligent digital health applications. This dissertation explores solutions to enhance EHR interoperability by applying classical and deep learning-based text classification methods, automating named entity recognition (NER) and medical concept normalization (MCN) using unsupervised techniques, and leveraging transformer models and large language models (LLMs) for mapping clinical narratives to terminologies such as SNOMED CT. Methods. This dissertation is divided into three main studies focused on enhancing EHR interoperability. The first study applies classical machine learning (ML) and deep learning models to classify EHR data related to oxygen supplementation, highlighting the effectiveness of ML methods in organizing large volumes of data, especially for COVID-19 research. The second study automates NER and MCN using unsupervised methods. A pipeline was developed to tokenize clinical narratives, detect entities, and map them to SNOMED CT, employing SapBERT for medical term embeddings and rule-based re-ranking for entity disambiguation. This approach reduces reliance on manually labelled data. The third study combines SapBERT and an LLM to enhance MCN, enabling exact mappings to SNOMED CT and UMLS. LLM-based data cleansing and re-ranking algorithms improved accuracy. Additional investigations include SapBERT-based term clustering, a comparison of contextual vs. non-contextual embeddings, hybrid approaches for mapping smoking status, and the development of Explainable AI (XAI) and visualization tools for patient data integration and navigation. Results. In the first of the three main studies mentioned, the text classification model achieved an F1 score of over 90% in categorizing oxygen supplementation records, demonstrating the effectiveness of the chosen ML approach for the domain- and task-specific challenge. The unsupervised learning methods in the second study showed promising performance in terms of precision and recall, significantly reducing the dependency on manually annotated data for this task in the future. The third study confirmed that BERT-based models outperformed traditional lexicon matching for the task of MCN, with a 91.8% improvement in detection rate, raising the F1 score from 0.297 to 0.568. Furthermore, the application of LLM for data cleaning and reranking enhanced the performance of BERT by 6.8% in F1 score for the task, refining the normalization process and improving alignment with standardized medical terminologies. Discussion. The collective results of this work underscore the critical role of advanced contextual ML methods and the application of LLM in supporting data interoperability within EHR. By improving the accuracy, consistency, and automation of MCN, this research contributes to the development of a standardized representation of health data in relation to international terminologies, particularly SNOMED CT. These advancements enable better utilization of such data in the context of intelligent healthcare applications, which can seamlessly exchange data, enhance clinical workflows, and optimize patient care. This work highlights the need for continuous innovation in health informatics. It provides a foundation for future research to bridge interoperability gaps in EHR systems through the creation of structured and standardized patient profiles.
Datum der Bewilligung2025
OriginalspracheEnglisch
Gradverleihende Hochschule
  • Medizinische Universität Graz
Betreuer/-inMarkus Eduard Kreuzthaler (Betreuer*in) & Stefan Schulz (Betreuer*in)

Dieses zitieren

'