서브메뉴
검색
Towards Scalable and Stable Machine Learning in Clinical Contexts: Addressing Computational Efficiency and Dataset Shift
Towards Scalable and Stable Machine Learning in Clinical Contexts: Addressing Computational Efficiency and Dataset Shift
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105218
- ISBN
- 9798291565995
- DDC
- 004
- 서명/저자
- Towards Scalable and Stable Machine Learning in Clinical Contexts: Addressing Computational Efficiency and Dataset Shift
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 183 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Wiens, Jenna.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약Clinical machine learning (ML) models have the potential to improve patient outcomes, but realizing this potential requires models that are both computationally efficient and robust to dataset shift. In terms of computational efficiency, simpler models that are memory and storage efficient can help increase the probability of adoption and the feasibility of running models at the bedside in health systems, which are often IT-resource constrained. Robustness to dataset shift ensures long-term model reliability since resources for model retraining are often unavailable. Yet, many existing clinical ML models fail on both fronts. This dissertation addresses these challenges by proposing methods that improve models' (i) computational efficiency and (ii) performance under dataset shift.To improve computational efficiency, we explore two domains. First, in genomic sequence classification, we show that standard approaches are inefficient because they rely on large reference databases at inference time. We introduce an ML-based method that removes this dependency, enabling accurate and memory-efficient classification. Second, we reduce unnecessary model complexity in multiple instance learning (MIL) for large-scale medical image classification. While transformers outperform simpler MIL approaches on tasks where relevant image regions are spatially aligned (e.g., detecting cardiac conditions), they are more computationally complex. We find their advantage stems from their inclusion of positional information. Based on this insight, we propose a positional encoding wrapper that boosts the accuracy of standard MIL models, without adding excessive computational overhead, enabling accurate and efficient image classification.To improve robustness to dataset shift, we examine two settings. The first involves developing models in the presence of unstable correlations-transient associations between features and outcomes that appear in training data but fail to generalize over time (e.g., due to changes in clinical practice). We show that standard model selection strategies, which average model performance across random or temporal splits, mask reliance on unstable correlations. In light of this shortcoming, we propose a new model selection approach that yields models with more stable performance over time. The second setting considers predicting the time to a medical event (e.g., spontaneous labor) when there is dataset shift due to changes in an individual's probability of being censored (e.g., undergoing a c-section) over time. This can lead to limited uncensored support during training compared to testing. Standard survival analysis methods often treat censoring times as lower bounds, causing them to overestimate event times for individuals resembling censored training data in this setting. We propose a method that identifies individuals likely censored close to their true event time and uses them as supervision during training, improving predictions for censored-like test cases without compromising accuracy on uncensored ones.Together, these contributions support the development of clinical ML models that are both efficient and robust. These contributions aim to support the development of models that are easier to integrate into clinical workflows and more reliable in real-world healthcare settings.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 키워드
- Machine learning
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359814
■00520260202105218
■006m o d
■007cr#unu||||||||
■020 ▼a9798291565995
■035 ▼a(MiAaPQ)AAI32271784
■035 ▼a(MiAaPQ)umichrackham006398
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aKrishnamoorthy, Meera.
■24510▼aTowards Scalable and Stable Machine Learning in Clinical Contexts: Addressing Computational Efficiency and Dataset Shift
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a183 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Wiens, Jenna.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aClinical machine learning (ML) models have the potential to improve patient outcomes, but realizing this potential requires models that are both computationally efficient and robust to dataset shift. In terms of computational efficiency, simpler models that are memory and storage efficient can help increase the probability of adoption and the feasibility of running models at the bedside in health systems, which are often IT-resource constrained. Robustness to dataset shift ensures long-term model reliability since resources for model retraining are often unavailable. Yet, many existing clinical ML models fail on both fronts. This dissertation addresses these challenges by proposing methods that improve models' (i) computational efficiency and (ii) performance under dataset shift.To improve computational efficiency, we explore two domains. First, in genomic sequence classification, we show that standard approaches are inefficient because they rely on large reference databases at inference time. We introduce an ML-based method that removes this dependency, enabling accurate and memory-efficient classification. Second, we reduce unnecessary model complexity in multiple instance learning (MIL) for large-scale medical image classification. While transformers outperform simpler MIL approaches on tasks where relevant image regions are spatially aligned (e.g., detecting cardiac conditions), they are more computationally complex. We find their advantage stems from their inclusion of positional information. Based on this insight, we propose a positional encoding wrapper that boosts the accuracy of standard MIL models, without adding excessive computational overhead, enabling accurate and efficient image classification.To improve robustness to dataset shift, we examine two settings. The first involves developing models in the presence of unstable correlations-transient associations between features and outcomes that appear in training data but fail to generalize over time (e.g., due to changes in clinical practice). We show that standard model selection strategies, which average model performance across random or temporal splits, mask reliance on unstable correlations. In light of this shortcoming, we propose a new model selection approach that yields models with more stable performance over time. The second setting considers predicting the time to a medical event (e.g., spontaneous labor) when there is dataset shift due to changes in an individual's probability of being censored (e.g., undergoing a c-section) over time. This can lead to limited uncensored support during training compared to testing. Standard survival analysis methods often treat censoring times as lower bounds, causing them to overestimate event times for individuals resembling censored training data in this setting. We propose a method that identifies individuals likely censored close to their true event time and uses them as supervision during training, improving predictions for censored-like test cases without compromising accuracy on uncensored ones.Together, these contributions support the development of clinical ML models that are both efficient and robust. These contributions aim to support the development of models that are easier to integrate into clinical workflows and more reliable in real-world healthcare settings.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aMachine learning
■653 ▼aMultiple instance learning
■653 ▼aComputational efficiency
■690 ▼a0984
■690 ▼a0464
■690 ▼a0800
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359814▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


