서브메뉴
검색
Toward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support Applications
Toward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support Applications
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103552
- ISBN
- 9798288862687
- DDC
- 621.3
- 저자명
- Hsu, Aliyah.
- 서명/저자
- Toward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support Applications
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 138 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Yu, Bin.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약As deep learning models are increasingly deployed in high-stakes domains like healthcare, understanding their decision-making processes has become essential. While numerous interpretation methods have been proposed in response, many remain unreliable (i.e., being sensitive to input perturbations, or misaligned with real-world reasoning) and struggle to scale effectively. This dissertation advances interpretability in deep learning through a structured investigation across three fronts: post-hoc explanations for black-box models, mechanistic insights into deep learning model internals, and interpretable real-world clinical applications guided by domain expertise. A central emphasis is placed on ensuring the trustworthiness of the developed methods through internal stability analyses and external validation in collaboration with domain experts on real-world tasks. First, we develop two black-box interpretation methods: one distills symbolic rules from concept bottleneck models, and the other uses prompt-based techniques to generate natural language explanations from text modules, both offering interpretable outputs without internal model access. Next, by extending the utility of contextual decomposition (a prior work proposed for local interpretations), we introduce a scalable, mathematically grounded method for mechanistic interpretability in transformers, efficiently identifying task-relevant computational subgraphs at fine granularity. Finally, we explore interpretability in real-world clinical decision support. In collaboration with clinicians, we develop a framework for analyzing fine-tuned transformer feature spaces to inform model suitability for tasks, and design a rule-based LLM system that autonomously applies clinical decision rules from unstructured notes to support emergency care, guided by expert feedback throughout development. These contributions collectively demonstrate how trustworthy interpretability can bridge the gap between model performance and trustworthy deployment in practice.
- 일반주제명
- Computer engineering
- 키워드
- Natural language
- 키워드
- Healthcare
- 기타저자
- University of California, Berkeley Electrical Engineering & Computer Sciences
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357731
■00520260202103552
■006m o d
■007cr#unu||||||||
■020 ▼a9798288862687
■035 ▼a(MiAaPQ)AAI32041910
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aHsu, Aliyah.
■24510▼aToward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support Applications
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a138 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Yu, Bin.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aAs deep learning models are increasingly deployed in high-stakes domains like healthcare, understanding their decision-making processes has become essential. While numerous interpretation methods have been proposed in response, many remain unreliable (i.e., being sensitive to input perturbations, or misaligned with real-world reasoning) and struggle to scale effectively. This dissertation advances interpretability in deep learning through a structured investigation across three fronts: post-hoc explanations for black-box models, mechanistic insights into deep learning model internals, and interpretable real-world clinical applications guided by domain expertise. A central emphasis is placed on ensuring the trustworthiness of the developed methods through internal stability analyses and external validation in collaboration with domain experts on real-world tasks. First, we develop two black-box interpretation methods: one distills symbolic rules from concept bottleneck models, and the other uses prompt-based techniques to generate natural language explanations from text modules, both offering interpretable outputs without internal model access. Next, by extending the utility of contextual decomposition (a prior work proposed for local interpretations), we introduce a scalable, mathematically grounded method for mechanistic interpretability in transformers, efficiently identifying task-relevant computational subgraphs at fine granularity. Finally, we explore interpretability in real-world clinical decision support. In collaboration with clinicians, we develop a framework for analyzing fine-tuned transformer feature spaces to inform model suitability for tasks, and design a rule-based LLM system that autonomously applies clinical decision rules from unstructured notes to support emergency care, guided by expert feedback throughout development. These contributions collectively demonstrate how trustworthy interpretability can bridge the gap between model performance and trustworthy deployment in practice.
■590 ▼aSchool code: 0028.
■650 4▼aComputer engineering
■653 ▼aNatural language
■653 ▼aDeep learning models
■653 ▼aBlack-box interpretation
■653 ▼aHealthcare
■690 ▼a0800
■690 ▼a0464
■690 ▼a0769
■71020▼aUniversity of California, Berkeley▼bElectrical Engineering & Computer Sciences.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357731▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


