본문

서브메뉴

Toward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support Applications
Toward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support A...
Toward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support Applications

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103552
ISBN  
9798288862687
DDC  
621.3
저자명  
Hsu, Aliyah.
서명/저자  
Toward Trustworthy Language Models: Interpretation Methods and Clinical Decision Support Applications
발행사항  
[Sl] : University of California, Berkeley, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
138 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
주기사항  
Advisor: Yu, Bin.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
초록/해제  
요약As deep learning models are increasingly deployed in high-stakes domains like healthcare, understanding their decision-making processes has become essential. While numerous interpretation methods have been proposed in response, many remain unreliable (i.e., being sensitive to input perturbations, or misaligned with real-world reasoning) and struggle to scale effectively. This dissertation advances interpretability in deep learning through a structured investigation across three fronts: post-hoc explanations for black-box models, mechanistic insights into deep learning model internals, and interpretable real-world clinical applications guided by domain expertise. A central emphasis is placed on ensuring the trustworthiness of the developed methods through internal stability analyses and external validation in collaboration with domain experts on real-world tasks. First, we develop two black-box interpretation methods: one distills symbolic rules from concept bottleneck models, and the other uses prompt-based techniques to generate natural language explanations from text modules, both offering interpretable outputs without internal model access. Next, by extending the utility of contextual decomposition (a prior work proposed for local interpretations), we introduce a scalable, mathematically grounded method for mechanistic interpretability in transformers, efficiently identifying task-relevant computational subgraphs at fine granularity. Finally, we explore interpretability in real-world clinical decision support. In collaboration with clinicians, we develop a framework for analyzing fine-tuned transformer feature spaces to inform model suitability for tasks, and design a rule-based LLM system that autonomously applies clinical decision rules from unstructured notes to support emergency care, guided by expert feedback throughout development. These contributions collectively demonstrate how trustworthy interpretability can bridge the gap between model performance and trustworthy deployment in practice.
일반주제명  
Computer engineering
키워드  
Natural language
키워드  
Deep learning models
키워드  
Black-box interpretation
키워드  
Healthcare
기타저자  
University of California, Berkeley Electrical Engineering & Computer Sciences
기본자료저록  
Dissertations Abstracts International. 87-01B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357731
■00520260202103552
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288862687
■035    ▼a(MiAaPQ)AAI32041910
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aHsu,  Aliyah.
■24510▼aToward  Trustworthy  Language  Models:  Interpretation  Methods  and  Clinical  Decision  Support  Applications
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a138  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  B.
■500    ▼aAdvisor:  Yu,  Bin.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aAs  deep  learning  models  are  increasingly  deployed  in  high-stakes  domains  like  healthcare,  understanding  their  decision-making  processes  has  become  essential.  While  numerous  interpretation  methods  have  been  proposed  in  response,  many  remain  unreliable  (i.e.,  being  sensitive  to  input  perturbations,  or  misaligned  with  real-world  reasoning)  and  struggle  to  scale  effectively.  This  dissertation  advances  interpretability  in  deep  learning  through  a  structured  investigation  across  three  fronts:  post-hoc  explanations  for  black-box  models,  mechanistic  insights  into  deep  learning  model  internals,  and  interpretable  real-world  clinical  applications  guided  by  domain  expertise.  A  central  emphasis  is  placed  on  ensuring  the  trustworthiness  of  the  developed  methods  through  internal  stability  analyses  and  external  validation  in  collaboration  with  domain  experts  on  real-world  tasks.  First,  we  develop  two  black-box  interpretation  methods:  one  distills  symbolic  rules  from  concept  bottleneck  models,  and  the  other  uses  prompt-based  techniques  to  generate  natural  language  explanations  from  text  modules,  both  offering  interpretable  outputs  without  internal  model  access.  Next,  by  extending  the  utility  of  contextual  decomposition  (a  prior  work  proposed  for  local  interpretations),  we  introduce  a  scalable,  mathematically  grounded  method  for  mechanistic  interpretability  in  transformers,  efficiently  identifying  task-relevant  computational  subgraphs  at  fine  granularity.  Finally,  we  explore  interpretability  in  real-world  clinical  decision  support.  In  collaboration  with  clinicians,  we  develop  a  framework  for  analyzing  fine-tuned  transformer  feature  spaces  to  inform  model  suitability  for  tasks,  and  design  a  rule-based  LLM  system  that  autonomously  applies  clinical  decision  rules  from  unstructured  notes  to  support  emergency  care,  guided  by  expert  feedback  throughout  development.  These  contributions  collectively  demonstrate  how  trustworthy  interpretability  can  bridge  the  gap  between  model  performance  and  trustworthy  deployment  in  practice.
■590    ▼aSchool  code:  0028.
■650  4▼aComputer  engineering
■653    ▼aNatural  language
■653    ▼aDeep  learning  models
■653    ▼aBlack-box  interpretation
■653    ▼aHealthcare
■690    ▼a0800
■690    ▼a0464
■690    ▼a0769
■71020▼aUniversity  of  California,  Berkeley▼bElectrical  Engineering  &  Computer  Sciences.
■7730  ▼tDissertations  Abstracts  International▼g87-01B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357731▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF19287 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.