본문

서브메뉴

Safety, Robustness, and Interpretability in Machine Learning
Safety, Robustness, and Interpretability in Machine Learning
Safety, Robustness, and Interpretability in Machine Learning

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103551
ISBN  
9798288862991
DDC  
004
저자명  
Pfrommer, Samuel Ian.
서명/저자  
Safety, Robustness, and Interpretability in Machine Learning
발행사항  
[Sl] : University of California, Berkeley, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
157 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
주기사항  
Advisor: Sojoudi, Somayeh.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
초록/해제  
요약Machine learning is poised to have a dramatic impact across many scientific, industrial, and social domains. While current Artificial Intelligence (AI) systems generally involve human supervision, future applications will demand significantly more autonomy. Such a transition will require us to trust the behavior of increasingly large models. This dissertation addresses three critical research areas towards this goal: safety, robustness, and interpretability.We first address safety concerns in Reinforcement Learning (RL) and Imitation Learning (IL). While learned policies have achieved impressive performance, they often exhibit unsafe behavior due to training-time exploration and test-time environmental shifts. We introduce a model predictive control-based safety guide which refines the actions of a base RL policy, conditioned on user-provided constraints. With an appropriate optimization formulation and loss function, we show theoretically that the final base policy is provably safe at optimality. IL suffers from a distinct causal confusion safety concern, where spurious correlations between observations and expert actions can lead to unsafe behavior upon deployment. We leverage tools from Structural Causal Models (SCMs) to identify and mask problematic observations. Whereas previous work requires access to a queryable expert or an expert reward function, our approach uses the typical ability of an experimenter to intervene on the initial state of an episode.The second part of this dissertation concerns robustifying machine learning classifiers against adversarial inputs. Classifiers are a critical component of many AI systems and have been shown to be highly sensitive to small input perturbations. We first extend randomized smoothing beyond traditional isotropic certification by projecting inputs into a data-manifold subspace, resulting in orders-of-magnitude improvements in certified volume. We then revisit the fundamental robustness problem by proposing asymmetric certification. This binary classification setting requires only certified robustness for one class, reflecting the fact that many real-world adversaries are strictly interested in producing false negatives. This more focused problem admits an interesting class of feature-convex architectures, which we leverage to provide efficient, deterministic, and closed-form certified radii.The third part of this dissertation discusses two distinct aspects of interpretability: how Large Language Models (LLMs) decide what to recommend to human users, and how we can build learned models which obey human-interpretable structures. We first analyze conversational search engines, in which we use LLMs to rank consumer products for a user query. Our results show that LLMs vary widely in prioritizing product names, associated website content, and input context position. Finally, we propose a new family of interpretable models in domains where latent embeddings carry mathematical structure: structural transport nets. Via a learned bijection to a carefully-designed mirrored algebra, we produce interpretable latent-space operations which respect the laws of the original input space. We demonstrate that respecting underlying algebraic laws is crucial for learning accurate and self-consistent operations.
일반주제명  
Computer science
일반주제명  
Computer engineering
키워드  
Interpretability
키워드  
Machine learning
키워드  
Robustness
키워드  
Safety
키워드  
Reinforcement Learning
기타저자  
University of California, Berkeley Electrical Engineering & Computer Sciences
기본자료저록  
Dissertations Abstracts International. 87-01B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357725
■00520260202103551
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288862991
■035    ▼a(MiAaPQ)AAI32041867
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aPfrommer,  Samuel  Ian.
■24510▼aSafety,  Robustness,  and  Interpretability  in  Machine  Learning
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a157  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  B.
■500    ▼aAdvisor:  Sojoudi,  Somayeh.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aMachine  learning  is  poised  to  have  a  dramatic  impact  across  many  scientific,  industrial,  and  social  domains.  While  current  Artificial  Intelligence  (AI)  systems  generally  involve  human  supervision,  future  applications  will  demand  significantly  more  autonomy.  Such  a  transition  will  require  us  to  trust  the  behavior  of  increasingly  large  models.  This  dissertation  addresses  three  critical  research  areas  towards  this  goal:  safety,  robustness,  and  interpretability.We  first  address  safety  concerns  in  Reinforcement  Learning  (RL)  and  Imitation  Learning  (IL).  While  learned  policies  have  achieved  impressive  performance,  they  often  exhibit  unsafe  behavior  due  to  training-time  exploration  and  test-time  environmental  shifts.  We  introduce  a  model  predictive  control-based  safety  guide  which  refines  the  actions  of  a  base  RL  policy,  conditioned  on  user-provided  constraints.  With  an  appropriate  optimization  formulation  and  loss  function,  we  show  theoretically  that  the  final  base  policy  is  provably  safe  at  optimality.  IL  suffers  from  a  distinct  causal  confusion  safety  concern,  where  spurious  correlations  between  observations  and  expert  actions  can  lead  to  unsafe  behavior  upon  deployment.  We  leverage  tools  from  Structural  Causal  Models  (SCMs)  to  identify  and  mask  problematic  observations.  Whereas  previous  work  requires  access  to  a  queryable  expert  or  an  expert  reward  function,  our  approach  uses  the  typical  ability  of  an  experimenter  to  intervene  on  the  initial  state  of  an  episode.The  second  part  of  this  dissertation  concerns  robustifying  machine  learning  classifiers  against  adversarial  inputs.  Classifiers  are  a  critical  component  of  many  AI  systems  and  have  been  shown  to  be  highly  sensitive  to  small  input  perturbations.  We  first  extend  randomized  smoothing  beyond  traditional  isotropic  certification  by  projecting  inputs  into  a  data-manifold  subspace,  resulting  in  orders-of-magnitude  improvements  in  certified  volume.  We  then  revisit  the  fundamental  robustness  problem  by  proposing  asymmetric  certification.  This  binary  classification  setting  requires  only  certified  robustness  for  one  class,  reflecting  the  fact  that  many  real-world  adversaries  are  strictly  interested  in  producing  false  negatives.  This  more  focused  problem  admits  an  interesting  class  of  feature-convex  architectures,  which  we  leverage  to  provide  efficient,  deterministic,  and  closed-form  certified  radii.The  third  part  of  this  dissertation  discusses  two  distinct  aspects  of  interpretability:  how  Large  Language  Models  (LLMs)  decide  what  to  recommend  to  human  users,  and  how  we  can  build  learned  models  which  obey  human-interpretable  structures.  We  first  analyze  conversational  search  engines,  in  which  we  use  LLMs  to  rank  consumer  products  for  a  user  query.  Our  results  show  that  LLMs  vary  widely  in  prioritizing  product  names,  associated  website  content,  and  input  context  position.  Finally,  we  propose  a  new  family  of  interpretable  models  in  domains  where  latent  embeddings  carry  mathematical  structure:  structural  transport  nets.  Via  a  learned  bijection  to  a  carefully-designed  mirrored  algebra,  we  produce  interpretable  latent-space  operations  which  respect  the  laws  of  the  original  input  space.  We  demonstrate  that  respecting  underlying  algebraic  laws  is  crucial  for  learning  accurate  and  self-consistent  operations.
■590    ▼aSchool  code:  0028.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■653    ▼aInterpretability
■653    ▼aMachine  learning
■653    ▼aRobustness
■653    ▼aSafety
■653    ▼aReinforcement  Learning
■690    ▼a0800
■690    ▼a0984
■690    ▼a0464
■71020▼aUniversity  of  California,  Berkeley▼bElectrical  Engineering  &  Computer  Sciences.
■7730  ▼tDissertations  Abstracts  International▼g87-01B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357725▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18954 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.