본문

서브메뉴

Natural Language Explanations of Dataset Patterns
Natural Language Explanations of Dataset Patterns
Natural Language Explanations of Dataset Patterns

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103351
ISBN  
9798288861659
DDC  
004
저자명  
Zhong, Ruiqi.
서명/저자  
Natural Language Explanations of Dataset Patterns
발행사항  
[Sl] : University of California, Berkeley, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
106 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: A.
주기사항  
Advisor: Steinhardt, Jacob.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
초록/해제  
요약Explaining patterns in large datasets is essential for empirical science, engineering, and business. For example, by analyzing a dataset of symptom descriptions, a doctor may discover that "tingling in the thumb" is a good explanatory variable for disease X. However, existing methods (e.g. regression) are primarily designed to analyze real-valued datasets and explain patterns in mathematical formulas (e.g. F=kx + b).This thesis proposes metrics and methods for discovering and explaining dataset patterns in structured modalities (text/images) using natural language strings such as "tingling in the thumb". We evaluate the explanations based on the predictive power they give to humans, which differs from common metrics based on human ratings or similarity to human demonstrations. We then generate dataset explanations by optimizing them against our evaluation metric, with the help of language models. Concretely, we sample candidate explanations from language models and select the highest-scoring one under our evaluation.Based on these principles, we build a general framework, "statistical models with natural language parameters", which allows us to explain distributional differences, clusters, and time-series in real-world datasets with structured modalities. Additionally, our metric can evaluate explanations of model decisions by treating them as explanations of datasets, which consist of the model's input-output behavior. Using this approach, we show that language models are still far from explaining themselves as of 2024.Our contribution paves the way for helping humans understand complex datasets and systems, thereby accelerating scientific discovery and advancing explainable AI systems.
일반주제명  
Computer science
일반주제명  
Engineering
일반주제명  
Linguistics
일반주제명  
Information technology
키워드  
Data mining
키워드  
Explainability
키워드  
Language models
키워드  
Machine learning
키워드  
Natural language processing
기타저자  
University of California, Berkeley Computer Science
기본자료저록  
Dissertations Abstracts International. 87-01A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357363
■00520260202103351
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288861659
■035    ▼a(MiAaPQ)AAI31995092
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aZhong,  Ruiqi.
■24510▼aNatural  Language  Explanations  of  Dataset  Patterns
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a106  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  A.
■500    ▼aAdvisor:  Steinhardt,  Jacob.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aExplaining  patterns  in  large  datasets  is  essential  for  empirical  science,  engineering,  and  business.  For  example,  by  analyzing  a  dataset  of  symptom  descriptions,  a  doctor  may  discover  that  "tingling  in  the  thumb"  is  a  good  explanatory  variable  for  disease  X.  However,  existing  methods  (e.g.  regression)  are  primarily  designed  to  analyze  real-valued  datasets  and  explain  patterns  in  mathematical  formulas  (e.g.  F=kx  +  b).This  thesis  proposes  metrics  and  methods  for  discovering  and  explaining  dataset  patterns  in  structured  modalities  (text/images)  using  natural  language  strings  such  as  "tingling  in  the  thumb".  We  evaluate  the  explanations  based  on  the  predictive  power  they  give  to  humans,  which  differs  from  common  metrics  based  on  human  ratings  or  similarity  to  human  demonstrations.  We  then  generate  dataset  explanations  by  optimizing  them  against  our  evaluation  metric,  with  the  help  of  language  models.  Concretely,  we  sample  candidate  explanations  from  language  models  and  select  the  highest-scoring  one  under  our  evaluation.Based  on  these  principles,  we  build  a  general  framework,  "statistical  models  with  natural  language  parameters",  which  allows  us  to  explain  distributional  differences,  clusters,  and  time-series  in  real-world  datasets  with  structured  modalities.  Additionally,  our  metric  can  evaluate  explanations  of  model  decisions  by  treating  them  as  explanations  of  datasets,  which  consist  of  the  model's  input-output  behavior.  Using  this  approach,  we  show  that  language  models  are  still  far  from  explaining  themselves  as  of  2024.Our  contribution  paves  the  way  for  helping  humans  understand  complex  datasets  and  systems,  thereby  accelerating  scientific  discovery  and  advancing  explainable  AI  systems.
■590    ▼aSchool  code:  0028.
■650  4▼aComputer  science
■650  4▼aEngineering
■650  4▼aLinguistics
■650  4▼aInformation  technology
■653    ▼aData  mining
■653    ▼aExplainability
■653    ▼aLanguage  models
■653    ▼aMachine  learning
■653    ▼aNatural  language  processing
■690    ▼a0984
■690    ▼a0489
■690    ▼a0800
■690    ▼a0537
■690    ▼a0290
■71020▼aUniversity  of  California,  Berkeley▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g87-01A.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357363▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17358 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.