서브메뉴
검색
Natural Language Explanations of Dataset Patterns
Natural Language Explanations of Dataset Patterns
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103351
- ISBN
- 9798288861659
- DDC
- 004
- 저자명
- Zhong, Ruiqi.
- 서명/저자
- Natural Language Explanations of Dataset Patterns
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 106 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: A.
- 주기사항
- Advisor: Steinhardt, Jacob.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Explaining patterns in large datasets is essential for empirical science, engineering, and business. For example, by analyzing a dataset of symptom descriptions, a doctor may discover that "tingling in the thumb" is a good explanatory variable for disease X. However, existing methods (e.g. regression) are primarily designed to analyze real-valued datasets and explain patterns in mathematical formulas (e.g. F=kx + b).This thesis proposes metrics and methods for discovering and explaining dataset patterns in structured modalities (text/images) using natural language strings such as "tingling in the thumb". We evaluate the explanations based on the predictive power they give to humans, which differs from common metrics based on human ratings or similarity to human demonstrations. We then generate dataset explanations by optimizing them against our evaluation metric, with the help of language models. Concretely, we sample candidate explanations from language models and select the highest-scoring one under our evaluation.Based on these principles, we build a general framework, "statistical models with natural language parameters", which allows us to explain distributional differences, clusters, and time-series in real-world datasets with structured modalities. Additionally, our metric can evaluate explanations of model decisions by treating them as explanations of datasets, which consist of the model's input-output behavior. Using this approach, we show that language models are still far from explaining themselves as of 2024.Our contribution paves the way for helping humans understand complex datasets and systems, thereby accelerating scientific discovery and advancing explainable AI systems.
- 일반주제명
- Computer science
- 일반주제명
- Engineering
- 일반주제명
- Linguistics
- 일반주제명
- Information technology
- 키워드
- Data mining
- 키워드
- Explainability
- 키워드
- Language models
- 키워드
- Machine learning
- 기타저자
- University of California, Berkeley Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-01A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357363
■00520260202103351
■006m o d
■007cr#unu||||||||
■020 ▼a9798288861659
■035 ▼a(MiAaPQ)AAI31995092
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aZhong, Ruiqi.
■24510▼aNatural Language Explanations of Dataset Patterns
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a106 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: A.
■500 ▼aAdvisor: Steinhardt, Jacob.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aExplaining patterns in large datasets is essential for empirical science, engineering, and business. For example, by analyzing a dataset of symptom descriptions, a doctor may discover that "tingling in the thumb" is a good explanatory variable for disease X. However, existing methods (e.g. regression) are primarily designed to analyze real-valued datasets and explain patterns in mathematical formulas (e.g. F=kx + b).This thesis proposes metrics and methods for discovering and explaining dataset patterns in structured modalities (text/images) using natural language strings such as "tingling in the thumb". We evaluate the explanations based on the predictive power they give to humans, which differs from common metrics based on human ratings or similarity to human demonstrations. We then generate dataset explanations by optimizing them against our evaluation metric, with the help of language models. Concretely, we sample candidate explanations from language models and select the highest-scoring one under our evaluation.Based on these principles, we build a general framework, "statistical models with natural language parameters", which allows us to explain distributional differences, clusters, and time-series in real-world datasets with structured modalities. Additionally, our metric can evaluate explanations of model decisions by treating them as explanations of datasets, which consist of the model's input-output behavior. Using this approach, we show that language models are still far from explaining themselves as of 2024.Our contribution paves the way for helping humans understand complex datasets and systems, thereby accelerating scientific discovery and advancing explainable AI systems.
■590 ▼aSchool code: 0028.
■650 4▼aComputer science
■650 4▼aEngineering
■650 4▼aLinguistics
■650 4▼aInformation technology
■653 ▼aData mining
■653 ▼aExplainability
■653 ▼aLanguage models
■653 ▼aMachine learning
■653 ▼aNatural language processing
■690 ▼a0984
■690 ▼a0489
■690 ▼a0800
■690 ▼a0537
■690 ▼a0290
■71020▼aUniversity of California, Berkeley▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g87-01A.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357363▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


