서브메뉴
검색
Conditional Guarantees in Model-Free Inference
Conditional Guarantees in Model-Free Inference
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104740
- ISBN
- 9798290651453
- DDC
- 305.8
- 저자명
- Cherian, John.
- 서명/저자
- Conditional Guarantees in Model-Free Inference
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 180 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Candès, Emmanuel.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약The deployment of machine learning in high-stakes settings has raised fundamental questions about the reliability and fairness of black-box models. For example, does a model treat different groups equitably, and can we quantify model uncertainty before taking action on each prediction? While numerous assumption-lean methods appear to address these questions, their guarantees often fall short in practice. For example, model-free guarantees in predictive inference are often marginal, meaning that the assessment of prediction error holds only on-average for a data point drawn exchangeably from the data used to calibrate the method. These guarantees, therefore, do not rule out the existence of substantial individual variation. The research program presented in this dissertation aims at the inherent tension of model-free statistical inference: the generic validity of such methods is appealing, but without a well-specified model, it is challenging to identify guarantees that are also practically useful.1.1 Fairness auditingWhile black-box models may demonstrate impressive accuracy on average, their performance can still vary substantially between subpopulations. For example, an algorithm deployed for recidivism prediction exhibits significantly higher false positive rates for African-American relative to Caucasian parolees [Angwin et al., 2016]. Similar performance disparities have been documented in other highstakes applications such as facial recognition and hiring [Buolamwini and Gebru, 2018, Dastin, 2018].Motivated by this concern, numerous stakeholders have solicited methods, often referred to as "fairness audits," that can discover and quantify such disparities [Brundage et al., 2020, Schaake and Clark, 2022]. Despite substantial prior work in this area [DiCiccio et al., 2020, Morina et al., 2019, Si et al., 2021, Taskesen et al., 2021, Tramer et al., 2017, von Zahn et al., 2023, Xue et al., 2020, Yan and Zhang, 2022], the definition of fairness auditing remains fraught. Fairness auditing is often framed as a single statistical test that rejects in the case of any performance disparity over a limited set of sensitive subpopulations [DiCiccio et al., 2020, Morina et al., 2019, Roy and Mohapatra, 2023, Si et al., 2021, Taskesen et al., 2021, Tramer et al., 2017, Xue et al., 2020]. While follow-up investigation to localize disparities is desired (and often performed), this task raises new challenges. For example, empirical parity across a limited collection of subgroups does not rule out substantial disparities among smaller subgroups [Kearns et al., 2018]. Further, if we consider multiple performance metrics over a rich collection of subpopulations, discovering some disparity between two subgroups is hardly surprising. Unfortunately, existing methods for identifying localized (dis-)parities are accompanied by few statistical guarantees [Schaake and Clark, 2022, von Zahn et al., 2023, Yan and Zhang, 2022].In Chapter 2, we develop a family of statistical methods that rigorously achieve two goals: (1) the "certification" of subpopulations for which the model performs adequately, and (2) the "flagging" of subpopulations that suffer harmful performance disparities. Formally, we approach these two tasks by allowing the auditor to define a "disparity" by comparing some measure of model performance on a subpopulation to a potentially data-dependent target. A certification audit then allows the auditor to identify subpopulations for which this disparity is acceptably low, while a flagging audit discovers subpopulations for which this disparity exceeds some prespecified threshold. Our proposed methods only require access to a so-called "audit trail," i.e., model predictions on a data set held out from training [Brundage et al., 2020], but not white-box access to the model itself.
- 일반주제명
- White people
- 일반주제명
- Statistical inference
- 일반주제명
- Conformity
- 일반주제명
- Large language models
- 일반주제명
- Hilbert space
- 일반주제명
- Certification
- 일반주제명
- Computer engineering
- 키워드
- Black-box models
- 키워드
- Machine learning
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358705
■00520260202104740
■006m o d
■007cr#unu||||||||
■020 ▼a9798290651453
■035 ▼a(MiAaPQ)AAI32149693
■035 ▼a(MiAaPQ)Stanfordmx648rd8955
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a305.8
■1001 ▼aCherian, John.
■24510▼aConditional Guarantees in Model-Free Inference
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a180 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Candès, Emmanuel.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼aThe deployment of machine learning in high-stakes settings has raised fundamental questions about the reliability and fairness of black-box models. For example, does a model treat different groups equitably, and can we quantify model uncertainty before taking action on each prediction? While numerous assumption-lean methods appear to address these questions, their guarantees often fall short in practice. For example, model-free guarantees in predictive inference are often marginal, meaning that the assessment of prediction error holds only on-average for a data point drawn exchangeably from the data used to calibrate the method. These guarantees, therefore, do not rule out the existence of substantial individual variation. The research program presented in this dissertation aims at the inherent tension of model-free statistical inference: the generic validity of such methods is appealing, but without a well-specified model, it is challenging to identify guarantees that are also practically useful.1.1 Fairness auditingWhile black-box models may demonstrate impressive accuracy on average, their performance can still vary substantially between subpopulations. For example, an algorithm deployed for recidivism prediction exhibits significantly higher false positive rates for African-American relative to Caucasian parolees [Angwin et al., 2016]. Similar performance disparities have been documented in other highstakes applications such as facial recognition and hiring [Buolamwini and Gebru, 2018, Dastin, 2018].Motivated by this concern, numerous stakeholders have solicited methods, often referred to as "fairness audits," that can discover and quantify such disparities [Brundage et al., 2020, Schaake and Clark, 2022]. Despite substantial prior work in this area [DiCiccio et al., 2020, Morina et al., 2019, Si et al., 2021, Taskesen et al., 2021, Tramer et al., 2017, von Zahn et al., 2023, Xue et al., 2020, Yan and Zhang, 2022], the definition of fairness auditing remains fraught. Fairness auditing is often framed as a single statistical test that rejects in the case of any performance disparity over a limited set of sensitive subpopulations [DiCiccio et al., 2020, Morina et al., 2019, Roy and Mohapatra, 2023, Si et al., 2021, Taskesen et al., 2021, Tramer et al., 2017, Xue et al., 2020]. While follow-up investigation to localize disparities is desired (and often performed), this task raises new challenges. For example, empirical parity across a limited collection of subgroups does not rule out substantial disparities among smaller subgroups [Kearns et al., 2018]. Further, if we consider multiple performance metrics over a rich collection of subpopulations, discovering some disparity between two subgroups is hardly surprising. Unfortunately, existing methods for identifying localized (dis-)parities are accompanied by few statistical guarantees [Schaake and Clark, 2022, von Zahn et al., 2023, Yan and Zhang, 2022].In Chapter 2, we develop a family of statistical methods that rigorously achieve two goals: (1) the "certification" of subpopulations for which the model performs adequately, and (2) the "flagging" of subpopulations that suffer harmful performance disparities. Formally, we approach these two tasks by allowing the auditor to define a "disparity" by comparing some measure of model performance on a subpopulation to a potentially data-dependent target. A certification audit then allows the auditor to identify subpopulations for which this disparity is acceptably low, while a flagging audit discovers subpopulations for which this disparity exceeds some prespecified threshold. Our proposed methods only require access to a so-called "audit trail," i.e., model predictions on a data set held out from training [Brundage et al., 2020], but not white-box access to the model itself.
■590 ▼aSchool code: 0212.
■650 4▼aWhite people
■650 4▼aStatistical inference
■650 4▼aConformity
■650 4▼aLarge language models
■650 4▼aHilbert space
■650 4▼aCertification
■650 4▼aComputer engineering
■653 ▼aBlack-box models
■653 ▼aMachine learning
■690 ▼a0464
■690 ▼a0800
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358705▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


