서브메뉴
검색
Advances in Multiple Testing and Variable Selection
Advances in Multiple Testing and Variable Selection
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103535
- ISBN
- 9798288866197
- DDC
- 310
- 저자명
- Luo, Yixiang.
- 서명/저자
- Advances in Multiple Testing and Variable Selection
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 151 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Fithian, William;Evans, Steven.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약In contemporary scientific research, variable selection from numerous candidates represents a fundamental challenge with diverse objectives: building predictive models that balance selection accuracy and predictive power, testing multiple hypotheses with controlled error rates, and quantifying the likelihood of each variable being a true signal. While variable selection can be formalized within multiple hypothesis testing frameworks, classical p-value-based methods often prove inadequate when complex modeling procedures are involved. This dissertation addresses the issues through three novel methodological contributions.Chapter 1-based on Luo et al. (2024)-introduces a conservative estimator for the false discovery rate (FDR) applicable to any variable selection procedure in common statistical modeling settings. Our estimator complements cross-validation by elucidating the trade-off between prediction error and variable selection accuracy as a function of model complexity. We prove that our estimator maintains conservative bias in finite samples under standard assumptions and provide a bootstrap methodology for standard error assessment.The knockoff filter of Barber and Candes (2015) provides a powerful framework for multiple testing with FDR control by leveraging supervised learning models, yet suffers from critical limitations in specific scenarios. Chapter 2-based on Luo et al. (2022)-develops the calibrated knockoff procedure, which uniformly improves the power of any knockoff procedure while preserving FDR control. Our theoretical and empirical analyses demonstrate particularly significant improvements in two scenarios where knockoff methods can be nearly powerless: when rejection sets are small, and when design matrix structures inhibit the construction of effective knockoff variables.Chapter 3 is as-yet unpublished work and presents a novel estimator for the frequentist local false discovery rate (lfdr), which addresses the question of how likely each variable is to represent noise rather than signal. While empirical Bayes methods effectively address large-scale testing problems, they rely on correct Bayesian model specifications. Our estimator, constructed within a purely frequentist framework, addresses this limitation. We establish that in asymptotic regimes characteristic of large-scale testing, our estimator converges to an upper bound of the true frequentist lfdr, enabling validation of empirical Bayes inference and conservative recalibration when necessary.
- 일반주제명
- Statistics
- 일반주제명
- Applied mathematics
- 키워드
- Knockoff
- 키워드
- Multiple testing
- 기타저자
- University of California, Berkeley Mathematics
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357605
■00520260202103535
■006m o d
■007cr#unu||||||||
■020 ▼a9798288866197
■035 ▼a(MiAaPQ)AAI32040414
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aLuo, Yixiang.
■24510▼aAdvances in Multiple Testing and Variable Selection
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a151 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Fithian, William;Evans, Steven.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aIn contemporary scientific research, variable selection from numerous candidates represents a fundamental challenge with diverse objectives: building predictive models that balance selection accuracy and predictive power, testing multiple hypotheses with controlled error rates, and quantifying the likelihood of each variable being a true signal. While variable selection can be formalized within multiple hypothesis testing frameworks, classical p-value-based methods often prove inadequate when complex modeling procedures are involved. This dissertation addresses the issues through three novel methodological contributions.Chapter 1-based on Luo et al. (2024)-introduces a conservative estimator for the false discovery rate (FDR) applicable to any variable selection procedure in common statistical modeling settings. Our estimator complements cross-validation by elucidating the trade-off between prediction error and variable selection accuracy as a function of model complexity. We prove that our estimator maintains conservative bias in finite samples under standard assumptions and provide a bootstrap methodology for standard error assessment.The knockoff filter of Barber and Candes (2015) provides a powerful framework for multiple testing with FDR control by leveraging supervised learning models, yet suffers from critical limitations in specific scenarios. Chapter 2-based on Luo et al. (2022)-develops the calibrated knockoff procedure, which uniformly improves the power of any knockoff procedure while preserving FDR control. Our theoretical and empirical analyses demonstrate particularly significant improvements in two scenarios where knockoff methods can be nearly powerless: when rejection sets are small, and when design matrix structures inhibit the construction of effective knockoff variables.Chapter 3 is as-yet unpublished work and presents a novel estimator for the frequentist local false discovery rate (lfdr), which addresses the question of how likely each variable is to represent noise rather than signal. While empirical Bayes methods effectively address large-scale testing problems, they rely on correct Bayesian model specifications. Our estimator, constructed within a purely frequentist framework, addresses this limitation. We establish that in asymptotic regimes characteristic of large-scale testing, our estimator converges to an upper bound of the true frequentist lfdr, enabling validation of empirical Bayes inference and conservative recalibration when necessary.
■590 ▼aSchool code: 0028.
■650 4▼aStatistics
■650 4▼aApplied mathematics
■653 ▼aFalse discovery rate
■653 ▼aKnockoff
■653 ▼aLocal false discovery rate
■653 ▼aMultiple testing
■653 ▼aVariable selection
■690 ▼a0463
■690 ▼a0364
■71020▼aUniversity of California, Berkeley▼bMathematics.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357605▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


