서브메뉴
검색
Detecting and Measuring Important Variables: Novel Methods With Statistical Guarantees
Detecting and Measuring Important Variables: Novel Methods With Statistical Guarantees
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104741
- ISBN
- 9798290651460
- DDC
- 519
- 저자명
- Gablenz, Paula.
- 서명/저자
- Detecting and Measuring Important Variables: Novel Methods With Statistical Guarantees
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 145 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Sabatti, Chiara.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약1.1 Background and motivationUnderstanding which explanatory variables, out of potentially millions, are important is a fundamental challenge in statistics. The eventual purpose is to help shed light on the mechanisms affecting the outcome of interest. The field of multiple testing addresses how to discover as many relevant variables as possible, while at the same time ensuring that the findings are replicable. In numerous applications the associated hypotheses are structured, but many conventional approaches do not take this structure into account. A compelling example is offered by genome-wide association studies (GWAS). The goal of genome-wide association studies is to identify which genetic variants influence a certain disease. The genetic variants have spatial structure and can be grouped at different levels of resolution. In addition, the outcomes of interest might be structured, such as the tree structure of the international classification of diseases. Finding the most precise rejections, ensuring consistency between the rejections at different levels of resolution, or generally filtering the rejection set thus represent salient problems. Moreover, there might be individual heterogeneity, where some genetic variants are only relevant for a certain subset of individuals. A challenge in multiple testing are dependencies between the tested hypotheses, which could invalidate the inferential procedure. In this dissertation, we develop new methodologies for (structured) hypotheses testing that are valid under any dependency structures. In addition, we present a novel application to reliably detect gene-environment interactions. Finally, we explore ideas on quantifying how important a variable is for the uncertainty and prediction of a particular model, as opposed to only testing whether it is important.1.2 Structure of the thesisChapter 2 addresses the specific problem of finding the most precise rejections possible, while ensuring error control. In particular, we consider problems where many, somewhat redundant, hypotheses are tested at different levels of resolution. We design a novel multiple comparison procedure that allows for an adaptive choice of resolution with false discovery rate (FDR) control, leveraging e-values.In chapter 3 we develop several additional approaches for structured hypotheses testing using e-values. Specifically, we consider the coordination of rejections across multiple resolutions to avoid conflicting rejections, filtering the rejection set more generally and testing partial conjunction hypotheses. We introduce e-value counterparts of p-value based procedures. Our methods provide error control under any dependency structures.Chapter 4 considers local conditional hypotheses which express how the relation between explanatory variables and outcomes change across different environments, described by covariates. We describe a practical implementation to GWAS-scale data and demonstrate its effectiveness through numerical experiments on the UK Biobank.Chapter 5 introduces a set of novel ideas that contribute to a deeper understanding of measuring variable importance. Building upon established frameworks for variable importance, we present an alternative perspective that offers new directions for understanding predictive uncertainty through the lens of conformal prediction intervals.
- 일반주제명
- Linear programming
- 일반주제명
- Genomes
- 일반주제명
- Biobanks
- 일반주제명
- Power
- 일반주제명
- Statistics
- 일반주제명
- Biostatistics
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358714
■00520260202104741
■006m o d
■007cr#unu||||||||
■020 ▼a9798290651460
■035 ▼a(MiAaPQ)AAI32149704
■035 ▼a(MiAaPQ)Stanfordpm479hw3823
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a519
■1001 ▼aGablenz, Paula.
■24510▼aDetecting and Measuring Important Variables: Novel Methods With Statistical Guarantees
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a145 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Sabatti, Chiara.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼a1.1 Background and motivationUnderstanding which explanatory variables, out of potentially millions, are important is a fundamental challenge in statistics. The eventual purpose is to help shed light on the mechanisms affecting the outcome of interest. The field of multiple testing addresses how to discover as many relevant variables as possible, while at the same time ensuring that the findings are replicable. In numerous applications the associated hypotheses are structured, but many conventional approaches do not take this structure into account. A compelling example is offered by genome-wide association studies (GWAS). The goal of genome-wide association studies is to identify which genetic variants influence a certain disease. The genetic variants have spatial structure and can be grouped at different levels of resolution. In addition, the outcomes of interest might be structured, such as the tree structure of the international classification of diseases. Finding the most precise rejections, ensuring consistency between the rejections at different levels of resolution, or generally filtering the rejection set thus represent salient problems. Moreover, there might be individual heterogeneity, where some genetic variants are only relevant for a certain subset of individuals. A challenge in multiple testing are dependencies between the tested hypotheses, which could invalidate the inferential procedure. In this dissertation, we develop new methodologies for (structured) hypotheses testing that are valid under any dependency structures. In addition, we present a novel application to reliably detect gene-environment interactions. Finally, we explore ideas on quantifying how important a variable is for the uncertainty and prediction of a particular model, as opposed to only testing whether it is important.1.2 Structure of the thesisChapter 2 addresses the specific problem of finding the most precise rejections possible, while ensuring error control. In particular, we consider problems where many, somewhat redundant, hypotheses are tested at different levels of resolution. We design a novel multiple comparison procedure that allows for an adaptive choice of resolution with false discovery rate (FDR) control, leveraging e-values.In chapter 3 we develop several additional approaches for structured hypotheses testing using e-values. Specifically, we consider the coordination of rejections across multiple resolutions to avoid conflicting rejections, filtering the rejection set more generally and testing partial conjunction hypotheses. We introduce e-value counterparts of p-value based procedures. Our methods provide error control under any dependency structures.Chapter 4 considers local conditional hypotheses which express how the relation between explanatory variables and outcomes change across different environments, described by covariates. We describe a practical implementation to GWAS-scale data and demonstrate its effectiveness through numerical experiments on the UK Biobank.Chapter 5 introduces a set of novel ideas that contribute to a deeper understanding of measuring variable importance. Building upon established frameworks for variable importance, we present an alternative perspective that offers new directions for understanding predictive uncertainty through the lens of conformal prediction intervals.
■590 ▼aSchool code: 0212.
■650 4▼aLinear programming
■650 4▼aGenomes
■650 4▼aBiobanks
■650 4▼aPower
■650 4▼aStatistics
■650 4▼aBiostatistics
■653 ▼aGenome-wide association studies
■653 ▼aFalse discovery rate
■690 ▼a0463
■690 ▼a0308
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358714▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


