서브메뉴
검색
Topics in Selective and Causal Inference
Topics in Selective and Causal Inference
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104743
- ISBN
- 9798290649481
- DDC
- 306
- 저자명
- Chen, Zhaomeng.
- 서명/저자
- Topics in Selective and Causal Inference
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 231 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Candès, Emmanuel.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약1.1 Controlled Variable Selection Based on Summary StatisticsModern scientific studies across fields such as genomics, neuroscience, and economics increasingly involve large-scale datasets with high-dimensional covariates and complex outcomes. A central goal in many of these applications is to identify variables that are meaningfully associated with a response of interest. For example, in genome-wide association studies (GWAS), researchers aim to discover genetic variants linked to complex traits and diseases by analyzing tens of millions of variants across a large number of individuals.These large-scale variable selection tasks present several practical challenges. First, the sheer number of candidate variables increases the risk of false discoveries, making error control essential for ensuring the reproducibility and reliability of scientific findings. Second, to protect privacy and enable broader data sharing, many applications release only summary statistics. For example, GWAS often provide summary data like marginal associations and estimated linkage disequilibrium (LD) patterns instead of individual-level data. This lack of access to individual-level data poses challenges for statistical inference.To address these challenges, the first part of this thesis develops novel methods for controlled variable selection using summary statistics, as presented in Chapters 2 and 3. Specifically, we frame variable selection as a multiple testing problem involving conditional independence hypotheses:H j0: Xj ⊥⊥ Y | X−j , 1 ≤ j ≤ p,Where X−j= (X1, . . . , Xj−1, Xj+1, . . . , Xp) denotes all variables except Xj. Under H j0, the variable Xjprovides no additional information about the response Ybeyond what is already captured by the other covariates. Our goal is to perform powerful testing for H j0, j= 1, . . . , p,while controlling the number of false discoveries. In this thesis, we consider two types of error control: the false discovery rate (FDR),which is the expected proportion of false positives among the selected variables, and the familywise error rate (FWER),which is the probability of making at least one false discovery. FDR control is less stringent and allows for greater power, making it suitable for exploratory analyses. In contrast, FWER control is more conservative and better suited for high-stakes applications where any false positive could be costly. In Chapter 2, we introduce novel GhostKnockoffmethods based on penalized regression that control the FDR in the aforementioned conditional independence testing problem using only summary statistics. In Chapter 3, we introduce a new filter for variable selection using summary statistics with rigorous FWER control. Together, these tools enable statistically principled variable selection in large-scale studies where only summary-level data are available.1.2 Localized Feature Selection and Causal InferenceWhile traditional statistical analysis often focuses on inference or exploratory insights at the population level, there is growing interest in drawing conclusions at a more localized or individual level to enable higher-resolution understanding and to better support personalized decision-making in practice. This has important implications for real-world applications such as personalized medicine, business analytics, and individualized education. For instance, in precision medicine, researchers often aim to identify genetic factors associated with disease while accounting for patient heterogeneity-such as differences in age or gender-to develop targeted treatments.
- 일반주제명
- Families & family life
- 일반주제명
- Feature selection
- 일반주제명
- Statistics
- 일반주제명
- Alzheimer's disease
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358725
■00520260202104743
■006m o d
■007cr#unu||||||||
■020 ▼a9798290649481
■035 ▼a(MiAaPQ)AAI32149722
■035 ▼a(MiAaPQ)Stanfordsq475hw3784
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a306
■1001 ▼aChen, Zhaomeng.
■24510▼aTopics in Selective and Causal Inference
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a231 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Candès, Emmanuel.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼a1.1 Controlled Variable Selection Based on Summary StatisticsModern scientific studies across fields such as genomics, neuroscience, and economics increasingly involve large-scale datasets with high-dimensional covariates and complex outcomes. A central goal in many of these applications is to identify variables that are meaningfully associated with a response of interest. For example, in genome-wide association studies (GWAS), researchers aim to discover genetic variants linked to complex traits and diseases by analyzing tens of millions of variants across a large number of individuals.These large-scale variable selection tasks present several practical challenges. First, the sheer number of candidate variables increases the risk of false discoveries, making error control essential for ensuring the reproducibility and reliability of scientific findings. Second, to protect privacy and enable broader data sharing, many applications release only summary statistics. For example, GWAS often provide summary data like marginal associations and estimated linkage disequilibrium (LD) patterns instead of individual-level data. This lack of access to individual-level data poses challenges for statistical inference.To address these challenges, the first part of this thesis develops novel methods for controlled variable selection using summary statistics, as presented in Chapters 2 and 3. Specifically, we frame variable selection as a multiple testing problem involving conditional independence hypotheses:H j0: Xj ⊥⊥ Y | X−j , 1 ≤ j ≤ p,Where X−j= (X1, . . . , Xj−1, Xj+1, . . . , Xp) denotes all variables except Xj. Under H j0, the variable Xjprovides no additional information about the response Ybeyond what is already captured by the other covariates. Our goal is to perform powerful testing for H j0, j= 1, . . . , p,while controlling the number of false discoveries. In this thesis, we consider two types of error control: the false discovery rate (FDR),which is the expected proportion of false positives among the selected variables, and the familywise error rate (FWER),which is the probability of making at least one false discovery. FDR control is less stringent and allows for greater power, making it suitable for exploratory analyses. In contrast, FWER control is more conservative and better suited for high-stakes applications where any false positive could be costly. In Chapter 2, we introduce novel GhostKnockoffmethods based on penalized regression that control the FDR in the aforementioned conditional independence testing problem using only summary statistics. In Chapter 3, we introduce a new filter for variable selection using summary statistics with rigorous FWER control. Together, these tools enable statistically principled variable selection in large-scale studies where only summary-level data are available.1.2 Localized Feature Selection and Causal InferenceWhile traditional statistical analysis often focuses on inference or exploratory insights at the population level, there is growing interest in drawing conclusions at a more localized or individual level to enable higher-resolution understanding and to better support personalized decision-making in practice. This has important implications for real-world applications such as personalized medicine, business analytics, and individualized education. For instance, in precision medicine, researchers often aim to identify genetic factors associated with disease while accounting for patient heterogeneity-such as differences in age or gender-to develop targeted treatments.
■590 ▼aSchool code: 0212.
■650 4▼aFamilies & family life
■650 4▼aFeature selection
■650 4▼aStatistics
■650 4▼aAlzheimer's disease
■690 ▼a0463
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358725▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


