본문

서브메뉴

Topics in Selective and Causal Inference
Topics in Selective and Causal Inference
Topics in Selective and Causal Inference

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104743
ISBN  
9798290649481
DDC  
306
저자명  
Chen, Zhaomeng.
서명/저자  
Topics in Selective and Causal Inference
발행사항  
[Sl] : Stanford University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
231 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
주기사항  
Advisor: Candès, Emmanuel.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2025.
초록/해제  
요약1.1 Controlled Variable Selection Based on Summary StatisticsModern scientific studies across fields such as genomics, neuroscience, and economics increasingly involve large-scale datasets with high-dimensional covariates and complex outcomes. A central goal in many of these applications is to identify variables that are meaningfully associated with a response of interest. For example, in genome-wide association studies (GWAS), researchers aim to discover genetic variants linked to complex traits and diseases by analyzing tens of millions of variants across a large number of individuals.These large-scale variable selection tasks present several practical challenges. First, the sheer number of candidate variables increases the risk of false discoveries, making error control essential for ensuring the reproducibility and reliability of scientific findings. Second, to protect privacy and enable broader data sharing, many applications release only summary statistics. For example, GWAS often provide summary data like marginal associations and estimated linkage disequilibrium (LD) patterns instead of individual-level data. This lack of access to individual-level data poses challenges for statistical inference.To address these challenges, the first part of this thesis develops novel methods for controlled variable selection using summary statistics, as presented in Chapters 2 and 3. Specifically, we frame variable selection as a multiple testing problem involving conditional independence hypotheses:H j0: Xj ⊥⊥ Y | X−j , 1 ≤ j ≤ p,Where X−j= (X1, . . . , Xj−1, Xj+1, . . . , Xp) denotes all variables except Xj. Under H j0, the variable Xjprovides no additional information about the response Ybeyond what is already captured by the other covariates. Our goal is to perform powerful testing for H j0, j= 1, . . . , p,while controlling the number of false discoveries. In this thesis, we consider two types of error control: the false discovery rate (FDR),which is the expected proportion of false positives among the selected variables, and the familywise error rate (FWER),which is the probability of making at least one false discovery. FDR control is less stringent and allows for greater power, making it suitable for exploratory analyses. In contrast, FWER control is more conservative and better suited for high-stakes applications where any false positive could be costly. In Chapter 2, we introduce novel GhostKnockoffmethods based on penalized regression that control the FDR in the aforementioned conditional independence testing problem using only summary statistics. In Chapter 3, we introduce a new filter for variable selection using summary statistics with rigorous FWER control. Together, these tools enable statistically principled variable selection in large-scale studies where only summary-level data are available.1.2 Localized Feature Selection and Causal InferenceWhile traditional statistical analysis often focuses on inference or exploratory insights at the population level, there is growing interest in drawing conclusions at a more localized or individual level to enable higher-resolution understanding and to better support personalized decision-making in practice. This has important implications for real-world applications such as personalized medicine, business analytics, and individualized education. For instance, in precision medicine, researchers often aim to identify genetic factors associated with disease while accounting for patient heterogeneity-such as differences in age or gender-to develop targeted treatments.
일반주제명  
Families & family life
일반주제명  
Feature selection
일반주제명  
Statistics
일반주제명  
Alzheimer's disease
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 87-01B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017358725
■00520260202104743
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798290649481
■035    ▼a(MiAaPQ)AAI32149722
■035    ▼a(MiAaPQ)Stanfordsq475hw3784
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a306
■1001  ▼aChen,  Zhaomeng.
■24510▼aTopics  in  Selective  and  Causal  Inference
■260    ▼a[Sl]▼bStanford  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a231  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  B.
■500    ▼aAdvisor:  Candès,  Emmanuel.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2025.
■520    ▼a1.1  Controlled  Variable  Selection  Based  on  Summary  StatisticsModern  scientific  studies  across  fields  such  as  genomics,  neuroscience,  and  economics  increasingly  involve  large-scale  datasets  with  high-dimensional  covariates  and  complex  outcomes.  A  central  goal  in  many  of  these  applications  is  to  identify  variables  that  are  meaningfully  associated  with  a  response  of  interest.  For  example,  in  genome-wide  association  studies  (GWAS),  researchers  aim  to  discover  genetic  variants  linked  to  complex  traits  and  diseases  by  analyzing  tens  of  millions  of  variants  across  a  large  number  of  individuals.These  large-scale  variable  selection  tasks  present  several  practical  challenges.  First,  the  sheer  number  of  candidate  variables  increases  the  risk  of  false  discoveries,  making  error  control  essential  for  ensuring  the  reproducibility  and  reliability  of  scientific  findings.  Second,  to  protect  privacy  and  enable  broader  data  sharing,  many  applications  release  only  summary  statistics.  For  example,  GWAS  often  provide  summary  data  like  marginal  associations  and  estimated  linkage  disequilibrium  (LD)  patterns  instead  of  individual-level  data.  This  lack  of  access  to  individual-level  data  poses  challenges  for  statistical  inference.To  address  these  challenges,  the  first  part  of  this  thesis  develops  novel  methods  for  controlled  variable  selection  using  summary  statistics,  as  presented  in  Chapters  2  and  3.  Specifically,  we  frame  variable  selection  as  a  multiple  testing  problem  involving  conditional  independence  hypotheses:H  j0:  Xj  ⊥⊥  Y  |  X−j  ,  1  ≤  j  ≤  p,Where  X−j=  (X1,  .  .  .  ,  Xj−1,  Xj+1,  .  .  .  ,  Xp)  denotes  all  variables  except  Xj.  Under  H  j0,  the  variable  Xjprovides  no  additional  information  about  the  response  Ybeyond  what  is  already  captured  by  the  other  covariates.  Our  goal  is  to  perform  powerful  testing  for  H  j0,  j=  1,  .  .  .  ,  p,while  controlling  the  number  of  false  discoveries.  In  this  thesis,  we  consider  two  types  of  error  control:  the  false  discovery  rate  (FDR),which  is  the  expected  proportion  of  false  positives  among  the  selected  variables,  and  the  familywise  error  rate  (FWER),which  is  the  probability  of  making  at  least  one  false  discovery.  FDR  control  is  less  stringent  and  allows  for  greater  power,  making  it  suitable  for  exploratory  analyses.  In  contrast,  FWER  control  is  more  conservative  and  better  suited  for  high-stakes  applications  where  any  false  positive  could  be  costly.  In  Chapter  2,  we  introduce  novel  GhostKnockoffmethods  based  on  penalized  regression  that  control  the  FDR  in  the  aforementioned  conditional  independence  testing  problem  using  only  summary  statistics.  In  Chapter  3,  we  introduce  a  new  filter  for  variable  selection  using  summary  statistics  with  rigorous  FWER  control.  Together,  these  tools  enable  statistically  principled  variable  selection  in  large-scale  studies  where  only  summary-level  data  are  available.1.2  Localized  Feature  Selection  and  Causal  InferenceWhile  traditional  statistical  analysis  often  focuses  on  inference  or  exploratory  insights  at  the  population  level,  there  is  growing  interest  in  drawing  conclusions  at  a  more  localized  or  individual  level  to  enable  higher-resolution  understanding  and  to  better  support  personalized  decision-making  in  practice.  This  has  important  implications  for  real-world  applications  such  as  personalized  medicine,  business  analytics,  and  individualized  education.  For  instance,  in  precision  medicine,  researchers  often  aim  to  identify  genetic  factors  associated  with  disease  while  accounting  for  patient  heterogeneity-such  as  differences  in  age  or  gender-to  develop  targeted  treatments.
■590    ▼aSchool  code:  0212.
■650  4▼aFamilies  &  family  life
■650  4▼aFeature  selection
■650  4▼aStatistics
■650  4▼aAlzheimer's  disease
■690    ▼a0463
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g87-01B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358725▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF19320 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.