본문

서브메뉴

Detecting and Measuring Important Variables: Novel Methods With Statistical Guarantees
Detecting and Measuring Important Variables: Novel Methods With Statistical Guarantees
Detecting and Measuring Important Variables: Novel Methods With Statistical Guarantees

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104741
ISBN  
9798290651460
DDC  
519
저자명  
Gablenz, Paula.
서명/저자  
Detecting and Measuring Important Variables: Novel Methods With Statistical Guarantees
발행사항  
[Sl] : Stanford University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
145 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
주기사항  
Advisor: Sabatti, Chiara.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2025.
초록/해제  
요약1.1 Background and motivationUnderstanding which explanatory variables, out of potentially millions, are important is a fundamental challenge in statistics. The eventual purpose is to help shed light on the mechanisms affecting the outcome of interest. The field of multiple testing addresses how to discover as many relevant variables as possible, while at the same time ensuring that the findings are replicable. In numerous applications the associated hypotheses are structured, but many conventional approaches do not take this structure into account. A compelling example is offered by genome-wide association studies (GWAS). The goal of genome-wide association studies is to identify which genetic variants influence a certain disease. The genetic variants have spatial structure and can be grouped at different levels of resolution. In addition, the outcomes of interest might be structured, such as the tree structure of the international classification of diseases. Finding the most precise rejections, ensuring consistency between the rejections at different levels of resolution, or generally filtering the rejection set thus represent salient problems. Moreover, there might be individual heterogeneity, where some genetic variants are only relevant for a certain subset of individuals. A challenge in multiple testing are dependencies between the tested hypotheses, which could invalidate the inferential procedure. In this dissertation, we develop new methodologies for (structured) hypotheses testing that are valid under any dependency structures. In addition, we present a novel application to reliably detect gene-environment interactions. Finally, we explore ideas on quantifying how important a variable is for the uncertainty and prediction of a particular model, as opposed to only testing whether it is important.1.2 Structure of the thesisChapter 2 addresses the specific problem of finding the most precise rejections possible, while ensuring error control. In particular, we consider problems where many, somewhat redundant, hypotheses are tested at different levels of resolution. We design a novel multiple comparison procedure that allows for an adaptive choice of resolution with false discovery rate (FDR) control, leveraging e-values.In chapter 3 we develop several additional approaches for structured hypotheses testing using e-values. Specifically, we consider the coordination of rejections across multiple resolutions to avoid conflicting rejections, filtering the rejection set more generally and testing partial conjunction hypotheses. We introduce e-value counterparts of p-value based procedures. Our methods provide error control under any dependency structures.Chapter 4 considers local conditional hypotheses which express how the relation between explanatory variables and outcomes change across different environments, described by covariates. We describe a practical implementation to GWAS-scale data and demonstrate its effectiveness through numerical experiments on the UK Biobank.Chapter 5 introduces a set of novel ideas that contribute to a deeper understanding of measuring variable importance. Building upon established frameworks for variable importance, we present an alternative perspective that offers new directions for understanding predictive uncertainty through the lens of conformal prediction intervals.
일반주제명  
Linear programming
일반주제명  
Genomes
일반주제명  
Biobanks
일반주제명  
Power
일반주제명  
Statistics
일반주제명  
Biostatistics
키워드  
Genome-wide association studies
키워드  
False discovery rate
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 87-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017358714
■00520260202104741
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798290651460
■035    ▼a(MiAaPQ)AAI32149704
■035    ▼a(MiAaPQ)Stanfordpm479hw3823
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a519
■1001  ▼aGablenz,  Paula.
■24510▼aDetecting  and  Measuring  Important  Variables:  Novel  Methods  With  Statistical  Guarantees
■260    ▼a[Sl]▼bStanford  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a145  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  B.
■500    ▼aAdvisor:  Sabatti,  Chiara.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2025.
■520    ▼a1.1  Background  and  motivationUnderstanding  which  explanatory  variables,  out  of  potentially  millions,  are  important  is  a  fundamental  challenge  in  statistics.  The  eventual  purpose  is  to  help  shed  light  on  the  mechanisms  affecting  the  outcome  of  interest.  The  field  of  multiple  testing  addresses  how  to  discover  as  many  relevant  variables  as  possible,  while  at  the  same  time  ensuring  that  the  findings  are  replicable.  In  numerous  applications  the  associated  hypotheses  are  structured,  but  many  conventional  approaches  do  not  take  this  structure  into  account.  A  compelling  example  is  offered  by  genome-wide  association  studies  (GWAS).  The  goal  of  genome-wide  association  studies  is  to  identify  which  genetic  variants  influence  a  certain  disease.  The  genetic  variants  have  spatial  structure  and  can  be  grouped  at  different  levels  of  resolution.  In  addition,  the  outcomes  of  interest  might  be  structured,  such  as  the  tree  structure  of  the  international  classification  of  diseases.  Finding  the  most  precise  rejections,  ensuring  consistency  between  the  rejections  at  different  levels  of  resolution,  or  generally  filtering  the  rejection  set  thus  represent  salient  problems.  Moreover,  there  might  be  individual  heterogeneity,  where  some  genetic  variants  are  only  relevant  for  a  certain  subset  of  individuals.  A  challenge  in  multiple  testing  are  dependencies  between  the  tested  hypotheses,  which  could  invalidate  the  inferential  procedure.  In  this  dissertation,  we  develop  new  methodologies  for  (structured)  hypotheses  testing  that  are  valid  under  any  dependency  structures.  In  addition,  we  present  a  novel  application  to  reliably  detect  gene-environment  interactions.  Finally,  we  explore  ideas  on  quantifying  how  important  a  variable  is  for  the  uncertainty  and  prediction  of  a  particular  model,  as  opposed  to  only  testing  whether  it  is  important.1.2  Structure  of  the  thesisChapter  2  addresses  the  specific  problem  of  finding  the  most  precise  rejections  possible,  while  ensuring  error  control.  In  particular,  we  consider  problems  where  many,  somewhat  redundant,  hypotheses  are  tested  at  different  levels  of  resolution.  We  design  a  novel  multiple  comparison  procedure  that  allows  for  an  adaptive  choice  of  resolution  with  false  discovery  rate  (FDR)  control,  leveraging  e-values.In  chapter  3  we  develop  several  additional  approaches  for  structured  hypotheses  testing  using  e-values.  Specifically,  we  consider  the  coordination  of  rejections  across  multiple  resolutions  to  avoid  conflicting  rejections,  filtering  the  rejection  set  more  generally  and  testing  partial  conjunction  hypotheses.  We  introduce  e-value  counterparts  of  p-value  based  procedures.  Our  methods  provide  error  control  under  any  dependency  structures.Chapter  4  considers  local  conditional  hypotheses  which  express  how  the  relation  between  explanatory  variables  and  outcomes  change  across  different  environments,  described  by  covariates.  We  describe  a  practical  implementation  to  GWAS-scale  data  and  demonstrate  its  effectiveness  through  numerical  experiments  on  the  UK  Biobank.Chapter  5  introduces  a  set  of  novel  ideas  that  contribute  to  a  deeper  understanding  of  measuring  variable  importance.  Building  upon  established  frameworks  for  variable  importance,  we  present  an  alternative  perspective  that  offers  new  directions  for  understanding  predictive  uncertainty  through  the  lens  of  conformal  prediction  intervals.
■590    ▼aSchool  code:  0212.
■650  4▼aLinear  programming
■650  4▼aGenomes
■650  4▼aBiobanks
■650  4▼aPower
■650  4▼aStatistics
■650  4▼aBiostatistics
■653    ▼aGenome-wide  association  studies
■653    ▼aFalse  discovery  rate
■690    ▼a0463
■690    ▼a0308
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g87-03B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358714▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF19258 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.