본문

서브메뉴

Large Scale Inference and Combinatorial Variable Selection for Complex Dataset
Large Scale Inference and Combinatorial Variable Selection for Complex Dataset
Large Scale Inference and Combinatorial Variable Selection for Complex Dataset

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151122
ISBN  
9798382775579
DDC  
310
저자명  
Liu, Yue.
서명/저자  
Large Scale Inference and Combinatorial Variable Selection for Complex Dataset
발행사항  
[Sl] : Harvard University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
297 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Liu, Jun.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2024.
초록/해제  
요약This dissertation advances the field of modern statistical theory and methodology by focusing on two primary areas: first, the quantification of uncertainty beyond mere estimation in combinatorial inference theory; and second, addressing the complexities and challenges inherent in electronic health records (EHR).Chapter 1 introduces a novel combinatorial inference framework to conduct general uncertainty quantification in ranking problems. By considering the Bradley-Terry-Luce model, we aim to infer both local and global ranking properties, and generalize the method to multi-tesing problem with false discovery rate (FDR) control.Chapter 2 focuses on the development of a semi-supervised approach that efficiently leverages sizable unlabeled samples with error-prone EHR surrogate outcomes from multiple local sites, to improve the learning accuracy of the small gold-labeled data. we apply our method to develop a high dimensional genetic risk model for type II diabetes using large-scale data sets from UK and Mass General Brigham biobanks, where only a small fraction of subjects in one site has been labeled via chart reviewing.Chapter 3 presents a novel inferential framework for general graphical models to select graph features with false discovery rate controlled. The proposed method is based on the maximum of p-values from single edges that comprise the topological feature of interest, thus is able to detect weak signals. Moreover, we introduce the K-dimensional persistent Homology Adaptive selectioN (KHAN) algorithm to select all the homological features within K dimensions with the uniform control of the false discovery rate over continuous filtration levels. The KHAN method applies a novel discrete Gram-Schmidt algorithm to select statistically significant generators from the homology group. We apply the structural screening method to identify the important residues of the SARS-CoV-2 spike protein during the binding process to the ACE2 receptors. We score the residues for all domains in the spike protein by the p-value weighted filtration level in the network persistent homology for the closed, partially open, and open states and identify the residues crucial for protein conformational changes and thus being potential targets for inhibition.
일반주제명  
Statistics
일반주제명  
Biostatistics
일반주제명  
Health sciences
키워드  
Electronic health records
키워드  
False discovery rate
키워드  
Statistical theory
키워드  
Genetic risk model
키워드  
Homology group
기타저자  
Harvard University Statistics
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017160826
■00520250211151122
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382775579
■035    ▼a(MiAaPQ)AAI31146719
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a310
■1001  ▼aLiu,  Yue.▼0(orcid)0000-0002-1897-4562
■24510▼aLarge  Scale  Inference  and  Combinatorial  Variable  Selection  for  Complex  Dataset
■260    ▼a[Sl]▼bHarvard  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a297  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Liu,  Jun.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2024.
■520    ▼aThis  dissertation  advances  the  field  of  modern  statistical  theory  and  methodology  by  focusing  on  two  primary  areas:  first,  the  quantification  of  uncertainty  beyond  mere  estimation  in  combinatorial  inference  theory;  and  second,  addressing  the  complexities  and  challenges  inherent  in  electronic  health  records  (EHR).Chapter  1  introduces  a  novel  combinatorial  inference  framework  to  conduct  general  uncertainty  quantification  in  ranking  problems.  By  considering  the  Bradley-Terry-Luce  model,  we  aim  to  infer  both  local  and  global  ranking  properties,  and  generalize  the  method  to  multi-tesing  problem  with  false  discovery  rate  (FDR)  control.Chapter  2  focuses  on  the  development  of  a  semi-supervised  approach  that  efficiently  leverages  sizable  unlabeled  samples  with  error-prone  EHR  surrogate  outcomes  from  multiple  local  sites,  to  improve  the  learning  accuracy  of  the  small  gold-labeled  data.  we  apply  our  method  to  develop  a  high  dimensional  genetic  risk  model  for  type  II  diabetes  using  large-scale  data  sets  from  UK  and  Mass  General  Brigham  biobanks,  where  only  a  small  fraction  of  subjects  in  one  site  has  been  labeled  via  chart  reviewing.Chapter  3  presents  a  novel  inferential  framework  for  general  graphical  models  to  select  graph  features  with  false  discovery  rate  controlled.  The  proposed  method  is  based  on  the  maximum  of  p-values  from  single  edges  that  comprise  the  topological  feature  of  interest,  thus  is  able  to  detect  weak  signals.  Moreover,  we  introduce  the  K-dimensional  persistent  Homology  Adaptive  selectioN  (KHAN)  algorithm  to  select  all  the  homological  features  within  K  dimensions  with  the  uniform  control  of  the  false  discovery  rate  over  continuous  filtration  levels.  The  KHAN  method  applies  a  novel  discrete  Gram-Schmidt  algorithm  to  select  statistically  significant  generators  from  the  homology  group.  We  apply  the  structural  screening  method  to  identify  the  important  residues  of  the  SARS-CoV-2  spike  protein  during  the  binding  process  to  the  ACE2  receptors.  We  score  the  residues  for  all  domains  in  the  spike  protein  by  the  p-value  weighted  filtration  level  in  the  network  persistent  homology  for  the  closed,  partially  open,  and  open  states  and  identify  the  residues  crucial  for  protein  conformational  changes  and  thus  being  potential  targets  for  inhibition.
■590    ▼aSchool  code:  0084.
■650  4▼aStatistics
■650  4▼aBiostatistics
■650  4▼aHealth  sciences
■653    ▼aElectronic  health  records
■653    ▼aFalse  discovery  rate
■653    ▼aStatistical  theory
■653    ▼aGenetic  risk  model
■653    ▼aHomology  group
■690    ▼a0463
■690    ▼a0308
■690    ▼a0566
■71020▼aHarvard  University▼bStatistics.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160826▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11709 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.