서브메뉴
검색
Large Scale Inference and Combinatorial Variable Selection for Complex Dataset
Large Scale Inference and Combinatorial Variable Selection for Complex Dataset
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151122
- ISBN
- 9798382775579
- DDC
- 310
- 저자명
- Liu, Yue.
- 서명/저자
- Large Scale Inference and Combinatorial Variable Selection for Complex Dataset
- 발행사항
- [Sl] : Harvard University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 297 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
- 주기사항
- Advisor: Liu, Jun.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2024.
- 초록/해제
- 요약This dissertation advances the field of modern statistical theory and methodology by focusing on two primary areas: first, the quantification of uncertainty beyond mere estimation in combinatorial inference theory; and second, addressing the complexities and challenges inherent in electronic health records (EHR).Chapter 1 introduces a novel combinatorial inference framework to conduct general uncertainty quantification in ranking problems. By considering the Bradley-Terry-Luce model, we aim to infer both local and global ranking properties, and generalize the method to multi-tesing problem with false discovery rate (FDR) control.Chapter 2 focuses on the development of a semi-supervised approach that efficiently leverages sizable unlabeled samples with error-prone EHR surrogate outcomes from multiple local sites, to improve the learning accuracy of the small gold-labeled data. we apply our method to develop a high dimensional genetic risk model for type II diabetes using large-scale data sets from UK and Mass General Brigham biobanks, where only a small fraction of subjects in one site has been labeled via chart reviewing.Chapter 3 presents a novel inferential framework for general graphical models to select graph features with false discovery rate controlled. The proposed method is based on the maximum of p-values from single edges that comprise the topological feature of interest, thus is able to detect weak signals. Moreover, we introduce the K-dimensional persistent Homology Adaptive selectioN (KHAN) algorithm to select all the homological features within K dimensions with the uniform control of the false discovery rate over continuous filtration levels. The KHAN method applies a novel discrete Gram-Schmidt algorithm to select statistically significant generators from the homology group. We apply the structural screening method to identify the important residues of the SARS-CoV-2 spike protein during the binding process to the ACE2 receptors. We score the residues for all domains in the spike protein by the p-value weighted filtration level in the network persistent homology for the closed, partially open, and open states and identify the residues crucial for protein conformational changes and thus being potential targets for inhibition.
- 일반주제명
- Statistics
- 일반주제명
- Biostatistics
- 일반주제명
- Health sciences
- 키워드
- Homology group
- 기타저자
- Harvard University Statistics
- 기본자료저록
- Dissertations Abstracts International. 85-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160826
■00520250211151122
■006m o d
■007cr#unu||||||||
■020 ▼a9798382775579
■035 ▼a(MiAaPQ)AAI31146719
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aLiu, Yue.▼0(orcid)0000-0002-1897-4562
■24510▼aLarge Scale Inference and Combinatorial Variable Selection for Complex Dataset
■260 ▼a[Sl]▼bHarvard University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a297 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-12, Section: B.
■500 ▼aAdvisor: Liu, Jun.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2024.
■520 ▼aThis dissertation advances the field of modern statistical theory and methodology by focusing on two primary areas: first, the quantification of uncertainty beyond mere estimation in combinatorial inference theory; and second, addressing the complexities and challenges inherent in electronic health records (EHR).Chapter 1 introduces a novel combinatorial inference framework to conduct general uncertainty quantification in ranking problems. By considering the Bradley-Terry-Luce model, we aim to infer both local and global ranking properties, and generalize the method to multi-tesing problem with false discovery rate (FDR) control.Chapter 2 focuses on the development of a semi-supervised approach that efficiently leverages sizable unlabeled samples with error-prone EHR surrogate outcomes from multiple local sites, to improve the learning accuracy of the small gold-labeled data. we apply our method to develop a high dimensional genetic risk model for type II diabetes using large-scale data sets from UK and Mass General Brigham biobanks, where only a small fraction of subjects in one site has been labeled via chart reviewing.Chapter 3 presents a novel inferential framework for general graphical models to select graph features with false discovery rate controlled. The proposed method is based on the maximum of p-values from single edges that comprise the topological feature of interest, thus is able to detect weak signals. Moreover, we introduce the K-dimensional persistent Homology Adaptive selectioN (KHAN) algorithm to select all the homological features within K dimensions with the uniform control of the false discovery rate over continuous filtration levels. The KHAN method applies a novel discrete Gram-Schmidt algorithm to select statistically significant generators from the homology group. We apply the structural screening method to identify the important residues of the SARS-CoV-2 spike protein during the binding process to the ACE2 receptors. We score the residues for all domains in the spike protein by the p-value weighted filtration level in the network persistent homology for the closed, partially open, and open states and identify the residues crucial for protein conformational changes and thus being potential targets for inhibition.
■590 ▼aSchool code: 0084.
■650 4▼aStatistics
■650 4▼aBiostatistics
■650 4▼aHealth sciences
■653 ▼aElectronic health records
■653 ▼aFalse discovery rate
■653 ▼aStatistical theory
■653 ▼aGenetic risk model
■653 ▼aHomology group
■690 ▼a0463
■690 ▼a0308
■690 ▼a0566
■71020▼aHarvard University▼bStatistics.
■7730 ▼tDissertations Abstracts International▼g85-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160826▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


