서브메뉴
검색
Reconsider Machine Learning Method for Variable Selection and Validation With High Dimensional Data
Reconsider Machine Learning Method for Variable Selection and Validation With High Dimensional Data
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152040
- ISBN
- 9798384093374
- DDC
- 574
- 저자명
- Liu, Lu.
- 서명/저자
- Reconsider Machine Learning Method for Variable Selection and Validation With High Dimensional Data
- 발행사항
- [Sl] : Duke University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 89 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: A.
- 주기사항
- Advisor: Jung, Sin-Ho.
- 학위논문주기
- Thesis (Ph.D.)--Duke University, 2024.
- 초록/해제
- 요약The big data tendency influences how people think and inspires potential research directions. Recent feats of machine learning have seized collective attention because of its profound performance in conducting big data analysis including text analysis and image processing. Machine learning is also a popular topic in clinical medicine to implement analysis on electronic health records and medical image data, which traditional statistics model is not adequate for. However, we realize that machine learning is not panacea and its defects such as loss of interpretability and excess selection may restrict its application. And we must also recognize that for many clinical prediction analyses, the simpler approach-generalized linear model is enough for what we need. In this dissertation, we propose to use standard regression methods, without any penalizing approach, combined with a stepwise variable selection procedure to overcome the over-selection issue of popular machine learning methods. For model validation, we propose a permutation approach to estimate the performance of various validation methods. Finally, we propose a repeated sieving approach, extending the standard regression methods with stepwise variable selection, to handle high dimensional modeling.
- 일반주제명
- Biostatistics
- 일반주제명
- Statistics
- 일반주제명
- Bioinformatics
- 일반주제명
- Information science
- 키워드
- Machine learning
- 기타저자
- Duke University Biostatistics and Bioinformatics Doctor of Philosophy
- 기본자료저록
- Dissertations Abstracts International. 86-03A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162676
■00520250211152040
■006m o d
■007cr#unu||||||||
■020 ▼a9798384093374
■035 ▼a(MiAaPQ)AAI31336592
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aLiu, Lu.
■24510▼aReconsider Machine Learning Method for Variable Selection and Validation With High Dimensional Data
■260 ▼a[Sl]▼bDuke University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a89 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: A.
■500 ▼aAdvisor: Jung, Sin-Ho.
■5021 ▼aThesis (Ph.D.)--Duke University, 2024.
■520 ▼aThe big data tendency influences how people think and inspires potential research directions. Recent feats of machine learning have seized collective attention because of its profound performance in conducting big data analysis including text analysis and image processing. Machine learning is also a popular topic in clinical medicine to implement analysis on electronic health records and medical image data, which traditional statistics model is not adequate for. However, we realize that machine learning is not panacea and its defects such as loss of interpretability and excess selection may restrict its application. And we must also recognize that for many clinical prediction analyses, the simpler approach-generalized linear model is enough for what we need. In this dissertation, we propose to use standard regression methods, without any penalizing approach, combined with a stepwise variable selection procedure to overcome the over-selection issue of popular machine learning methods. For model validation, we propose a permutation approach to estimate the performance of various validation methods. Finally, we propose a repeated sieving approach, extending the standard regression methods with stepwise variable selection, to handle high dimensional modeling.
■590 ▼aSchool code: 0066.
■650 4▼aBiostatistics
■650 4▼aStatistics
■650 4▼aBioinformatics
■650 4▼aInformation science
■653 ▼aLogistic regression
■653 ▼aMachine learning
■653 ▼aPermutation approach
■653 ▼aVariable selection
■653 ▼aValidation methods
■690 ▼a0308
■690 ▼a0723
■690 ▼a0715
■690 ▼a0463
■71020▼aDuke University▼bBiostatistics and Bioinformatics Doctor of Philosophy.
■7730 ▼tDissertations Abstracts International▼g86-03A.
■790 ▼a0066
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162676▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


