본문

서브메뉴

Reconsider Machine Learning Method for Variable Selection and Validation With High Dimensional Data
Reconsider Machine Learning Method for Variable Selection and Validation With High Dimensi...
Reconsider Machine Learning Method for Variable Selection and Validation With High Dimensional Data

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152040
ISBN  
9798384093374
DDC  
574
저자명  
Liu, Lu.
서명/저자  
Reconsider Machine Learning Method for Variable Selection and Validation With High Dimensional Data
발행사항  
[Sl] : Duke University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
89 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: A.
주기사항  
Advisor: Jung, Sin-Ho.
학위논문주기  
Thesis (Ph.D.)--Duke University, 2024.
초록/해제  
요약The big data tendency influences how people think and inspires potential research directions. Recent feats of machine learning have seized collective attention because of its profound performance in conducting big data analysis including text analysis and image processing. Machine learning is also a popular topic in clinical medicine to implement analysis on electronic health records and medical image data, which traditional statistics model is not adequate for. However, we realize that machine learning is not panacea and its defects such as loss of interpretability and excess selection may restrict its application. And we must also recognize that for many clinical prediction analyses, the simpler approach-generalized linear model is enough for what we need. In this dissertation, we propose to use standard regression methods, without any penalizing approach, combined with a stepwise variable selection procedure to overcome the over-selection issue of popular machine learning methods. For model validation, we propose a permutation approach to estimate the performance of various validation methods. Finally, we propose a repeated sieving approach, extending the standard regression methods with stepwise variable selection, to handle high dimensional modeling.
일반주제명  
Biostatistics
일반주제명  
Statistics
일반주제명  
Bioinformatics
일반주제명  
Information science
키워드  
Logistic regression
키워드  
Machine learning
키워드  
Permutation approach
키워드  
Variable selection
키워드  
Validation methods
기타저자  
Duke University Biostatistics and Bioinformatics Doctor of Philosophy
기본자료저록  
Dissertations Abstracts International. 86-03A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017162676
■00520250211152040
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384093374
■035    ▼a(MiAaPQ)AAI31336592
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aLiu,  Lu.
■24510▼aReconsider  Machine  Learning  Method  for  Variable  Selection  and  Validation  With  High  Dimensional  Data
■260    ▼a[Sl]▼bDuke  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a89  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  A.
■500    ▼aAdvisor:  Jung,  Sin-Ho.
■5021  ▼aThesis  (Ph.D.)--Duke  University,  2024.
■520    ▼aThe  big  data  tendency  influences  how  people  think  and  inspires  potential  research  directions.  Recent  feats  of  machine  learning  have  seized  collective  attention  because  of  its  profound  performance  in  conducting  big  data  analysis  including  text  analysis  and  image  processing.  Machine  learning  is  also  a  popular  topic  in  clinical  medicine  to  implement  analysis  on  electronic  health  records  and  medical  image  data,  which  traditional  statistics  model  is  not  adequate  for.  However,  we  realize  that  machine  learning  is  not  panacea  and  its  defects  such  as  loss  of  interpretability  and  excess  selection  may  restrict  its  application.  And  we  must  also  recognize  that  for  many  clinical  prediction  analyses,  the  simpler  approach-generalized  linear  model  is  enough  for  what  we  need.  In  this  dissertation,  we  propose  to  use  standard  regression  methods,  without  any  penalizing  approach,  combined  with  a  stepwise  variable  selection  procedure  to  overcome  the  over-selection  issue  of  popular  machine  learning  methods.  For  model  validation,  we  propose  a  permutation  approach  to  estimate  the  performance  of  various  validation  methods.  Finally,  we  propose  a  repeated  sieving  approach,  extending  the  standard  regression  methods  with  stepwise  variable  selection,  to  handle  high  dimensional  modeling.
■590    ▼aSchool  code:  0066.
■650  4▼aBiostatistics
■650  4▼aStatistics
■650  4▼aBioinformatics
■650  4▼aInformation  science
■653    ▼aLogistic  regression
■653    ▼aMachine  learning
■653    ▼aPermutation  approach
■653    ▼aVariable  selection
■653    ▼aValidation  methods
■690    ▼a0308
■690    ▼a0723
■690    ▼a0715
■690    ▼a0463
■71020▼aDuke  University▼bBiostatistics  and  Bioinformatics  Doctor  of  Philosophy.
■7730  ▼tDissertations  Abstracts  International▼g86-03A.
■790    ▼a0066
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162676▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF09961 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.