본문

서브메뉴

A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods
A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Meth...
A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20250211153126
ISBN  
9798346859673
DDC  
151
저자명  
Tibbe, Tristan Dale.
서명/저자  
A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods
발행사항  
[Sl] : University of California, Los Angeles, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
151 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-06, Section: A.
주기사항  
Advisor: Montoya, Amanda K.
학위논문주기  
Thesis (Ph.D.)--University of California, Los Angeles, 2024.
초록/해제  
요약As machine learning techniques increase in popularity among psychology researchers, decision trees---and their offshoots such as random forests and qualitative interaction trees (QUINT)---have received special attention due to their interpretability and ease of use. Although these tree-based methods are versatile in terms of the types and number of relationships they can model, missingness still needs to be addressed before they can produce predictions. The current literature comparing missing data methods available for tree-based models is fragmented, with only subsets of conditions or methods examined in each study. Furthermore, the application of missing data methods to the specialized tree-based algorithm QUINT has largely gone uninvestigated in previous research. Thus, there is a great need to clarify which missing data methods are best to use in which situations, especially with niche tree-based models like QUINT. Since tree-based methods are quick to set up and easy to interpret, it is particularly important that users who may not be experienced in machine learning or statistics receive guidance on how they should manage missingness in their data before they apply such methods. In order to consolidate the research that has already been done on missing data methods with tree-based models, this dissertation provides a summary of the existing literature on the topic, introducing the methods and factors that are important to consider when choosing how to deal with missingness in datasets for tree models. Also, to understand which missing data methods are currently applied to tree-based models in psychological studies and under which conditions, recent substantive research articles were reviewed and information about their datasets and methodologies were recorded. Finally, to extend the knowledge accumulated in the literature, a plethora of popular/modern missing data methods for tree models were applied to the QUINT algorithm and compared in a simulation study, varying factors such as the amount of missingness, the type of missingness, and where the missingness appeared in the data to cover a variety of scenarios that may arise in the real world. The results reveal that, in terms of both prediction accuracy and variable selection, imputation methods---specifically regression and hot deck imputation---and missingness incorporated in attributes (MIA) are able to address missing data and produce QUINT models of consistently high quality. The complete case method, on the other hand, should be avoided due to its highly variable performance and inconsistent nature, leading to models that differ greatly from would have been produced had the data been complete.
일반주제명  
Quantitative psychology
일반주제명  
Statistics
일반주제명  
Psychology
일반주제명  
Information science
키워드  
Decision trees
키워드  
Missing data
키워드  
Qualitative interaction trees
키워드  
Tree-based methods
키워드  
Machine learning
기타저자  
University of California, Los Angeles Psychology 0780
기본자료저록  
Dissertations Abstracts International. 86-06A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017165121
■00520250211153126
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798346859673
■035    ▼a(MiAaPQ)AAI31765212
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a151
■1001  ▼aTibbe,  Tristan  Dale.
■24512▼aA  Comprehensive  Comparison  of  Missing  Data  Procedures  for  Tree-Based  Machine  Learning  Methods
■260    ▼a[Sl]▼bUniversity  of  California,  Los  Angeles▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a151  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-06,  Section:  A.
■500    ▼aAdvisor:  Montoya,  Amanda  K.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Los  Angeles,  2024.
■520    ▼aAs  machine  learning  techniques  increase  in  popularity  among  psychology  researchers,  decision  trees---and  their  offshoots  such  as  random  forests  and  qualitative  interaction  trees  (QUINT)---have  received  special  attention  due  to  their  interpretability  and  ease  of  use.  Although  these  tree-based  methods  are  versatile  in  terms  of  the  types  and  number  of  relationships  they  can  model,  missingness  still  needs  to  be  addressed  before  they  can  produce  predictions.  The  current  literature  comparing  missing  data  methods  available  for  tree-based  models  is  fragmented,  with  only  subsets  of  conditions  or  methods  examined  in  each  study.  Furthermore,  the  application  of  missing  data  methods  to  the  specialized  tree-based  algorithm  QUINT  has  largely  gone  uninvestigated  in  previous  research.  Thus,  there  is  a  great  need  to  clarify  which  missing  data  methods  are  best  to  use  in  which  situations,  especially  with  niche  tree-based  models  like  QUINT.  Since  tree-based  methods  are  quick  to  set  up  and  easy  to  interpret,  it  is  particularly  important  that  users  who  may  not  be  experienced  in  machine  learning  or  statistics  receive  guidance  on  how  they  should  manage  missingness  in  their  data  before  they  apply  such  methods.  In  order  to  consolidate  the  research  that  has  already  been  done  on  missing  data  methods  with  tree-based  models,  this  dissertation  provides  a  summary  of  the  existing  literature  on  the  topic,  introducing  the  methods  and  factors  that  are  important  to  consider  when  choosing  how  to  deal  with  missingness  in  datasets  for  tree  models.  Also,  to  understand  which  missing  data  methods  are  currently  applied  to  tree-based  models  in  psychological  studies  and  under  which  conditions,  recent  substantive  research  articles  were  reviewed  and  information  about  their  datasets  and  methodologies  were  recorded.  Finally,  to  extend  the  knowledge  accumulated  in  the  literature,  a  plethora  of  popular/modern  missing  data  methods  for  tree  models  were  applied  to  the  QUINT  algorithm  and  compared  in  a  simulation  study,  varying  factors  such  as  the  amount  of  missingness,  the  type  of  missingness,  and  where  the  missingness  appeared  in  the  data  to  cover  a  variety  of  scenarios  that  may  arise  in  the  real  world.  The  results  reveal  that,  in  terms  of  both  prediction  accuracy  and  variable  selection,  imputation  methods---specifically  regression  and  hot  deck  imputation---and  missingness  incorporated  in  attributes  (MIA)  are  able  to  address  missing  data  and  produce  QUINT  models  of  consistently  high  quality.  The  complete  case  method,  on  the  other  hand,  should  be  avoided  due  to  its  highly  variable  performance  and  inconsistent  nature,  leading  to  models  that  differ  greatly  from  would  have  been  produced  had  the  data  been  complete.
■590    ▼aSchool  code:  0031.
■650  4▼aQuantitative  psychology
■650  4▼aStatistics
■650  4▼aPsychology
■650  4▼aInformation  science
■653    ▼aDecision  trees
■653    ▼aMissing  data
■653    ▼aQualitative  interaction  trees
■653    ▼aTree-based  methods
■653    ▼aMachine  learning
■690    ▼a0632
■690    ▼a0621
■690    ▼a0723
■690    ▼a0463
■71020▼aUniversity  of  California,  Los  Angeles▼bPsychology  0780.
■7730  ▼tDissertations  Abstracts  International▼g86-06A.
■790    ▼a0031
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17165121▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF13264 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.