서브메뉴
검색
A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods
A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153126
- ISBN
- 9798346859673
- DDC
- 151
- 서명/저자
- A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods
- 발행사항
- [Sl] : University of California, Los Angeles, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 151 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-06, Section: A.
- 주기사항
- Advisor: Montoya, Amanda K.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2024.
- 초록/해제
- 요약As machine learning techniques increase in popularity among psychology researchers, decision trees---and their offshoots such as random forests and qualitative interaction trees (QUINT)---have received special attention due to their interpretability and ease of use. Although these tree-based methods are versatile in terms of the types and number of relationships they can model, missingness still needs to be addressed before they can produce predictions. The current literature comparing missing data methods available for tree-based models is fragmented, with only subsets of conditions or methods examined in each study. Furthermore, the application of missing data methods to the specialized tree-based algorithm QUINT has largely gone uninvestigated in previous research. Thus, there is a great need to clarify which missing data methods are best to use in which situations, especially with niche tree-based models like QUINT. Since tree-based methods are quick to set up and easy to interpret, it is particularly important that users who may not be experienced in machine learning or statistics receive guidance on how they should manage missingness in their data before they apply such methods. In order to consolidate the research that has already been done on missing data methods with tree-based models, this dissertation provides a summary of the existing literature on the topic, introducing the methods and factors that are important to consider when choosing how to deal with missingness in datasets for tree models. Also, to understand which missing data methods are currently applied to tree-based models in psychological studies and under which conditions, recent substantive research articles were reviewed and information about their datasets and methodologies were recorded. Finally, to extend the knowledge accumulated in the literature, a plethora of popular/modern missing data methods for tree models were applied to the QUINT algorithm and compared in a simulation study, varying factors such as the amount of missingness, the type of missingness, and where the missingness appeared in the data to cover a variety of scenarios that may arise in the real world. The results reveal that, in terms of both prediction accuracy and variable selection, imputation methods---specifically regression and hot deck imputation---and missingness incorporated in attributes (MIA) are able to address missing data and produce QUINT models of consistently high quality. The complete case method, on the other hand, should be avoided due to its highly variable performance and inconsistent nature, leading to models that differ greatly from would have been produced had the data been complete.
- 일반주제명
- Quantitative psychology
- 일반주제명
- Statistics
- 일반주제명
- Psychology
- 일반주제명
- Information science
- 키워드
- Decision trees
- 키워드
- Missing data
- 키워드
- Machine learning
- 기타저자
- University of California, Los Angeles Psychology 0780
- 기본자료저록
- Dissertations Abstracts International. 86-06A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017165121
■00520250211153126
■006m o d
■007cr#unu||||||||
■020 ▼a9798346859673
■035 ▼a(MiAaPQ)AAI31765212
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a151
■1001 ▼aTibbe, Tristan Dale.
■24512▼aA Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a151 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-06, Section: A.
■500 ▼aAdvisor: Montoya, Amanda K.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2024.
■520 ▼aAs machine learning techniques increase in popularity among psychology researchers, decision trees---and their offshoots such as random forests and qualitative interaction trees (QUINT)---have received special attention due to their interpretability and ease of use. Although these tree-based methods are versatile in terms of the types and number of relationships they can model, missingness still needs to be addressed before they can produce predictions. The current literature comparing missing data methods available for tree-based models is fragmented, with only subsets of conditions or methods examined in each study. Furthermore, the application of missing data methods to the specialized tree-based algorithm QUINT has largely gone uninvestigated in previous research. Thus, there is a great need to clarify which missing data methods are best to use in which situations, especially with niche tree-based models like QUINT. Since tree-based methods are quick to set up and easy to interpret, it is particularly important that users who may not be experienced in machine learning or statistics receive guidance on how they should manage missingness in their data before they apply such methods. In order to consolidate the research that has already been done on missing data methods with tree-based models, this dissertation provides a summary of the existing literature on the topic, introducing the methods and factors that are important to consider when choosing how to deal with missingness in datasets for tree models. Also, to understand which missing data methods are currently applied to tree-based models in psychological studies and under which conditions, recent substantive research articles were reviewed and information about their datasets and methodologies were recorded. Finally, to extend the knowledge accumulated in the literature, a plethora of popular/modern missing data methods for tree models were applied to the QUINT algorithm and compared in a simulation study, varying factors such as the amount of missingness, the type of missingness, and where the missingness appeared in the data to cover a variety of scenarios that may arise in the real world. The results reveal that, in terms of both prediction accuracy and variable selection, imputation methods---specifically regression and hot deck imputation---and missingness incorporated in attributes (MIA) are able to address missing data and produce QUINT models of consistently high quality. The complete case method, on the other hand, should be avoided due to its highly variable performance and inconsistent nature, leading to models that differ greatly from would have been produced had the data been complete.
■590 ▼aSchool code: 0031.
■650 4▼aQuantitative psychology
■650 4▼aStatistics
■650 4▼aPsychology
■650 4▼aInformation science
■653 ▼aDecision trees
■653 ▼aMissing data
■653 ▼aQualitative interaction trees
■653 ▼aTree-based methods
■653 ▼aMachine learning
■690 ▼a0632
■690 ▼a0621
■690 ▼a0723
■690 ▼a0463
■71020▼aUniversity of California, Los Angeles▼bPsychology 0780.
■7730 ▼tDissertations Abstracts International▼g86-06A.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17165121▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


