서브메뉴
검색
The Use of Large Language Models to Predict Item Properties
The Use of Large Language Models to Predict Item Properties
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153034
- ISBN
- 9798346851141
- DDC
- 379.1
- 저자명
- Smart, Francis.
- 서명/저자
- The Use of Large Language Models to Predict Item Properties
- 발행사항
- [Sl] : Michigan State University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 162 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-06, Section: B.
- 주기사항
- Advisor: Kelly, Kimberly.
- 학위논문주기
- Thesis (Ph.D.)--Michigan State University, 2024.
- 초록/해제
- 요약Calibrating items is a crucial yet costly requirement for both new tests and existing ones as items become outdated due to changing relevance or overexposure. Traditionally, this calibration involves giving items to a large number of participants, a process that requires substantial time and resources. To reduce these costs, researchers have sought alternative calibration methods. Before the emergence of Large Language Models (LLMs), these methods mainly relied on expert opinions or computational analysis of item features. Yet, the accuracy of experts in predicting item performance has varied, and computational approaches often struggle to capture the intricate semantic details of test items.The emergence of LLMs might offer a new avenue of addressing the need for item calibration. These models, popularized by OpenAI (like the GPT series), have shown remarkable abilities in mimicking complex human thought processes, and performing advanced reasoning tasks. Their achievements in passing sophisticated exams and executing cross-language translations underline their potential. However, their capacity for predicting item properties in test calibration has not been thoroughly investigated. Traditional calibration relies heavily on direct human interaction, such as pretesting and expert assessment, or on statistical modeling of item features through resource intensive machine learning algorithms. This dissertation explores the potential of LLMs to predict item characteristics, tasks that have traditionally required human insight or complex statistical models. With the increasing accessibility of high-performance LLMs from organizations like OpenAI, Meta, and Google, and through open-source platforms such as HuggingFace.com, there is promising ground for investigation. This study examines whether LLMs could replace human efforts in item calibration tasks.To evaluate the effectiveness of LLMs in predicting item properties, this dissertation implements a training and testing framework, focusing on assessing both the relative and absolute difficulties of items. It undertakes three theoretical investigations: firstly, examining the ability of LLMs to predict the relative difficulty of items; secondly, assessing the feasibility of using multiple LLMs as substitutes for test-takers and attempts to use their responses predictors of item difficulty; and thirdly, applying a search algorithm, guided by LLM predictions of relative difficulty, to ascertain absolute difficulties.The findings indicate that the models have statistical significance in predicting relative item difficulty, limited by modest explanatory power - with adjusted R-squared values around 5-10%. However, the application of LLMs in predicting relative item difficulties through pairwise comparisons proves to be more promising, achieving a pairwise accuracy of about 62% and demonstrating predicted correlations with item difficulty ranging between 0.36 and 0.42.This suggests that whereas LLMs show potential in certain aspects of item calibration, their effectiveness varies depending on the specific task. This demonstrates a potential promising result that warrants further exploration into the capabilities of LLMs for item calibration, potentially leading to more efficient and cost-effective methods in the field of test development and maintenance.
- 일반주제명
- Educational evaluation
- 키워드
- Item calibration
- 키워드
- Pre-calibration
- 키워드
- Test development
- 기타저자
- Michigan State University Measurement and Quantitative Methods - Doctor of Philosophy
- 기본자료저록
- Dissertations Abstracts International. 86-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164708
■00520250211153034
■006m o d
■007cr#unu||||||||
■020 ▼a9798346851141
■035 ▼a(MiAaPQ)AAI31637273
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a379.1
■1001 ▼aSmart, Francis.▼0(orcid)0000-0002-3414-7786
■24510▼aThe Use of Large Language Models to Predict Item Properties
■260 ▼a[Sl]▼bMichigan State University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a162 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-06, Section: B.
■500 ▼aAdvisor: Kelly, Kimberly.
■5021 ▼aThesis (Ph.D.)--Michigan State University, 2024.
■520 ▼aCalibrating items is a crucial yet costly requirement for both new tests and existing ones as items become outdated due to changing relevance or overexposure. Traditionally, this calibration involves giving items to a large number of participants, a process that requires substantial time and resources. To reduce these costs, researchers have sought alternative calibration methods. Before the emergence of Large Language Models (LLMs), these methods mainly relied on expert opinions or computational analysis of item features. Yet, the accuracy of experts in predicting item performance has varied, and computational approaches often struggle to capture the intricate semantic details of test items.The emergence of LLMs might offer a new avenue of addressing the need for item calibration. These models, popularized by OpenAI (like the GPT series), have shown remarkable abilities in mimicking complex human thought processes, and performing advanced reasoning tasks. Their achievements in passing sophisticated exams and executing cross-language translations underline their potential. However, their capacity for predicting item properties in test calibration has not been thoroughly investigated. Traditional calibration relies heavily on direct human interaction, such as pretesting and expert assessment, or on statistical modeling of item features through resource intensive machine learning algorithms. This dissertation explores the potential of LLMs to predict item characteristics, tasks that have traditionally required human insight or complex statistical models. With the increasing accessibility of high-performance LLMs from organizations like OpenAI, Meta, and Google, and through open-source platforms such as HuggingFace.com, there is promising ground for investigation. This study examines whether LLMs could replace human efforts in item calibration tasks.To evaluate the effectiveness of LLMs in predicting item properties, this dissertation implements a training and testing framework, focusing on assessing both the relative and absolute difficulties of items. It undertakes three theoretical investigations: firstly, examining the ability of LLMs to predict the relative difficulty of items; secondly, assessing the feasibility of using multiple LLMs as substitutes for test-takers and attempts to use their responses predictors of item difficulty; and thirdly, applying a search algorithm, guided by LLM predictions of relative difficulty, to ascertain absolute difficulties.The findings indicate that the models have statistical significance in predicting relative item difficulty, limited by modest explanatory power - with adjusted R-squared values around 5-10%. However, the application of LLMs in predicting relative item difficulties through pairwise comparisons proves to be more promising, achieving a pairwise accuracy of about 62% and demonstrating predicted correlations with item difficulty ranging between 0.36 and 0.42.This suggests that whereas LLMs show potential in certain aspects of item calibration, their effectiveness varies depending on the specific task. This demonstrates a potential promising result that warrants further exploration into the capabilities of LLMs for item calibration, potentially leading to more efficient and cost-effective methods in the field of test development and maintenance.
■590 ▼aSchool code: 0128.
■650 4▼aEducational evaluation
■650 4▼aEducational tests & measurements
■653 ▼aExpert assessment
■653 ▼aItem calibration
■653 ▼aLarge language models
■653 ▼aPre-calibration
■653 ▼aTest development
■690 ▼a0443
■690 ▼a0288
■690 ▼a0800
■71020▼aMichigan State University▼bMeasurement and Quantitative Methods - Doctor of Philosophy.
■7730 ▼tDissertations Abstracts International▼g86-06B.
■790 ▼a0128
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164708▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


