서브메뉴
검색
Data Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining
Data Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103110
- ISBN
- 9798280756434
- DDC
- 370
- 저자명
- Slater, Stefan.
- 서명/저자
- Data Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining
- 발행사항
- [Sl] : University of Pennsylvania, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 99 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
- 주기사항
- Advisor: Baker, Ryan S.
- 학위논문주기
- Thesis (Ph.D.)--University of Pennsylvania, 2025.
- 초록/해제
- 요약The purpose of this dissertation was to determine the necessary amount of data to generate reliable, generalizable, and replicable machine learning models for educational contexts. Algorithms are ubiquitous across a range of educational settings and used to detect or predict an increasing number of student performance metrics, conceptualizations, and behaviors. But determining the correct amount of data to use for the construction and use of these algorithms often comes down to 'rules of thumb' rather than empirically generated benchmarks. The first study explored the amount of data necessary to generate stable predictions of student knowledge using the Bayesian Knowledge Tracing (BKT) algorithm, while the second study explored the differences in algorithm performance on predicting student stopout behavior in real data from Massive Open Online Courses (MOOCs). In both studies, subsets of data of varying sizes were taken from a larger overall dataset, and model performance on these subsets was compared to the performance of a model that used all available data. Findings from Study 1 showed that BKT is able to generate good predictions of student mastery at sample sizes as low as 25 and assessments as short as three problems, while findings from Study 2 showed that sample sizes of around 500 are suitable for more complex prediction tasks like stopout behaviors. In discussing future avenues for research, this dissertation explained how particular characteristics of a dataset, such as its homogeneity or the quality of features used for the analysis, could further influence the performance of models alongside sample size exclusively. The conclusion also examined whether the sample size requirements in this work can be applied to under-studied populations and demographics, in order to ensure that a dataset has sufficient representation for these populations when modeling and prediction tasks are undertaken.
- 일반주제명
- Education
- 일반주제명
- Educational technology
- 일반주제명
- Information science
- 키워드
- Data science
- 키워드
- Machine learning
- 기타저자
- University of Pennsylvania Education
- 기본자료저록
- Dissertations Abstracts International. 86-12A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356972
■00520260202103110
■006m o d
■007cr#unu||||||||
■020 ▼a9798280756434
■035 ▼a(MiAaPQ)AAI31936028
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a370
■1001 ▼aSlater, Stefan.
■24510▼aData Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining
■260 ▼a[Sl]▼bUniversity of Pennsylvania▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a99 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: A.
■500 ▼aAdvisor: Baker, Ryan S.
■5021 ▼aThesis (Ph.D.)--University of Pennsylvania, 2025.
■520 ▼aThe purpose of this dissertation was to determine the necessary amount of data to generate reliable, generalizable, and replicable machine learning models for educational contexts. Algorithms are ubiquitous across a range of educational settings and used to detect or predict an increasing number of student performance metrics, conceptualizations, and behaviors. But determining the correct amount of data to use for the construction and use of these algorithms often comes down to 'rules of thumb' rather than empirically generated benchmarks. The first study explored the amount of data necessary to generate stable predictions of student knowledge using the Bayesian Knowledge Tracing (BKT) algorithm, while the second study explored the differences in algorithm performance on predicting student stopout behavior in real data from Massive Open Online Courses (MOOCs). In both studies, subsets of data of varying sizes were taken from a larger overall dataset, and model performance on these subsets was compared to the performance of a model that used all available data. Findings from Study 1 showed that BKT is able to generate good predictions of student mastery at sample sizes as low as 25 and assessments as short as three problems, while findings from Study 2 showed that sample sizes of around 500 are suitable for more complex prediction tasks like stopout behaviors. In discussing future avenues for research, this dissertation explained how particular characteristics of a dataset, such as its homogeneity or the quality of features used for the analysis, could further influence the performance of models alongside sample size exclusively. The conclusion also examined whether the sample size requirements in this work can be applied to under-studied populations and demographics, in order to ensure that a dataset has sufficient representation for these populations when modeling and prediction tasks are undertaken.
■590 ▼aSchool code: 0175.
■650 4▼aEducation
■650 4▼aEducational technology
■650 4▼aInformation science
■653 ▼aData science
■653 ▼aEducational Data Mining
■653 ▼aBayesian Knowledge Tracing
■653 ▼aLearning analytics
■653 ▼aMachine learning
■653 ▼aPredictive modeling
■690 ▼a0515
■690 ▼a0710
■690 ▼a0723
■71020▼aUniversity of Pennsylvania▼bEducation.
■7730 ▼tDissertations Abstracts International▼g86-12A.
■790 ▼a0175
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356972▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


