본문

서브메뉴

Data Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining
Data Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining
Data Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103110
ISBN  
9798280756434
DDC  
370
저자명  
Slater, Stefan.
서명/저자  
Data Quantity and Data Characteristics for Modeling Approaches in Educational Data Mining
발행사항  
[Sl] : University of Pennsylvania, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
99 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
주기사항  
Advisor: Baker, Ryan S.
학위논문주기  
Thesis (Ph.D.)--University of Pennsylvania, 2025.
초록/해제  
요약The purpose of this dissertation was to determine the necessary amount of data to generate reliable, generalizable, and replicable machine learning models for educational contexts. Algorithms are ubiquitous across a range of educational settings and used to detect or predict an increasing number of student performance metrics, conceptualizations, and behaviors. But determining the correct amount of data to use for the construction and use of these algorithms often comes down to 'rules of thumb' rather than empirically generated benchmarks. The first study explored the amount of data necessary to generate stable predictions of student knowledge using the Bayesian Knowledge Tracing (BKT) algorithm, while the second study explored the differences in algorithm performance on predicting student stopout behavior in real data from Massive Open Online Courses (MOOCs). In both studies, subsets of data of varying sizes were taken from a larger overall dataset, and model performance on these subsets was compared to the performance of a model that used all available data. Findings from Study 1 showed that BKT is able to generate good predictions of student mastery at sample sizes as low as 25 and assessments as short as three problems, while findings from Study 2 showed that sample sizes of around 500 are suitable for more complex prediction tasks like stopout behaviors. In discussing future avenues for research, this dissertation explained how particular characteristics of a dataset, such as its homogeneity or the quality of features used for the analysis, could further influence the performance of models alongside sample size exclusively. The conclusion also examined whether the sample size requirements in this work can be applied to under-studied populations and demographics, in order to ensure that a dataset has sufficient representation for these populations when modeling and prediction tasks are undertaken.
일반주제명  
Education
일반주제명  
Educational technology
일반주제명  
Information science
키워드  
Data science
키워드  
Educational Data Mining
키워드  
Bayesian Knowledge Tracing
키워드  
Learning analytics
키워드  
Machine learning
키워드  
Predictive modeling
기타저자  
University of Pennsylvania Education
기본자료저록  
Dissertations Abstracts International. 86-12A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017356972
■00520260202103110
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280756434
■035    ▼a(MiAaPQ)AAI31936028
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a370
■1001  ▼aSlater,  Stefan.
■24510▼aData  Quantity  and  Data  Characteristics  for  Modeling  Approaches  in  Educational  Data  Mining
■260    ▼a[Sl]▼bUniversity  of  Pennsylvania▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a99  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  A.
■500    ▼aAdvisor:  Baker,  Ryan  S.
■5021  ▼aThesis  (Ph.D.)--University  of  Pennsylvania,  2025.
■520    ▼aThe  purpose  of  this  dissertation  was  to  determine  the  necessary  amount  of  data  to  generate  reliable,  generalizable,  and  replicable  machine  learning  models  for  educational  contexts.  Algorithms  are  ubiquitous  across  a  range  of  educational  settings  and  used  to  detect  or  predict  an  increasing  number  of  student  performance  metrics,  conceptualizations,  and  behaviors.  But  determining  the  correct  amount  of  data  to  use  for  the  construction  and  use  of  these  algorithms  often  comes  down  to  'rules  of  thumb'  rather  than  empirically  generated  benchmarks.  The  first  study  explored  the  amount  of  data  necessary  to  generate  stable  predictions  of  student  knowledge  using  the  Bayesian  Knowledge  Tracing  (BKT)  algorithm,  while  the  second  study  explored  the  differences  in  algorithm  performance  on  predicting  student  stopout  behavior  in  real  data  from  Massive  Open  Online  Courses  (MOOCs).  In  both  studies,  subsets  of  data  of  varying  sizes  were  taken  from  a  larger  overall  dataset,  and  model  performance  on  these  subsets  was  compared  to  the  performance  of  a  model  that  used  all  available  data.  Findings  from  Study  1  showed  that  BKT  is  able  to  generate  good  predictions  of  student  mastery  at  sample  sizes  as  low  as  25  and  assessments  as  short  as  three  problems,  while  findings  from  Study  2  showed  that  sample  sizes  of  around  500  are  suitable  for  more  complex  prediction  tasks  like  stopout  behaviors.  In  discussing  future  avenues  for  research,  this  dissertation  explained  how  particular  characteristics  of  a  dataset,  such  as  its  homogeneity  or  the  quality  of  features  used  for  the  analysis,  could  further  influence  the  performance  of  models  alongside  sample  size  exclusively.  The  conclusion  also  examined  whether  the  sample  size  requirements  in  this  work  can  be  applied  to  under-studied  populations  and  demographics,  in  order  to  ensure  that  a  dataset  has  sufficient  representation  for  these  populations  when  modeling  and  prediction  tasks  are  undertaken.
■590    ▼aSchool  code:  0175.
■650  4▼aEducation
■650  4▼aEducational  technology
■650  4▼aInformation  science
■653    ▼aData  science
■653    ▼aEducational  Data  Mining
■653    ▼aBayesian  Knowledge  Tracing
■653    ▼aLearning  analytics
■653    ▼aMachine  learning
■653    ▼aPredictive  modeling
■690    ▼a0515
■690    ▼a0710
■690    ▼a0723
■71020▼aUniversity  of  Pennsylvania▼bEducation.
■7730  ▼tDissertations  Abstracts  International▼g86-12A.
■790    ▼a0175
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356972▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF14845 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.