본문

서브메뉴

Essays on Data Science: Computational Measurement for Learning and Teaching
Essays on Data Science: Computational Measurement for Learning and Teaching
Essays on Data Science: Computational Measurement for Learning and Teaching

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103600
ISBN  
9798280715646
DDC  
370
저자명  
Himmelsbach, Zachary.
서명/저자  
Essays on Data Science: Computational Measurement for Learning and Teaching
발행사항  
[Sl] : Harvard University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
140 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Advisor: Miratrix, Luke;Galvez, Sebastian Munoz-Najar.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2025.
초록/해제  
요약To study teaching and learning at a large scale, we must introduce new methods for the analysis of rich, unstructured data - such as audio, video, and transcribed text - from classrooms. In this dissertation, I develop and apply computational and statistical methods to measure teaching and learning processes captured in two unstructured sources: student writing and transcripts of teacher speech from classroom lessons. My first essay introduces Coupled Likelihood Estimation (CLE), a method that improves the precision of parameter estimates in models of unstructured data features while requiring fewer expert-labeled observations. It combines information from limited samples of expert-labeled data with larger samples of data with machine-predicted labels. CLE leverages the geometric structure of the joint likelihood from both identifying (labeled) and non-identifying (unlabeled) data, constraining parameter estimates to, approximately, a surface defined by the unlabeled data's likelihood. Simulations demonstrate that CLE is unbiased, reduces root mean squared error, and yields narrower confidence intervals compared to existing methods, in some cases effectively achieving average efficiency gains equivalent to doubling the expert-labeled sample size. An application estimating the effect of an educational intervention on student writing quality illustrates CLE's practical utility, producing estimates closer to an oracle benchmark using only 18% of the expert-labeled data. By amplifying the value of limited labeled data, CLE lowers barriers to high-quality inference in resource-constrained domains such as healthcare, education, and policy evaluation. The method's broad applicability, theoretical guarantees, and computational approach offer a pathway to cost-effective, reliable analyses in settings where researchers face high labeling costs.My second essay leverages natural language processing techniques to study the use of mathematical vocabulary in elementary math classrooms. My collaborators and I develop a rules-based computational measure of mathematical vocabulary use. We find that teachers differ substantially in the amount of mathematical vocabulary they model for their students. Students of teachers in the 75th percentile were exposed to 28 more mathematical terms per lesson (4,480 per year) than students of a teacher in the 25th percentile. Observed characteristics explain very little of this variation in teachers' mathematical vocabulary use. Finally, students randomly assigned to teachers' who used more mathematical vocabulary in previous years scored higher on standardized tests of mathematics. This implies that teachers who expose their students to more mathematical vocabulary are more effective teachers of mathematics. Across value-added studies, a teacher one standard deviation above the mean in effectiveness raises math scores by between .10 and .15 (Bacher-Hicks & Koedel, 2023); our estimate of the effect of being assigned to a teacher who uses one standard deviation more mathematical language accounts for roughly half of this variation, indicating that our measure is a powerful predictor of teacher effectiveness. In my third essay, I develop Contextual Value Separation (CVS), a general method for identifying words used differently between pre-specified subsets of documents in large text corpora. CVS achieves this by combining contextual embeddings with machine learning classifiers, permutation testing, and statistical adjustments for multiple comparisons. Whereas current methods identify words that predict membership within a given class of documents, CVS reveals cases where separate classes of the documents use the same word in differing ways. For example, experienced and novice math teachers may use a mathematical vocabulary term with similar frequency but in markedly different ways or contexts. This approach can search over a specified set of target words or over the entire vocabulary of the corpus. For each target word, CVS infers how consistently its contextual embeddings differ by subset. Because vocabularies are large, the method includes multiple testing correction to control the false-discovery rate, typically yielding a small set of words whose usage varies between the document classes. After identifying the words whose usage most consistently differs, example usages from each subset are extracted for qualitative examination. CVS easily extends to other forms of unstructured data represented by embeddings, such as video and audio. The method can be used as an exploratory tool for hypothesis generation, to test a priori hypotheses, or to detect treatment effects on textual outcomes in experimental settings. To demonstrate the method, I analyze a set of transcripts from upper elementary mathematics lessons and identify two ways that teachers with larger impacts on math scores use mathematical vocabulary differently: more use of the mathematical meanings of polysemous terms and more requests that students engage with questions related to the terms. The method can be easily extended to other forms of unstructured data can be encoded into vectors, e.g., audio and video. As a collection, these three essays reveal the promise of computational methods for enabling the analysis of text data (and other rich, unstructured data sources). They contribute several novel findings in the field of education regarding mathematical vocabulary and effective teaching. From a statistical point of view, CLE introduces a new way to leverage large amounts of machine labeled data, which, in addition to its value for educational research, can lower the cost of research in several domains, such as phenotyping electronic health records.
일반주제명  
Education
일반주제명  
Statistics
일반주제명  
Mathematics education
일반주제명  
Educational leadership
키워드  
Contextual Value Separation
키워드  
Mathematical vocabulary
키워드  
Expert-labeled observations
키워드  
Policy evaluation
키워드  
Natural language processing
기타저자  
Harvard University Education
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357785
■00520260202103600
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280715646
■035    ▼a(MiAaPQ)AAI32042401
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a370
■1001  ▼aHimmelsbach,  Zachary.▼0(orcid)0000-0002-5444-0648
■24510▼aEssays  on  Data  Science:  Computational  Measurement  for  Learning  and  Teaching
■260    ▼a[Sl]▼bHarvard  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a140  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aAdvisor:  Miratrix,  Luke;Galvez,  Sebastian  Munoz-Najar.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2025.
■520    ▼aTo  study  teaching  and  learning  at  a  large  scale,  we  must  introduce  new  methods  for  the  analysis  of  rich,  unstructured  data  -  such  as  audio,  video,  and  transcribed  text  -  from  classrooms.  In  this  dissertation,  I  develop  and  apply  computational  and  statistical  methods  to  measure  teaching  and  learning  processes  captured  in  two  unstructured  sources:  student  writing  and  transcripts  of  teacher  speech  from  classroom  lessons. My  first  essay  introduces  Coupled  Likelihood  Estimation  (CLE),  a  method  that  improves  the  precision  of  parameter  estimates  in  models  of  unstructured  data  features  while  requiring  fewer  expert-labeled  observations.  It  combines  information  from  limited  samples  of  expert-labeled  data  with  larger  samples  of  data  with  machine-predicted  labels.  CLE  leverages  the  geometric  structure  of  the  joint  likelihood  from  both  identifying  (labeled)  and  non-identifying  (unlabeled)  data,  constraining  parameter  estimates  to,  approximately,  a  surface  defined  by  the  unlabeled  data's  likelihood.  Simulations  demonstrate  that  CLE  is  unbiased,  reduces  root  mean  squared  error,  and  yields  narrower  confidence  intervals  compared  to  existing  methods,  in  some  cases  effectively  achieving  average  efficiency  gains  equivalent  to  doubling  the  expert-labeled  sample  size.  An  application  estimating  the  effect  of  an  educational  intervention  on  student  writing  quality  illustrates  CLE's  practical  utility,  producing  estimates  closer  to  an  oracle  benchmark  using  only  18%  of  the  expert-labeled  data.  By  amplifying  the  value  of  limited  labeled  data,  CLE  lowers  barriers  to  high-quality  inference  in  resource-constrained  domains  such  as  healthcare,  education,  and  policy  evaluation.  The  method's  broad  applicability,  theoretical  guarantees,  and  computational  approach  offer  a  pathway  to  cost-effective,  reliable  analyses  in  settings  where  researchers  face  high  labeling  costs.My  second  essay  leverages  natural  language  processing  techniques  to  study  the  use  of  mathematical  vocabulary  in  elementary  math  classrooms.  My  collaborators  and  I  develop  a  rules-based  computational  measure  of  mathematical  vocabulary  use.  We  find  that  teachers  differ  substantially  in  the  amount  of  mathematical  vocabulary  they  model  for  their  students.  Students  of  teachers  in  the  75th  percentile  were  exposed  to  28  more  mathematical  terms  per  lesson  (4,480  per  year)  than  students  of  a  teacher  in  the  25th  percentile.  Observed  characteristics  explain  very  little  of  this  variation  in  teachers'  mathematical  vocabulary  use.  Finally,  students  randomly  assigned  to  teachers'  who  used  more  mathematical  vocabulary  in  previous  years  scored  higher  on  standardized  tests  of  mathematics.  This  implies  that  teachers  who  expose  their  students  to  more  mathematical  vocabulary  are  more  effective  teachers  of  mathematics.  Across  value-added  studies,  a  teacher  one  standard  deviation  above  the  mean  in  effectiveness  raises  math  scores  by  between  .10  and  .15  (Bacher-Hicks  &  Koedel,  2023);  our  estimate  of  the  effect  of  being  assigned  to  a  teacher  who  uses  one  standard  deviation  more  mathematical  language  accounts  for  roughly  half  of  this  variation,  indicating  that  our  measure  is  a  powerful  predictor  of  teacher  effectiveness. In  my  third  essay,  I  develop  Contextual  Value  Separation  (CVS),  a  general  method  for  identifying  words  used  differently  between  pre-specified  subsets  of  documents  in  large  text  corpora.  CVS  achieves  this  by  combining  contextual  embeddings  with  machine  learning  classifiers,  permutation  testing,  and  statistical  adjustments  for  multiple  comparisons.  Whereas  current  methods  identify  words  that  predict  membership  within  a  given  class  of  documents,  CVS  reveals  cases  where  separate  classes  of  the  documents  use  the  same  word  in  differing  ways.  For  example,  experienced  and  novice  math  teachers  may  use  a  mathematical  vocabulary  term  with  similar  frequency  but  in  markedly  different  ways  or  contexts.  This  approach  can  search  over  a  specified  set  of  target  words  or  over  the  entire  vocabulary  of  the  corpus.  For  each  target  word,  CVS  infers  how  consistently  its  contextual  embeddings  differ  by  subset.  Because  vocabularies  are  large,  the  method  includes  multiple  testing  correction  to  control  the  false-discovery  rate,  typically  yielding  a  small  set  of  words  whose  usage  varies  between  the  document  classes.  After  identifying  the  words  whose  usage  most  consistently  differs,  example  usages  from  each  subset  are  extracted  for  qualitative  examination.  CVS  easily  extends  to  other  forms  of  unstructured  data  represented  by  embeddings,  such  as  video  and  audio.  The  method  can  be  used  as  an  exploratory  tool  for  hypothesis  generation,  to  test  a  priori  hypotheses,  or  to  detect  treatment  effects  on  textual  outcomes  in  experimental  settings.  To  demonstrate  the  method,  I  analyze  a  set  of  transcripts  from  upper  elementary  mathematics  lessons  and  identify  two  ways  that  teachers  with  larger  impacts  on  math  scores  use  mathematical  vocabulary  differently:  more  use  of  the  mathematical  meanings  of  polysemous  terms  and  more  requests  that  students  engage  with  questions  related  to  the  terms.  The  method  can  be  easily  extended  to  other  forms  of  unstructured  data  can  be  encoded  into  vectors,  e.g.,  audio  and  video. As  a  collection,  these  three  essays  reveal  the  promise  of  computational  methods  for  enabling  the  analysis  of  text  data  (and  other  rich,  unstructured  data  sources).  They  contribute  several  novel  findings  in  the  field  of  education  regarding  mathematical  vocabulary  and  effective  teaching.  From  a  statistical  point  of  view,  CLE  introduces  a  new  way  to  leverage  large  amounts  of  machine  labeled  data,  which,  in  addition  to  its  value  for  educational  research,  can  lower  the  cost  of  research  in  several  domains,  such  as  phenotyping  electronic  health  records.
■590    ▼aSchool  code:  0084.
■650  4▼aEducation
■650  4▼aStatistics
■650  4▼aMathematics  education
■650  4▼aEducational  leadership
■653    ▼aContextual  Value  Separation  
■653    ▼aMathematical  vocabulary
■653    ▼aExpert-labeled  observations
■653    ▼aPolicy  evaluation
■653    ▼aNatural  language  processing
■690    ▼a0515
■690    ▼a0463
■690    ▼a0449
■690    ▼a0280
■71020▼aHarvard  University▼bEducation.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357785▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF19374 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.