본문

서브메뉴

Towards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems for African American English-Speaking Children
Towards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems fo...
Towards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems for African American English-Speaking Children

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151038
ISBN  
9798381953404
DDC  
621.3
저자명  
Johnson, Alexander.
서명/저자  
Towards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems for African American English-Speaking Children
발행사항  
[Sl] : University of California, Los Angeles, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
119 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-09, Section: A.
주기사항  
Advisor: Alwan, Abeer A.
학위논문주기  
Thesis (Ph.D.)--University of California, Los Angeles, 2024.
초록/해제  
요약The potential of speech technology to improve educational outcomes has been a topic of great interest in recent years. For example, automatic speech recognition (ASR) systems could be employed to provide kindergarten-aged children with real-time feedback on their literacy and pronunciation as they practice reading aloud. Within these systems, speaker identification (SID) technology could additionally be used to identify the user's speaker characteristics in order to ensure that they receive age, language, and dialect-appropriate feedback. While these technologies are more established for well-represented groups in STEM (ie. able-bodied, adult, first-language speakers of mainstream dialects), they give much worse performance for underrepresented groups (young children, speakers of non-mainstream dialects, people with speech-related disabilities, etc.). This work focuses on improving speech technology performance for children's speech and African American English (AAE) dialect speech with the goal of creating more equitable outcomes in early education. The contributions of this work span three primary areas: 1) Dialect identification and density scoring, 2) data augmentation for speech recognition, and 3) Natural Language Processing for fair and inclusive automatic speech assessment.First, we create a robust system for dialect identification of African American English for both children and adult's speech. This system aims to take an input utterance from a speaker of either African American English or Mainstream American English and determine which of the two dialects the utterance belongs. The system fuses features from paralinguistics, self-supervised learning representations, automatic speech recognition system outputs, prosodic contours, and other descriptors of the speech signal in order to learn a mapping from the input acoustic information to a dialect classification decision. We further explore this architecture in automatic dialect density estimation, a task we create and develop. In dialect density scoring, we train a system to automatically predict a speaker's frequency of usage of dialect-specific patterns. This information can then be passed to a speech recognition system for more dialect-informed processing.Second, we develop a data augmentation algorithm to improve zero-shot and few-shot speech recognition of low-resource dialects. The algorithm, named LPCAugment, deconstructs an input speech signal into a source and filter representation using linear predictive coding (LPC) analysis. The poles of the filter representation can then be perturbed independently of the source representation in order to model formant shifts that may be seen across accents and dialects. We use this perturbation method to artificially generate speech samples with shifted formant locations to serve as additional training data for a speech recognition system. This speech recognition system is then evaluated on children's speech for child speakers of a Southern California dialect and child speakers of an Atlanta, Georgia, area dialect.Third, we explore automatic analysis and scoring of speech recognition transcripts for educational assessments. Given information about a student's spoken dialect and automatically generated transcripts of their oral response to an assessment prompt, we train a system to automatically grade the quality of the response with respect to a pre-determined criterion. This system uses language modeling and spoken information retrieval to identify key features in the spoken response and holistically decide if the response aligns with the grading criteria. Combined, the steps in this work form a framework for inclusive spoken language understanding technology that can be used to perform provide students with dialect-appropriate language training or language assessment.
일반주제명  
Electrical engineering
일반주제명  
Educational technology
일반주제명  
Linguistics
일반주제명  
African American studies
일반주제명  
Communication
키워드  
Speech technology
키워드  
Automatic speech recognition
키워드  
Speaker identification
키워드  
Linear predictive coding
키워드  
African American English
기타저자  
University of California, Los Angeles Electrical and Computer Engineering 0333
기본자료저록  
Dissertations Abstracts International. 85-09A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017160550
■00520250211151038
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798381953404
■035    ▼a(MiAaPQ)AAI31139575
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aJohnson,  Alexander.
■24510▼aTowards  Inclusive  Low-Resource  Speech  Technologies:  A  Case  Study  of  Educational  Systems  for  African  American  English-Speaking  Children
■260    ▼a[Sl]▼bUniversity  of  California,  Los  Angeles▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a119  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-09,  Section:  A.
■500    ▼aAdvisor:  Alwan,  Abeer  A.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Los  Angeles,  2024.
■520    ▼aThe  potential  of  speech  technology  to  improve  educational  outcomes  has  been  a  topic  of  great  interest  in  recent  years.    For  example,  automatic  speech  recognition  (ASR)  systems  could  be  employed  to  provide  kindergarten-aged  children  with  real-time  feedback  on  their  literacy  and  pronunciation  as  they  practice  reading  aloud.    Within  these  systems,  speaker  identification  (SID)  technology  could  additionally  be  used  to  identify  the  user's  speaker  characteristics  in  order  to  ensure  that  they  receive  age,  language,  and  dialect-appropriate  feedback.    While  these  technologies  are  more  established  for  well-represented  groups  in  STEM  (ie.  able-bodied,  adult,  first-language  speakers  of  mainstream  dialects),  they  give  much  worse  performance  for  underrepresented  groups  (young  children,  speakers  of  non-mainstream  dialects,  people  with  speech-related  disabilities,  etc.).    This  work  focuses  on  improving  speech  technology  performance  for  children's  speech  and  African  American  English  (AAE)  dialect  speech  with  the  goal  of  creating  more  equitable  outcomes  in  early  education.    The  contributions  of  this  work  span  three  primary  areas:  1)  Dialect  identification  and  density  scoring,  2)  data  augmentation  for  speech  recognition,  and  3)  Natural  Language  Processing  for  fair  and  inclusive  automatic  speech  assessment.First,  we  create  a  robust  system  for  dialect  identification  of  African  American  English  for  both  children  and  adult's  speech.    This  system  aims  to  take  an  input  utterance  from  a  speaker  of  either  African  American  English  or  Mainstream  American  English  and  determine  which  of  the  two  dialects  the  utterance  belongs.    The  system  fuses  features  from  paralinguistics,  self-supervised  learning  representations,  automatic  speech  recognition  system  outputs,  prosodic  contours,  and  other  descriptors  of  the  speech  signal  in  order  to  learn  a  mapping  from  the  input  acoustic  information  to  a  dialect  classification  decision.    We  further  explore  this  architecture  in  automatic  dialect  density  estimation,  a  task  we  create  and  develop.    In  dialect  density  scoring,  we  train  a  system  to  automatically  predict  a  speaker's  frequency  of  usage  of  dialect-specific  patterns.    This  information  can  then  be  passed  to  a  speech  recognition  system  for  more  dialect-informed  processing.Second,  we  develop  a  data  augmentation  algorithm  to  improve  zero-shot  and  few-shot  speech  recognition  of  low-resource  dialects.    The  algorithm,  named  LPCAugment,  deconstructs  an  input  speech  signal  into  a  source  and  filter  representation  using  linear  predictive  coding  (LPC)  analysis.    The  poles  of  the  filter  representation  can  then  be  perturbed  independently  of  the  source  representation  in  order  to  model  formant  shifts  that  may  be  seen  across  accents  and  dialects.    We  use  this  perturbation  method  to  artificially  generate  speech  samples  with  shifted  formant  locations  to  serve  as  additional  training  data  for  a  speech  recognition  system.    This  speech  recognition  system  is  then  evaluated  on  children's  speech  for  child  speakers  of  a  Southern  California  dialect  and  child  speakers  of  an  Atlanta,  Georgia,  area  dialect.Third,  we  explore  automatic  analysis  and  scoring  of  speech  recognition  transcripts  for  educational  assessments.    Given  information  about  a  student's  spoken  dialect  and  automatically  generated  transcripts  of  their  oral  response  to  an  assessment  prompt,  we  train  a  system  to  automatically  grade  the  quality  of  the  response  with  respect  to  a  pre-determined  criterion.    This  system  uses  language  modeling  and  spoken  information  retrieval  to  identify  key  features  in  the  spoken  response  and  holistically  decide  if  the  response  aligns  with  the  grading  criteria.    Combined,  the  steps  in  this  work  form  a  framework  for  inclusive  spoken  language  understanding  technology  that  can  be  used  to  perform  provide  students  with  dialect-appropriate  language  training  or  language  assessment.
■590    ▼aSchool  code:  0031.
■650  4▼aElectrical  engineering
■650  4▼aEducational  technology
■650  4▼aLinguistics
■650  4▼aAfrican  American  studies
■650  4▼aCommunication
■653    ▼aSpeech  technology
■653    ▼aAutomatic  speech  recognition
■653    ▼aSpeaker  identification
■653    ▼aLinear  predictive  coding
■653    ▼aAfrican  American  English
■690    ▼a0544
■690    ▼a0459
■690    ▼a0290
■690    ▼a0710
■690    ▼a0296
■71020▼aUniversity  of  California,  Los  Angeles▼bElectrical  and  Computer  Engineering  0333.
■7730  ▼tDissertations  Abstracts  International▼g85-09A.
■790    ▼a0031
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160550▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11197 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.