서브메뉴
검색
Towards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems for African American English-Speaking Children
Towards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems for African American English-Speaking Children
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151038
- ISBN
- 9798381953404
- DDC
- 621.3
- 서명/저자
- Towards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems for African American English-Speaking Children
- 발행사항
- [Sl] : University of California, Los Angeles, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 119 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-09, Section: A.
- 주기사항
- Advisor: Alwan, Abeer A.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2024.
- 초록/해제
- 요약The potential of speech technology to improve educational outcomes has been a topic of great interest in recent years. For example, automatic speech recognition (ASR) systems could be employed to provide kindergarten-aged children with real-time feedback on their literacy and pronunciation as they practice reading aloud. Within these systems, speaker identification (SID) technology could additionally be used to identify the user's speaker characteristics in order to ensure that they receive age, language, and dialect-appropriate feedback. While these technologies are more established for well-represented groups in STEM (ie. able-bodied, adult, first-language speakers of mainstream dialects), they give much worse performance for underrepresented groups (young children, speakers of non-mainstream dialects, people with speech-related disabilities, etc.). This work focuses on improving speech technology performance for children's speech and African American English (AAE) dialect speech with the goal of creating more equitable outcomes in early education. The contributions of this work span three primary areas: 1) Dialect identification and density scoring, 2) data augmentation for speech recognition, and 3) Natural Language Processing for fair and inclusive automatic speech assessment.First, we create a robust system for dialect identification of African American English for both children and adult's speech. This system aims to take an input utterance from a speaker of either African American English or Mainstream American English and determine which of the two dialects the utterance belongs. The system fuses features from paralinguistics, self-supervised learning representations, automatic speech recognition system outputs, prosodic contours, and other descriptors of the speech signal in order to learn a mapping from the input acoustic information to a dialect classification decision. We further explore this architecture in automatic dialect density estimation, a task we create and develop. In dialect density scoring, we train a system to automatically predict a speaker's frequency of usage of dialect-specific patterns. This information can then be passed to a speech recognition system for more dialect-informed processing.Second, we develop a data augmentation algorithm to improve zero-shot and few-shot speech recognition of low-resource dialects. The algorithm, named LPCAugment, deconstructs an input speech signal into a source and filter representation using linear predictive coding (LPC) analysis. The poles of the filter representation can then be perturbed independently of the source representation in order to model formant shifts that may be seen across accents and dialects. We use this perturbation method to artificially generate speech samples with shifted formant locations to serve as additional training data for a speech recognition system. This speech recognition system is then evaluated on children's speech for child speakers of a Southern California dialect and child speakers of an Atlanta, Georgia, area dialect.Third, we explore automatic analysis and scoring of speech recognition transcripts for educational assessments. Given information about a student's spoken dialect and automatically generated transcripts of their oral response to an assessment prompt, we train a system to automatically grade the quality of the response with respect to a pre-determined criterion. This system uses language modeling and spoken information retrieval to identify key features in the spoken response and holistically decide if the response aligns with the grading criteria. Combined, the steps in this work form a framework for inclusive spoken language understanding technology that can be used to perform provide students with dialect-appropriate language training or language assessment.
- 일반주제명
- Electrical engineering
- 일반주제명
- Educational technology
- 일반주제명
- Linguistics
- 일반주제명
- African American studies
- 일반주제명
- Communication
- 기타저자
- University of California, Los Angeles Electrical and Computer Engineering 0333
- 기본자료저록
- Dissertations Abstracts International. 85-09A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160550
■00520250211151038
■006m o d
■007cr#unu||||||||
■020 ▼a9798381953404
■035 ▼a(MiAaPQ)AAI31139575
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aJohnson, Alexander.
■24510▼aTowards Inclusive Low-Resource Speech Technologies: A Case Study of Educational Systems for African American English-Speaking Children
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a119 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-09, Section: A.
■500 ▼aAdvisor: Alwan, Abeer A.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2024.
■520 ▼aThe potential of speech technology to improve educational outcomes has been a topic of great interest in recent years. For example, automatic speech recognition (ASR) systems could be employed to provide kindergarten-aged children with real-time feedback on their literacy and pronunciation as they practice reading aloud. Within these systems, speaker identification (SID) technology could additionally be used to identify the user's speaker characteristics in order to ensure that they receive age, language, and dialect-appropriate feedback. While these technologies are more established for well-represented groups in STEM (ie. able-bodied, adult, first-language speakers of mainstream dialects), they give much worse performance for underrepresented groups (young children, speakers of non-mainstream dialects, people with speech-related disabilities, etc.). This work focuses on improving speech technology performance for children's speech and African American English (AAE) dialect speech with the goal of creating more equitable outcomes in early education. The contributions of this work span three primary areas: 1) Dialect identification and density scoring, 2) data augmentation for speech recognition, and 3) Natural Language Processing for fair and inclusive automatic speech assessment.First, we create a robust system for dialect identification of African American English for both children and adult's speech. This system aims to take an input utterance from a speaker of either African American English or Mainstream American English and determine which of the two dialects the utterance belongs. The system fuses features from paralinguistics, self-supervised learning representations, automatic speech recognition system outputs, prosodic contours, and other descriptors of the speech signal in order to learn a mapping from the input acoustic information to a dialect classification decision. We further explore this architecture in automatic dialect density estimation, a task we create and develop. In dialect density scoring, we train a system to automatically predict a speaker's frequency of usage of dialect-specific patterns. This information can then be passed to a speech recognition system for more dialect-informed processing.Second, we develop a data augmentation algorithm to improve zero-shot and few-shot speech recognition of low-resource dialects. The algorithm, named LPCAugment, deconstructs an input speech signal into a source and filter representation using linear predictive coding (LPC) analysis. The poles of the filter representation can then be perturbed independently of the source representation in order to model formant shifts that may be seen across accents and dialects. We use this perturbation method to artificially generate speech samples with shifted formant locations to serve as additional training data for a speech recognition system. This speech recognition system is then evaluated on children's speech for child speakers of a Southern California dialect and child speakers of an Atlanta, Georgia, area dialect.Third, we explore automatic analysis and scoring of speech recognition transcripts for educational assessments. Given information about a student's spoken dialect and automatically generated transcripts of their oral response to an assessment prompt, we train a system to automatically grade the quality of the response with respect to a pre-determined criterion. This system uses language modeling and spoken information retrieval to identify key features in the spoken response and holistically decide if the response aligns with the grading criteria. Combined, the steps in this work form a framework for inclusive spoken language understanding technology that can be used to perform provide students with dialect-appropriate language training or language assessment.
■590 ▼aSchool code: 0031.
■650 4▼aElectrical engineering
■650 4▼aEducational technology
■650 4▼aLinguistics
■650 4▼aAfrican American studies
■650 4▼aCommunication
■653 ▼aSpeech technology
■653 ▼aAutomatic speech recognition
■653 ▼aSpeaker identification
■653 ▼aLinear predictive coding
■653 ▼aAfrican American English
■690 ▼a0544
■690 ▼a0459
■690 ▼a0290
■690 ▼a0710
■690 ▼a0296
■71020▼aUniversity of California, Los Angeles▼bElectrical and Computer Engineering 0333.
■7730 ▼tDissertations Abstracts International▼g85-09A.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160550▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


