서브메뉴
검색
Transforming Disfluency Detection: Integrating Large Language and Acoustic Models
Transforming Disfluency Detection: Integrating Large Language and Acoustic Models
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153005
- ISBN
- 9798384044567
- DDC
- 004
- 저자명
- Romana, Amrit.
- 서명/저자
- Transforming Disfluency Detection: Integrating Large Language and Acoustic Models
- 발행사항
- [Sl] : University of Michigan, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 136 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Mower Provost, Emily.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2024.
- 초록/해제
- 요약Speech disfluencies, such as filled pauses or revisions, are disruptions in the typical flow of speech. While all speakers experience some disfluencies, the frequency of disfluencies increases with certain speaker and environmental characteristics, such as a speech disorder or heightened cognitive load. Modeling disfluent events has been shown to be helpful for a range of downstream tasks, and as a result, disfluency detection and categorization has gained traction as a research area. However, variability across speakers and disfluency types make it difficult to develop scalable and generalizable methods for this task.In this thesis, I begin by exploring disfluencies as a predictor of cognitive impairment. I use these findings to motivate my investigation into models for automatic disfluency detection and categorization. I find that a fine-tuned transformer-based large language model, namely BERT, outperforms previously used neural networks that rely on hand-crafted features. With scalability in mind, I then address the challenge of performing this task without manually transcribed text. I evaluate the potential of detecting disfluencies from automatic speech recognition (ASR) transcripts, including those generated by Whisper, but I find that ASR errors limit performance of the downstream task.As an alternative approach, I fine-tune the acoustic ASR models directly for disfluency detection from audio, eliminating the intermediate transcription step. I then propose a multimodal model that combines language and acoustic representations. This multimodal approach effectively compensates for ASR errors and outperforms the unimodal models. Lastly, I consider a multi-task training objective, with disfluency detection as the primary task and ASR as an auxiliary task. I find that this multi-task training results in a model that performs similarly to the multimodal model when evaluated in-domain. Importantly, I also find that multi-task training results in a model that generalizes best out-of-domain.The overarching goal of this thesis is to develop robust and scalable methods for automatically detecting and categorizing disfluencies. These advancements will lay the groundwork for future research to explore disfluencies as potential signals of speaker or environmental characteristics.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Acoustics
- 키워드
- Disfluencies
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164463
■00520250211153005
■006m o d
■007cr#unu||||||||
■020 ▼a9798384044567
■035 ▼a(MiAaPQ)AAI31631383
■035 ▼a(MiAaPQ)umichrackham005594
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aRomana, Amrit.
■24510▼aTransforming Disfluency Detection: Integrating Large Language and Acoustic Models
■260 ▼a[Sl]▼bUniversity of Michigan▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a136 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Mower Provost, Emily.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2024.
■520 ▼aSpeech disfluencies, such as filled pauses or revisions, are disruptions in the typical flow of speech. While all speakers experience some disfluencies, the frequency of disfluencies increases with certain speaker and environmental characteristics, such as a speech disorder or heightened cognitive load. Modeling disfluent events has been shown to be helpful for a range of downstream tasks, and as a result, disfluency detection and categorization has gained traction as a research area. However, variability across speakers and disfluency types make it difficult to develop scalable and generalizable methods for this task.In this thesis, I begin by exploring disfluencies as a predictor of cognitive impairment. I use these findings to motivate my investigation into models for automatic disfluency detection and categorization. I find that a fine-tuned transformer-based large language model, namely BERT, outperforms previously used neural networks that rely on hand-crafted features. With scalability in mind, I then address the challenge of performing this task without manually transcribed text. I evaluate the potential of detecting disfluencies from automatic speech recognition (ASR) transcripts, including those generated by Whisper, but I find that ASR errors limit performance of the downstream task.As an alternative approach, I fine-tune the acoustic ASR models directly for disfluency detection from audio, eliminating the intermediate transcription step. I then propose a multimodal model that combines language and acoustic representations. This multimodal approach effectively compensates for ASR errors and outperforms the unimodal models. Lastly, I consider a multi-task training objective, with disfluency detection as the primary task and ASR as an auxiliary task. I find that this multi-task training results in a model that performs similarly to the multimodal model when evaluated in-domain. Importantly, I also find that multi-task training results in a model that generalizes best out-of-domain.The overarching goal of this thesis is to develop robust and scalable methods for automatically detecting and categorizing disfluencies. These advancements will lay the groundwork for future research to explore disfluencies as potential signals of speaker or environmental characteristics.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aAcoustics
■653 ▼aSpeech processing
■653 ▼aDisfluencies
■653 ▼aAutomatic speech recognition
■653 ▼aLarge language models
■653 ▼aCognitive impairment
■690 ▼a0984
■690 ▼a0464
■690 ▼a0986
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164463▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
ค้นหาข้อมูลรายละเอียด
- จองห้องพัก
- ไม่อยู่
- โฟลเดอร์ของฉัน
- ขอดูแรก
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


