본문

서브메뉴

Transforming Disfluency Detection: Integrating Large Language and Acoustic Models
Transforming Disfluency Detection: Integrating Large Language and Acoustic Models
Transforming Disfluency Detection: Integrating Large Language and Acoustic Models

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20250211153005
ISBN  
9798384044567
DDC  
004
저자명  
Romana, Amrit.
서명/저자  
Transforming Disfluency Detection: Integrating Large Language and Acoustic Models
발행사항  
[Sl] : University of Michigan, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
136 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
주기사항  
Advisor: Mower Provost, Emily.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2024.
초록/해제  
요약Speech disfluencies, such as filled pauses or revisions, are disruptions in the typical flow of speech. While all speakers experience some disfluencies, the frequency of disfluencies increases with certain speaker and environmental characteristics, such as a speech disorder or heightened cognitive load. Modeling disfluent events has been shown to be helpful for a range of downstream tasks, and as a result, disfluency detection and categorization has gained traction as a research area. However, variability across speakers and disfluency types make it difficult to develop scalable and generalizable methods for this task.In this thesis, I begin by exploring disfluencies as a predictor of cognitive impairment. I use these findings to motivate my investigation into models for automatic disfluency detection and categorization. I find that a fine-tuned transformer-based large language model, namely BERT, outperforms previously used neural networks that rely on hand-crafted features. With scalability in mind, I then address the challenge of performing this task without manually transcribed text. I evaluate the potential of detecting disfluencies from automatic speech recognition (ASR) transcripts, including those generated by Whisper, but I find that ASR errors limit performance of the downstream task.As an alternative approach, I fine-tune the acoustic ASR models directly for disfluency detection from audio, eliminating the intermediate transcription step. I then propose a multimodal model that combines language and acoustic representations. This multimodal approach effectively compensates for ASR errors and outperforms the unimodal models. Lastly, I consider a multi-task training objective, with disfluency detection as the primary task and ASR as an auxiliary task. I find that this multi-task training results in a model that performs similarly to the multimodal model when evaluated in-domain. Importantly, I also find that multi-task training results in a model that generalizes best out-of-domain.The overarching goal of this thesis is to develop robust and scalable methods for automatically detecting and categorizing disfluencies. These advancements will lay the groundwork for future research to explore disfluencies as potential signals of speaker or environmental characteristics.
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Acoustics
키워드  
Speech processing
키워드  
Disfluencies
키워드  
Automatic speech recognition
키워드  
Large language models
키워드  
Cognitive impairment
기타저자  
University of Michigan Computer Science & Engineering
기본자료저록  
Dissertations Abstracts International. 86-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164463
■00520250211153005
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384044567
■035    ▼a(MiAaPQ)AAI31631383
■035    ▼a(MiAaPQ)umichrackham005594
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aRomana,  Amrit.
■24510▼aTransforming  Disfluency  Detection:  Integrating  Large  Language  and  Acoustic  Models
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a136  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  B.
■500    ▼aAdvisor:  Mower  Provost,  Emily.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2024.
■520    ▼aSpeech  disfluencies,  such  as  filled  pauses  or  revisions,  are  disruptions  in  the  typical  flow  of  speech.  While  all  speakers  experience  some  disfluencies,  the  frequency  of  disfluencies  increases  with  certain  speaker  and  environmental  characteristics,  such  as  a  speech  disorder  or  heightened  cognitive  load.  Modeling  disfluent  events  has  been  shown  to  be  helpful  for  a  range  of  downstream  tasks,  and  as  a  result,  disfluency  detection  and  categorization  has  gained  traction  as  a  research  area.  However,  variability  across  speakers  and  disfluency  types  make  it  difficult  to  develop  scalable  and  generalizable  methods  for  this  task.In  this  thesis,  I  begin  by  exploring  disfluencies  as  a  predictor  of  cognitive  impairment.  I  use  these  findings  to  motivate  my  investigation  into  models  for  automatic  disfluency  detection  and  categorization.  I  find  that  a  fine-tuned  transformer-based  large  language  model,  namely  BERT,  outperforms  previously  used  neural  networks  that  rely  on  hand-crafted  features.  With  scalability  in  mind,  I  then  address  the  challenge  of  performing  this  task  without  manually  transcribed  text.  I  evaluate  the  potential  of  detecting  disfluencies  from  automatic  speech  recognition  (ASR)  transcripts,  including  those  generated  by  Whisper,  but  I  find  that  ASR  errors  limit  performance  of  the  downstream  task.As  an  alternative  approach,  I  fine-tune  the  acoustic  ASR  models  directly  for  disfluency  detection  from  audio,  eliminating  the  intermediate  transcription  step.  I  then  propose  a  multimodal  model  that  combines  language  and  acoustic  representations.  This  multimodal  approach  effectively  compensates  for  ASR  errors  and  outperforms  the  unimodal  models.  Lastly,  I  consider  a  multi-task  training  objective,  with  disfluency  detection  as  the  primary  task  and  ASR  as  an  auxiliary  task.  I  find  that  this  multi-task  training  results  in  a  model  that  performs  similarly  to  the  multimodal  model  when  evaluated  in-domain.  Importantly,  I  also  find  that  multi-task  training  results  in  a  model  that  generalizes  best  out-of-domain.The  overarching  goal  of  this  thesis  is  to  develop  robust  and  scalable  methods  for  automatically  detecting  and  categorizing  disfluencies.  These  advancements  will  lay  the  groundwork  for  future  research  to  explore  disfluencies  as  potential  signals  of  speaker  or  environmental  characteristics.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aAcoustics
■653    ▼aSpeech  processing
■653    ▼aDisfluencies
■653    ▼aAutomatic  speech  recognition
■653    ▼aLarge  language  models
■653    ▼aCognitive  impairment
■690    ▼a0984
■690    ▼a0464
■690    ▼a0986
■71020▼aUniversity  of  Michigan▼bComputer  Science  &  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-03B.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164463▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    ค้นหาข้อมูลรายละเอียด

    • จองห้องพัก
    • ไม่อยู่
    • โฟลเดอร์ของฉัน
    • ขอดูแรก
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    วัสดุ
    Reg No. Call No. ตำแหน่งที่ตั้ง สถานะ ยืมข้อมูล
    TF11371 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * จองมีอยู่ในหนังสือยืม เพื่อให้การสำรองที่นั่งคลิกที่ปุ่มจองห้องพัก

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.