본문

서브메뉴

Towards Effective and Efficient Open Speech Foundation Models
Towards Effective and Efficient Open Speech Foundation Models
Towards Effective and Efficient Open Speech Foundation Models

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103143
ISBN  
9798314868096
DDC  
004
저자명  
Peng, Yifan.
서명/저자  
Towards Effective and Efficient Open Speech Foundation Models
발행사항  
[Sl] : Carnegie Mellon University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
222 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
주기사항  
Advisor: Watanabe, Shinji.
학위논문주기  
Thesis (Ph.D.)--Carnegie Mellon University, 2025.
초록/해제  
요약Speech is a key modality for human-computer interaction, enabling a wide range of speech processing applications. Traditionally, these applications have relied on separate models for each task, limiting scalability and impeding cross-task knowledge sharing. Recently, speech foundation models (SFMs) have emerged as a unifying framework for speech-related tasks. We define SFMs as models trained on broad data that can be adapted to various tasks or languages related to speech processing. Common types of SFMs include self-supervised learning (SSL) speech representation models, task-specific SFMs, and general instruction-following SFMs. This thesis primarily focuses on the latter two categories, as SSL models cannot directly perform downstream tasks and are typically used as feature extractors or tokenizers via additional fine-tuning.Despite their impressive performance, most existing SFMs-primarily developed by large corporations-lack openness and reproducibility, hindering scientific transparency and slowing broader research progress. This thesis addresses these limitations by developing SFMs with full transparency, architectural innovations, and improved efficiency.We begin with task-specific SFMs trained via large-scale supervised learning, following the paradigm of models like Whisper. We present the Open Whisper-style Speech Models (OWSM), a reproducible framework for large-scale speech model training using only publicly available data and open-source toolkits. To improve speech modeling capabilities, we propose novel encoder architectures, including Branchformer and E-Branchformer, which achieve state-of-the-art results across diverse speech tasks. Through architectural improvements, data scaling, and systematic data cleaning, the OWSM models match or exceed the performance of leading proprietary systems in several benchmarks.Building on this foundation, we extend to instruction-following SFMs capable of handling unseen tasks via natural language prompts. We introduce VoiceTextBlender, a spoken language model (SLM) that integrates speech and language understanding through a novel single-stage, joint speech-text supervised fine-tuning approach.To address the computational demands of SFMs, we further propose model compression techniques ranging from static pruning to dynamic architectures, significantly reducing inference costs while maintaining high performance.Finally, this thesis offers new insights into the behavior of SFMs, including emergent capabilities and scaling trends, contributing to a deeper understanding of their design and potential. By tackling the key challenges of openness, efficiency, and architecture, this work aims to democratize access to advanced speech technologies and catalyze further innovation in the field.
일반주제명  
Computer science
일반주제명  
Electrical engineering
일반주제명  
Computer engineering
일반주제명  
Speech therapy
키워드  
Natural language processing
키워드  
Speech foundation models
키워드  
Speech language models
키워드  
Speech processing
키워드  
Speech recognition
키워드  
Speech translation
기타저자  
Carnegie Mellon University Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 86-11B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357172
■00520260202103143
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798314868096
■035    ▼a(MiAaPQ)AAI31994585
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aPeng,  Yifan.▼0(orcid)0000-0002-8581-8674
■24510▼aTowards  Effective  and  Efficient  Open  Speech  Foundation  Models
■260    ▼a[Sl]▼bCarnegie  Mellon  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a222  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-11,  Section:  B.
■500    ▼aAdvisor:  Watanabe,  Shinji.
■5021  ▼aThesis  (Ph.D.)--Carnegie  Mellon  University,  2025.
■520    ▼aSpeech  is  a  key  modality  for  human-computer  interaction,  enabling  a  wide  range  of  speech  processing  applications.  Traditionally,  these  applications  have  relied  on  separate  models  for  each  task,  limiting  scalability  and  impeding  cross-task  knowledge  sharing.  Recently,  speech  foundation  models  (SFMs)  have  emerged  as  a  unifying  framework  for  speech-related  tasks.  We  define  SFMs  as  models  trained  on  broad  data  that  can  be  adapted  to  various  tasks  or  languages  related  to  speech  processing.  Common  types  of  SFMs  include  self-supervised  learning  (SSL)  speech  representation  models,  task-specific  SFMs,  and  general  instruction-following  SFMs.  This  thesis  primarily  focuses  on  the  latter  two  categories,  as  SSL  models  cannot  directly  perform  downstream  tasks  and  are  typically  used  as  feature  extractors  or  tokenizers  via  additional  fine-tuning.Despite  their  impressive  performance,  most  existing  SFMs-primarily  developed  by  large  corporations-lack  openness  and  reproducibility,  hindering  scientific  transparency  and  slowing  broader  research  progress.  This  thesis  addresses  these  limitations  by  developing  SFMs  with  full  transparency,  architectural  innovations,  and  improved  efficiency.We  begin  with  task-specific  SFMs  trained  via  large-scale  supervised  learning,  following  the  paradigm  of  models  like  Whisper.  We  present  the  Open  Whisper-style  Speech  Models  (OWSM),  a  reproducible  framework  for  large-scale  speech  model  training  using  only  publicly  available  data  and  open-source  toolkits.  To  improve  speech  modeling  capabilities,  we  propose  novel  encoder  architectures,  including  Branchformer  and  E-Branchformer,  which  achieve  state-of-the-art  results  across  diverse  speech  tasks.  Through  architectural  improvements,  data  scaling,  and  systematic  data  cleaning,  the  OWSM  models  match  or  exceed  the  performance  of  leading  proprietary  systems  in  several  benchmarks.Building  on  this  foundation,  we  extend  to  instruction-following  SFMs  capable  of  handling  unseen  tasks  via  natural  language  prompts.  We  introduce  VoiceTextBlender,  a  spoken  language  model  (SLM)  that  integrates  speech  and  language  understanding  through  a  novel  single-stage,  joint  speech-text  supervised  fine-tuning  approach.To  address  the  computational  demands  of  SFMs,  we  further  propose  model  compression  techniques  ranging  from  static  pruning  to  dynamic  architectures,  significantly  reducing  inference  costs  while  maintaining  high  performance.Finally,  this  thesis  offers  new  insights  into  the  behavior  of  SFMs,  including  emergent  capabilities  and  scaling  trends,  contributing  to  a  deeper  understanding  of  their  design  and  potential.  By  tackling  the  key  challenges  of  openness,  efficiency,  and  architecture,  this  work  aims  to  democratize  access  to  advanced  speech  technologies  and  catalyze  further  innovation  in  the  field.
■590    ▼aSchool  code:  0041.
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■650  4▼aComputer  engineering
■650  4▼aSpeech  therapy
■653    ▼aNatural  language  processing
■653    ▼aSpeech  foundation  models
■653    ▼aSpeech  language  models
■653    ▼aSpeech  processing
■653    ▼aSpeech  recognition
■653    ▼aSpeech  translation
■690    ▼a0800
■690    ▼a0984
■690    ▼a0544
■690    ▼a0464
■690    ▼a0460
■71020▼aCarnegie  Mellon  University▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-11B.
■790    ▼a0041
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357172▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18525 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.