서브메뉴
검색
Towards Effective and Efficient Open Speech Foundation Models
Towards Effective and Efficient Open Speech Foundation Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103143
- ISBN
- 9798314868096
- DDC
- 004
- 저자명
- Peng, Yifan.
- 서명/저자
- Towards Effective and Efficient Open Speech Foundation Models
- 발행사항
- [Sl] : Carnegie Mellon University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 222 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
- 주기사항
- Advisor: Watanabe, Shinji.
- 학위논문주기
- Thesis (Ph.D.)--Carnegie Mellon University, 2025.
- 초록/해제
- 요약Speech is a key modality for human-computer interaction, enabling a wide range of speech processing applications. Traditionally, these applications have relied on separate models for each task, limiting scalability and impeding cross-task knowledge sharing. Recently, speech foundation models (SFMs) have emerged as a unifying framework for speech-related tasks. We define SFMs as models trained on broad data that can be adapted to various tasks or languages related to speech processing. Common types of SFMs include self-supervised learning (SSL) speech representation models, task-specific SFMs, and general instruction-following SFMs. This thesis primarily focuses on the latter two categories, as SSL models cannot directly perform downstream tasks and are typically used as feature extractors or tokenizers via additional fine-tuning.Despite their impressive performance, most existing SFMs-primarily developed by large corporations-lack openness and reproducibility, hindering scientific transparency and slowing broader research progress. This thesis addresses these limitations by developing SFMs with full transparency, architectural innovations, and improved efficiency.We begin with task-specific SFMs trained via large-scale supervised learning, following the paradigm of models like Whisper. We present the Open Whisper-style Speech Models (OWSM), a reproducible framework for large-scale speech model training using only publicly available data and open-source toolkits. To improve speech modeling capabilities, we propose novel encoder architectures, including Branchformer and E-Branchformer, which achieve state-of-the-art results across diverse speech tasks. Through architectural improvements, data scaling, and systematic data cleaning, the OWSM models match or exceed the performance of leading proprietary systems in several benchmarks.Building on this foundation, we extend to instruction-following SFMs capable of handling unseen tasks via natural language prompts. We introduce VoiceTextBlender, a spoken language model (SLM) that integrates speech and language understanding through a novel single-stage, joint speech-text supervised fine-tuning approach.To address the computational demands of SFMs, we further propose model compression techniques ranging from static pruning to dynamic architectures, significantly reducing inference costs while maintaining high performance.Finally, this thesis offers new insights into the behavior of SFMs, including emergent capabilities and scaling trends, contributing to a deeper understanding of their design and potential. By tackling the key challenges of openness, efficiency, and architecture, this work aims to democratize access to advanced speech technologies and catalyze further innovation in the field.
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 일반주제명
- Computer engineering
- 일반주제명
- Speech therapy
- 기타저자
- Carnegie Mellon University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357172
■00520260202103143
■006m o d
■007cr#unu||||||||
■020 ▼a9798314868096
■035 ▼a(MiAaPQ)AAI31994585
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aPeng, Yifan.▼0(orcid)0000-0002-8581-8674
■24510▼aTowards Effective and Efficient Open Speech Foundation Models
■260 ▼a[Sl]▼bCarnegie Mellon University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a222 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-11, Section: B.
■500 ▼aAdvisor: Watanabe, Shinji.
■5021 ▼aThesis (Ph.D.)--Carnegie Mellon University, 2025.
■520 ▼aSpeech is a key modality for human-computer interaction, enabling a wide range of speech processing applications. Traditionally, these applications have relied on separate models for each task, limiting scalability and impeding cross-task knowledge sharing. Recently, speech foundation models (SFMs) have emerged as a unifying framework for speech-related tasks. We define SFMs as models trained on broad data that can be adapted to various tasks or languages related to speech processing. Common types of SFMs include self-supervised learning (SSL) speech representation models, task-specific SFMs, and general instruction-following SFMs. This thesis primarily focuses on the latter two categories, as SSL models cannot directly perform downstream tasks and are typically used as feature extractors or tokenizers via additional fine-tuning.Despite their impressive performance, most existing SFMs-primarily developed by large corporations-lack openness and reproducibility, hindering scientific transparency and slowing broader research progress. This thesis addresses these limitations by developing SFMs with full transparency, architectural innovations, and improved efficiency.We begin with task-specific SFMs trained via large-scale supervised learning, following the paradigm of models like Whisper. We present the Open Whisper-style Speech Models (OWSM), a reproducible framework for large-scale speech model training using only publicly available data and open-source toolkits. To improve speech modeling capabilities, we propose novel encoder architectures, including Branchformer and E-Branchformer, which achieve state-of-the-art results across diverse speech tasks. Through architectural improvements, data scaling, and systematic data cleaning, the OWSM models match or exceed the performance of leading proprietary systems in several benchmarks.Building on this foundation, we extend to instruction-following SFMs capable of handling unseen tasks via natural language prompts. We introduce VoiceTextBlender, a spoken language model (SLM) that integrates speech and language understanding through a novel single-stage, joint speech-text supervised fine-tuning approach.To address the computational demands of SFMs, we further propose model compression techniques ranging from static pruning to dynamic architectures, significantly reducing inference costs while maintaining high performance.Finally, this thesis offers new insights into the behavior of SFMs, including emergent capabilities and scaling trends, contributing to a deeper understanding of their design and potential. By tackling the key challenges of openness, efficiency, and architecture, this work aims to democratize access to advanced speech technologies and catalyze further innovation in the field.
■590 ▼aSchool code: 0041.
■650 4▼aComputer science
■650 4▼aElectrical engineering
■650 4▼aComputer engineering
■650 4▼aSpeech therapy
■653 ▼aNatural language processing
■653 ▼aSpeech foundation models
■653 ▼aSpeech language models
■653 ▼aSpeech processing
■653 ▼aSpeech recognition
■653 ▼aSpeech translation
■690 ▼a0800
■690 ▼a0984
■690 ▼a0544
■690 ▼a0464
■690 ▼a0460
■71020▼aCarnegie Mellon University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g86-11B.
■790 ▼a0041
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357172▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


