서브메뉴
검색
Speech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Transfer
Speech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Transfer
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105659
- ISBN
- 9798263307721
- DDC
- 621.3
- 저자명
- Harvill, John.
- 서명/저자
- Speech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Transfer
- 발행사항
- [Sl] : University of Illinois at Urbana-Champaign, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 132 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Hasegawa-Johnson, Mark.
- 학위논문주기
- Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
- 초록/해제
- 요약The field of speech and natural language processing has experienced dramatic progress over the past decade due to a major paradigm shift. Instead of using training data for a target task only, modern speech and text applications rely on pretraining as the first step. After pretraining a model on a certain task, or potentially multiple tasks, the knowledge that was learned can be transferred to a downstream task and lead to large performance gains. Given that labeled data is much more challenging to collect than raw speech or text, the most explosive growth in the field has come from discovering effective ways to perform pretraining in a self-supervised fashion. By cleverly manipulating a raw speech waveform or raw text, it is possible to learn an immense amount of information without requiring annotations from humans. In this dissertation, I explore several speech and text tasks that benefit from self-supervision and knowledge transfer. For speech, I demonstrate that for both the stutter detection and device arbitration problems, tailored self-supervised pretraining schemes can be developed that lead to significant performance gains compared to relying on labeled data only. For stutter detection, I propose the idea of creating artificial stuttered speech from healthy speech and using it for pretraining. I also show that knowledge of whether stuttering occurs somewhere within a window of several seconds of speech audio can be used to learn the location of stuttering to a much finer degree via multiple instance learning. For device arbitration, I show that contrastive learning and autoencoding can both create useful representations of acoustic information that improve the ability of an arbitration system to determine which voice assistant is closest to a user. In the text domain, I explore lexical semantic modeling, exemplification modeling, and Automatic Speech Recognition (ASR) error detection and correction. Similar to the speech tasks, I find that all text-based tasks can be improved via knowledge transfer, self-supervision, or a combination of the two. For lexical semantic modeling, I propose a graph-based solution and find that knowledge from many languages is required to perform well on any single language. For exemplification modeling, I propose an autoencoding technique that can effectively isolate information related to contextual meaning of a target polysemous word and generate new, diverse sentences using that word with the intended meaning. For ASR error detection and correction, I show that significant errors can be detected to a high degree of accuracy by combining knowledge from both a sentence-level semantic encoder and Large Language Model (LLM) and highlight the existence of statistical bias within correction and detection models.
- 일반주제명
- Electrical engineering
- 일반주제명
- Engineering
- 일반주제명
- Computer science
- 일반주제명
- Acoustics
- 키워드
- Self-supervision
- 기타저자
- University of Illinois at Urbana-Champaign Electrical & Computer Eng
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017361061
■00520260202105659
■006m o d
■007cr#unu||||||||
■020 ▼a9798263307721
■035 ▼a(MiAaPQ)AAI32409849
■035 ▼a(MiAaPQ)124335
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aHarvill, John.
■24510▼aSpeech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Transfer
■260 ▼a[Sl]▼bUniversity of Illinois at Urbana-Champaign▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a132 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Hasegawa-Johnson, Mark.
■5021 ▼aThesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
■520 ▼aThe field of speech and natural language processing has experienced dramatic progress over the past decade due to a major paradigm shift. Instead of using training data for a target task only, modern speech and text applications rely on pretraining as the first step. After pretraining a model on a certain task, or potentially multiple tasks, the knowledge that was learned can be transferred to a downstream task and lead to large performance gains. Given that labeled data is much more challenging to collect than raw speech or text, the most explosive growth in the field has come from discovering effective ways to perform pretraining in a self-supervised fashion. By cleverly manipulating a raw speech waveform or raw text, it is possible to learn an immense amount of information without requiring annotations from humans. In this dissertation, I explore several speech and text tasks that benefit from self-supervision and knowledge transfer. For speech, I demonstrate that for both the stutter detection and device arbitration problems, tailored self-supervised pretraining schemes can be developed that lead to significant performance gains compared to relying on labeled data only. For stutter detection, I propose the idea of creating artificial stuttered speech from healthy speech and using it for pretraining. I also show that knowledge of whether stuttering occurs somewhere within a window of several seconds of speech audio can be used to learn the location of stuttering to a much finer degree via multiple instance learning. For device arbitration, I show that contrastive learning and autoencoding can both create useful representations of acoustic information that improve the ability of an arbitration system to determine which voice assistant is closest to a user. In the text domain, I explore lexical semantic modeling, exemplification modeling, and Automatic Speech Recognition (ASR) error detection and correction. Similar to the speech tasks, I find that all text-based tasks can be improved via knowledge transfer, self-supervision, or a combination of the two. For lexical semantic modeling, I propose a graph-based solution and find that knowledge from many languages is required to perform well on any single language. For exemplification modeling, I propose an autoencoding technique that can effectively isolate information related to contextual meaning of a target polysemous word and generate new, diverse sentences using that word with the intended meaning. For ASR error detection and correction, I show that significant errors can be detected to a high degree of accuracy by combining knowledge from both a sentence-level semantic encoder and Large Language Model (LLM) and highlight the existence of statistical bias within correction and detection models.
■590 ▼aSchool code: 0090.
■650 4▼aElectrical engineering
■650 4▼aEngineering
■650 4▼aComputer science
■650 4▼aAcoustics
■653 ▼aSelf-supervision
■653 ▼aKnowledge transfer
■653 ▼aSpeech classification
■653 ▼aLexical semantics
■690 ▼a0544
■690 ▼a0984
■690 ▼a0537
■690 ▼a0986
■71020▼aUniversity of Illinois at Urbana-Champaign▼bElectrical & Computer Eng.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0090
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361061▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


