서브메뉴
검색
Prosody in Human Communication and Machine Understanding
Prosody in Human Communication and Machine Understanding
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152809
- ISBN
- 9798384094975
- DDC
- 401
- 저자명
- Ng, Sara B.
- 서명/저자
- Prosody in Human Communication and Machine Understanding
- 발행사항
- [Sl] : University of Washington, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 91 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Wright, Richard A.;Ostendorf, Mari.
- 학위논문주기
- Thesis (Ph.D.)--University of Washington, 2024.
- 초록/해제
- 요약Speech technology is a ubiquitous part of the modern world, from the voice-enabled assistants in smartphones to bespoke tools used by language researchers. Technological advances and the curation of large speech datasets have enabled these systems to identify words with remarkable quality. However, the black-box nature of large commercial speech understanding systems brings into question the extent to which they can take advantage of cues from prosody.Prosody has great potential as an untapped source of linguistic information for speech understanding that is not surfaced in other aspects of language. Previous work has shown that prosodic information can be exploited computationally to resolve ambiguity for linguistic structures in computational models, and to perform tasks which are considered prosodically significant, such as sarcasm detection. However, computational systems do not benefit from the same social and conversational context that humans have in processing this kind of communication, making such tasks more challenging and further motivating the careful study of prosodic input.This work investigates the hypothesis that explicit encoding of acoustic-prosodic features is a benefit to speech understanding technology. From the domain of punctuation prediction in automatic speech recognition, I show that adding acoustic-prosodic measures can improve the performance of punctuation prediction models for speech transcripts compared to a system that uses only the word sequence. I provide a potential use case for prosodic modeling in the domain of speech entrainment. Finally, I show how computational methods can be used to understand human behavior in prosodically marked speech within the domains of speech timing and regions of presumed hyper-articulation.This work bridges the gap between linguistic questions about prosody, and computational questions about the use of or need for linguistically-motivated acoustic features. Understanding how prosody influences the quality of speech understanding systems is vital in enhancing their utility across various domains and for diverse speakers. The synthesis of the these research strands provides a bird's eye view of the methodologies and challenges that can be involved in computational processing of prosody.
- 일반주제명
- Linguistics
- 일반주제명
- Communication
- 일반주제명
- Speech therapy
- 키워드
- Entrainment
- 키워드
- Prosody
- 키워드
- Punctuation
- 키워드
- Stance
- 기타저자
- University of Washington Linguistics
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163913
■00520250211152809
■006m o d
■007cr#unu||||||||
■020 ▼a9798384094975
■035 ▼a(MiAaPQ)AAI31557585
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a401
■1001 ▼aNg, Sara B.
■24510▼aProsody in Human Communication and Machine Understanding
■260 ▼a[Sl]▼bUniversity of Washington▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a91 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Wright, Richard A.;Ostendorf, Mari.
■5021 ▼aThesis (Ph.D.)--University of Washington, 2024.
■520 ▼aSpeech technology is a ubiquitous part of the modern world, from the voice-enabled assistants in smartphones to bespoke tools used by language researchers. Technological advances and the curation of large speech datasets have enabled these systems to identify words with remarkable quality. However, the black-box nature of large commercial speech understanding systems brings into question the extent to which they can take advantage of cues from prosody.Prosody has great potential as an untapped source of linguistic information for speech understanding that is not surfaced in other aspects of language. Previous work has shown that prosodic information can be exploited computationally to resolve ambiguity for linguistic structures in computational models, and to perform tasks which are considered prosodically significant, such as sarcasm detection. However, computational systems do not benefit from the same social and conversational context that humans have in processing this kind of communication, making such tasks more challenging and further motivating the careful study of prosodic input.This work investigates the hypothesis that explicit encoding of acoustic-prosodic features is a benefit to speech understanding technology. From the domain of punctuation prediction in automatic speech recognition, I show that adding acoustic-prosodic measures can improve the performance of punctuation prediction models for speech transcripts compared to a system that uses only the word sequence. I provide a potential use case for prosodic modeling in the domain of speech entrainment. Finally, I show how computational methods can be used to understand human behavior in prosodically marked speech within the domains of speech timing and regions of presumed hyper-articulation.This work bridges the gap between linguistic questions about prosody, and computational questions about the use of or need for linguistically-motivated acoustic features. Understanding how prosody influences the quality of speech understanding systems is vital in enhancing their utility across various domains and for diverse speakers. The synthesis of the these research strands provides a bird's eye view of the methodologies and challenges that can be involved in computational processing of prosody.
■590 ▼aSchool code: 0250.
■650 4▼aLinguistics
■650 4▼aCommunication
■650 4▼aSpeech therapy
■653 ▼aEntrainment
■653 ▼aHyperarticulation
■653 ▼aProsody
■653 ▼aPunctuation
■653 ▼aSpeech recognition
■653 ▼aStance
■690 ▼a0290
■690 ▼a0800
■690 ▼a0459
■690 ▼a0460
■71020▼aUniversity of Washington▼bLinguistics.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0250
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163913▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


