서브메뉴
검색
Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data
Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data
Detailed Information
- Material Type
- 단행본
- 0017358320
- Date and Time of Latest Transaction
- 20260202104643
- ISBN
- 9798288822049
- DDC
- 401
- Author
- Proch Ahn, Emily.
- Title/Author
- Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data
- Publish Info
- [Sl] : University of Washington, 2025
- Publish Info
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- Material Info
- 132 p
- General Note
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- General Note
- Advisor: Levow, Gina-Anne.
- 학위논문주기
- Thesis (Ph.D.)--University of Washington, 2025.
- Abstracts/Etc
- 요약Corpus phonetics research has become increasingly large-scale as both data and automated tools have become more plentiful and available. Now that there are resources to study more kinds of data, what are some best practices in using these resources, especially when the data is diverse? This dissertation addresses the following research questions: How do we process diverse speech data, and how much can we rely on automated tools to conduct corpus phonetics research? The types of diversity covered in this work include multilingual and fieldwork corpora covering styles including read, spontaneous, and code-switched speech. Across four studies, we show that automated systems in the corpus phonetics pipeline are viable on multilingual and low-resource datasets. We first propose a pipeline that utilizes automated systems that convert orthography to phonemes, model the acoustics and align audio to those phonemes, and extract features for phonetic analysis. We apply this pipeline to a large, multilingual corpus and show both the utility and limitations of this derivative corpus in a careful study of outlying phonetic features. Then, we apply novel techniques to improve the phonetic forced alignment of low-resource field data, a challenging yet important process in language documentation. We encourage the research community to continue developing tools to aid in language documentation and cross-linguistic research. In doing so, it is important to include manual audits and to examine whether or not the tools are genuinely modeling the data.
- Subject Added Entry-Topical Term
- Linguistics
- Subject Added Entry-Topical Term
- Computer science
- Subject Added Entry-Topical Term
- Language arts
- Subject Added Entry-Topical Term
- Information technology
- Index Term-Uncontrolled
- Corpus phonetics
- Index Term-Uncontrolled
- Forced alignment
- Index Term-Uncontrolled
- Language documentation
- Index Term-Uncontrolled
- Speech technology
- Index Term-Uncontrolled
- Diverse speech data
- Added Entry-Corporate Name
- University of Washington Linguistics
- Host Item Entry
- Dissertations Abstracts International. 87-01B.
- Electronic Location and Access
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358320
■00520260202104643
■006m o d
■007cr#unu||||||||
■020 ▼a9798288822049
■035 ▼a(MiAaPQ)AAI32114404
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a401
■1001 ▼aProch Ahn, Emily.
■24510▼aInvestigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data
■260 ▼a[Sl]▼bUniversity of Washington▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a132 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Levow, Gina-Anne.
■5021 ▼aThesis (Ph.D.)--University of Washington, 2025.
■520 ▼aCorpus phonetics research has become increasingly large-scale as both data and automated tools have become more plentiful and available. Now that there are resources to study more kinds of data, what are some best practices in using these resources, especially when the data is diverse? This dissertation addresses the following research questions: How do we process diverse speech data, and how much can we rely on automated tools to conduct corpus phonetics research? The types of diversity covered in this work include multilingual and fieldwork corpora covering styles including read, spontaneous, and code-switched speech. Across four studies, we show that automated systems in the corpus phonetics pipeline are viable on multilingual and low-resource datasets. We first propose a pipeline that utilizes automated systems that convert orthography to phonemes, model the acoustics and align audio to those phonemes, and extract features for phonetic analysis. We apply this pipeline to a large, multilingual corpus and show both the utility and limitations of this derivative corpus in a careful study of outlying phonetic features. Then, we apply novel techniques to improve the phonetic forced alignment of low-resource field data, a challenging yet important process in language documentation. We encourage the research community to continue developing tools to aid in language documentation and cross-linguistic research. In doing so, it is important to include manual audits and to examine whether or not the tools are genuinely modeling the data.
■590 ▼aSchool code: 0250.
■650 4▼aLinguistics
■650 4▼aComputer science
■650 4▼aLanguage arts
■650 4▼aInformation technology
■653 ▼aCorpus phonetics
■653 ▼aForced alignment
■653 ▼aLanguage documentation
■653 ▼aSpeech technology
■653 ▼aDiverse speech data
■690 ▼a0290
■690 ▼a0984
■690 ▼a0279
■690 ▼a0489
■71020▼aUniversity of Washington▼bLinguistics.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0250
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358320▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Detail Info.
- Reservation
- Not Exist
- My Folder
- First Request
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


