본문

서브메뉴

Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data
Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data
Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104643
ISBN  
9798288822049
DDC  
401
저자명  
Proch Ahn, Emily.
서명/저자  
Investigating the Corpus Phonetics Pipeline Applied to Diverse Speech Data
발행사항  
[Sl] : University of Washington, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
132 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
주기사항  
Advisor: Levow, Gina-Anne.
학위논문주기  
Thesis (Ph.D.)--University of Washington, 2025.
초록/해제  
요약Corpus phonetics research has become increasingly large-scale as both data and automated tools have become more plentiful and available. Now that there are resources to study more kinds of data, what are some best practices in using these resources, especially when the data is diverse? This dissertation addresses the following research questions: How do we process diverse speech data, and how much can we rely on automated tools to conduct corpus phonetics research? The types of diversity covered in this work include multilingual and fieldwork corpora covering styles including read, spontaneous, and code-switched speech. Across four studies, we show that automated systems in the corpus phonetics pipeline are viable on multilingual and low-resource datasets. We first propose a pipeline that utilizes automated systems that convert orthography to phonemes, model the acoustics and align audio to those phonemes, and extract features for phonetic analysis. We apply this pipeline to a large, multilingual corpus and show both the utility and limitations of this derivative corpus in a careful study of outlying phonetic features. Then, we apply novel techniques to improve the phonetic forced alignment of low-resource field data, a challenging yet important process in language documentation. We encourage the research community to continue developing tools to aid in language documentation and cross-linguistic research. In doing so, it is important to include manual audits and to examine whether or not the tools are genuinely modeling the data.
일반주제명  
Linguistics
일반주제명  
Computer science
일반주제명  
Language arts
일반주제명  
Information technology
키워드  
Corpus phonetics
키워드  
Forced alignment
키워드  
Language documentation
키워드  
Speech technology
키워드  
Diverse speech data
기타저자  
University of Washington Linguistics
기본자료저록  
Dissertations Abstracts International. 87-01B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017358320
■00520260202104643
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288822049
■035    ▼a(MiAaPQ)AAI32114404
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a401
■1001  ▼aProch  Ahn,  Emily.
■24510▼aInvestigating  the  Corpus  Phonetics  Pipeline  Applied  to  Diverse  Speech  Data
■260    ▼a[Sl]▼bUniversity  of  Washington▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a132  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  B.
■500    ▼aAdvisor:  Levow,  Gina-Anne.
■5021  ▼aThesis  (Ph.D.)--University  of  Washington,  2025.
■520    ▼aCorpus  phonetics  research  has  become  increasingly  large-scale  as  both  data  and  automated  tools  have  become  more  plentiful  and  available.  Now  that  there  are  resources  to  study  more  kinds  of  data,  what  are  some  best  practices  in  using  these  resources,  especially  when  the  data  is  diverse?  This  dissertation  addresses  the  following  research  questions:  How  do  we  process  diverse  speech  data,  and  how  much  can  we  rely  on  automated  tools  to  conduct  corpus  phonetics  research?  The  types  of  diversity  covered  in  this  work  include  multilingual  and  fieldwork  corpora  covering  styles  including  read,  spontaneous,  and  code-switched  speech.  Across  four  studies,  we  show  that  automated  systems  in  the  corpus  phonetics  pipeline  are  viable  on  multilingual  and  low-resource  datasets.  We  first  propose  a  pipeline  that  utilizes  automated  systems  that  convert  orthography  to  phonemes,  model  the  acoustics  and  align  audio  to  those  phonemes,  and  extract  features  for  phonetic  analysis.  We  apply  this  pipeline  to  a  large,  multilingual  corpus  and  show  both  the  utility  and  limitations  of  this  derivative  corpus  in  a  careful  study  of  outlying  phonetic  features.  Then,  we  apply  novel  techniques  to  improve  the  phonetic  forced  alignment  of  low-resource  field  data,  a  challenging  yet  important  process  in  language  documentation.  We  encourage  the  research  community  to  continue  developing  tools  to  aid  in  language  documentation  and  cross-linguistic  research.  In  doing  so,  it  is  important  to  include  manual  audits  and  to  examine  whether  or  not  the  tools  are  genuinely  modeling  the  data.
■590    ▼aSchool  code:  0250.
■650  4▼aLinguistics
■650  4▼aComputer  science
■650  4▼aLanguage  arts
■650  4▼aInformation  technology
■653    ▼aCorpus  phonetics
■653    ▼aForced  alignment
■653    ▼aLanguage  documentation
■653    ▼aSpeech  technology
■653    ▼aDiverse  speech  data
■690    ▼a0290
■690    ▼a0984
■690    ▼a0279
■690    ▼a0489
■71020▼aUniversity  of  Washington▼bLinguistics.
■7730  ▼tDissertations  Abstracts  International▼g87-01B.
■790    ▼a0250
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358320▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17881 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.