본문

서브메뉴

Speech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Transfer
Speech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Tra...
Speech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Transfer

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105659
ISBN  
9798263307721
DDC  
621.3
저자명  
Harvill, John.
서명/저자  
Speech Classification and Lexical Semantic Modeling via Self-Supervision and Knowledge Transfer
발행사항  
[Sl] : University of Illinois at Urbana-Champaign, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
132 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Hasegawa-Johnson, Mark.
학위논문주기  
Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
초록/해제  
요약The field of speech and natural language processing has experienced dramatic progress over the past decade due to a major paradigm shift. Instead of using training data for a target task only, modern speech and text applications rely on pretraining as the first step. After pretraining a model on a certain task, or potentially multiple tasks, the knowledge that was learned can be transferred to a downstream task and lead to large performance gains. Given that labeled data is much more challenging to collect than raw speech or text, the most explosive growth in the field has come from discovering effective ways to perform pretraining in a self-supervised fashion. By cleverly manipulating a raw speech waveform or raw text, it is possible to learn an immense amount of information without requiring annotations from humans. In this dissertation, I explore several speech and text tasks that benefit from self-supervision and knowledge transfer. For speech, I demonstrate that for both the stutter detection and device arbitration problems, tailored self-supervised pretraining schemes can be developed that lead to significant performance gains compared to relying on labeled data only. For stutter detection, I propose the idea of creating artificial stuttered speech from healthy speech and using it for pretraining. I also show that knowledge of whether stuttering occurs somewhere within a window of several seconds of speech audio can be used to learn the location of stuttering to a much finer degree via multiple instance learning. For device arbitration, I show that contrastive learning and autoencoding can both create useful representations of acoustic information that improve the ability of an arbitration system to determine which voice assistant is closest to a user. In the text domain, I explore lexical semantic modeling, exemplification modeling, and Automatic Speech Recognition (ASR) error detection and correction. Similar to the speech tasks, I find that all text-based tasks can be improved via knowledge transfer, self-supervision, or a combination of the two. For lexical semantic modeling, I propose a graph-based solution and find that knowledge from many languages is required to perform well on any single language. For exemplification modeling, I propose an autoencoding technique that can effectively isolate information related to contextual meaning of a target polysemous word and generate new, diverse sentences using that word with the intended meaning. For ASR error detection and correction, I show that significant errors can be detected to a high degree of accuracy by combining knowledge from both a sentence-level semantic encoder and Large Language Model (LLM) and highlight the existence of statistical bias within correction and detection models.
일반주제명  
Electrical engineering
일반주제명  
Engineering
일반주제명  
Computer science
일반주제명  
Acoustics
키워드  
Self-supervision
키워드  
Knowledge transfer
키워드  
Speech classification
키워드  
Lexical semantics
기타저자  
University of Illinois at Urbana-Champaign Electrical & Computer Eng
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017361061
■00520260202105659
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263307721
■035    ▼a(MiAaPQ)AAI32409849
■035    ▼a(MiAaPQ)124335
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aHarvill,  John.
■24510▼aSpeech  Classification  and  Lexical  Semantic  Modeling  via  Self-Supervision  and  Knowledge  Transfer
■260    ▼a[Sl]▼bUniversity  of  Illinois  at  Urbana-Champaign▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a132  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Hasegawa-Johnson,  Mark.
■5021  ▼aThesis  (Ph.D.)--University  of  Illinois  at  Urbana-Champaign,  2024.
■520    ▼aThe  field  of  speech  and  natural  language  processing  has  experienced  dramatic  progress  over  the  past  decade  due  to  a  major  paradigm  shift.  Instead  of  using  training  data  for  a  target  task  only,  modern  speech  and  text  applications  rely  on  pretraining  as  the  first  step.  After  pretraining  a  model  on  a  certain  task,  or  potentially  multiple  tasks,  the  knowledge  that  was  learned  can  be  transferred  to  a  downstream  task  and  lead  to  large  performance  gains.  Given  that  labeled  data  is  much  more  challenging  to  collect  than  raw  speech  or  text,  the  most  explosive  growth  in  the  field  has  come  from  discovering  effective  ways  to  perform  pretraining  in  a  self-supervised  fashion.  By  cleverly  manipulating  a  raw  speech  waveform  or  raw  text,  it  is  possible  to  learn  an  immense  amount  of  information  without  requiring  annotations  from  humans.                        In  this  dissertation,  I  explore  several  speech  and  text  tasks  that  benefit  from  self-supervision  and  knowledge  transfer.  For  speech,  I  demonstrate  that  for  both  the  stutter  detection  and  device  arbitration  problems,  tailored  self-supervised  pretraining  schemes  can  be  developed  that  lead  to  significant  performance  gains  compared  to  relying  on  labeled  data  only.  For  stutter  detection,  I  propose  the  idea  of  creating  artificial  stuttered  speech  from  healthy  speech  and  using  it  for  pretraining.  I  also  show  that  knowledge  of  whether  stuttering  occurs  somewhere  within  a  window  of  several  seconds  of  speech  audio  can  be  used  to  learn  the  location  of  stuttering  to  a  much  finer  degree  via  multiple  instance  learning.  For  device  arbitration,  I  show  that  contrastive  learning  and  autoencoding  can  both  create  useful  representations  of  acoustic  information  that  improve  the  ability  of  an  arbitration  system  to  determine  which  voice  assistant  is  closest  to  a  user.                        In  the  text  domain,  I  explore  lexical  semantic  modeling,  exemplification  modeling,  and  Automatic  Speech  Recognition  (ASR)  error  detection  and  correction.  Similar  to  the  speech  tasks,  I  find  that  all  text-based  tasks  can  be  improved  via  knowledge  transfer,  self-supervision,  or  a  combination  of  the  two.  For  lexical  semantic  modeling,  I  propose  a  graph-based  solution  and  find  that  knowledge  from  many  languages  is  required  to  perform  well  on  any  single  language.  For  exemplification  modeling,  I  propose  an  autoencoding  technique  that  can  effectively  isolate  information  related  to  contextual  meaning  of  a  target  polysemous  word  and  generate  new,  diverse  sentences  using  that  word  with  the  intended  meaning.  For  ASR  error  detection  and  correction,  I  show  that  significant  errors  can  be  detected  to  a  high  degree  of  accuracy  by  combining  knowledge  from  both  a  sentence-level  semantic  encoder  and  Large  Language  Model  (LLM)  and  highlight  the  existence  of  statistical  bias  within  correction  and  detection  models.
■590    ▼aSchool  code:  0090.
■650  4▼aElectrical  engineering
■650  4▼aEngineering
■650  4▼aComputer  science
■650  4▼aAcoustics
■653    ▼aSelf-supervision
■653    ▼aKnowledge  transfer
■653    ▼aSpeech  classification
■653    ▼aLexical  semantics
■690    ▼a0544
■690    ▼a0984
■690    ▼a0537
■690    ▼a0986
■71020▼aUniversity  of  Illinois  at  Urbana-Champaign▼bElectrical  &  Computer  Eng.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0090
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361061▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF16738 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.