본문

서브메뉴

Continual Learning on Speech and Audio: Towards Data, Model and Metrics
Continual Learning on Speech and Audio: Towards Data, Model and Metrics
Continual Learning on Speech and Audio: Towards Data, Model and Metrics

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152005
ISBN  
9798382832937
DDC  
004
저자명  
Yang, Muqiao.
서명/저자  
Continual Learning on Speech and Audio: Towards Data, Model and Metrics
발행사항  
[Sl] : Carnegie Mellon University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
99 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Ramakrishnan, Bhiksha.
학위논문주기  
Thesis (Ph.D.)--Carnegie Mellon University, 2024.
초록/해제  
요약In recent years, the community has witnessed the enormous progress of deep neural network models in matching or even surpassing human performance on a variety of speech and audio tasks, including Automatic Speech Recognition (ASR), Spoken Language Understanding (SLU), Text-to-Speech (TTS), etc. However, their impressive and powerful achievement is predominantly dependent on training with a large set of data defined by a particular and rigid task. In such a paradigm, the model is expected to learn universal knowledge from a static entity of data and stationary environments. In contrast, the real world is inherently ever-changing and non-stationary. New data is often generated and collected every second in a stream format, and novel classes may also emerge from time to time. Without proper adaptation techniques, the knowledge learned in the past might be erased easily when the model is learning subsequent tasks, thus resulting in overall performance degradation. Such a phenomenon is called catastrophic forgetting, which limits the practical use and expansion of many deep neural network models.Continual learning has emerged as a new machine learning paradigm that enables artificial intelligence (AI) systems to learn from a continuous stream of data and incrementally improve their performance over time. By adapting to changing environments and user needs, continual learning aims to address the catastrophic forgetting effect, so that the model can gradually extend the knowledge it acquires without drastically forgetting the knowledge that has been learned in the past. Such a property is crucial in practical applications to enable artificial systems to learn from the infinite streams of data of the changing world in a lifelong manner.This thesis mainly focuses on the underexplored area of how continual learning techniques can be effective in speech and audio tasks via three perspectives: data, model, and metrics. We will introduce the background and formulations of multiple continual learning scenarios, including data-incremental, class-incremental, and task-incremental settings. Then we will present how different categories of continual learning scenarios and methods can be applied to different modules of the modeling pipeline. Starting from the taxonomy of methods, we propose to improve continual learning towards the three perspectives. First, we demonstrate how to address data sampling, selection, and imbalance to help with continual learning on different audio tasks. Second, we show how the joint use of model architecture and data with different learning strategies could benefit continual learning processes. Lastly, we propose new continual evaluation metrics to give us a comprehensive and deeper understanding of the general continual learning behaviors. We believe that this thesis provides an overall exploration of continual learning scenarios in various speech and audio tasks, and makes an important step towards realizing lifelong learning of speech interfaces.
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Electrical engineering
키워드  
Deep neural network
키워드  
Continual learning
키워드  
Machine learning
키워드  
Speech interfaces
키워드  
Speech recognition
키워드  
Lifelong learning
기타저자  
Carnegie Mellon University Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017162381
■00520250211152005
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382832937
■035    ▼a(MiAaPQ)AAI31330553
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aYang,  Muqiao.▼0(orcid)0000-0001-6273-0138
■24510▼aContinual  Learning  on  Speech  and  Audio:  Towards  Data,  Model  and  Metrics
■260    ▼a[Sl]▼bCarnegie  Mellon  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a99  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Ramakrishnan,  Bhiksha.
■5021  ▼aThesis  (Ph.D.)--Carnegie  Mellon  University,  2024.
■520    ▼aIn  recent  years,  the  community  has  witnessed  the  enormous  progress  of  deep  neural  network  models  in  matching  or  even  surpassing  human  performance  on  a  variety  of  speech  and  audio  tasks,  including  Automatic  Speech  Recognition  (ASR),  Spoken  Language  Understanding  (SLU),  Text-to-Speech  (TTS),  etc.  However,  their  impressive  and  powerful  achievement  is  predominantly  dependent  on  training  with  a  large  set  of  data  defined  by  a  particular  and  rigid  task.  In  such  a  paradigm,  the  model  is  expected  to  learn  universal  knowledge  from  a  static  entity  of  data  and  stationary  environments.  In  contrast,  the  real  world  is  inherently  ever-changing  and  non-stationary.  New  data  is  often  generated  and  collected  every  second  in  a  stream  format,  and  novel  classes  may  also  emerge  from  time  to  time.  Without  proper  adaptation  techniques,  the  knowledge  learned  in  the  past  might  be  erased  easily  when  the  model  is  learning  subsequent  tasks,  thus  resulting  in  overall  performance  degradation.  Such  a  phenomenon  is  called  catastrophic  forgetting,  which  limits  the  practical  use  and  expansion  of  many  deep  neural  network  models.Continual  learning  has  emerged  as  a  new  machine  learning  paradigm  that  enables  artificial  intelligence  (AI)  systems  to  learn  from  a  continuous  stream  of  data  and  incrementally  improve  their  performance  over  time.  By  adapting  to  changing  environments  and  user  needs,  continual  learning  aims  to  address  the  catastrophic  forgetting  effect,  so  that  the  model  can  gradually  extend  the  knowledge  it  acquires  without  drastically  forgetting  the  knowledge  that  has  been  learned  in  the  past.  Such  a  property  is  crucial  in  practical  applications  to  enable  artificial  systems  to  learn  from  the  infinite  streams  of  data  of  the  changing  world  in  a  lifelong  manner.This  thesis  mainly  focuses  on  the  underexplored  area  of  how  continual  learning  techniques  can  be  effective  in  speech  and  audio  tasks  via  three  perspectives:  data,  model,  and  metrics.  We  will  introduce  the  background  and  formulations  of  multiple  continual  learning  scenarios,  including  data-incremental,  class-incremental,  and  task-incremental  settings.  Then  we  will  present  how  different  categories  of  continual  learning  scenarios  and  methods  can  be  applied  to  different  modules  of  the  modeling  pipeline.  Starting  from  the  taxonomy  of  methods,  we  propose  to  improve  continual  learning  towards  the  three  perspectives.  First,  we  demonstrate  how  to  address  data  sampling,  selection,  and  imbalance  to  help  with  continual  learning  on  different  audio  tasks.  Second,  we  show  how  the  joint  use  of  model  architecture  and  data  with  different  learning  strategies  could  benefit  continual  learning  processes.  Lastly,  we  propose  new  continual  evaluation  metrics  to  give  us  a  comprehensive  and  deeper  understanding  of  the  general  continual  learning  behaviors.  We  believe  that  this  thesis  provides  an  overall  exploration  of  continual  learning  scenarios  in  various  speech  and  audio  tasks,  and  makes  an  important  step  towards  realizing  lifelong  learning  of  speech  interfaces.
■590    ▼aSchool  code:  0041.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aElectrical  engineering
■653    ▼aDeep  neural  network
■653    ▼aContinual  learning
■653    ▼aMachine  learning
■653    ▼aSpeech  interfaces
■653    ▼aSpeech  recognition
■653    ▼aLifelong  learning
■690    ▼a0984
■690    ▼a0464
■690    ▼a0800
■690    ▼a0544
■71020▼aCarnegie  Mellon  University▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0041
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162381▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13675 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.