본문

서브메뉴

Data-Efficient Approaches for Audio Classification and Separation
Data-Efficient Approaches for Audio Classification and Separation
Data-Efficient Approaches for Audio Classification and Separation

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260209102855
ISBN  
9798291574317
DDC  
004
저자명  
Wang, Zhepei.
서명/저자  
Data-Efficient Approaches for Audio Classification and Separation
발행사항  
[Sl] : University of Illinois at Urbana-Champaign, 2023
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2023
형태사항  
115 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
주기사항  
Advisor: Smaragdis, Paris.
학위논문주기  
Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
초록/해제  
요약Recent advances in deep learning for computational audio processing are established upon sufficient annotated audio data. However, obtaining a substantial volume of high-quality annotations from in-the-wild audio remains a significant challenge. In this thesis, we propose and analyze data-efficient approaches for modeling audio signals to perform sound classification and separation. First, we present neural network architectures based on multi-dimensional unrolling of recurrent neural networks that allow the model to perform sound event detection with efficient usage of training data. Equipped with adaptive computation, the model further learns to intelligently adjust the amount of computation and enables processing when only partial information is available. Next, we propose approaches for recognizing sound classes under a time-varying distribution. We investigate continual learning techniques to train a classifier that can efficiently learn new sound classes without forgetting the past using generative replay. We further extend our approach to an unsupervised learning setup, where the model progressively learns representations for an indefinite number of sound classes with few labels presented. Last but not least, we investigate learning with limited annotated data using semi-supervised learning. We demonstrate the effectiveness of the proposed teacher-student framework on tasks including cross-modal audio-text representation learning, singing voice separation, and personalized speech enhancement. To this end, our proposed data-efficient algorithms for audio classification and source separation show high potential for reducing the labor cost for collecting high-quality annotated data, improving computational and storage efficiency, and enabling processing on memory-limited edge devices.
일반주제명  
Computer science
일반주제명  
Electrical engineering
일반주제명  
Computer engineering
키워드  
Deep learning
키워드  
Sound classification
키워드  
Source separation
키워드  
Self-supervised learning
키워드  
Semi-supervised learning
키워드  
Continual learning
기타저자  
University of Illinois at Urbana-Champaign Computer Science
기본자료저록  
Dissertations Abstracts International. 87-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260203s2023        us                              c    eng  d
■001000017365921
■00520260209102855
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291574317
■035    ▼a(MiAaPQ)AAI32272133
■035    ▼a(MiAaPQ)httphdlhandlenet2142121935
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aWang,  Zhepei.
■24510▼aData-Efficient  Approaches  for  Audio  Classification  and  Separation
■260    ▼a[Sl]▼bUniversity  of  Illinois  at  Urbana-Champaign▼c2023
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2023
■300    ▼a115  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  B.
■500    ▼aAdvisor:  Smaragdis,  Paris.
■5021  ▼aThesis  (Ph.D.)--University  of  Illinois  at  Urbana-Champaign,  2023.
■520    ▼aRecent  advances  in  deep  learning  for  computational  audio  processing  are  established  upon  sufficient  annotated  audio  data.  However,  obtaining  a  substantial  volume  of  high-quality  annotations  from  in-the-wild  audio  remains  a  significant  challenge.  In  this  thesis,  we  propose  and  analyze  data-efficient  approaches  for  modeling  audio  signals  to  perform  sound  classification  and  separation.  First,  we  present  neural  network  architectures  based  on  multi-dimensional  unrolling  of  recurrent  neural  networks  that  allow  the  model  to  perform  sound  event  detection  with  efficient  usage  of  training  data.  Equipped  with  adaptive  computation,  the  model  further  learns  to  intelligently  adjust  the  amount  of  computation  and  enables  processing  when  only  partial  information  is  available.  Next,  we  propose  approaches  for  recognizing  sound  classes  under  a  time-varying  distribution.  We  investigate  continual  learning  techniques  to  train  a  classifier  that  can  efficiently  learn  new  sound  classes  without  forgetting  the  past  using  generative  replay.  We  further  extend  our  approach  to  an  unsupervised  learning  setup,  where  the  model  progressively  learns  representations  for  an  indefinite  number  of  sound  classes  with  few  labels  presented.  Last  but  not  least,  we  investigate  learning  with  limited  annotated  data  using  semi-supervised  learning.  We  demonstrate  the  effectiveness  of  the  proposed  teacher-student  framework  on  tasks  including  cross-modal  audio-text  representation  learning,  singing  voice  separation,  and  personalized  speech  enhancement.  To  this  end,  our  proposed  data-efficient  algorithms  for  audio  classification  and  source  separation  show  high  potential  for  reducing  the  labor  cost  for  collecting  high-quality  annotated  data,  improving  computational  and  storage  efficiency,  and  enabling  processing  on  memory-limited  edge  devices.
■590    ▼aSchool  code:  0090.
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■650  4▼aComputer  engineering
■653    ▼aDeep  learning
■653    ▼aSound  classification
■653    ▼aSource  separation
■653    ▼aSelf-supervised  learning
■653    ▼aSemi-supervised  learning
■653    ▼aContinual  learning
■690    ▼a0984
■690    ▼a0544
■690    ▼a0464
■71020▼aUniversity  of  Illinois  at  Urbana-Champaign▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g87-03B.
■790    ▼a0090
■791    ▼aPh.D.
■792    ▼a2023
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365921▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15721 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.