서브메뉴
검색
Data-Efficient Approaches for Audio Classification and Separation
Data-Efficient Approaches for Audio Classification and Separation
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260209102855
- ISBN
- 9798291574317
- DDC
- 004
- 저자명
- Wang, Zhepei.
- 서명/저자
- Data-Efficient Approaches for Audio Classification and Separation
- 발행사항
- [Sl] : University of Illinois at Urbana-Champaign, 2023
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2023
- 형태사항
- 115 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Smaragdis, Paris.
- 학위논문주기
- Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
- 초록/해제
- 요약Recent advances in deep learning for computational audio processing are established upon sufficient annotated audio data. However, obtaining a substantial volume of high-quality annotations from in-the-wild audio remains a significant challenge. In this thesis, we propose and analyze data-efficient approaches for modeling audio signals to perform sound classification and separation. First, we present neural network architectures based on multi-dimensional unrolling of recurrent neural networks that allow the model to perform sound event detection with efficient usage of training data. Equipped with adaptive computation, the model further learns to intelligently adjust the amount of computation and enables processing when only partial information is available. Next, we propose approaches for recognizing sound classes under a time-varying distribution. We investigate continual learning techniques to train a classifier that can efficiently learn new sound classes without forgetting the past using generative replay. We further extend our approach to an unsupervised learning setup, where the model progressively learns representations for an indefinite number of sound classes with few labels presented. Last but not least, we investigate learning with limited annotated data using semi-supervised learning. We demonstrate the effectiveness of the proposed teacher-student framework on tasks including cross-modal audio-text representation learning, singing voice separation, and personalized speech enhancement. To this end, our proposed data-efficient algorithms for audio classification and source separation show high potential for reducing the labor cost for collecting high-quality annotated data, improving computational and storage efficiency, and enabling processing on memory-limited edge devices.
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 일반주제명
- Computer engineering
- 키워드
- Deep learning
- 기타저자
- University of Illinois at Urbana-Champaign Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260203s2023 us c eng d■001000017365921
■00520260209102855
■006m o d
■007cr#unu||||||||
■020 ▼a9798291574317
■035 ▼a(MiAaPQ)AAI32272133
■035 ▼a(MiAaPQ)httphdlhandlenet2142121935
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWang, Zhepei.
■24510▼aData-Efficient Approaches for Audio Classification and Separation
■260 ▼a[Sl]▼bUniversity of Illinois at Urbana-Champaign▼c2023
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2023
■300 ▼a115 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Smaragdis, Paris.
■5021 ▼aThesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
■520 ▼aRecent advances in deep learning for computational audio processing are established upon sufficient annotated audio data. However, obtaining a substantial volume of high-quality annotations from in-the-wild audio remains a significant challenge. In this thesis, we propose and analyze data-efficient approaches for modeling audio signals to perform sound classification and separation. First, we present neural network architectures based on multi-dimensional unrolling of recurrent neural networks that allow the model to perform sound event detection with efficient usage of training data. Equipped with adaptive computation, the model further learns to intelligently adjust the amount of computation and enables processing when only partial information is available. Next, we propose approaches for recognizing sound classes under a time-varying distribution. We investigate continual learning techniques to train a classifier that can efficiently learn new sound classes without forgetting the past using generative replay. We further extend our approach to an unsupervised learning setup, where the model progressively learns representations for an indefinite number of sound classes with few labels presented. Last but not least, we investigate learning with limited annotated data using semi-supervised learning. We demonstrate the effectiveness of the proposed teacher-student framework on tasks including cross-modal audio-text representation learning, singing voice separation, and personalized speech enhancement. To this end, our proposed data-efficient algorithms for audio classification and source separation show high potential for reducing the labor cost for collecting high-quality annotated data, improving computational and storage efficiency, and enabling processing on memory-limited edge devices.
■590 ▼aSchool code: 0090.
■650 4▼aComputer science
■650 4▼aElectrical engineering
■650 4▼aComputer engineering
■653 ▼aDeep learning
■653 ▼aSound classification
■653 ▼aSource separation
■653 ▼aSelf-supervised learning
■653 ▼aSemi-supervised learning
■653 ▼aContinual learning
■690 ▼a0984
■690 ▼a0544
■690 ▼a0464
■71020▼aUniversity of Illinois at Urbana-Champaign▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0090
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365921▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


