서브메뉴
검색
Continual Learning on Speech and Audio: Towards Data, Model and Metrics
Continual Learning on Speech and Audio: Towards Data, Model and Metrics
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152005
- ISBN
- 9798382832937
- DDC
- 004
- 저자명
- Yang, Muqiao.
- 서명/저자
- Continual Learning on Speech and Audio: Towards Data, Model and Metrics
- 발행사항
- [Sl] : Carnegie Mellon University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 99 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
- 주기사항
- Advisor: Ramakrishnan, Bhiksha.
- 학위논문주기
- Thesis (Ph.D.)--Carnegie Mellon University, 2024.
- 초록/해제
- 요약In recent years, the community has witnessed the enormous progress of deep neural network models in matching or even surpassing human performance on a variety of speech and audio tasks, including Automatic Speech Recognition (ASR), Spoken Language Understanding (SLU), Text-to-Speech (TTS), etc. However, their impressive and powerful achievement is predominantly dependent on training with a large set of data defined by a particular and rigid task. In such a paradigm, the model is expected to learn universal knowledge from a static entity of data and stationary environments. In contrast, the real world is inherently ever-changing and non-stationary. New data is often generated and collected every second in a stream format, and novel classes may also emerge from time to time. Without proper adaptation techniques, the knowledge learned in the past might be erased easily when the model is learning subsequent tasks, thus resulting in overall performance degradation. Such a phenomenon is called catastrophic forgetting, which limits the practical use and expansion of many deep neural network models.Continual learning has emerged as a new machine learning paradigm that enables artificial intelligence (AI) systems to learn from a continuous stream of data and incrementally improve their performance over time. By adapting to changing environments and user needs, continual learning aims to address the catastrophic forgetting effect, so that the model can gradually extend the knowledge it acquires without drastically forgetting the knowledge that has been learned in the past. Such a property is crucial in practical applications to enable artificial systems to learn from the infinite streams of data of the changing world in a lifelong manner.This thesis mainly focuses on the underexplored area of how continual learning techniques can be effective in speech and audio tasks via three perspectives: data, model, and metrics. We will introduce the background and formulations of multiple continual learning scenarios, including data-incremental, class-incremental, and task-incremental settings. Then we will present how different categories of continual learning scenarios and methods can be applied to different modules of the modeling pipeline. Starting from the taxonomy of methods, we propose to improve continual learning towards the three perspectives. First, we demonstrate how to address data sampling, selection, and imbalance to help with continual learning on different audio tasks. Second, we show how the joint use of model architecture and data with different learning strategies could benefit continual learning processes. Lastly, we propose new continual evaluation metrics to give us a comprehensive and deeper understanding of the general continual learning behaviors. We believe that this thesis provides an overall exploration of continual learning scenarios in various speech and audio tasks, and makes an important step towards realizing lifelong learning of speech interfaces.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Electrical engineering
- 키워드
- Machine learning
- 기타저자
- Carnegie Mellon University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 85-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162381
■00520250211152005
■006m o d
■007cr#unu||||||||
■020 ▼a9798382832937
■035 ▼a(MiAaPQ)AAI31330553
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aYang, Muqiao.▼0(orcid)0000-0001-6273-0138
■24510▼aContinual Learning on Speech and Audio: Towards Data, Model and Metrics
■260 ▼a[Sl]▼bCarnegie Mellon University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a99 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-12, Section: B.
■500 ▼aAdvisor: Ramakrishnan, Bhiksha.
■5021 ▼aThesis (Ph.D.)--Carnegie Mellon University, 2024.
■520 ▼aIn recent years, the community has witnessed the enormous progress of deep neural network models in matching or even surpassing human performance on a variety of speech and audio tasks, including Automatic Speech Recognition (ASR), Spoken Language Understanding (SLU), Text-to-Speech (TTS), etc. However, their impressive and powerful achievement is predominantly dependent on training with a large set of data defined by a particular and rigid task. In such a paradigm, the model is expected to learn universal knowledge from a static entity of data and stationary environments. In contrast, the real world is inherently ever-changing and non-stationary. New data is often generated and collected every second in a stream format, and novel classes may also emerge from time to time. Without proper adaptation techniques, the knowledge learned in the past might be erased easily when the model is learning subsequent tasks, thus resulting in overall performance degradation. Such a phenomenon is called catastrophic forgetting, which limits the practical use and expansion of many deep neural network models.Continual learning has emerged as a new machine learning paradigm that enables artificial intelligence (AI) systems to learn from a continuous stream of data and incrementally improve their performance over time. By adapting to changing environments and user needs, continual learning aims to address the catastrophic forgetting effect, so that the model can gradually extend the knowledge it acquires without drastically forgetting the knowledge that has been learned in the past. Such a property is crucial in practical applications to enable artificial systems to learn from the infinite streams of data of the changing world in a lifelong manner.This thesis mainly focuses on the underexplored area of how continual learning techniques can be effective in speech and audio tasks via three perspectives: data, model, and metrics. We will introduce the background and formulations of multiple continual learning scenarios, including data-incremental, class-incremental, and task-incremental settings. Then we will present how different categories of continual learning scenarios and methods can be applied to different modules of the modeling pipeline. Starting from the taxonomy of methods, we propose to improve continual learning towards the three perspectives. First, we demonstrate how to address data sampling, selection, and imbalance to help with continual learning on different audio tasks. Second, we show how the joint use of model architecture and data with different learning strategies could benefit continual learning processes. Lastly, we propose new continual evaluation metrics to give us a comprehensive and deeper understanding of the general continual learning behaviors. We believe that this thesis provides an overall exploration of continual learning scenarios in various speech and audio tasks, and makes an important step towards realizing lifelong learning of speech interfaces.
■590 ▼aSchool code: 0041.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aElectrical engineering
■653 ▼aDeep neural network
■653 ▼aContinual learning
■653 ▼aMachine learning
■653 ▼aSpeech interfaces
■653 ▼aSpeech recognition
■653 ▼aLifelong learning
■690 ▼a0984
■690 ▼a0464
■690 ▼a0800
■690 ▼a0544
■71020▼aCarnegie Mellon University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g85-12B.
■790 ▼a0041
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162381▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


