서브메뉴
검색
Accuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches Through Data Augmentation, and Speaker Identity Disentanglement
Accuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches Through Data Augmentation, and Speaker Identity Disentanglement
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153120
- ISBN
- 9798346852056
- DDC
- 621.3
- 저자명
- Ravi, Vijay.
- 서명/저자
- Accuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches Through Data Augmentation, and Speaker Identity Disentanglement
- 발행사항
- [Sl] : University of California, Los Angeles, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 140 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-06, Section: B.
- 주기사항
- Advisor: Alwan, Abeer.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2024.
- 초록/해제
- 요약Major Depressive Disorder (MDD) is a prevalent mental illness that affects a significant portion of the global population. Despite its severity, traditional diagnostic methods often fail to identify and treat MDD effectively, highlighting the need for automated diagnostic tools. Recent research has identified speech signals as promising biomarkers for objectively detecting depression. However, the development of speech-based depression detection systems faces several challenges including data scarcity and privacy preservation. The sensitive nature of mental health data makes it difficult to collect large datasets required for training robust models. Moreover, many current approaches rely on features that can compromise patient confidentiality, hindering the adoption of these systems in clinical settings. This thesis presents novel methods to address these challenges and to enhance the performance and privacy of speech-based depression detection. The contributions include a frame rate-based data augmentation technique (FrAUG) to increase training data while preserving depression-related acoustic information. Additionally, five speaker identity disentanglement methods are proposed: adversarial loss maximization, loss equalization via Cross-Entropy, Variance, and KL Divergence, and unsupervised speaker disentanglement via cosine similarity minimization. These methods aim to reduce the reliance on speaker identity during depression detection. The proposed techniques are evaluated on multiple datasets in two languages - English (DAIC-WoZ dataset) and Mandarin (EATD and CONVERGE datasets), demonstrating improved depression detection accuracy and reduced speaker separability compared to state-of-the-art approaches. Furthermore, the privacy preservation capabilities of these methods are quantified using gain of voice distinctiveness and de-identification scores, showcasing their potential for safeguarding patient privacy. By advancing speech-based depression detection in terms of accuracy and privacy, this thesis aims to facilitate the development of effective and secure diagnostic tools that can be readily adopted in clinical settings.
- 일반주제명
- Computer engineering
- 일반주제명
- Mental health
- 일반주제명
- Computer science
- 일반주제명
- Information technology
- 기타저자
- University of California, Los Angeles Electrical and Computer Engineering 0333
- 기본자료저록
- Dissertations Abstracts International. 86-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017165074
■00520250211153120
■006m o d
■007cr#unu||||||||
■020 ▼a9798346852056
■035 ▼a(MiAaPQ)AAI31761642
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aRavi, Vijay.
■24510▼aAccuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches Through Data Augmentation, and Speaker Identity Disentanglement
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a140 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-06, Section: B.
■500 ▼aAdvisor: Alwan, Abeer.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2024.
■520 ▼aMajor Depressive Disorder (MDD) is a prevalent mental illness that affects a significant portion of the global population. Despite its severity, traditional diagnostic methods often fail to identify and treat MDD effectively, highlighting the need for automated diagnostic tools. Recent research has identified speech signals as promising biomarkers for objectively detecting depression. However, the development of speech-based depression detection systems faces several challenges including data scarcity and privacy preservation. The sensitive nature of mental health data makes it difficult to collect large datasets required for training robust models. Moreover, many current approaches rely on features that can compromise patient confidentiality, hindering the adoption of these systems in clinical settings. This thesis presents novel methods to address these challenges and to enhance the performance and privacy of speech-based depression detection. The contributions include a frame rate-based data augmentation technique (FrAUG) to increase training data while preserving depression-related acoustic information. Additionally, five speaker identity disentanglement methods are proposed: adversarial loss maximization, loss equalization via Cross-Entropy, Variance, and KL Divergence, and unsupervised speaker disentanglement via cosine similarity minimization. These methods aim to reduce the reliance on speaker identity during depression detection. The proposed techniques are evaluated on multiple datasets in two languages - English (DAIC-WoZ dataset) and Mandarin (EATD and CONVERGE datasets), demonstrating improved depression detection accuracy and reduced speaker separability compared to state-of-the-art approaches. Furthermore, the privacy preservation capabilities of these methods are quantified using gain of voice distinctiveness and de-identification scores, showcasing their potential for safeguarding patient privacy. By advancing speech-based depression detection in terms of accuracy and privacy, this thesis aims to facilitate the development of effective and secure diagnostic tools that can be readily adopted in clinical settings.
■590 ▼aSchool code: 0031.
■650 4▼aComputer engineering
■650 4▼aMental health
■650 4▼aComputer science
■650 4▼aInformation technology
■653 ▼aDepression detection
■653 ▼aPrivacy-preserving
■653 ▼aSpeaker disentanglement
■653 ▼aSpeech processing
■653 ▼aData augmentation
■653 ▼aMajor Depressive Disorder
■690 ▼a0984
■690 ▼a0489
■690 ▼a0464
■690 ▼a0347
■71020▼aUniversity of California, Los Angeles▼bElectrical and Computer Engineering 0333.
■7730 ▼tDissertations Abstracts International▼g86-06B.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17165074▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


