서브메뉴
검색
Towards Better and Privacy-Preserving Speech Modeling for Depression Detection
Towards Better and Privacy-Preserving Speech Modeling for Depression Detection
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152758
- ISBN
- 9798383684481
- DDC
- 534
- 저자명
- Wang, Jinhan.
- 서명/저자
- Towards Better and Privacy-Preserving Speech Modeling for Depression Detection
- 발행사항
- [Sl] : University of California, Los Angeles, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 103 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-02, Section: B.
- 주기사항
- Advisor: Alwan, Abeer A.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2024.
- 초록/해제
- 요약Automatic depression detection systems based on speech signals have recently garnered significant attention. Depression modeling from speech signals, however, faces three challenges. The first challenge is data scarcity. The second, is the risk of privacy exposure in depression detection systems. The third, is the lack of consideration of non-uniformly distributed depression patterns within speech signals. In this dissertation, we address these challenges so that better and privacy-preserving speech-based depression detection systems are built.To address the data scarcity issue, we propose a modified Instance Discriminative Learning (IDL) pre-training method to enable the model to extract augment-invariant and instance-spread-out embeddings from pre-training tasks using out-of-domain unlabeled data. The pre-trained model is then used for initialization of the downstream model, DepAudioNet, and fine-tuned for depression detection tasks. We investigate different augmentation techniques and instance sampling strategies in the pre-training stage. Specifically, we propose a novel sampling strategy, Pseudo-Instance-based Sampling (PIS), to further reveal the correlation between depression characteristics with the underlying acoustic units.Second, to address the privacy-preservation issue, we propose a novel non-uniform speaker disentanglement (NUSD) adversarial learning framework to disentangle speaker-identity information from depression characteristics. The approach utilizes idiosyncratic behaviors of different layers of detection models, and varies the adversarial disentanglement strength of different model components. The method shows that depression detection can be done without an over-reliance on speaker-identity features. More importantly, we found that attenuating more speaker information in the Feature Extraction (FE) module yields better performance than assigning the disentanglement weights uniformly.Third, to address the non-uniformity of depression patterns in speech signals, we propose a novel framework. The framework, Speechformer-CTC, models dynamically varying depression characteristics within speech segments using a Connectionist Temporal Classification (CTC) objective function without the necessity of input-output alignment. Two novel CTC-label generation policies, namely the Expectation-One-Hot and the HuBERT policies, are proposed and incorporated in objectives at various granularities. Additionally, experiments using Automatic Speech Recognition (ASR) features are conducted to demonstrate the compatibility of the proposed method with content-based features. Our findings show that depression detection can benefit from modeling non-uniformly distributed depression patterns and the proposed framework can potentially be used to determine significant depressive regions in speech utterances.Experiments show that the proposed techniques achieve state-of-the-art performance and are validated for both English and Mandarin Chinese.
- 일반주제명
- Acoustics
- 일반주제명
- Mental health
- 일반주제명
- Speech therapy
- 일반주제명
- Electrical engineering
- 키워드
- Depression
- 키워드
- HuBERT policies
- 기타저자
- University of California, Los Angeles Electrical and Computer Engineering 0333
- 기본자료저록
- Dissertations Abstracts International. 86-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163826
■00520250211152758
■006m o d
■007cr#unu||||||||
■020 ▼a9798383684481
■035 ▼a(MiAaPQ)AAI31556052
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a534
■1001 ▼aWang, Jinhan.
■24510▼aTowards Better and Privacy-Preserving Speech Modeling for Depression Detection
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a103 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-02, Section: B.
■500 ▼aAdvisor: Alwan, Abeer A.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2024.
■520 ▼aAutomatic depression detection systems based on speech signals have recently garnered significant attention. Depression modeling from speech signals, however, faces three challenges. The first challenge is data scarcity. The second, is the risk of privacy exposure in depression detection systems. The third, is the lack of consideration of non-uniformly distributed depression patterns within speech signals. In this dissertation, we address these challenges so that better and privacy-preserving speech-based depression detection systems are built.To address the data scarcity issue, we propose a modified Instance Discriminative Learning (IDL) pre-training method to enable the model to extract augment-invariant and instance-spread-out embeddings from pre-training tasks using out-of-domain unlabeled data. The pre-trained model is then used for initialization of the downstream model, DepAudioNet, and fine-tuned for depression detection tasks. We investigate different augmentation techniques and instance sampling strategies in the pre-training stage. Specifically, we propose a novel sampling strategy, Pseudo-Instance-based Sampling (PIS), to further reveal the correlation between depression characteristics with the underlying acoustic units.Second, to address the privacy-preservation issue, we propose a novel non-uniform speaker disentanglement (NUSD) adversarial learning framework to disentangle speaker-identity information from depression characteristics. The approach utilizes idiosyncratic behaviors of different layers of detection models, and varies the adversarial disentanglement strength of different model components. The method shows that depression detection can be done without an over-reliance on speaker-identity features. More importantly, we found that attenuating more speaker information in the Feature Extraction (FE) module yields better performance than assigning the disentanglement weights uniformly.Third, to address the non-uniformity of depression patterns in speech signals, we propose a novel framework. The framework, Speechformer-CTC, models dynamically varying depression characteristics within speech segments using a Connectionist Temporal Classification (CTC) objective function without the necessity of input-output alignment. Two novel CTC-label generation policies, namely the Expectation-One-Hot and the HuBERT policies, are proposed and incorporated in objectives at various granularities. Additionally, experiments using Automatic Speech Recognition (ASR) features are conducted to demonstrate the compatibility of the proposed method with content-based features. Our findings show that depression detection can benefit from modeling non-uniformly distributed depression patterns and the proposed framework can potentially be used to determine significant depressive regions in speech utterances.Experiments show that the proposed techniques achieve state-of-the-art performance and are validated for both English and Mandarin Chinese.
■590 ▼aSchool code: 0031.
■650 4▼aAcoustics
■650 4▼aMental health
■650 4▼aSpeech therapy
■650 4▼aElectrical engineering
■653 ▼aDepression
■653 ▼aInstance Discriminative Learning
■653 ▼aHuBERT policies
■690 ▼a0986
■690 ▼a0544
■690 ▼a0347
■690 ▼a0460
■71020▼aUniversity of California, Los Angeles▼bElectrical and Computer Engineering 0333.
■7730 ▼tDissertations Abstracts International▼g86-02B.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163826▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


