본문

서브메뉴

Towards Better and Privacy-Preserving Speech Modeling for Depression Detection
Towards Better and Privacy-Preserving Speech Modeling for Depression Detection
Towards Better and Privacy-Preserving Speech Modeling for Depression Detection

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152758
ISBN  
9798383684481
DDC  
534
저자명  
Wang, Jinhan.
서명/저자  
Towards Better and Privacy-Preserving Speech Modeling for Depression Detection
발행사항  
[Sl] : University of California, Los Angeles, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
103 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-02, Section: B.
주기사항  
Advisor: Alwan, Abeer A.
학위논문주기  
Thesis (Ph.D.)--University of California, Los Angeles, 2024.
초록/해제  
요약Automatic depression detection systems based on speech signals have recently garnered significant attention. Depression modeling from speech signals, however, faces three challenges. The first challenge is data scarcity. The second, is the risk of privacy exposure in depression detection systems. The third, is the lack of consideration of non-uniformly distributed depression patterns within speech signals. In this dissertation, we address these challenges so that better and privacy-preserving speech-based depression detection systems are built.To address the data scarcity issue, we propose a modified Instance Discriminative Learning (IDL) pre-training method to enable the model to extract augment-invariant and instance-spread-out embeddings from pre-training tasks using out-of-domain unlabeled data. The pre-trained model is then used for initialization of the downstream model, DepAudioNet, and fine-tuned for depression detection tasks. We investigate different augmentation techniques and instance sampling strategies in the pre-training stage. Specifically, we propose a novel sampling strategy, Pseudo-Instance-based Sampling (PIS), to further reveal the correlation between depression characteristics with the underlying acoustic units.Second, to address the privacy-preservation issue, we propose a novel non-uniform speaker disentanglement (NUSD) adversarial learning framework to disentangle speaker-identity information from depression characteristics. The approach utilizes idiosyncratic behaviors of different layers of detection models, and varies the adversarial disentanglement strength of different model components. The method shows that depression detection can be done without an over-reliance on speaker-identity features. More importantly, we found that attenuating more speaker information in the Feature Extraction (FE) module yields better performance than assigning the disentanglement weights uniformly.Third, to address the non-uniformity of depression patterns in speech signals, we propose a novel framework. The framework, Speechformer-CTC, models dynamically varying depression characteristics within speech segments using a Connectionist Temporal Classification (CTC) objective function without the necessity of input-output alignment. Two novel CTC-label generation policies, namely the Expectation-One-Hot and the HuBERT policies, are proposed and incorporated in objectives at various granularities. Additionally, experiments using Automatic Speech Recognition (ASR) features are conducted to demonstrate the compatibility of the proposed method with content-based features. Our findings show that depression detection can benefit from modeling non-uniformly distributed depression patterns and the proposed framework can potentially be used to determine significant depressive regions in speech utterances.Experiments show that the proposed techniques achieve state-of-the-art performance and are validated for both English and Mandarin Chinese.
일반주제명  
Acoustics
일반주제명  
Mental health
일반주제명  
Speech therapy
일반주제명  
Electrical engineering
키워드  
Depression
키워드  
Instance Discriminative Learning
키워드  
HuBERT policies
기타저자  
University of California, Los Angeles Electrical and Computer Engineering 0333
기본자료저록  
Dissertations Abstracts International. 86-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017163826
■00520250211152758
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798383684481
■035    ▼a(MiAaPQ)AAI31556052
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a534
■1001  ▼aWang,  Jinhan.
■24510▼aTowards  Better  and  Privacy-Preserving  Speech  Modeling  for  Depression  Detection
■260    ▼a[Sl]▼bUniversity  of  California,  Los  Angeles▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a103  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-02,  Section:  B.
■500    ▼aAdvisor:  Alwan,  Abeer  A.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Los  Angeles,  2024.
■520    ▼aAutomatic  depression  detection  systems  based  on  speech  signals  have  recently  garnered  significant  attention.  Depression  modeling  from  speech  signals,  however,  faces  three  challenges.  The  first  challenge  is  data  scarcity.  The  second,  is  the  risk  of  privacy  exposure  in  depression  detection  systems.  The  third,  is  the  lack  of  consideration  of  non-uniformly  distributed  depression  patterns  within  speech  signals.  In  this  dissertation,  we  address  these  challenges  so  that  better  and  privacy-preserving  speech-based  depression  detection  systems  are  built.To  address  the  data  scarcity  issue,  we  propose  a  modified  Instance  Discriminative  Learning  (IDL)  pre-training  method  to  enable  the  model  to  extract  augment-invariant  and  instance-spread-out  embeddings  from  pre-training  tasks  using  out-of-domain  unlabeled  data.  The  pre-trained  model  is  then  used  for  initialization  of  the  downstream  model,  DepAudioNet,  and  fine-tuned  for  depression  detection  tasks.  We  investigate  different  augmentation  techniques  and  instance  sampling  strategies  in  the  pre-training  stage.  Specifically,  we  propose  a  novel  sampling  strategy,  Pseudo-Instance-based  Sampling  (PIS),  to  further  reveal  the  correlation  between  depression  characteristics  with  the  underlying  acoustic  units.Second,  to  address  the  privacy-preservation  issue,  we  propose  a  novel  non-uniform  speaker  disentanglement  (NUSD)  adversarial  learning  framework  to  disentangle  speaker-identity  information  from  depression  characteristics.  The  approach  utilizes  idiosyncratic  behaviors  of  different  layers  of  detection  models,  and  varies  the  adversarial  disentanglement  strength  of  different  model  components.  The  method  shows  that  depression  detection  can  be  done  without  an  over-reliance  on  speaker-identity  features.  More  importantly,  we  found  that  attenuating  more  speaker  information  in  the  Feature  Extraction  (FE)  module  yields  better  performance  than  assigning  the  disentanglement  weights  uniformly.Third,  to  address  the  non-uniformity  of  depression  patterns  in  speech  signals,  we  propose  a  novel  framework.  The  framework,  Speechformer-CTC,  models  dynamically  varying  depression  characteristics  within  speech  segments  using  a  Connectionist  Temporal  Classification  (CTC)  objective  function  without  the  necessity  of  input-output  alignment.  Two  novel  CTC-label  generation  policies,  namely  the  Expectation-One-Hot  and  the  HuBERT  policies,  are  proposed  and  incorporated  in  objectives  at  various  granularities.  Additionally,  experiments  using  Automatic  Speech  Recognition  (ASR)  features  are  conducted  to  demonstrate  the  compatibility  of  the  proposed  method  with  content-based  features.  Our  findings  show  that  depression  detection  can  benefit  from  modeling  non-uniformly  distributed  depression  patterns  and  the  proposed  framework  can  potentially  be  used  to  determine  significant  depressive  regions  in  speech  utterances.Experiments  show  that  the  proposed  techniques  achieve  state-of-the-art  performance  and  are  validated  for  both  English  and  Mandarin  Chinese.
■590    ▼aSchool  code:  0031.
■650  4▼aAcoustics
■650  4▼aMental  health
■650  4▼aSpeech  therapy
■650  4▼aElectrical  engineering
■653    ▼aDepression
■653    ▼aInstance  Discriminative  Learning
■653    ▼aHuBERT  policies
■690    ▼a0986
■690    ▼a0544
■690    ▼a0347
■690    ▼a0460
■71020▼aUniversity  of  California,  Los  Angeles▼bElectrical  and  Computer  Engineering  0333.
■7730  ▼tDissertations  Abstracts  International▼g86-02B.
■790    ▼a0031
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163826▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF09815 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.