본문

서브메뉴

Accuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches Through Data Augmentation, and Speaker Identity Disentanglement
Accuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches T...
Accuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches Through Data Augmentation, and Speaker Identity Disentanglement

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211153120
ISBN  
9798346852056
DDC  
621.3
저자명  
Ravi, Vijay.
서명/저자  
Accuracy and Privacy in Speech-Based Modeling of Major Depression: Innovative Approaches Through Data Augmentation, and Speaker Identity Disentanglement
발행사항  
[Sl] : University of California, Los Angeles, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
140 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-06, Section: B.
주기사항  
Advisor: Alwan, Abeer.
학위논문주기  
Thesis (Ph.D.)--University of California, Los Angeles, 2024.
초록/해제  
요약Major Depressive Disorder (MDD) is a prevalent mental illness that affects a significant portion of the global population. Despite its severity, traditional diagnostic methods often fail to identify and treat MDD effectively, highlighting the need for automated diagnostic tools. Recent research has identified speech signals as promising biomarkers for objectively detecting depression. However, the development of speech-based depression detection systems faces several challenges including data scarcity and privacy preservation. The sensitive nature of mental health data makes it difficult to collect large datasets required for training robust models. Moreover, many current approaches rely on features that can compromise patient confidentiality, hindering the adoption of these systems in clinical settings. This thesis presents novel methods to address these challenges and to enhance the performance and privacy of speech-based depression detection. The contributions include a frame rate-based data augmentation technique (FrAUG) to increase training data while preserving depression-related acoustic information. Additionally, five speaker identity disentanglement methods are proposed: adversarial loss maximization, loss equalization via Cross-Entropy, Variance, and KL Divergence, and unsupervised speaker disentanglement via cosine similarity minimization. These methods aim to reduce the reliance on speaker identity during depression detection. The proposed techniques are evaluated on multiple datasets in two languages - English (DAIC-WoZ dataset) and Mandarin (EATD and CONVERGE datasets), demonstrating improved depression detection accuracy and reduced speaker separability compared to state-of-the-art approaches. Furthermore, the privacy preservation capabilities of these methods are quantified using gain of voice distinctiveness and de-identification scores, showcasing their potential for safeguarding patient privacy. By advancing speech-based depression detection in terms of accuracy and privacy, this thesis aims to facilitate the development of effective and secure diagnostic tools that can be readily adopted in clinical settings.
일반주제명  
Computer engineering
일반주제명  
Mental health
일반주제명  
Computer science
일반주제명  
Information technology
키워드  
Depression detection
키워드  
Privacy-preserving
키워드  
Speaker disentanglement
키워드  
Speech processing
키워드  
Data augmentation
키워드  
Major Depressive Disorder
기타저자  
University of California, Los Angeles Electrical and Computer Engineering 0333
기본자료저록  
Dissertations Abstracts International. 86-06B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017165074
■00520250211153120
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798346852056
■035    ▼a(MiAaPQ)AAI31761642
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aRavi,  Vijay.
■24510▼aAccuracy  and  Privacy  in  Speech-Based  Modeling  of  Major  Depression:  Innovative  Approaches  Through  Data  Augmentation,  and  Speaker  Identity  Disentanglement
■260    ▼a[Sl]▼bUniversity  of  California,  Los  Angeles▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a140  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-06,  Section:  B.
■500    ▼aAdvisor:  Alwan,  Abeer.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Los  Angeles,  2024.
■520    ▼aMajor  Depressive  Disorder  (MDD)  is  a  prevalent  mental  illness  that  affects  a  significant  portion  of  the  global  population.  Despite  its  severity,  traditional  diagnostic  methods  often  fail  to  identify  and  treat  MDD  effectively,  highlighting  the  need  for  automated  diagnostic  tools.  Recent  research  has  identified  speech  signals  as  promising  biomarkers  for  objectively  detecting  depression.  However,  the  development  of  speech-based  depression  detection  systems  faces  several  challenges  including  data  scarcity  and  privacy  preservation.  The  sensitive  nature  of  mental  health  data  makes  it  difficult  to  collect  large  datasets  required  for  training  robust  models.  Moreover,  many  current  approaches  rely  on  features  that  can  compromise  patient  confidentiality,  hindering  the  adoption  of  these  systems  in  clinical  settings.  This  thesis  presents  novel  methods  to  address  these  challenges  and  to  enhance  the  performance  and  privacy  of  speech-based  depression  detection.  The  contributions  include  a  frame  rate-based  data  augmentation  technique  (FrAUG)  to  increase  training  data  while  preserving  depression-related  acoustic  information.  Additionally,  five  speaker  identity  disentanglement  methods  are  proposed:  adversarial  loss  maximization,  loss  equalization  via  Cross-Entropy,  Variance,  and  KL  Divergence,  and  unsupervised  speaker  disentanglement  via  cosine  similarity  minimization.  These  methods  aim  to  reduce  the  reliance  on  speaker  identity  during  depression  detection.  The  proposed  techniques  are  evaluated  on  multiple  datasets  in  two  languages  -  English  (DAIC-WoZ  dataset)  and  Mandarin  (EATD  and  CONVERGE  datasets),  demonstrating  improved  depression  detection  accuracy  and  reduced  speaker  separability  compared  to  state-of-the-art  approaches.  Furthermore,  the  privacy  preservation  capabilities  of  these  methods  are  quantified  using  gain  of  voice  distinctiveness  and  de-identification  scores,  showcasing  their  potential  for  safeguarding  patient  privacy.  By  advancing  speech-based  depression  detection  in  terms  of  accuracy  and  privacy,  this  thesis  aims  to  facilitate  the  development  of  effective  and  secure  diagnostic  tools  that  can  be  readily  adopted  in  clinical  settings.
■590    ▼aSchool  code:  0031.
■650  4▼aComputer  engineering
■650  4▼aMental  health
■650  4▼aComputer  science
■650  4▼aInformation  technology
■653    ▼aDepression  detection
■653    ▼aPrivacy-preserving
■653    ▼aSpeaker  disentanglement
■653    ▼aSpeech  processing
■653    ▼aData  augmentation
■653    ▼aMajor  Depressive  Disorder
■690    ▼a0984
■690    ▼a0489
■690    ▼a0464
■690    ▼a0347
■71020▼aUniversity  of  California,  Los  Angeles▼bElectrical  and  Computer  Engineering  0333.
■7730  ▼tDissertations  Abstracts  International▼g86-06B.
■790    ▼a0031
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17165074▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF12672 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.