본문

서브메뉴

Modeling Affect in Speech and Language in the Presence of Natural Inconsistency
Modeling Affect in Speech and Language in the Presence of Natural Inconsistency
Modeling Affect in Speech and Language in the Presence of Natural Inconsistency

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105218
ISBN  
9798291565940
DDC  
621.3
저자명  
Niu, Minxue.
서명/저자  
Modeling Affect in Speech and Language in the Presence of Natural Inconsistency
발행사항  
[Sl] : University of Michigan, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
128 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
주기사항  
Advisor: Mower Provost, Emily.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2025.
초록/해제  
요약Understanding human affect, including emotions and moods, is crucial for a wide range of applications, from improving mental health monitoring to enhancing user experiences in human-computer interaction. The complex, ambiguous, and subjective nature of human affect poses many challenges in developing affect recognition models, such as disagreements in human labels and misaligned signals across modalities (e.g., text and voice). While these inconsistencies present obstacles, they reflect fundamental aspects of human affect and can offer valuable signals for affect modeling if properly harnessed. By addressing these challenges and effectively making use of the information embedded in these inconsistencies, we can work toward building more reliable, interpretable, and human-aligned affect recognition systems.This dissertation investigates the causes and effects of inconsistencies in affect models and explores strategies to mitigate their impact or to use them as informative signals. We focus on emotions expressed through speech and language, the most common modes of communication, particularly in interactions with intelligent systems. First, we examine modality inconsistency: emotions conveyed through text and vocal modalities do not always align. In the mental health domain, we demonstrate that their mismatch carries important information about people's mood and can serve as signals for mood disorder monitoring. We extend this finding and more broadly show that modeling text and acoustic modalities separately allows for disentangling their respective information, yielding richer and more robust speech representations for various speech understanding tasks. Second, we examine annotation inconsistency observed in human emotion labels of text data, highlighting their variability and sensitivity to annotation study designs. We compare human and Large Language Models (LLMs) generated annotations and find that LLMs achieve strong performance. Building on this insight, we propose integrating LLMs into the human annotation workflow, which improves both the annotators' experience and the quality of the labels. We then explore interpersonal inconsistency caused by the subjectivity of emotion perception across individuals. We show that emotion perception significantly differs across demographic and personality groups, and incorporating annotator-level information can improve personalized speech emotion recognition models. Finally, drawing on insights from annotation and interpersonal inconsistencies, we propose a contrastive distillation framework that transfers the generalizable emotion understanding capabilities of LLMs into a lightweight text emotion recognition model. Leveraging the breadth of LLMs' pretraining, the distilled model learns an efficient, emotion-salient embedding space and can seamlessly handle unseen emotion label spaces without extra training.Together, these findings offer a new perspective on the role of inconsistencies in affective modeling. It is important to understand their origins and implications, develop strategies to mitigate unintended ones, and leverage useful signals from them, to achieve a more comprehensive understanding and modeling of human affect.
일반주제명  
Computer engineering
일반주제명  
Computer science
키워드  
Affective computing
키워드  
Emotion recognition
키워드  
Large Language Models
키워드  
Modality inconsistency
기타저자  
University of Michigan Computer Science & Engineering
기본자료저록  
Dissertations Abstracts International. 87-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359811
■00520260202105218
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291565940
■035    ▼a(MiAaPQ)AAI32271782
■035    ▼a(MiAaPQ)umichrackham006462
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aNiu,  Minxue.
■24510▼aModeling  Affect  in  Speech  and  Language  in  the  Presence  of  Natural  Inconsistency
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a128  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  B.
■500    ▼aAdvisor:  Mower  Provost,  Emily.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2025.
■520    ▼aUnderstanding  human  affect,  including  emotions  and  moods,  is  crucial  for  a  wide  range  of  applications,  from  improving  mental  health  monitoring  to  enhancing  user  experiences  in  human-computer  interaction.  The  complex,  ambiguous,  and  subjective  nature  of  human  affect  poses  many  challenges  in  developing  affect  recognition  models,  such  as  disagreements  in  human  labels  and  misaligned  signals  across  modalities  (e.g.,  text  and  voice).  While  these  inconsistencies  present  obstacles,  they  reflect  fundamental  aspects  of  human  affect  and  can  offer  valuable  signals  for  affect  modeling  if  properly  harnessed.  By  addressing  these  challenges  and  effectively  making  use  of  the  information  embedded  in  these  inconsistencies,  we  can  work  toward  building  more  reliable,  interpretable,  and  human-aligned  affect  recognition  systems.This  dissertation  investigates  the  causes  and  effects  of  inconsistencies  in  affect  models  and  explores  strategies  to  mitigate  their  impact  or  to  use  them  as  informative  signals.  We  focus  on  emotions  expressed  through  speech  and  language,  the  most  common  modes  of  communication,  particularly  in  interactions  with  intelligent  systems.  First,  we  examine  modality  inconsistency:  emotions  conveyed  through  text  and  vocal  modalities  do  not  always  align.  In  the  mental  health  domain,  we  demonstrate  that  their  mismatch  carries  important  information  about  people's  mood  and  can  serve  as  signals  for  mood  disorder  monitoring.  We  extend  this  finding  and  more  broadly  show  that  modeling  text  and  acoustic  modalities  separately  allows  for  disentangling  their  respective  information,  yielding  richer  and  more  robust  speech  representations  for  various  speech  understanding  tasks.  Second,  we  examine  annotation  inconsistency  observed  in  human  emotion  labels  of  text  data,  highlighting  their  variability  and  sensitivity  to  annotation  study  designs.  We  compare  human  and  Large  Language  Models  (LLMs)  generated  annotations  and  find  that  LLMs  achieve  strong  performance.  Building  on  this  insight,  we  propose  integrating  LLMs  into  the  human  annotation  workflow,  which  improves  both  the  annotators'  experience  and  the  quality  of  the  labels.  We  then  explore  interpersonal  inconsistency  caused  by  the  subjectivity  of  emotion  perception  across  individuals.  We  show  that  emotion  perception  significantly  differs  across  demographic  and  personality  groups,  and  incorporating  annotator-level  information  can  improve  personalized  speech  emotion  recognition  models.  Finally,  drawing  on  insights  from  annotation  and  interpersonal  inconsistencies,  we  propose  a  contrastive  distillation  framework  that  transfers  the  generalizable  emotion  understanding  capabilities  of  LLMs  into  a  lightweight  text  emotion  recognition  model.  Leveraging  the  breadth  of  LLMs'  pretraining,  the  distilled  model  learns  an  efficient,  emotion-salient  embedding  space  and  can  seamlessly  handle  unseen  emotion  label  spaces  without  extra  training.Together,  these  findings  offer  a  new  perspective  on  the  role  of  inconsistencies  in  affective  modeling.  It  is  important  to  understand  their  origins  and  implications,  develop  strategies  to  mitigate  unintended  ones,  and  leverage  useful  signals  from  them,  to  achieve  a  more  comprehensive  understanding  and  modeling  of  human  affect.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  engineering
■650  4▼aComputer  science
■653    ▼aAffective  computing
■653    ▼aEmotion  recognition
■653    ▼aLarge  Language  Models
■653    ▼aModality  inconsistency
■690    ▼a0464
■690    ▼a0984
■690    ▼a0800
■71020▼aUniversity  of  Michigan▼bComputer  Science  &  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g87-03B.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359811▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15486 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.