서브메뉴
검색
Modeling Affect in Speech and Language in the Presence of Natural Inconsistency
Modeling Affect in Speech and Language in the Presence of Natural Inconsistency
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105218
- ISBN
- 9798291565940
- DDC
- 621.3
- 저자명
- Niu, Minxue.
- 서명/저자
- Modeling Affect in Speech and Language in the Presence of Natural Inconsistency
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 128 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Mower Provost, Emily.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약Understanding human affect, including emotions and moods, is crucial for a wide range of applications, from improving mental health monitoring to enhancing user experiences in human-computer interaction. The complex, ambiguous, and subjective nature of human affect poses many challenges in developing affect recognition models, such as disagreements in human labels and misaligned signals across modalities (e.g., text and voice). While these inconsistencies present obstacles, they reflect fundamental aspects of human affect and can offer valuable signals for affect modeling if properly harnessed. By addressing these challenges and effectively making use of the information embedded in these inconsistencies, we can work toward building more reliable, interpretable, and human-aligned affect recognition systems.This dissertation investigates the causes and effects of inconsistencies in affect models and explores strategies to mitigate their impact or to use them as informative signals. We focus on emotions expressed through speech and language, the most common modes of communication, particularly in interactions with intelligent systems. First, we examine modality inconsistency: emotions conveyed through text and vocal modalities do not always align. In the mental health domain, we demonstrate that their mismatch carries important information about people's mood and can serve as signals for mood disorder monitoring. We extend this finding and more broadly show that modeling text and acoustic modalities separately allows for disentangling their respective information, yielding richer and more robust speech representations for various speech understanding tasks. Second, we examine annotation inconsistency observed in human emotion labels of text data, highlighting their variability and sensitivity to annotation study designs. We compare human and Large Language Models (LLMs) generated annotations and find that LLMs achieve strong performance. Building on this insight, we propose integrating LLMs into the human annotation workflow, which improves both the annotators' experience and the quality of the labels. We then explore interpersonal inconsistency caused by the subjectivity of emotion perception across individuals. We show that emotion perception significantly differs across demographic and personality groups, and incorporating annotator-level information can improve personalized speech emotion recognition models. Finally, drawing on insights from annotation and interpersonal inconsistencies, we propose a contrastive distillation framework that transfers the generalizable emotion understanding capabilities of LLMs into a lightweight text emotion recognition model. Leveraging the breadth of LLMs' pretraining, the distilled model learns an efficient, emotion-salient embedding space and can seamlessly handle unseen emotion label spaces without extra training.Together, these findings offer a new perspective on the role of inconsistencies in affective modeling. It is important to understand their origins and implications, develop strategies to mitigate unintended ones, and leverage useful signals from them, to achieve a more comprehensive understanding and modeling of human affect.
- 일반주제명
- Computer engineering
- 일반주제명
- Computer science
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359811
■00520260202105218
■006m o d
■007cr#unu||||||||
■020 ▼a9798291565940
■035 ▼a(MiAaPQ)AAI32271782
■035 ▼a(MiAaPQ)umichrackham006462
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aNiu, Minxue.
■24510▼aModeling Affect in Speech and Language in the Presence of Natural Inconsistency
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a128 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Mower Provost, Emily.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aUnderstanding human affect, including emotions and moods, is crucial for a wide range of applications, from improving mental health monitoring to enhancing user experiences in human-computer interaction. The complex, ambiguous, and subjective nature of human affect poses many challenges in developing affect recognition models, such as disagreements in human labels and misaligned signals across modalities (e.g., text and voice). While these inconsistencies present obstacles, they reflect fundamental aspects of human affect and can offer valuable signals for affect modeling if properly harnessed. By addressing these challenges and effectively making use of the information embedded in these inconsistencies, we can work toward building more reliable, interpretable, and human-aligned affect recognition systems.This dissertation investigates the causes and effects of inconsistencies in affect models and explores strategies to mitigate their impact or to use them as informative signals. We focus on emotions expressed through speech and language, the most common modes of communication, particularly in interactions with intelligent systems. First, we examine modality inconsistency: emotions conveyed through text and vocal modalities do not always align. In the mental health domain, we demonstrate that their mismatch carries important information about people's mood and can serve as signals for mood disorder monitoring. We extend this finding and more broadly show that modeling text and acoustic modalities separately allows for disentangling their respective information, yielding richer and more robust speech representations for various speech understanding tasks. Second, we examine annotation inconsistency observed in human emotion labels of text data, highlighting their variability and sensitivity to annotation study designs. We compare human and Large Language Models (LLMs) generated annotations and find that LLMs achieve strong performance. Building on this insight, we propose integrating LLMs into the human annotation workflow, which improves both the annotators' experience and the quality of the labels. We then explore interpersonal inconsistency caused by the subjectivity of emotion perception across individuals. We show that emotion perception significantly differs across demographic and personality groups, and incorporating annotator-level information can improve personalized speech emotion recognition models. Finally, drawing on insights from annotation and interpersonal inconsistencies, we propose a contrastive distillation framework that transfers the generalizable emotion understanding capabilities of LLMs into a lightweight text emotion recognition model. Leveraging the breadth of LLMs' pretraining, the distilled model learns an efficient, emotion-salient embedding space and can seamlessly handle unseen emotion label spaces without extra training.Together, these findings offer a new perspective on the role of inconsistencies in affective modeling. It is important to understand their origins and implications, develop strategies to mitigate unintended ones, and leverage useful signals from them, to achieve a more comprehensive understanding and modeling of human affect.
■590 ▼aSchool code: 0127.
■650 4▼aComputer engineering
■650 4▼aComputer science
■653 ▼aAffective computing
■653 ▼aEmotion recognition
■653 ▼aLarge Language Models
■653 ▼aModality inconsistency
■690 ▼a0464
■690 ▼a0984
■690 ▼a0800
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359811▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


