서브메뉴
검색
"Seeing Red" or "Tickled Pink"?: Investigating the Power of Language and Vision Models Through Color, Emotion, and Metaphor- [electronic resource]
"Seeing Red" or "Tickled Pink"?: Investigating the Power of Language and Vision Models Through Color, Emotion, and Metaphor- [electronic resource]
상세정보
- 자료유형
- 학위논문파일 국외
- 최종처리일시
- 20240214101257
- ISBN
- 9798380091695
- DDC
- 621.3
- 저자명
- Winn, Olivia.
- 서명/저자
- Seeing Red or Tickled Pink?: Investigating the Power of Language and Vision Models Through Color, Emotion, and Metaphor - [electronic resource]
- 발행사항
- [S.l.]: : Columbia University., 2023
- 발행사항
- Ann Arbor : : ProQuest Dissertations & Theses,, 2023
- 형태사항
- 1 online resource(142 p.)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-02, Section: B.
- 주기사항
- Advisor: Muresan, Smaranda.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2023.
- 사용제한주기
- This item must not be sold to any third party vendors.
- 초록/해제
- 요약Multimodal NLP is an approach to language understanding that incorporates data from nontextual in order to enhance our linguistic understanding through additional contextual information. In particular, incorporating visual data has allowed for great strides in our ability to model language related to physical phenomena. The performance of these models has so far been contingent upon access to large datasets, focusing on classification problems without relative information, and constraining the problem space to literal descriptions and interpretations. In this thesis, we examine these limitations by investigating how types of data previously unused in these models can be reconfigured and worked with intelligently and on a small scale to enhance our understanding of the pragmatics of language.We contribute to comparative language grounding, emotional interpretation, and metaphoric understanding by releasing multiple annotated datasets, developing a new paradigm for modeling relative data, creating a new task in examining the generation of emotional descriptions for image information, and demonstrating a novel approach to working with figurative text for image generation.We start by examining how traditional grounding models could be adapted to incorporate relative information. As no previous work has ever utilized relative textual description for imageunderstanding, we first constrain the problem by focusing on the language of color. We create a new dataset of comparative color terms with associated RGB datapoints, and use this data to develop a novel paradigm of grounding comparative color terms in RGB space, providing the first avenue towards utilizing relative information in a multimodal setting.Continuing our study of color, we then turn to examining the relationship between color and emotion. In order to further our understanding of this relationship, we define a new task, called Justified Affect Transformation, in which an image is recolored specifically to alter its emotional evocation and text is generated to explain the recoloring from an emotional perspective. We create a dataset of abstract art with contiguous emotion labels and textual rationales for the emotional evocation of multiple images, and using our new dataset for training, introduce a new unified model that recolors an image and provides a textual rationale explaining the recoloring with respect to the specified emotion. We use this model to examine the relationship between color and emotion devoid of confounding factors.Finally, we turn to figurative language as a resource, examining the pragmatics of visualizing metaphoric phrases. We demonstrate a novel approach to generating visual metaphor through the collaboration of large language models and diffusion-based text-to-image models, and in doing so create a novel dataset of visual metaphor with both literal and figurative captions. We then develop an evaluation framework using human-AI collaboration to examine the efficacy of the model collaboration, and choose a downstream task of visual entailment to evaluate the human-AI collaboration.
- 일반주제명
- Computer engineering.
- 키워드
- Computer vision
- 키워드
- Image generation
- 키워드
- Text generation
- 기타저자
- Columbia University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-02B.
- 기본자료저록
- Dissertation Abstract International
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008240612s2023 us |||||||||||||||c||eng d■001000016933526
■00520240214101257
■006m o d
■007cr#unu||||||||
■020 ▼a9798380091695
■035 ▼a(MiAaPQ)AAI30530597
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aWinn, Olivia.
■24510▼a"Seeing Red" or "Tickled Pink"?: Investigating the Power of Language and Vision Models Through Color, Emotion, and Metaphor▼h[electronic resource]
■260 ▼a[S.l.]:▼bColumbia University. ▼c2023
■260 1▼aAnn Arbor :▼bProQuest Dissertations & Theses, ▼c2023
■300 ▼a1 online resource(142 p.)
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-02, Section: B.
■500 ▼aAdvisor: Muresan, Smaranda.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2023.
■506 ▼aThis item must not be sold to any third party vendors.
■520 ▼aMultimodal NLP is an approach to language understanding that incorporates data from nontextual in order to enhance our linguistic understanding through additional contextual information. In particular, incorporating visual data has allowed for great strides in our ability to model language related to physical phenomena. The performance of these models has so far been contingent upon access to large datasets, focusing on classification problems without relative information, and constraining the problem space to literal descriptions and interpretations. In this thesis, we examine these limitations by investigating how types of data previously unused in these models can be reconfigured and worked with intelligently and on a small scale to enhance our understanding of the pragmatics of language.We contribute to comparative language grounding, emotional interpretation, and metaphoric understanding by releasing multiple annotated datasets, developing a new paradigm for modeling relative data, creating a new task in examining the generation of emotional descriptions for image information, and demonstrating a novel approach to working with figurative text for image generation.We start by examining how traditional grounding models could be adapted to incorporate relative information. As no previous work has ever utilized relative textual description for imageunderstanding, we first constrain the problem by focusing on the language of color. We create a new dataset of comparative color terms with associated RGB datapoints, and use this data to develop a novel paradigm of grounding comparative color terms in RGB space, providing the first avenue towards utilizing relative information in a multimodal setting.Continuing our study of color, we then turn to examining the relationship between color and emotion. In order to further our understanding of this relationship, we define a new task, called Justified Affect Transformation, in which an image is recolored specifically to alter its emotional evocation and text is generated to explain the recoloring from an emotional perspective. We create a dataset of abstract art with contiguous emotion labels and textual rationales for the emotional evocation of multiple images, and using our new dataset for training, introduce a new unified model that recolors an image and provides a textual rationale explaining the recoloring with respect to the specified emotion. We use this model to examine the relationship between color and emotion devoid of confounding factors.Finally, we turn to figurative language as a resource, examining the pragmatics of visualizing metaphoric phrases. We demonstrate a novel approach to generating visual metaphor through the collaboration of large language models and diffusion-based text-to-image models, and in doing so create a novel dataset of visual metaphor with both literal and figurative captions. We then develop an evaluation framework using human-AI collaboration to examine the efficacy of the model collaboration, and choose a downstream task of visual entailment to evaluate the human-AI collaboration.
■590 ▼aSchool code: 0054.
■650 4▼aComputer engineering.
■653 ▼aComputational linguistics
■653 ▼aComputer vision
■653 ▼aImage generation
■653 ▼aLanguage grounding
■653 ▼aText generation
■690 ▼a0800
■690 ▼a0464
■71020▼aColumbia University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-02B.
■773 ▼tDissertation Abstract International
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T16933526▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
■980 ▼a202402▼f2024


