본문

서브메뉴

"Seeing Red" or "Tickled Pink"?: Investigating the Power of Language and Vision Models Through Color, Emotion, and Metaphor- [electronic resource]
"Seeing Red" or "Tickled Pink"?: Investigating the Power of Language and Vision Models Thr...
"Seeing Red" or "Tickled Pink"?: Investigating the Power of Language and Vision Models Through Color, Emotion, and Metaphor- [electronic resource]

상세정보

자료유형  
 학위논문파일 국외
최종처리일시  
20240214101257
ISBN  
9798380091695
DDC  
621.3
저자명  
Winn, Olivia.
서명/저자  
Seeing Red or Tickled Pink?: Investigating the Power of Language and Vision Models Through Color, Emotion, and Metaphor - [electronic resource]
발행사항  
[S.l.]: : Columbia University., 2023
발행사항  
Ann Arbor : : ProQuest Dissertations & Theses,, 2023
형태사항  
1 online resource(142 p.)
주기사항  
Source: Dissertations Abstracts International, Volume: 85-02, Section: B.
주기사항  
Advisor: Muresan, Smaranda.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2023.
사용제한주기  
This item must not be sold to any third party vendors.
초록/해제  
요약Multimodal NLP is an approach to language understanding that incorporates data from nontextual in order to enhance our linguistic understanding through additional contextual information. In particular, incorporating visual data has allowed for great strides in our ability to model language related to physical phenomena. The performance of these models has so far been contingent upon access to large datasets, focusing on classification problems without relative information, and constraining the problem space to literal descriptions and interpretations. In this thesis, we examine these limitations by investigating how types of data previously unused in these models can be reconfigured and worked with intelligently and on a small scale to enhance our understanding of the pragmatics of language.We contribute to comparative language grounding, emotional interpretation, and metaphoric understanding by releasing multiple annotated datasets, developing a new paradigm for modeling relative data, creating a new task in examining the generation of emotional descriptions for image information, and demonstrating a novel approach to working with figurative text for image generation.We start by examining how traditional grounding models could be adapted to incorporate relative information. As no previous work has ever utilized relative textual description for imageunderstanding, we first constrain the problem by focusing on the language of color. We create a new dataset of comparative color terms with associated RGB datapoints, and use this data to develop a novel paradigm of grounding comparative color terms in RGB space, providing the first avenue towards utilizing relative information in a multimodal setting.Continuing our study of color, we then turn to examining the relationship between color and emotion. In order to further our understanding of this relationship, we define a new task, called Justified Affect Transformation, in which an image is recolored specifically to alter its emotional evocation and text is generated to explain the recoloring from an emotional perspective. We create a dataset of abstract art with contiguous emotion labels and textual rationales for the emotional evocation of multiple images, and using our new dataset for training, introduce a new unified model that recolors an image and provides a textual rationale explaining the recoloring with respect to the specified emotion. We use this model to examine the relationship between color and emotion devoid of confounding factors.Finally, we turn to figurative language as a resource, examining the pragmatics of visualizing metaphoric phrases. We demonstrate a novel approach to generating visual metaphor through the collaboration of large language models and diffusion-based text-to-image models, and in doing so create a novel dataset of visual metaphor with both literal and figurative captions. We then develop an evaluation framework using human-AI collaboration to examine the efficacy of the model collaboration, and choose a downstream task of visual entailment to evaluate the human-AI collaboration.
일반주제명  
Computer engineering.
키워드  
Computational linguistics
키워드  
Computer vision
키워드  
Image generation
키워드  
Language grounding
키워드  
Text generation
기타저자  
Columbia University Computer Science
기본자료저록  
Dissertations Abstracts International. 85-02B.
기본자료저록  
Dissertation Abstract International
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008240612s2023      us  |||||||||||||||c||eng  d
■001000016933526
■00520240214101257
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798380091695
■035    ▼a(MiAaPQ)AAI30530597
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aWinn,  Olivia.
■24510▼a"Seeing  Red"  or  "Tickled  Pink"?:  Investigating  the  Power  of  Language  and  Vision  Models  Through  Color,  Emotion,  and  Metaphor▼h[electronic  resource]
■260    ▼a[S.l.]:▼bColumbia  University.  ▼c2023
■260  1▼aAnn  Arbor  :▼bProQuest  Dissertations  &  Theses,  ▼c2023
■300    ▼a1  online  resource(142  p.)
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-02,  Section:  B.
■500    ▼aAdvisor:  Muresan,  Smaranda.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2023.
■506    ▼aThis  item  must  not  be  sold  to  any  third  party  vendors.
■520    ▼aMultimodal  NLP  is  an  approach  to  language  understanding  that  incorporates  data  from  nontextual  in  order  to  enhance  our  linguistic  understanding  through  additional  contextual  information.  In  particular,  incorporating  visual  data  has  allowed  for  great  strides  in  our  ability  to  model  language  related  to  physical  phenomena.  The  performance  of  these  models  has  so  far  been  contingent  upon  access  to  large  datasets,  focusing  on  classification  problems  without  relative  information,  and  constraining  the  problem  space  to  literal  descriptions  and  interpretations.  In  this  thesis,  we  examine  these  limitations  by  investigating  how  types  of  data  previously  unused  in  these  models  can  be  reconfigured  and  worked  with  intelligently  and  on  a  small  scale  to  enhance  our  understanding  of  the  pragmatics  of  language.We  contribute  to  comparative  language  grounding,  emotional  interpretation,  and  metaphoric  understanding  by  releasing  multiple  annotated  datasets,  developing  a  new  paradigm  for  modeling  relative  data,  creating  a  new  task  in  examining  the  generation  of  emotional  descriptions  for  image  information,  and  demonstrating  a  novel  approach  to  working  with  figurative  text  for  image  generation.We  start  by  examining  how  traditional  grounding  models  could  be  adapted  to  incorporate  relative  information.  As  no  previous  work  has  ever  utilized  relative  textual  description  for  imageunderstanding,  we  first  constrain  the  problem  by  focusing  on  the  language  of  color.  We  create  a  new  dataset  of  comparative  color  terms  with  associated  RGB  datapoints,  and  use  this  data  to  develop  a  novel  paradigm  of  grounding  comparative  color  terms  in  RGB  space,  providing  the  first  avenue  towards  utilizing  relative  information  in  a  multimodal  setting.Continuing  our  study  of  color,  we  then  turn  to  examining  the  relationship  between  color  and  emotion.  In  order  to  further  our  understanding  of  this  relationship,  we  define  a  new  task,  called  Justified  Affect  Transformation,  in  which  an  image  is  recolored  specifically  to  alter  its  emotional  evocation  and  text  is  generated  to  explain  the  recoloring  from  an  emotional  perspective.  We  create  a  dataset  of  abstract  art  with  contiguous  emotion  labels  and  textual  rationales  for  the  emotional  evocation  of  multiple  images,  and  using  our  new  dataset  for  training,  introduce  a  new  unified  model  that  recolors  an  image  and  provides  a  textual  rationale  explaining  the  recoloring  with  respect  to  the  specified  emotion.  We  use  this  model  to  examine  the  relationship  between  color  and  emotion  devoid  of  confounding  factors.Finally,  we  turn  to  figurative  language  as  a  resource,  examining  the  pragmatics  of  visualizing  metaphoric  phrases.  We  demonstrate  a  novel  approach  to  generating  visual  metaphor  through  the  collaboration  of  large  language  models  and  diffusion-based  text-to-image  models,  and  in  doing  so  create  a  novel  dataset  of  visual  metaphor  with  both  literal  and  figurative  captions.  We  then  develop  an  evaluation  framework  using  human-AI  collaboration  to  examine  the  efficacy  of  the  model  collaboration,  and  choose  a  downstream  task  of  visual  entailment  to  evaluate  the  human-AI  collaboration.
■590    ▼aSchool  code:  0054.
■650  4▼aComputer  engineering.
■653    ▼aComputational  linguistics
■653    ▼aComputer  vision
■653    ▼aImage  generation
■653    ▼aLanguage  grounding
■653    ▼aText  generation
■690    ▼a0800
■690    ▼a0464
■71020▼aColumbia  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g85-02B.
■773    ▼tDissertation  Abstract  International
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2023
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T16933526▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.
■980    ▼a202402▼f2024

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF09266 전자도서 마이폴더 부재도서신고 비도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.