본문

서브메뉴

Event-Centric Multimodal Knowledge Acquisition
Event-Centric Multimodal Knowledge Acquisition
Event-Centric Multimodal Knowledge Acquisition

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260209102848
ISBN  
9798291563670
DDC  
004
저자명  
Li, Manling.
서명/저자  
Event-Centric Multimodal Knowledge Acquisition
발행사항  
[Sl] : University of Illinois at Urbana-Champaign, 2023
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2023
형태사항  
161 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
주기사항  
Advisor: Ji, Heng.
학위논문주기  
Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
초록/해제  
요약What happened? Who? When? Where? Why? What will happen next? are the fundamental questions asked to comprehend the overwhelming amount of information. Answers to these questions are the core knowledge communicated through multiple forms of information, regardless of whether presented as text, images, videos, audio, or other modalities.To obtain such knowledge from multimodal data, this dissertation focuses on Multimodal Information Extraction (IE), and propose Event-Centric Multimodal Knowledge Acquisition to evolve traditional Entity-centric Single-modality knowledge into Event-centric Multi-modality knowledge. Traditional entity-centric approaches to consuming multimodal information focus on concrete concepts (such as objects, object types, physical relations, e.g., a person in a car), while this dissertation endows machines to understand complex abstract semantic structures that are difficult to ground into image regions but are essential knowledge (such as events and semantic roles of objects, e.g., driver, passenger, passerby, salesperson). It is able to consolidate complex semantic structures of multiple modalities, providing a major benefit over recent research advances in single-modality (text-only or vision-only) knowledge.Such a transformation poses significant challenges in terms of understanding multimodal semantic structures (such as semantic roles) and temporal dynamics (such as future participants and their roles):Understanding Multimodal Semantic Structures to answer What happened?, Who?, Where?, and When? (Knowledge Extraction): Due to the structural nature and lack of anchoring in a specific image region, abstract semantic structures are difficult to synthesize between text and vision modalities through general large-scale pretraining. We introduce complex event semantic structures into vision-language pretraining (CLIP-Event), and propose a zero-shot cross-modal transfer of semantic understanding abilities from language to vision, which resolves the poor portability issue of IE and supports Zero-shot Multimodal Event Extraction (M2E2) for the first time. We also release an open-source Multimodal IE system GAIA to serve as an off-the-shelf tool for the research community.Understanding Temporal Dynamics to answer What will happen next?, Who will participant? and Why? (Knowledge Reasoning): The significance of capturing temporal dynamics has led to recent advances in script knowledge learning, however, which has been overly simplified to be local and sequential. We propose Event Graph Schema, which open doors to a global event graph context to enable alternative predictions, along with structural justifications including location-, attribute-, and participant-specific details.Generating truthfully with Event-Centric Knowledge Facts (Knowledge Driven Applications): Our work has shown positive results on long-standing open problems, such as Timeline Summarization, Meeting Summarization, and Multimedia News Question Answering, Report Generation, etc.This work on Multimedia Event Knowledge Graphs aims to open doors to the next generation of information access, in order to equip machines with factual knowledge discovery and reasoning from diverse sources of information, so that we can lay a foundation for promoting factuality and truthfulness in information access, through a structured knowledge view that is easily explainable, highly compositional, and capable of long-horizon reasoning.
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Systems science
키워드  
Multimodal semantic structures
키워드  
Timeline Summarization
키워드  
Event-Centric Knowledge
키워드  
Event Graph Schema
키워드  
Multimodal information
기타저자  
University of Illinois at Urbana-Champaign Computer Science
기본자료저록  
Dissertations Abstracts International. 87-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260203s2023        us                              c    eng  d
■001000017365885
■00520260209102848
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291563670
■035    ▼a(MiAaPQ)AAI32271344
■035    ▼a(MiAaPQ)httphdlhandlenet2142121435
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aLi,  Manling.
■24510▼aEvent-Centric  Multimodal  Knowledge  Acquisition
■260    ▼a[Sl]▼bUniversity  of  Illinois  at  Urbana-Champaign▼c2023
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2023
■300    ▼a161  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-02,  Section:  B.
■500    ▼aAdvisor:  Ji,  Heng.
■5021  ▼aThesis  (Ph.D.)--University  of  Illinois  at  Urbana-Champaign,  2023.
■520    ▼aWhat  happened?  Who?  When?  Where?  Why?  What  will  happen  next?  are  the  fundamental  questions  asked  to  comprehend  the  overwhelming  amount  of  information.  Answers  to  these  questions  are  the  core  knowledge  communicated  through  multiple  forms  of  information,  regardless  of  whether  presented  as  text,  images,  videos,  audio,  or  other  modalities.To  obtain  such  knowledge  from  multimodal  data,  this  dissertation  focuses  on  Multimodal  Information  Extraction  (IE),  and  propose  Event-Centric  Multimodal  Knowledge  Acquisition  to  evolve  traditional  Entity-centric  Single-modality  knowledge  into  Event-centric  Multi-modality  knowledge.  Traditional  entity-centric  approaches  to  consuming  multimodal  information  focus  on  concrete  concepts  (such  as  objects,  object  types,  physical  relations,  e.g.,  a  person  in  a  car),  while  this  dissertation  endows  machines  to  understand  complex  abstract  semantic  structures  that  are  difficult  to  ground  into  image  regions  but  are  essential  knowledge  (such  as  events  and  semantic  roles  of  objects,  e.g.,  driver,  passenger,  passerby,  salesperson).  It  is  able  to  consolidate  complex  semantic  structures  of  multiple  modalities,  providing  a  major  benefit  over  recent  research  advances  in  single-modality  (text-only  or  vision-only)  knowledge.Such  a  transformation  poses  significant  challenges  in  terms  of  understanding  multimodal  semantic  structures  (such  as  semantic  roles)  and  temporal  dynamics  (such  as  future  participants  and  their  roles):Understanding  Multimodal  Semantic  Structures  to  answer  What  happened?,  Who?,  Where?,  and  When?  (Knowledge  Extraction):  Due  to  the  structural  nature  and  lack  of  anchoring  in  a  specific  image  region,  abstract  semantic  structures  are  difficult  to  synthesize  between  text  and  vision  modalities  through  general  large-scale  pretraining.  We  introduce  complex  event  semantic  structures  into  vision-language  pretraining  (CLIP-Event),  and  propose  a  zero-shot  cross-modal  transfer  of  semantic  understanding  abilities  from  language  to  vision,  which  resolves  the  poor  portability  issue  of  IE  and  supports  Zero-shot  Multimodal  Event  Extraction  (M2E2)  for  the  first  time.  We  also  release  an  open-source  Multimodal  IE  system  GAIA  to  serve  as  an  off-the-shelf  tool  for  the  research  community.Understanding  Temporal  Dynamics  to  answer  What  will  happen  next?,  Who  will  participant?  and  Why?  (Knowledge  Reasoning):  The  significance  of  capturing  temporal  dynamics  has  led  to  recent  advances  in  script  knowledge  learning,  however,  which  has  been  overly  simplified  to  be  local  and  sequential.  We  propose  Event  Graph  Schema,  which  open  doors  to  a  global  event  graph  context  to  enable  alternative  predictions,  along  with  structural  justifications  including  location-,  attribute-,  and  participant-specific  details.Generating  truthfully  with  Event-Centric  Knowledge  Facts  (Knowledge  Driven  Applications):  Our  work  has  shown  positive  results  on  long-standing  open  problems,  such  as  Timeline  Summarization,  Meeting  Summarization,  and  Multimedia  News  Question  Answering,  Report  Generation,  etc.This  work  on  Multimedia  Event  Knowledge  Graphs  aims  to  open  doors  to  the  next  generation  of  information  access,  in  order  to  equip  machines  with  factual  knowledge  discovery  and  reasoning  from  diverse  sources  of  information,  so  that  we  can  lay  a  foundation  for  promoting  factuality  and  truthfulness  in  information  access,  through  a  structured  knowledge  view  that  is  easily  explainable,  highly  compositional,  and  capable  of  long-horizon  reasoning.
■590    ▼aSchool  code:  0090.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aSystems  science
■653    ▼aMultimodal  semantic  structures
■653    ▼aTimeline  Summarization
■653    ▼aEvent-Centric  Knowledge
■653    ▼aEvent  Graph  Schema
■653    ▼aMultimodal  information
■690    ▼a0984
■690    ▼a0464
■690    ▼a0790
■71020▼aUniversity  of  Illinois  at  Urbana-Champaign▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g87-02B.
■790    ▼a0090
■791    ▼aPh.D.
■792    ▼a2023
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365885▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18991 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.