서브메뉴
검색
Event-Centric Multimodal Knowledge Acquisition
Event-Centric Multimodal Knowledge Acquisition
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260209102848
- ISBN
- 9798291563670
- DDC
- 004
- 저자명
- Li, Manling.
- 서명/저자
- Event-Centric Multimodal Knowledge Acquisition
- 발행사항
- [Sl] : University of Illinois at Urbana-Champaign, 2023
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2023
- 형태사항
- 161 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Ji, Heng.
- 학위논문주기
- Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
- 초록/해제
- 요약What happened? Who? When? Where? Why? What will happen next? are the fundamental questions asked to comprehend the overwhelming amount of information. Answers to these questions are the core knowledge communicated through multiple forms of information, regardless of whether presented as text, images, videos, audio, or other modalities.To obtain such knowledge from multimodal data, this dissertation focuses on Multimodal Information Extraction (IE), and propose Event-Centric Multimodal Knowledge Acquisition to evolve traditional Entity-centric Single-modality knowledge into Event-centric Multi-modality knowledge. Traditional entity-centric approaches to consuming multimodal information focus on concrete concepts (such as objects, object types, physical relations, e.g., a person in a car), while this dissertation endows machines to understand complex abstract semantic structures that are difficult to ground into image regions but are essential knowledge (such as events and semantic roles of objects, e.g., driver, passenger, passerby, salesperson). It is able to consolidate complex semantic structures of multiple modalities, providing a major benefit over recent research advances in single-modality (text-only or vision-only) knowledge.Such a transformation poses significant challenges in terms of understanding multimodal semantic structures (such as semantic roles) and temporal dynamics (such as future participants and their roles):Understanding Multimodal Semantic Structures to answer What happened?, Who?, Where?, and When? (Knowledge Extraction): Due to the structural nature and lack of anchoring in a specific image region, abstract semantic structures are difficult to synthesize between text and vision modalities through general large-scale pretraining. We introduce complex event semantic structures into vision-language pretraining (CLIP-Event), and propose a zero-shot cross-modal transfer of semantic understanding abilities from language to vision, which resolves the poor portability issue of IE and supports Zero-shot Multimodal Event Extraction (M2E2) for the first time. We also release an open-source Multimodal IE system GAIA to serve as an off-the-shelf tool for the research community.Understanding Temporal Dynamics to answer What will happen next?, Who will participant? and Why? (Knowledge Reasoning): The significance of capturing temporal dynamics has led to recent advances in script knowledge learning, however, which has been overly simplified to be local and sequential. We propose Event Graph Schema, which open doors to a global event graph context to enable alternative predictions, along with structural justifications including location-, attribute-, and participant-specific details.Generating truthfully with Event-Centric Knowledge Facts (Knowledge Driven Applications): Our work has shown positive results on long-standing open problems, such as Timeline Summarization, Meeting Summarization, and Multimedia News Question Answering, Report Generation, etc.This work on Multimedia Event Knowledge Graphs aims to open doors to the next generation of information access, in order to equip machines with factual knowledge discovery and reasoning from diverse sources of information, so that we can lay a foundation for promoting factuality and truthfulness in information access, through a structured knowledge view that is easily explainable, highly compositional, and capable of long-horizon reasoning.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Systems science
- 기타저자
- University of Illinois at Urbana-Champaign Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260203s2023 us c eng d■001000017365885
■00520260209102848
■006m o d
■007cr#unu||||||||
■020 ▼a9798291563670
■035 ▼a(MiAaPQ)AAI32271344
■035 ▼a(MiAaPQ)httphdlhandlenet2142121435
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aLi, Manling.
■24510▼aEvent-Centric Multimodal Knowledge Acquisition
■260 ▼a[Sl]▼bUniversity of Illinois at Urbana-Champaign▼c2023
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2023
■300 ▼a161 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Ji, Heng.
■5021 ▼aThesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
■520 ▼aWhat happened? Who? When? Where? Why? What will happen next? are the fundamental questions asked to comprehend the overwhelming amount of information. Answers to these questions are the core knowledge communicated through multiple forms of information, regardless of whether presented as text, images, videos, audio, or other modalities.To obtain such knowledge from multimodal data, this dissertation focuses on Multimodal Information Extraction (IE), and propose Event-Centric Multimodal Knowledge Acquisition to evolve traditional Entity-centric Single-modality knowledge into Event-centric Multi-modality knowledge. Traditional entity-centric approaches to consuming multimodal information focus on concrete concepts (such as objects, object types, physical relations, e.g., a person in a car), while this dissertation endows machines to understand complex abstract semantic structures that are difficult to ground into image regions but are essential knowledge (such as events and semantic roles of objects, e.g., driver, passenger, passerby, salesperson). It is able to consolidate complex semantic structures of multiple modalities, providing a major benefit over recent research advances in single-modality (text-only or vision-only) knowledge.Such a transformation poses significant challenges in terms of understanding multimodal semantic structures (such as semantic roles) and temporal dynamics (such as future participants and their roles):Understanding Multimodal Semantic Structures to answer What happened?, Who?, Where?, and When? (Knowledge Extraction): Due to the structural nature and lack of anchoring in a specific image region, abstract semantic structures are difficult to synthesize between text and vision modalities through general large-scale pretraining. We introduce complex event semantic structures into vision-language pretraining (CLIP-Event), and propose a zero-shot cross-modal transfer of semantic understanding abilities from language to vision, which resolves the poor portability issue of IE and supports Zero-shot Multimodal Event Extraction (M2E2) for the first time. We also release an open-source Multimodal IE system GAIA to serve as an off-the-shelf tool for the research community.Understanding Temporal Dynamics to answer What will happen next?, Who will participant? and Why? (Knowledge Reasoning): The significance of capturing temporal dynamics has led to recent advances in script knowledge learning, however, which has been overly simplified to be local and sequential. We propose Event Graph Schema, which open doors to a global event graph context to enable alternative predictions, along with structural justifications including location-, attribute-, and participant-specific details.Generating truthfully with Event-Centric Knowledge Facts (Knowledge Driven Applications): Our work has shown positive results on long-standing open problems, such as Timeline Summarization, Meeting Summarization, and Multimedia News Question Answering, Report Generation, etc.This work on Multimedia Event Knowledge Graphs aims to open doors to the next generation of information access, in order to equip machines with factual knowledge discovery and reasoning from diverse sources of information, so that we can lay a foundation for promoting factuality and truthfulness in information access, through a structured knowledge view that is easily explainable, highly compositional, and capable of long-horizon reasoning.
■590 ▼aSchool code: 0090.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aSystems science
■653 ▼aMultimodal semantic structures
■653 ▼aTimeline Summarization
■653 ▼aEvent-Centric Knowledge
■653 ▼aEvent Graph Schema
■653 ▼aMultimodal information
■690 ▼a0984
■690 ▼a0464
■690 ▼a0790
■71020▼aUniversity of Illinois at Urbana-Champaign▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0090
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365885▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


