서브메뉴
검색
Grounding Language in Images and Videos
Grounding Language in Images and Videos
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151509
- ISBN
- 9798382652443
- DDC
- 004
- 저자명
- Sadhu, Arka.
- 서명/저자
- Grounding Language in Images and Videos
- 발행사항
- [Sl] : University of Southern California, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 244 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-11, Section: A.
- 주기사항
- Advisor: Nevatia, Ramakant.
- 학위논문주기
- Thesis (Ph.D.)--University of Southern California, 2024.
- 초록/해제
- 요약While machine learning research has traditionally explored image, video and text understanding as separate fields, the surge in multi-modal content in today's digital landscape underscores the importance of computation models that adeptly navigate complex interactions between text, images and videos. This dissertation addresses this challenge of grounding language in visual media - the task of associating linguistic symbols with perceptual experiences and actions. The overarching goal of this dissertation is to bridge the gap between language and vision as a means to a "deeper understanding" of images and videos to allow developing models capable of reasoning over longer-time horizons such as hour-long movies, or a collection of images, or even multiple videos.A pivotal contribution of my work is the use of Semantic Roles for images, videos and text. Unlike previous works that primarily focused on recognizing single entities or generating holistic captions, the use of Semantic Roles facilitates a fine-grained understanding of "who did what to whom" in a structured format. It maintains the advantages of having free-form language phrases and at the same time also being comprehensive and complete like entity recognition, thus enriching the model's interpretive capabilities.In this thesis, we will introduce the various vision-language tasks developed during my Ph.D. This includes grounding unseen words, spatio-temporal localization of entities in a video, video question answering, visual semantic role labeling in videos, reasoning across more than one image or a video, and finally, weakly-supervised open-vocabulary object detection. Each task is accompanied by the creation and development of dedicated datasets, evaluation protocols, and model frameworks. These tasks aim to investigate a particular phenomenon inherent in image or video understanding in isolation, develop corresponding datasets and model frameworks, and outline evaluation protocols robust to data priors.The resulting models can be used for other downstream tasks like obtaining common-sense knowledge graphs from instructional videos or drive end-user applications like Retrieval, Question Answering, and Captioning. By facilitating the deeper integration of language and vision, this dissertation represents a step-forward in machine learning models capable of finer-understanding of the world around us.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Linguistics
- 키워드
- Computer vision
- 키워드
- Machine learning
- 기타저자
- University of Southern California Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-11A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017161970
■00520250211151509
■006m o d
■007cr#unu||||||||
■020 ▼a9798382652443
■035 ▼a(MiAaPQ)AAI31299441
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aSadhu, Arka.
■24510▼aGrounding Language in Images and Videos
■260 ▼a[Sl]▼bUniversity of Southern California▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a244 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-11, Section: A.
■500 ▼aAdvisor: Nevatia, Ramakant.
■5021 ▼aThesis (Ph.D.)--University of Southern California, 2024.
■520 ▼aWhile machine learning research has traditionally explored image, video and text understanding as separate fields, the surge in multi-modal content in today's digital landscape underscores the importance of computation models that adeptly navigate complex interactions between text, images and videos. This dissertation addresses this challenge of grounding language in visual media - the task of associating linguistic symbols with perceptual experiences and actions. The overarching goal of this dissertation is to bridge the gap between language and vision as a means to a "deeper understanding" of images and videos to allow developing models capable of reasoning over longer-time horizons such as hour-long movies, or a collection of images, or even multiple videos.A pivotal contribution of my work is the use of Semantic Roles for images, videos and text. Unlike previous works that primarily focused on recognizing single entities or generating holistic captions, the use of Semantic Roles facilitates a fine-grained understanding of "who did what to whom" in a structured format. It maintains the advantages of having free-form language phrases and at the same time also being comprehensive and complete like entity recognition, thus enriching the model's interpretive capabilities.In this thesis, we will introduce the various vision-language tasks developed during my Ph.D. This includes grounding unseen words, spatio-temporal localization of entities in a video, video question answering, visual semantic role labeling in videos, reasoning across more than one image or a video, and finally, weakly-supervised open-vocabulary object detection. Each task is accompanied by the creation and development of dedicated datasets, evaluation protocols, and model frameworks. These tasks aim to investigate a particular phenomenon inherent in image or video understanding in isolation, develop corresponding datasets and model frameworks, and outline evaluation protocols robust to data priors.The resulting models can be used for other downstream tasks like obtaining common-sense knowledge graphs from instructional videos or drive end-user applications like Retrieval, Question Answering, and Captioning. By facilitating the deeper integration of language and vision, this dissertation represents a step-forward in machine learning models capable of finer-understanding of the world around us.
■590 ▼aSchool code: 0208.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aLinguistics
■653 ▼aComputer vision
■653 ▼aImage understanding
■653 ▼aMachine learning
■653 ▼aNatural language processing
■653 ▼aVideo understanding
■690 ▼a0984
■690 ▼a0464
■690 ▼a0800
■690 ▼a0290
■71020▼aUniversity of Southern California▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-11A.
■790 ▼a0208
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161970▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


