서브메뉴
검색
Learning Video Representations With Limited Supervision
Learning Video Representations With Limited Supervision
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260209102859
- ISBN
- 9798291578162
- DDC
- 004
- 저자명
- McKee, Daniel.
- 서명/저자
- Learning Video Representations With Limited Supervision
- 발행사항
- [Sl] : University of Illinois at Urbana-Champaign, 2023
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2023
- 형태사항
- 95 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Lazebnik, Svetlana.
- 학위논문주기
- Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
- 초록/해제
- 요약With the rapid growth of deep computer vision models, demand for large quantities of annotated data has risen higher than ever. Obtaining visual annotations, especially dense annotations requiring fine-grained localization of objects, is a costly and intensive process. Dense video tasks like tracking or video object segmentation provide an even greater annotation challenge due to the steep cost increase associated with labeling many individual frames. As a result, datasets for these tasks often lack the scale and diversity of samples in annotated image datasets. To combat such limitations, we investigate how we can take advantage of unlabeled videos, image annotations, and transfer of large-scale pretrained models to achieve effective performance on dense video tasks. First, we study representations for dense label propagation tasks in video, focusing on self-supervised approaches to learning temporal correspondence and comparing how image-trained models might be adapted for these tasks. Second, we investigate how to train a multi-object tracking model in the absence of tracking annotations. In place of fully supervised annotations, we demonstrate how to learn from unlabeled videos and videos that are hallucinated from annotated images using data augmentation techniques. Lastly, we explore a multi-modal problem setting where we wish to automatically recommend an audio soundtrack for an input video and text description of desired music. In this setting, we explore adapting large scale models like CLIP for joint modeling of video, text, and audio. We also investigate mechanisms for generating text pseudo-label descriptions for training using recent large language models.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 키워드
- Computer vision
- 키워드
- Deep learning
- 키워드
- Language models
- 키워드
- Object tracking
- 기타저자
- University of Illinois at Urbana-Champaign Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260203s2023 us c eng d■001000017365942
■00520260209102859
■006m o d
■007cr#unu||||||||
■020 ▼a9798291578162
■035 ▼a(MiAaPQ)AAI32272183
■035 ▼a(MiAaPQ)httphdlhandlenet2142121988
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aMcKee, Daniel.
■24510▼aLearning Video Representations With Limited Supervision
■260 ▼a[Sl]▼bUniversity of Illinois at Urbana-Champaign▼c2023
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2023
■300 ▼a95 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Lazebnik, Svetlana.
■5021 ▼aThesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
■520 ▼aWith the rapid growth of deep computer vision models, demand for large quantities of annotated data has risen higher than ever. Obtaining visual annotations, especially dense annotations requiring fine-grained localization of objects, is a costly and intensive process. Dense video tasks like tracking or video object segmentation provide an even greater annotation challenge due to the steep cost increase associated with labeling many individual frames. As a result, datasets for these tasks often lack the scale and diversity of samples in annotated image datasets. To combat such limitations, we investigate how we can take advantage of unlabeled videos, image annotations, and transfer of large-scale pretrained models to achieve effective performance on dense video tasks. First, we study representations for dense label propagation tasks in video, focusing on self-supervised approaches to learning temporal correspondence and comparing how image-trained models might be adapted for these tasks. Second, we investigate how to train a multi-object tracking model in the absence of tracking annotations. In place of fully supervised annotations, we demonstrate how to learn from unlabeled videos and videos that are hallucinated from annotated images using data augmentation techniques. Lastly, we explore a multi-modal problem setting where we wish to automatically recommend an audio soundtrack for an input video and text description of desired music. In this setting, we explore adapting large scale models like CLIP for joint modeling of video, text, and audio. We also investigate mechanisms for generating text pseudo-label descriptions for training using recent large language models.
■590 ▼aSchool code: 0090.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aComputer vision
■653 ▼aDeep learning
■653 ▼aSelf-supervised learning
■653 ▼aWeakly supervised learning
■653 ▼aLanguage models
■653 ▼aMulti-modal models
■653 ▼aObject tracking
■690 ▼a0984
■690 ▼a0464
■71020▼aUniversity of Illinois at Urbana-Champaign▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0090
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365942▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


