서브메뉴
검색
Language Supervision for Computer Vision
Language Supervision for Computer Vision
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152057
- ISBN
- 9798382739380
- DDC
- 004
- 저자명
- Desai, Karan P.
- 서명/저자
- Language Supervision for Computer Vision
- 발행사항
- [Sl] : University of Michigan, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 179 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
- 주기사항
- Advisor: Johnson, Justin C.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2024.
- 초록/해제
- 요약Representation learning lies at the core of modern Artificial Intelligence. In computer vision, labeled image datasets like ImageNet have been the standard choice for representation learning. Despite being empirically successful, this approach is expensive to scale due to labeling costs. Moreover, the representation quality is limited by the size and diversity of datasets and their associated label ontologies.My research explores using natural language supervision for computer vision. Using natural language allows us to go beyond fixed label ontologies and scale up to more general sources such as internet data. Toward this goal, my dissertation explores four problems - (1) Learning representations: I propose one of the first methods for language-supervised visual learning that uses image captioning as the training objective, showing its efficacy compared to ImageNet-trained methods on downstream tasks like object detection and segmentation. (2) Scaling data: I explore social media as a rich source of high-quality image descriptions and curate a dataset of 12 million image-text pairs while ensuring responsible curation practices. (3) Understanding data: It is difficult to comprehend the diversity of visual concepts present in millions of image-text pairs. I posit that images and text naturally organize into a tree-like hierarchy and propose an approach for learning representations that capture this hierarchy using tools from hyperbolic geometry. (4) Transfer to downstream tasks: Large vision-language models show impressive zero-shot transfer capabilities on image-level tasks like classification and retrieval. However, their transferability to pixel-level tasks like object detection and segmentation has relied on expensive labeled mask annotations. I propose an object detector to efficiently transfer pre-trained vision models to segment and classify visual objects without any fine-tuning, unlike existing detectors that train using orders of magnitude more labeled masks to achieve high performance.In summary, my research affirms that using language supervision can drive the next leap of progress in computer vision and has immense utility in practical applications.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 키워드
- Computer vision
- 키워드
- Natural language
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 85-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162812
■00520250211152057
■006m o d
■007cr#unu||||||||
■020 ▼a9798382739380
■035 ▼a(MiAaPQ)AAI31348980
■035 ▼a(MiAaPQ)umichrackham005447
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aDesai, Karan P.
■24510▼aLanguage Supervision for Computer Vision
■260 ▼a[Sl]▼bUniversity of Michigan▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a179 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-12, Section: B.
■500 ▼aAdvisor: Johnson, Justin C.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2024.
■520 ▼aRepresentation learning lies at the core of modern Artificial Intelligence. In computer vision, labeled image datasets like ImageNet have been the standard choice for representation learning. Despite being empirically successful, this approach is expensive to scale due to labeling costs. Moreover, the representation quality is limited by the size and diversity of datasets and their associated label ontologies.My research explores using natural language supervision for computer vision. Using natural language allows us to go beyond fixed label ontologies and scale up to more general sources such as internet data. Toward this goal, my dissertation explores four problems - (1) Learning representations: I propose one of the first methods for language-supervised visual learning that uses image captioning as the training objective, showing its efficacy compared to ImageNet-trained methods on downstream tasks like object detection and segmentation. (2) Scaling data: I explore social media as a rich source of high-quality image descriptions and curate a dataset of 12 million image-text pairs while ensuring responsible curation practices. (3) Understanding data: It is difficult to comprehend the diversity of visual concepts present in millions of image-text pairs. I posit that images and text naturally organize into a tree-like hierarchy and propose an approach for learning representations that capture this hierarchy using tools from hyperbolic geometry. (4) Transfer to downstream tasks: Large vision-language models show impressive zero-shot transfer capabilities on image-level tasks like classification and retrieval. However, their transferability to pixel-level tasks like object detection and segmentation has relied on expensive labeled mask annotations. I propose an object detector to efficiently transfer pre-trained vision models to segment and classify visual objects without any fine-tuning, unlike existing detectors that train using orders of magnitude more labeled masks to achieve high performance.In summary, my research affirms that using language supervision can drive the next leap of progress in computer vision and has immense utility in practical applications.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aComputer vision
■653 ▼aRepresentation learning
■653 ▼aHyperbolic geometry
■653 ▼aNatural language
■690 ▼a0984
■690 ▼a0800
■690 ▼a0464
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g85-12B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162812▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


