서브메뉴
검색
Learning Visual Concepts
Learning Visual Concepts
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152037
- ISBN
- 9798384023104
- DDC
- 004
- 서명/저자
- Learning Visual Concepts
- 발행사항
- [Sl] : University of Pennsylvania, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 198 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-02, Section: A.
- 주기사항
- Advisor: Taylor, Camillo J.
- 학위논문주기
- Thesis (Ph.D.)--University of Pennsylvania, 2024.
- 초록/해제
- 요약We propose a framework to use off-the-shelf pre-trained object detection models and extend them for use on unseen datasets in a manner requiring little to no modification of the original architecture, and by adding only a few additional components to the overall pipeline. Motivated by the role of attributes in zero-shot-learning paradigms, we define conceptual groups by using positive and negative exemplars retroactively, and evaluate the feasibility of recognizing a variety of these proposed conceptual groups in a corpus of previously unseen data, including unseen categories. We conduct experiments with networks trained on the COCO dataset, and utilize Open-Images-V7 as our held out unseen dataset. Our analysis suggests that existing off-the-shelf object detection networks such as Faster-RCNN can be leveraged to extract useful information beyond the scope of a straightforward category prediction framework. This information can be used to operationalize the idea of concept learning through a set of positive and negative exemplars and a simple linear SVM operating on the features produced by the deep network. We compare this approach to vision enabled large language models such as LLaVA, CogVLM and GPT4V, and show a strong baseline performance with lower resource requirements. Additionally, we illustrate that this method can be scaled to larger concept sets by validating this approach on a larger set of concepts in the LVIS dataset. We illustrate a few approaches to better understand the semantic topology of their learned feature space, and we measure the feasibility of using these features for the identification of the proposed conceptual groups. We propose strategies to leverage this information to predict these conceptual groups on previously unseen samples containing unseen class categories.
- 일반주제명
- Computer science
- 일반주제명
- Robotics
- 일반주제명
- Information science
- 키워드
- Classification
- 키워드
- Computer vision
- 키워드
- Object detection
- 기타저자
- University of Pennsylvania Computer and Information Science
- 기본자료저록
- Dissertations Abstracts International. 86-02A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162643
■00520250211152037
■006m o d
■007cr#unu||||||||
■020 ▼a9798384023104
■035 ▼a(MiAaPQ)AAI31335541
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aShivakumar, Shreyas Skandan.
■24510▼aLearning Visual Concepts
■260 ▼a[Sl]▼bUniversity of Pennsylvania▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a198 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-02, Section: A.
■500 ▼aAdvisor: Taylor, Camillo J.
■5021 ▼aThesis (Ph.D.)--University of Pennsylvania, 2024.
■520 ▼aWe propose a framework to use off-the-shelf pre-trained object detection models and extend them for use on unseen datasets in a manner requiring little to no modification of the original architecture, and by adding only a few additional components to the overall pipeline. Motivated by the role of attributes in zero-shot-learning paradigms, we define conceptual groups by using positive and negative exemplars retroactively, and evaluate the feasibility of recognizing a variety of these proposed conceptual groups in a corpus of previously unseen data, including unseen categories. We conduct experiments with networks trained on the COCO dataset, and utilize Open-Images-V7 as our held out unseen dataset. Our analysis suggests that existing off-the-shelf object detection networks such as Faster-RCNN can be leveraged to extract useful information beyond the scope of a straightforward category prediction framework. This information can be used to operationalize the idea of concept learning through a set of positive and negative exemplars and a simple linear SVM operating on the features produced by the deep network. We compare this approach to vision enabled large language models such as LLaVA, CogVLM and GPT4V, and show a strong baseline performance with lower resource requirements. Additionally, we illustrate that this method can be scaled to larger concept sets by validating this approach on a larger set of concepts in the LVIS dataset. We illustrate a few approaches to better understand the semantic topology of their learned feature space, and we measure the feasibility of using these features for the identification of the proposed conceptual groups. We propose strategies to leverage this information to predict these conceptual groups on previously unseen samples containing unseen class categories.
■590 ▼aSchool code: 0175.
■650 4▼aComputer science
■650 4▼aRobotics
■650 4▼aInformation science
■653 ▼aClassification
■653 ▼aComputer vision
■653 ▼aMachine perception
■653 ▼aObject detection
■653 ▼aObject recognition
■690 ▼a0800
■690 ▼a0984
■690 ▼a0771
■690 ▼a0723
■71020▼aUniversity of Pennsylvania▼bComputer and Information Science.
■7730 ▼tDissertations Abstracts International▼g86-02A.
■790 ▼a0175
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162643▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


