서브메뉴
검색
Learning Pose and State-Invariant Object Representations for Fine-Grained Recognition and Retrieval
Learning Pose and State-Invariant Object Representations for Fine-Grained Recognition and Retrieval
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152751
- ISBN
- 9798342122955
- DDC
- 375
- 저자명
- Sarkar, Rohan.
- 서명/저자
- Learning Pose and State-Invariant Object Representations for Fine-Grained Recognition and Retrieval
- 발행사항
- [Sl] : Purdue University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 165 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
- 주기사항
- Advisor: Kak, Avinash.
- 학위논문주기
- Thesis (Ph.D.)--Purdue University, 2024.
- 초록/해제
- 요약Object Recognition and Retrieval is a fundamental problem in Computer Vision that involves recognizing objects and retrieving similar object images through visual queries. While deep metric learning is commonly employed to learn image embeddings for solving such problems, the representations learned using existing methods are not robust to changes in viewpoint, pose, and object state, especially for fine-grained recognition and retrieval tasks. To overcome these limitations, this dissertation aims to learn robust object representations that remain invariant to such transformations for fine-grained tasks. First, it focuses on learning dual pose-invariant embeddings to facilitate recognition and retrieval at both the category and finer object-identity levels by learning category and object-identity specific representations in separate embedding spaces simultaneously. For this, the PiRO framework is introduced that utilizes an attention-based dual encoder architecture and novel pose-invariant ranking losses for each embedding space to disentangle the category and object representations while learning pose-invariant features. Second, the dissertation introduces ranking losses that cluster multi-view images of an object together in both the embedding spaces while simultaneously pulling the embeddings of two objects from the same category closer in the category embedding space to learn fundamental category-specific attributes and pushing them apart in the object embedding space to learn discriminative features to distinguish between them. Third, the dissertation addresses state-invariance and introduces a novel Objects: With State: Change dataset to facilitate research in recognizing fine-grained objects with state changes involving structural transformations in addition to pose and viewpoint changes. Fourth, it proposes a curriculum learning strategyto progressively sample object images that are harder to distinguish for training the model, enhancing its ability to capture discriminative features for fine-grained tasks amidst state changes and other transformations. Experimental evaluations demonstrate significant improvements in object recognition and retrieval performance compared to previous methods, validating the effectiveness of the proposed approaches across several challenging datasets under various transformations.
- 일반주제명
- Curricula
- 일반주제명
- Computer science
- 기타저자
- Purdue University.
- 기본자료저록
- Dissertations Abstracts International. 86-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163774
■00520250211152751
■006m o d
■007cr#unu||||||||
■020 ▼a9798342122955
■035 ▼a(MiAaPQ)AAI31532247
■035 ▼a(MiAaPQ)Purdue26236742
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a375
■1001 ▼aSarkar, Rohan.
■24510▼aLearning Pose and State-Invariant Object Representations for Fine-Grained Recognition and Retrieval
■260 ▼a[Sl]▼bPurdue University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a165 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-05, Section: B.
■500 ▼aAdvisor: Kak, Avinash.
■5021 ▼aThesis (Ph.D.)--Purdue University, 2024.
■520 ▼aObject Recognition and Retrieval is a fundamental problem in Computer Vision that involves recognizing objects and retrieving similar object images through visual queries. While deep metric learning is commonly employed to learn image embeddings for solving such problems, the representations learned using existing methods are not robust to changes in viewpoint, pose, and object state, especially for fine-grained recognition and retrieval tasks. To overcome these limitations, this dissertation aims to learn robust object representations that remain invariant to such transformations for fine-grained tasks. First, it focuses on learning dual pose-invariant embeddings to facilitate recognition and retrieval at both the category and finer object-identity levels by learning category and object-identity specific representations in separate embedding spaces simultaneously. For this, the PiRO framework is introduced that utilizes an attention-based dual encoder architecture and novel pose-invariant ranking losses for each embedding space to disentangle the category and object representations while learning pose-invariant features. Second, the dissertation introduces ranking losses that cluster multi-view images of an object together in both the embedding spaces while simultaneously pulling the embeddings of two objects from the same category closer in the category embedding space to learn fundamental category-specific attributes and pushing them apart in the object embedding space to learn discriminative features to distinguish between them. Third, the dissertation addresses state-invariance and introduces a novel Objects: With State: Change dataset to facilitate research in recognizing fine-grained objects with state changes involving structural transformations in addition to pose and viewpoint changes. Fourth, it proposes a curriculum learning strategyto progressively sample object images that are harder to distinguish for training the model, enhancing its ability to capture discriminative features for fine-grained tasks amidst state changes and other transformations. Experimental evaluations demonstrate significant improvements in object recognition and retrieval performance compared to previous methods, validating the effectiveness of the proposed approaches across several challenging datasets under various transformations.
■590 ▼aSchool code: 0183.
■650 4▼aCurricula
■650 4▼aObject linking & embedding
■650 4▼aComputer science
■690 ▼a0984
■71020▼aPurdue University.
■7730 ▼tDissertations Abstracts International▼g86-05B.
■790 ▼a0183
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163774▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


