서브메뉴
검색
Domain Adapted Visual Representation Learning for Machine Perception- [electronic resource]
Domain Adapted Visual Representation Learning for Machine Perception- [electronic resource]
Detailed Information
- 자료유형
- 학위논문파일 국외
- 최종처리일시
- 20240214101645
- ISBN
- 9798380394345
- DDC
- 621.3
- 저자명
- Li, Yu-Jhe.
- 서명/저자
- Domain Adapted Visual Representation Learning for Machine Perception - [electronic resource]
- 발행사항
- [S.l.]: : Carnegie Mellon University., 2023
- 발행사항
- Ann Arbor : : ProQuest Dissertations & Theses,, 2023
- 형태사항
- 1 online resource(194 p.)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-03, Section: B.
- 주기사항
- Advisor: Kitani, Kris.
- 학위논문주기
- Thesis (Ph.D.)--Carnegie Mellon University, 2023.
- 사용제한주기
- This item must not be sold to any third party vendors.
- 초록/해제
- 요약Our objective is to enhance the generalization capabilities of existing machine perception models and achieve diverse domain alignments through adept representation learning. Many established approaches for perception tasks, encompassing object classification, detection, tracking, and rendering, often confront diverse domain changes that curtail their adaptability to novel domains. We categorize these changes into three types: 1) alterations in pose and viewpoint, 2) variations in visual capture conditions, and 3) diversity in modalities. Initially, models trained on specific viewpoints may falter when faced with viewpoints outside their training range. Second, changes in visual data capture conditions, encompassing changes in illumination or image resolution, can erode the generalization of trained models. Third, employing pre-trained models across distinct modalities, such as RGB, Lidar point clouds, Radar maps, or text embeddings, can lead to performance degradation. In this thesis, we propose to perform domain alignment to handle the aforementioned domain changes.The first segment of this thesis outlines our approach to performing domain alignment without the need for arduously training extensive models across multiple domains. We advocate for efficient handling of each type of change through visual representation learning techniques, utilizing models with minimal network parameters and judicious training data. This process, known as domain adaptation, unfolds in three stages. Initially, for pose and viewpoint variation, we propose acquiring viewpoint-invariant or pose-invariant representations, relevant to tasks like Re-ID, object tracking, and 3D face rendering. Subsequently, to mitigate the impact of changes in visual capture conditions, we harness semi-supervised and adversarial learning methods for tasks such as object detection and Re-ID. Lastly, to address cross-modal domain changes, we leverage self-training strategies to cultivate modality-agnostic representations for object detection.The second part of this thesis extends our domain-aligning framework to manage scenarios involving more than two forms of domain changes. To concurrently handle viewpoint variation and diverse modalities, we devise models capable of learning view-invariant representations for multiple modalities within the realm of 3D human pose estimation and rendering. Moreover, to combat changes arising from changes in resolution and diverse modalities in physical devices (e.g., ADC signals and Radar's RGB images), we advocate for the acquisition of super-resolution representations using models featuring complex values. Broadly, this thesis delves into the intricacies of perception tasks affected by domain changes and provides pragmatic solutions to address these challenges in real-world contexts.
- 일반주제명
- Computer engineering.
- 일반주제명
- Computer science.
- 키워드
- Deep learning
- 키워드
- Perception tasks
- 기타저자
- Carnegie Mellon University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 85-03B.
- 기본자료저록
- Dissertation Abstract International
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008240612s2023 us |||||||||||||||c||eng d■001000016934714
■00520240214101645
■006m o d
■007cr#unu||||||||
■020 ▼a9798380394345
■035 ▼a(MiAaPQ)AAI30633473
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aLi, Yu-Jhe.▼0(orcid)0000-0002-0912-4742
■24510▼aDomain Adapted Visual Representation Learning for Machine Perception▼h[electronic resource]
■260 ▼a[S.l.]:▼bCarnegie Mellon University. ▼c2023
■260 1▼aAnn Arbor :▼bProQuest Dissertations & Theses, ▼c2023
■300 ▼a1 online resource(194 p.)
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-03, Section: B.
■500 ▼aAdvisor: Kitani, Kris.
■5021 ▼aThesis (Ph.D.)--Carnegie Mellon University, 2023.
■506 ▼aThis item must not be sold to any third party vendors.
■520 ▼aOur objective is to enhance the generalization capabilities of existing machine perception models and achieve diverse domain alignments through adept representation learning. Many established approaches for perception tasks, encompassing object classification, detection, tracking, and rendering, often confront diverse domain changes that curtail their adaptability to novel domains. We categorize these changes into three types: 1) alterations in pose and viewpoint, 2) variations in visual capture conditions, and 3) diversity in modalities. Initially, models trained on specific viewpoints may falter when faced with viewpoints outside their training range. Second, changes in visual data capture conditions, encompassing changes in illumination or image resolution, can erode the generalization of trained models. Third, employing pre-trained models across distinct modalities, such as RGB, Lidar point clouds, Radar maps, or text embeddings, can lead to performance degradation. In this thesis, we propose to perform domain alignment to handle the aforementioned domain changes.The first segment of this thesis outlines our approach to performing domain alignment without the need for arduously training extensive models across multiple domains. We advocate for efficient handling of each type of change through visual representation learning techniques, utilizing models with minimal network parameters and judicious training data. This process, known as domain adaptation, unfolds in three stages. Initially, for pose and viewpoint variation, we propose acquiring viewpoint-invariant or pose-invariant representations, relevant to tasks like Re-ID, object tracking, and 3D face rendering. Subsequently, to mitigate the impact of changes in visual capture conditions, we harness semi-supervised and adversarial learning methods for tasks such as object detection and Re-ID. Lastly, to address cross-modal domain changes, we leverage self-training strategies to cultivate modality-agnostic representations for object detection.The second part of this thesis extends our domain-aligning framework to manage scenarios involving more than two forms of domain changes. To concurrently handle viewpoint variation and diverse modalities, we devise models capable of learning view-invariant representations for multiple modalities within the realm of 3D human pose estimation and rendering. Moreover, to combat changes arising from changes in resolution and diverse modalities in physical devices (e.g., ADC signals and Radar's RGB images), we advocate for the acquisition of super-resolution representations using models featuring complex values. Broadly, this thesis delves into the intricacies of perception tasks affected by domain changes and provides pragmatic solutions to address these challenges in real-world contexts.
■590 ▼aSchool code: 0041.
■650 4▼aComputer engineering.
■650 4▼aComputer science.
■653 ▼aDeep learning
■653 ▼aDomain adaptation
■653 ▼aMulti-modality learning
■653 ▼aPerception tasks
■653 ▼aRepresentation learning
■690 ▼a0464
■690 ▼a0984
■690 ▼a0800
■71020▼aCarnegie Mellon University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g85-03B.
■773 ▼tDissertation Abstract International
■790 ▼a0041
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T16934714▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
■980 ▼a202402▼f2024
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


