본문

서브메뉴

Domain Adapted Visual Representation Learning for Machine Perception- [electronic resource]
Domain Adapted Visual Representation Learning for Machine Perception - [electronic resourc...
Domain Adapted Visual Representation Learning for Machine Perception- [electronic resource]

Detailed Information

자료유형  
 학위논문파일 국외
최종처리일시  
20240214101645
ISBN  
9798380394345
DDC  
621.3
저자명  
Li, Yu-Jhe.
서명/저자  
Domain Adapted Visual Representation Learning for Machine Perception - [electronic resource]
발행사항  
[S.l.]: : Carnegie Mellon University., 2023
발행사항  
Ann Arbor : : ProQuest Dissertations & Theses,, 2023
형태사항  
1 online resource(194 p.)
주기사항  
Source: Dissertations Abstracts International, Volume: 85-03, Section: B.
주기사항  
Advisor: Kitani, Kris.
학위논문주기  
Thesis (Ph.D.)--Carnegie Mellon University, 2023.
사용제한주기  
This item must not be sold to any third party vendors.
초록/해제  
요약Our objective is to enhance the generalization capabilities of existing machine perception models and achieve diverse domain alignments through adept representation learning. Many established approaches for perception tasks, encompassing object classification, detection, tracking, and rendering, often confront diverse domain changes that curtail their adaptability to novel domains. We categorize these changes into three types: 1) alterations in pose and viewpoint, 2) variations in visual capture conditions, and 3) diversity in modalities. Initially, models trained on specific viewpoints may falter when faced with viewpoints outside their training range. Second, changes in visual data capture conditions, encompassing changes in illumination or image resolution, can erode the generalization of trained models. Third, employing pre-trained models across distinct modalities, such as RGB, Lidar point clouds, Radar maps, or text embeddings, can lead to performance degradation. In this thesis, we propose to perform domain alignment to handle the aforementioned domain changes.The first segment of this thesis outlines our approach to performing domain alignment without the need for arduously training extensive models across multiple domains. We advocate for efficient handling of each type of change through visual representation learning techniques, utilizing models with minimal network parameters and judicious training data. This process, known as domain adaptation, unfolds in three stages. Initially, for pose and viewpoint variation, we propose acquiring viewpoint-invariant or pose-invariant representations, relevant to tasks like Re-ID, object tracking, and 3D face rendering. Subsequently, to mitigate the impact of changes in visual capture conditions, we harness semi-supervised and adversarial learning methods for tasks such as object detection and Re-ID. Lastly, to address cross-modal domain changes, we leverage self-training strategies to cultivate modality-agnostic representations for object detection.The second part of this thesis extends our domain-aligning framework to manage scenarios involving more than two forms of domain changes. To concurrently handle viewpoint variation and diverse modalities, we devise models capable of learning view-invariant representations for multiple modalities within the realm of 3D human pose estimation and rendering. Moreover, to combat changes arising from changes in resolution and diverse modalities in physical devices (e.g., ADC signals and Radar's RGB images), we advocate for the acquisition of super-resolution representations using models featuring complex values. Broadly, this thesis delves into the intricacies of perception tasks affected by domain changes and provides pragmatic solutions to address these challenges in real-world contexts.
일반주제명  
Computer engineering.
일반주제명  
Computer science.
키워드  
Deep learning
키워드  
Domain adaptation
키워드  
Multi-modality learning
키워드  
Perception tasks
키워드  
Representation learning
기타저자  
Carnegie Mellon University Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 85-03B.
기본자료저록  
Dissertation Abstract International
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008240612s2023      us  |||||||||||||||c||eng  d
■001000016934714
■00520240214101645
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798380394345
■035    ▼a(MiAaPQ)AAI30633473
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aLi,  Yu-Jhe.▼0(orcid)0000-0002-0912-4742
■24510▼aDomain  Adapted  Visual  Representation  Learning  for  Machine  Perception▼h[electronic  resource]
■260    ▼a[S.l.]:▼bCarnegie  Mellon  University.  ▼c2023
■260  1▼aAnn  Arbor  :▼bProQuest  Dissertations  &  Theses,  ▼c2023
■300    ▼a1  online  resource(194  p.)
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-03,  Section:  B.
■500    ▼aAdvisor:  Kitani,  Kris.
■5021  ▼aThesis  (Ph.D.)--Carnegie  Mellon  University,  2023.
■506    ▼aThis  item  must  not  be  sold  to  any  third  party  vendors.
■520    ▼aOur  objective  is  to  enhance  the  generalization  capabilities  of  existing  machine  perception  models  and  achieve  diverse  domain  alignments  through  adept  representation  learning.  Many  established  approaches  for  perception  tasks,  encompassing  object  classification,  detection,  tracking,  and  rendering,  often  confront  diverse  domain  changes  that  curtail  their  adaptability  to  novel  domains.  We  categorize  these  changes  into  three  types:  1)  alterations  in  pose  and  viewpoint,  2)  variations  in  visual  capture  conditions,  and  3)  diversity  in  modalities.  Initially,  models  trained  on  specific  viewpoints  may  falter  when  faced  with  viewpoints  outside  their  training  range.  Second,  changes  in  visual  data  capture  conditions,  encompassing  changes  in  illumination  or  image  resolution,  can  erode  the  generalization  of  trained  models.  Third,  employing  pre-trained  models  across  distinct  modalities,  such  as  RGB,  Lidar  point  clouds,  Radar  maps,  or  text  embeddings,  can  lead  to  performance  degradation.  In  this  thesis,  we  propose  to  perform  domain  alignment  to  handle  the  aforementioned  domain  changes.The  first  segment  of  this  thesis  outlines  our  approach  to  performing  domain  alignment  without  the  need  for  arduously  training  extensive  models  across  multiple  domains.  We  advocate  for  efficient  handling  of  each  type  of  change  through  visual  representation  learning  techniques,  utilizing  models  with  minimal  network  parameters  and  judicious  training  data.  This  process,  known  as  domain  adaptation,  unfolds  in  three  stages.  Initially,  for  pose  and  viewpoint  variation,  we  propose  acquiring  viewpoint-invariant  or  pose-invariant  representations,  relevant  to  tasks  like  Re-ID,  object  tracking,  and  3D  face  rendering.  Subsequently,  to  mitigate  the  impact  of  changes  in  visual  capture  conditions,  we  harness  semi-supervised  and  adversarial  learning  methods  for  tasks  such  as  object  detection  and  Re-ID.  Lastly,  to  address  cross-modal  domain  changes,  we  leverage  self-training  strategies  to  cultivate  modality-agnostic  representations  for  object  detection.The  second  part  of  this  thesis  extends  our  domain-aligning  framework  to  manage  scenarios  involving  more  than  two  forms  of  domain  changes.  To  concurrently  handle  viewpoint  variation  and  diverse  modalities,  we  devise  models  capable  of  learning  view-invariant  representations  for  multiple  modalities  within  the  realm  of  3D  human  pose  estimation  and  rendering.  Moreover,  to  combat  changes  arising  from  changes  in  resolution  and  diverse  modalities  in  physical  devices  (e.g.,  ADC  signals  and  Radar's  RGB  images),  we  advocate  for  the  acquisition  of  super-resolution  representations  using  models  featuring  complex  values.  Broadly,  this  thesis  delves  into  the  intricacies  of  perception  tasks  affected  by  domain  changes  and  provides  pragmatic  solutions  to  address  these  challenges  in  real-world  contexts.
■590    ▼aSchool  code:  0041.
■650  4▼aComputer  engineering.
■650  4▼aComputer  science.
■653    ▼aDeep  learning
■653    ▼aDomain  adaptation
■653    ▼aMulti-modality  learning
■653    ▼aPerception  tasks
■653    ▼aRepresentation  learning
■690    ▼a0464
■690    ▼a0984
■690    ▼a0800
■71020▼aCarnegie  Mellon  University▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g85-03B.
■773    ▼tDissertation  Abstract  International
■790    ▼a0041
■791    ▼aPh.D.
■792    ▼a2023
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T16934714▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.
■980    ▼a202402▼f2024

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF09443 전자도서 My Folder 부재도서신고 비도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.