서브메뉴
검색
Multi-Modal Perception With Vision, Language, and Touch for Robot Manipulation
Multi-Modal Perception With Vision, Language, and Touch for Robot Manipulation
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103017
- ISBN
- 9798288861758
- DDC
- 629.8
- 저자명
- Huang, Huang.
- 서명/저자
- Multi-Modal Perception With Vision, Language, and Touch for Robot Manipulation
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 251 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Goldberg, Ken.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Humans use multiple senses to interact with the environments in daily life. Visions are used for perceiving and understanding the environments. Body awareness is used for positioning. Language is used for communicating and semantic understanding, and touch is used for con- tact feedback. Similarly, robots need the same kind of sensory integration for manipulation tasks in unstructured, real-world environments. This thesis explores how combining these various sensory inputs can improve robots' ability to manipulate objects in the real world. By integrating vision (which gives the robot detailed spatial info), proprioception (providing feedback on body positioning), language (for understanding and following instructions), and touch (for precise contact details), I develop safe, efficient, and generalizable robot systems. I present a series of contributions spanning sensorimotor control, motion planning, imitation learning, mechanical search, contact-rich manipulation, and multi-modal alignment, with a focus on improving robots' ability to perceive, reason, and act beyond the limitations of a single sensory modality.I start by exploring vision and proprioception integration to enhance control under distribution shifts and improve planning efficiency through diffusion-based trajectory generation. I introduce in-context imitation learning via next-token prediction, enabling prompt-driven adaptation to novel tasks. Vision and language is then incorporated to improve mechanical search for occluded objects and general manipulation. By leveraging large vision-language models, I demonstrate improved semantic reasoning, resulting in more effective manipulation strategies. I then investigate touch sensing in precise manipulation tasks such as industrial insertion and garment handling, enabling self-supervised policy learning and visuo-tactile pretraining that significantly improve task success rates. Finally, I introduce a novel dataset aligning vision, touch, and language to facilitate multi-modal learning for robotics.Through theoretical analysis, simulation, and real-world experiments, this thesis demonstrates how multi-modal perception enhances generalization, adaptability, and safety in robot manipulation.
- 일반주제명
- Robotics
- 일반주제명
- Computer engineering
- 키워드
- Robot learning
- 기타저자
- University of California, Berkeley Electrical Engineering & Computer Sciences
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356686
■00520260202103017
■006m o d
■007cr#unu||||||||
■020 ▼a9798288861758
■035 ▼a(MiAaPQ)AAI31843883
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a629.8
■1001 ▼aHuang, Huang.
■24510▼aMulti-Modal Perception With Vision, Language, and Touch for Robot Manipulation
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a251 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Goldberg, Ken.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aHumans use multiple senses to interact with the environments in daily life. Visions are used for perceiving and understanding the environments. Body awareness is used for positioning. Language is used for communicating and semantic understanding, and touch is used for con- tact feedback. Similarly, robots need the same kind of sensory integration for manipulation tasks in unstructured, real-world environments. This thesis explores how combining these various sensory inputs can improve robots' ability to manipulate objects in the real world. By integrating vision (which gives the robot detailed spatial info), proprioception (providing feedback on body positioning), language (for understanding and following instructions), and touch (for precise contact details), I develop safe, efficient, and generalizable robot systems. I present a series of contributions spanning sensorimotor control, motion planning, imitation learning, mechanical search, contact-rich manipulation, and multi-modal alignment, with a focus on improving robots' ability to perceive, reason, and act beyond the limitations of a single sensory modality.I start by exploring vision and proprioception integration to enhance control under distribution shifts and improve planning efficiency through diffusion-based trajectory generation. I introduce in-context imitation learning via next-token prediction, enabling prompt-driven adaptation to novel tasks. Vision and language is then incorporated to improve mechanical search for occluded objects and general manipulation. By leveraging large vision-language models, I demonstrate improved semantic reasoning, resulting in more effective manipulation strategies. I then investigate touch sensing in precise manipulation tasks such as industrial insertion and garment handling, enabling self-supervised policy learning and visuo-tactile pretraining that significantly improve task success rates. Finally, I introduce a novel dataset aligning vision, touch, and language to facilitate multi-modal learning for robotics.Through theoretical analysis, simulation, and real-world experiments, this thesis demonstrates how multi-modal perception enhances generalization, adaptability, and safety in robot manipulation.
■590 ▼aSchool code: 0028.
■650 4▼aRobotics
■650 4▼aComputer engineering
■653 ▼aMulti-modal perception
■653 ▼aRobot learning
■653 ▼aVision language action models
■690 ▼a0771
■690 ▼a0464
■71020▼aUniversity of California, Berkeley▼bElectrical Engineering & Computer Sciences.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356686▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


