서브메뉴
검색
Perception and Reasoning With Visual Relations in Humans and Machines
Perception and Reasoning With Visual Relations in Humans and Machines
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104644
- ISBN
- 9798280755574
- DDC
- 153
- 저자명
- Fu, Shuhao.
- 서명/저자
- Perception and Reasoning With Visual Relations in Humans and Machines
- 발행사항
- [Sl] : University of California, Los Angeles, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 193 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
- 주기사항
- Advisor: Lu, Hongjing.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2025.
- 초록/해제
- 요약The world consists of objects, and narratives are built from words. Rather than perceiving the world as a list of entities, humans represent the world in a more cohesive way by apprehending and expressing the relations between entities. Equipped with the ability to represent and process relations, human thinking and creativity are deeply rooted in analogy: the ability to identify and utilize resemblances based on relations between entities. Despite decades of research on relation perception and analogy, a fundamental question remains: how do relational representations arise from linguistic and visual inputs? This dissertation investigates the cognitive and computational mechanisms underlying relation perception and reasoning in humans, and examines the capacities of advanced AI models across a wide range of relation tasks. Through a combination of behavioral experiments, computational modeling, and AI model evaluation, this work bridges insights from cognitive science with innovations in deep learning.Chapter 2 investigates relation perception in humans using both realistic and synthetic stimuli, and compares human performance with that of vision-only and vision-language models, highlighting key areas where current AI falls short in accounting for human relation perception. Chapter 3 systematically evaluates the capacity for relation understanding and compositionality in multimodal generative models, revealing fundamental limitations in their ability to ground spatial and agentic relations. Chapter 4 focuses on spatial relations in 3D object recognition, and demonstrates that hierarchical abstraction mechanisms, commonly known as local-to-global visual processing, are crucial for enabling AI models to achieve human-like robustness for 3D object recognition. Chapter 5 examines visual analogy with realistic car stimuli, showing that a part-based comparison model more closely aligns with human reasoning performance than neural networks trained specifically on analogy tasks. Chapter 6 introduces VisiPAM, a vision-based probabilistic analogical mapping model that requires no analogy-specific training, yet best accounts for human performance in a novel mapping task.By integrating experimental and modeling approaches, this work offers novel benchmarks, cognitively-inspired design principles, and empirical evidence that reveals both the promise and limitations of current AI systems in relational tasks. These findings establish a foundation for developing models with greater generalizability, interpretability, and alignment with human relation perception and reasoning.
- 일반주제명
- Cognitive psychology
- 일반주제명
- Linguistics
- 일반주제명
- Computer science
- 키워드
- Machine learning
- 키워드
- Perception
- 키워드
- Analogy
- 기타저자
- University of California, Los Angeles Psychology 0780
- 기본자료저록
- Dissertations Abstracts International. 86-12A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358325
■00520260202104644
■006m o d
■007cr#unu||||||||
■020 ▼a9798280755574
■035 ▼a(MiAaPQ)AAI32114485
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a153
■1001 ▼aFu, Shuhao.
■24510▼aPerception and Reasoning With Visual Relations in Humans and Machines
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a193 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: A.
■500 ▼aAdvisor: Lu, Hongjing.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2025.
■520 ▼aThe world consists of objects, and narratives are built from words. Rather than perceiving the world as a list of entities, humans represent the world in a more cohesive way by apprehending and expressing the relations between entities. Equipped with the ability to represent and process relations, human thinking and creativity are deeply rooted in analogy: the ability to identify and utilize resemblances based on relations between entities. Despite decades of research on relation perception and analogy, a fundamental question remains: how do relational representations arise from linguistic and visual inputs? This dissertation investigates the cognitive and computational mechanisms underlying relation perception and reasoning in humans, and examines the capacities of advanced AI models across a wide range of relation tasks. Through a combination of behavioral experiments, computational modeling, and AI model evaluation, this work bridges insights from cognitive science with innovations in deep learning.Chapter 2 investigates relation perception in humans using both realistic and synthetic stimuli, and compares human performance with that of vision-only and vision-language models, highlighting key areas where current AI falls short in accounting for human relation perception. Chapter 3 systematically evaluates the capacity for relation understanding and compositionality in multimodal generative models, revealing fundamental limitations in their ability to ground spatial and agentic relations. Chapter 4 focuses on spatial relations in 3D object recognition, and demonstrates that hierarchical abstraction mechanisms, commonly known as local-to-global visual processing, are crucial for enabling AI models to achieve human-like robustness for 3D object recognition. Chapter 5 examines visual analogy with realistic car stimuli, showing that a part-based comparison model more closely aligns with human reasoning performance than neural networks trained specifically on analogy tasks. Chapter 6 introduces VisiPAM, a vision-based probabilistic analogical mapping model that requires no analogy-specific training, yet best accounts for human performance in a novel mapping task.By integrating experimental and modeling approaches, this work offers novel benchmarks, cognitively-inspired design principles, and empirical evidence that reveals both the promise and limitations of current AI systems in relational tasks. These findings establish a foundation for developing models with greater generalizability, interpretability, and alignment with human relation perception and reasoning.
■590 ▼aSchool code: 0031.
■650 4▼aCognitive psychology
■650 4▼aLinguistics
■650 4▼aComputer science
■653 ▼aCognitive science
■653 ▼aMachine learning
■653 ▼aPerception
■653 ▼aRelational reasoning
■653 ▼aAnalogy
■690 ▼a0633
■690 ▼a0984
■690 ▼a0800
■690 ▼a0290
■71020▼aUniversity of California, Los Angeles▼bPsychology 0780.
■7730 ▼tDissertations Abstracts International▼g86-12A.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358325▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


