서브메뉴
검색
Visual Intelligence Beyond Human Supervision
Visual Intelligence Beyond Human Supervision
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103608
- ISBN
- 9798293892341
- DDC
- 004
- 저자명
- Wang, XuDong.
- 서명/저자
- Visual Intelligence Beyond Human Supervision
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 231 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
- 주기사항
- Advisor: Darrell, Trevor.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Achieving artificial general intelligence requires developing models capable of perceiving, understanding, and interacting with the world across diverse sensory modalities-beyond the confines of language alone. While self-supervised learning has enabled remarkable advances in large language models (LLMs), replicating this success in the visual domain remains a significant challenge, largely due to the continued reliance on human-annotated data. This dissertation explores how self-supervised learning can unlock visual intelligence beyond human supervision, enabling models to learn directly from the inherent structure and regularities of the visual world.The thesis presents a series of efforts aimed at advancing this vision. First, it investigates self-supervised visual world understanding, demonstrating that models can achieve strong segmentation performance without the billions of labeled masks used in supervised approaches such as the Segment Anything Model (SAM). Instead, our work shows that models can "segment anything'' by leveraging the rich semantics present in unlabeled data. Second, it introduces methods that unify generative and discriminative visual models through self-supervision and synthetic data, allowing these systems to complement one another and improve both visual understanding and generation. Third, the dissertation examines how to build robust visual models through self-supervised debiased learning, proposing techniques that mitigate bias and enhance generalization under imperfect data conditions, within a data-centric representation learning framework.Together, these contributions serve a common goal: building scalable, multi-modality visual intelligence that learn not by mimicking human annotations, but by discovering the latent structure of the world itself!
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 기타저자
- University of California, Berkeley Electrical Engineering & Computer Sciences
- 기본자료저록
- Dissertations Abstracts International. 87-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357849
■00520260202103608
■006m o d
■007cr#unu||||||||
■020 ▼a9798293892341
■035 ▼a(MiAaPQ)AAI32043020
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWang, XuDong.
■24510▼aVisual Intelligence Beyond Human Supervision
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a231 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-04, Section: B.
■500 ▼aAdvisor: Darrell, Trevor.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aAchieving artificial general intelligence requires developing models capable of perceiving, understanding, and interacting with the world across diverse sensory modalities-beyond the confines of language alone. While self-supervised learning has enabled remarkable advances in large language models (LLMs), replicating this success in the visual domain remains a significant challenge, largely due to the continued reliance on human-annotated data. This dissertation explores how self-supervised learning can unlock visual intelligence beyond human supervision, enabling models to learn directly from the inherent structure and regularities of the visual world.The thesis presents a series of efforts aimed at advancing this vision. First, it investigates self-supervised visual world understanding, demonstrating that models can achieve strong segmentation performance without the billions of labeled masks used in supervised approaches such as the Segment Anything Model (SAM). Instead, our work shows that models can "segment anything'' by leveraging the rich semantics present in unlabeled data. Second, it introduces methods that unify generative and discriminative visual models through self-supervision and synthetic data, allowing these systems to complement one another and improve both visual understanding and generation. Third, the dissertation examines how to build robust visual models through self-supervised debiased learning, proposing techniques that mitigate bias and enhance generalization under imperfect data conditions, within a data-centric representation learning framework.Together, these contributions serve a common goal: building scalable, multi-modality visual intelligence that learn not by mimicking human annotations, but by discovering the latent structure of the world itself!
■590 ▼aSchool code: 0028.
■650 4▼aComputer science
■650 4▼aElectrical engineering
■653 ▼aLarge language models
■653 ▼aSegment Anything Model
■653 ▼aHuman-annotated data
■690 ▼a0800
■690 ▼a0984
■690 ▼a0544
■71020▼aUniversity of California, Berkeley▼bElectrical Engineering & Computer Sciences.
■7730 ▼tDissertations Abstracts International▼g87-04B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357849▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


