서브메뉴
검색
Advancing Vision-Language and Language Models in Low-Resource Settings
Advancing Vision-Language and Language Models in Low-Resource Settings
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152827
- ISBN
- 9798384095958
- DDC
- 004
- 서명/저자
- Advancing Vision-Language and Language Models in Low-Resource Settings
- 발행사항
- [Sl] : University of California, Los Angeles, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 86 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Chang, Kai-Wei;Yang, Lin.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2024.
- 초록/해제
- 요약Vision-language modeling is a crucial subfield of AI that focuses on jointly learning and representing image and text data, often using one modality to enhance understanding of the other. In cognitive science, humans use their visual system to grasp deep aspects of a concept, such as shape and size, while language helps them understand its semantics. Similarly, a machine can gain a better understanding of the world by utilizing multiple modalities, providing deeper insights compared to learning from a single modality. VL modeling is widely explored in the general domain, thanks to the vast image-text data available online and extensive annotated VL datasets. There are several strong VL models in the general domain, such as CLIP, which perform well on various tasks. However, in low-resource domains with limited data or dense knowledge areas, like the medical field, the data shortage hinders the development of robust multimodal models with reliable performance, especially where model reliability is critical. My research goal is to study the underlying capability of Vision-Language and Language Models and to develop innovative approaches to enhance their usage for low-resource domains such as the medical domain.
- 일반주제명
- Computer science
- 일반주제명
- Biomedical engineering
- 일반주제명
- Information technology
- 기타저자
- University of California, Los Angeles Electrical and Computer Engineering 0333
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164059
■00520250211152827
■006m o d
■007cr#unu||||||||
■020 ▼a9798384095958
■035 ▼a(MiAaPQ)AAI31560219
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aMonajatipoor, Masoud.
■24510▼aAdvancing Vision-Language and Language Models in Low-Resource Settings
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a86 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Chang, Kai-Wei;Yang, Lin.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2024.
■520 ▼aVision-language modeling is a crucial subfield of AI that focuses on jointly learning and representing image and text data, often using one modality to enhance understanding of the other. In cognitive science, humans use their visual system to grasp deep aspects of a concept, such as shape and size, while language helps them understand its semantics. Similarly, a machine can gain a better understanding of the world by utilizing multiple modalities, providing deeper insights compared to learning from a single modality. VL modeling is widely explored in the general domain, thanks to the vast image-text data available online and extensive annotated VL datasets. There are several strong VL models in the general domain, such as CLIP, which perform well on various tasks. However, in low-resource domains with limited data or dense knowledge areas, like the medical field, the data shortage hinders the development of robust multimodal models with reliable performance, especially where model reliability is critical. My research goal is to study the underlying capability of Vision-Language and Language Models and to develop innovative approaches to enhance their usage for low-resource domains such as the medical domain.
■590 ▼aSchool code: 0031.
■650 4▼aComputer science
■650 4▼aBiomedical engineering
■650 4▼aInformation technology
■653 ▼aMultimodal models
■653 ▼aNatural language processing
■653 ▼aVision-language modeling
■653 ▼aLarge language models
■653 ▼aLow-resource domains
■690 ▼a0984
■690 ▼a0489
■690 ▼a0541
■690 ▼a0800
■71020▼aUniversity of California, Los Angeles▼bElectrical and Computer Engineering 0333.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164059▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


