본문

서브메뉴

Advancing Vision-Language and Language Models in Low-Resource Settings
Advancing Vision-Language and Language Models in Low-Resource Settings
Advancing Vision-Language and Language Models in Low-Resource Settings

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152827
ISBN  
9798384095958
DDC  
004
저자명  
Monajatipoor, Masoud.
서명/저자  
Advancing Vision-Language and Language Models in Low-Resource Settings
발행사항  
[Sl] : University of California, Los Angeles, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
86 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
주기사항  
Advisor: Chang, Kai-Wei;Yang, Lin.
학위논문주기  
Thesis (Ph.D.)--University of California, Los Angeles, 2024.
초록/해제  
요약Vision-language modeling is a crucial subfield of AI that focuses on jointly learning and representing image and text data, often using one modality to enhance understanding of the other. In cognitive science, humans use their visual system to grasp deep aspects of a concept, such as shape and size, while language helps them understand its semantics. Similarly, a machine can gain a better understanding of the world by utilizing multiple modalities, providing deeper insights compared to learning from a single modality. VL modeling is widely explored in the general domain, thanks to the vast image-text data available online and extensive annotated VL datasets. There are several strong VL models in the general domain, such as CLIP, which perform well on various tasks. However, in low-resource domains with limited data or dense knowledge areas, like the medical field, the data shortage hinders the development of robust multimodal models with reliable performance, especially where model reliability is critical. My research goal is to study the underlying capability of Vision-Language and Language Models and to develop innovative approaches to enhance their usage for low-resource domains such as the medical domain.
일반주제명  
Computer science
일반주제명  
Biomedical engineering
일반주제명  
Information technology
키워드  
Multimodal models
키워드  
Natural language processing
키워드  
Vision-language modeling
키워드  
Large language models
키워드  
Low-resource domains
기타저자  
University of California, Los Angeles Electrical and Computer Engineering 0333
기본자료저록  
Dissertations Abstracts International. 86-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164059
■00520250211152827
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384095958
■035    ▼a(MiAaPQ)AAI31560219
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aMonajatipoor,  Masoud.
■24510▼aAdvancing  Vision-Language  and  Language  Models  in  Low-Resource  Settings
■260    ▼a[Sl]▼bUniversity  of  California,  Los  Angeles▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a86  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  B.
■500    ▼aAdvisor:  Chang,  Kai-Wei;Yang,  Lin.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Los  Angeles,  2024.
■520    ▼aVision-language  modeling  is  a  crucial  subfield  of  AI  that  focuses  on  jointly  learning  and  representing  image  and  text  data,  often  using  one  modality  to  enhance  understanding  of  the  other.  In  cognitive  science,  humans  use  their  visual  system  to  grasp  deep  aspects  of  a  concept,  such  as  shape  and  size,  while  language  helps  them  understand  its  semantics.  Similarly,  a  machine  can  gain  a  better  understanding  of  the  world  by  utilizing  multiple  modalities,  providing  deeper  insights  compared  to  learning  from  a  single  modality.  VL  modeling  is  widely  explored  in  the  general  domain,  thanks  to  the  vast  image-text  data  available  online  and  extensive  annotated  VL  datasets.  There  are  several  strong  VL  models  in  the  general  domain,  such  as  CLIP,  which  perform  well  on  various  tasks.  However,  in  low-resource  domains  with  limited  data  or  dense  knowledge  areas,  like  the  medical  field,  the  data  shortage  hinders  the  development  of  robust  multimodal  models  with  reliable  performance,  especially  where  model  reliability  is  critical.  My  research  goal  is  to  study  the  underlying  capability  of  Vision-Language  and  Language  Models  and  to  develop  innovative  approaches  to  enhance  their  usage  for  low-resource  domains  such  as  the  medical  domain.
■590    ▼aSchool  code:  0031.
■650  4▼aComputer  science
■650  4▼aBiomedical  engineering
■650  4▼aInformation  technology
■653    ▼aMultimodal  models
■653    ▼aNatural  language  processing
■653    ▼aVision-language  modeling
■653    ▼aLarge  language  models
■653    ▼aLow-resource  domains
■690    ▼a0984
■690    ▼a0489
■690    ▼a0541
■690    ▼a0800
■71020▼aUniversity  of  California,  Los  Angeles▼bElectrical  and  Computer  Engineering  0333.
■7730  ▼tDissertations  Abstracts  International▼g86-03B.
■790    ▼a0031
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164059▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11992 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.