서브메뉴
검색
Extracting Knowledge with Multimodal and Multilingual Intelligent Systems
Extracting Knowledge with Multimodal and Multilingual Intelligent Systems
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105550
- ISBN
- 9798265401205
- DDC
- 496
- 저자명
- Chen, Yang.
- 서명/저자
- Extracting Knowledge with Multimodal and Multilingual Intelligent Systems
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 196 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Ritter, Alan;Xu, Wei.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약Recent advancements in large language models (LLMs) have revolutionized natural language processing and vision-language tasks. While demonstrating emergent capabilities, these models present challenges in responsible development, particularly in visual world knowledge, privacy concerns, and multilingual capabilities. This thesis addresses these challenges through three main contributions. First, we introduce InfoSeek, a vision-language benchmark assessing models' ability to leverage world knowledge for answering queries about visual entities. To improve performance on this challenging task, we develop multimodal retrieval-augmented generation systems to acquire knowledge from external resources. Inspired by the visual knowledge we found from InfoSeek, we raise an emergent privacy concern of multimodal LLM to reveal geolocation information of user posted images. We then present PrivQA, a benchmark evaluating models' ability to follow access control instructions and prevent private information disclosure. Our findings reveal biases and vulnerabilities in current privacy protection mechanisms, especially in adversarial settings. Third, we propose three approaches to enhance multilingual capabilities and improve information extraction (IE) in low-resource languages: TransFusion, a framework leveraging English translations to enhance multilingual performance; EasyProject, a simplified method for creating synthetic multilingual IE data; and a model selection algorithm predicting multilingual model performance on unseen languages. These contributions aim to develop and benchmark methods for extracting knowledge with multimodal and multilingual intelligent systems, addressing key challenges in emergent visual knowledge, privacy concerns, and multilingual capabilities.
- 일반주제명
- African languages
- 일반주제명
- Multilingualism
- 일반주제명
- Privacy
- 일반주제명
- Large language models
- 일반주제명
- Bilingualism
- 일반주제명
- Access control
- 일반주제명
- Labeling
- 일반주제명
- Bilingual education
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360580
■00520260202105550
■006m o d
■007cr#unu||||||||
■020 ▼a9798265401205
■035 ▼a(MiAaPQ)AAI32315679
■035 ▼a(MiAaPQ)GeorgiaTech75688
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a496
■1001 ▼aChen, Yang.
■24510▼aExtracting Knowledge with Multimodal and Multilingual Intelligent Systems
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a196 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Ritter, Alan;Xu, Wei.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aRecent advancements in large language models (LLMs) have revolutionized natural language processing and vision-language tasks. While demonstrating emergent capabilities, these models present challenges in responsible development, particularly in visual world knowledge, privacy concerns, and multilingual capabilities. This thesis addresses these challenges through three main contributions. First, we introduce InfoSeek, a vision-language benchmark assessing models' ability to leverage world knowledge for answering queries about visual entities. To improve performance on this challenging task, we develop multimodal retrieval-augmented generation systems to acquire knowledge from external resources. Inspired by the visual knowledge we found from InfoSeek, we raise an emergent privacy concern of multimodal LLM to reveal geolocation information of user posted images. We then present PrivQA, a benchmark evaluating models' ability to follow access control instructions and prevent private information disclosure. Our findings reveal biases and vulnerabilities in current privacy protection mechanisms, especially in adversarial settings. Third, we propose three approaches to enhance multilingual capabilities and improve information extraction (IE) in low-resource languages: TransFusion, a framework leveraging English translations to enhance multilingual performance; EasyProject, a simplified method for creating synthetic multilingual IE data; and a model selection algorithm predicting multilingual model performance on unseen languages. These contributions aim to develop and benchmark methods for extracting knowledge with multimodal and multilingual intelligent systems, addressing key challenges in emergent visual knowledge, privacy concerns, and multilingual capabilities.
■590 ▼aSchool code: 0078.
■650 4▼aAfrican languages
■650 4▼aMultilingualism
■650 4▼aPrivacy
■650 4▼aLarge language models
■650 4▼aBilingualism
■650 4▼aAccess control
■650 4▼aLabeling
■650 4▼aBilingual education
■690 ▼a0800
■690 ▼a0282
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360580▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


