서브메뉴
검색
Robust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applications
Robust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applications
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105529
- ISBN
- 9798263343934
- DDC
- 616.8522
- 저자명
- Verma, Gaurav.
- 서명/저자
- Robust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applications
- 발행사항
- [Sl] : Georgia Institute of Technology, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 213 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: A.
- 주기사항
- Advisor: Kumar, Srijan.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2025.
- 초록/해제
- 요약Large artificial intelligence (AI) models have attracted widespread attention for theirimpressive, sometimes "superhuman," performance on standardized benchmarks. Yet,a decade of research applying these models in domains such as healthcare, web safety,and education has revealed persistent challenges. For instance, many popular AI modelsare brittle to small input variations, large language models (LLMs) exhibit sensitivity toprompt formatting, and their performance deteriorates in highly specialized settings. Theseshortcomings are often amplified when AI systems struggle to effectively support diverseuser groups, such as users with a lower Need For Cognition. Consequently, the successfuladoption of large AI models in specialized verticals requires strategic efforts.The key contributions made in this thesis - along with the thesis statement, overview, andimpact - are presented in Chapter 1. To address the aforementioned challenges, this thesisstarts by presenting a framework that helps researchers and developers optimize large-modeldevelopment and deployment (Part 1; Chapter 2). The framework articulates four modularlayers: (i) starting with large AI models at the bottom, (ii) vertical-agnostic properties, (iii)vertical-specific applications, and (iv) finally, vertical-user interfacing. Beyond capturingthe modular steps in developing practical systems that deliver real value to end users, theframework also serves as the foundation for further contributions made in this thesis.Large AI models that process, understand, and generate multimodal data-spanningvision and language-form the backbone of human-like AI interactions and richer worldunderstanding than unimodal systems. To this end, this thesis advances the use of largemultimodal models in multiple verticals by addressing three vertical-agnostic properties(Part 2): (a) robustness to realistic data variations (Chapter 3), (b) efficient cross-modalmapping (Chapter 4), and (c) adaptability to new tasks and domains (Chapter 5). First, weassess how current multimodal models respond to cross-modally grounded input variationsand expose their brittleness. Next, we propose a straightforward yet effective methodto capture text "visualness," thereby making text-to-image retrieval and generation moreefficient. Finally, we show how to adapt multimodal agents to custom workflows withminimal human demonstrations. Addressing these issues of robustness, efficiency, andadaptability reduces barriers to integrating multimodal AI across verticals, enabling moreeffective remediation techniques.Figure 1: Overview of the thesis; read from bottom to top. We address vertical-agnosticproperties as well as vertical-specific challenges of applying large multimodal models.Building on these foundational elements, the thesis transitions from vertical-agnosticproperties to vertical-specific applications (Part 3). We show that delivering value inspecific verticals requires tailored data, models, and evaluation methods. Focusing onweb safety and well-being, we collaborate with domain experts to (a) characterize anddetect violence-provoking speech (Chapter 6), and (b) use large language models (LLMs) touncover personal well-being insights that can inform policymaking (Chapter 7). Throughthese focused efforts, we develop a nuanced view of current AI strengths and limitations.Concluding these explorations, we demonstrate how vertical-specific insights can loop backto improve large AI models, by spotlighting inequitable outcomes across languages andshowing that multimodal training mitigates these disparities (Chapter 8).In summary, while large multimodal models have the potential to revolutionize verticalsystems and deliver tangible value to end users, they remain challenging to adopt acrossspecialized domains. This thesis offers a modular framework for addressing both verticalagnostic properties - robustness, efficiency, and adaptability - and vertical-specificchallenges, exemplified by web safety and well-being. Beyond these contributions, weemphasize that effective user interfacing stands as a major open challenge, providingopportunities for future work (Chapter 9).
- 일반주제명
- Anxiety
- 일반주제명
- Adaptability
- 일반주제명
- Humanitarianism
- 일반주제명
- Adaptation
- 일반주제명
- Violence
- 일반주제명
- Community
- 일반주제명
- Likert scale
- 일반주제명
- Keywords
- 일반주제명
- Image retrieval
- 일반주제명
- False information
- 일반주제명
- Multilingualism
- 일반주제명
- Speech
- 일반주제명
- Large language models
- 일반주제명
- Bilingual education
- 일반주제명
- Clinical psychology
- 일반주제명
- Web studies
- 기본자료저록
- Dissertations Abstracts International. 87-05A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017360456
■00520260202105529
■006m o d
■007cr#unu||||||||
■020 ▼a9798263343934
■035 ▼a(MiAaPQ)AAI32309806
■035 ▼a(MiAaPQ)GeorgiaTech77829
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a616.8522
■1001 ▼aVerma, Gaurav.
■24510▼aRobust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applications
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a213 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: A.
■500 ▼aAdvisor: Kumar, Srijan.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2025.
■520 ▼aLarge artificial intelligence (AI) models have attracted widespread attention for theirimpressive, sometimes "superhuman," performance on standardized benchmarks. Yet,a decade of research applying these models in domains such as healthcare, web safety,and education has revealed persistent challenges. For instance, many popular AI modelsare brittle to small input variations, large language models (LLMs) exhibit sensitivity toprompt formatting, and their performance deteriorates in highly specialized settings. Theseshortcomings are often amplified when AI systems struggle to effectively support diverseuser groups, such as users with a lower Need For Cognition. Consequently, the successfuladoption of large AI models in specialized verticals requires strategic efforts.The key contributions made in this thesis - along with the thesis statement, overview, andimpact - are presented in Chapter 1. To address the aforementioned challenges, this thesisstarts by presenting a framework that helps researchers and developers optimize large-modeldevelopment and deployment (Part 1; Chapter 2). The framework articulates four modularlayers: (i) starting with large AI models at the bottom, (ii) vertical-agnostic properties, (iii)vertical-specific applications, and (iv) finally, vertical-user interfacing. Beyond capturingthe modular steps in developing practical systems that deliver real value to end users, theframework also serves as the foundation for further contributions made in this thesis.Large AI models that process, understand, and generate multimodal data-spanningvision and language-form the backbone of human-like AI interactions and richer worldunderstanding than unimodal systems. To this end, this thesis advances the use of largemultimodal models in multiple verticals by addressing three vertical-agnostic properties(Part 2): (a) robustness to realistic data variations (Chapter 3), (b) efficient cross-modalmapping (Chapter 4), and (c) adaptability to new tasks and domains (Chapter 5). First, weassess how current multimodal models respond to cross-modally grounded input variationsand expose their brittleness. Next, we propose a straightforward yet effective methodto capture text "visualness," thereby making text-to-image retrieval and generation moreefficient. Finally, we show how to adapt multimodal agents to custom workflows withminimal human demonstrations. Addressing these issues of robustness, efficiency, andadaptability reduces barriers to integrating multimodal AI across verticals, enabling moreeffective remediation techniques.Figure 1: Overview of the thesis; read from bottom to top. We address vertical-agnosticproperties as well as vertical-specific challenges of applying large multimodal models.Building on these foundational elements, the thesis transitions from vertical-agnosticproperties to vertical-specific applications (Part 3). We show that delivering value inspecific verticals requires tailored data, models, and evaluation methods. Focusing onweb safety and well-being, we collaborate with domain experts to (a) characterize anddetect violence-provoking speech (Chapter 6), and (b) use large language models (LLMs) touncover personal well-being insights that can inform policymaking (Chapter 7). Throughthese focused efforts, we develop a nuanced view of current AI strengths and limitations.Concluding these explorations, we demonstrate how vertical-specific insights can loop backto improve large AI models, by spotlighting inequitable outcomes across languages andshowing that multimodal training mitigates these disparities (Chapter 8).In summary, while large multimodal models have the potential to revolutionize verticalsystems and deliver tangible value to end users, they remain challenging to adopt acrossspecialized domains. This thesis offers a modular framework for addressing both verticalagnostic properties - robustness, efficiency, and adaptability - and vertical-specificchallenges, exemplified by web safety and well-being. Beyond these contributions, weemphasize that effective user interfacing stands as a major open challenge, providingopportunities for future work (Chapter 9).
■590 ▼aSchool code: 0078.
■650 4▼aAnxiety
■650 4▼aAdaptability
■650 4▼aHumanitarianism
■650 4▼aAdaptation
■650 4▼aViolence
■650 4▼aCommunity
■650 4▼aLikert scale
■650 4▼aKeywords
■650 4▼aImage retrieval
■650 4▼aFalse information
■650 4▼aMultilingualism
■650 4▼aSpeech
■650 4▼aLarge language models
■650 4▼aBilingual education
■650 4▼aClinical psychology
■650 4▼aWeb studies
■690 ▼a0800
■690 ▼a0282
■690 ▼a0622
■690 ▼a0646
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05A.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360456▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


