본문

서브메뉴

Robust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applications
Robust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applicati...
Robust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applications

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105529
ISBN  
9798263343934
DDC  
616.8522
저자명  
Verma, Gaurav.
서명/저자  
Robust, Efficient, and Adaptable Multimodal Artificial Intelligence for Vertical Applications
발행사항  
[Sl] : Georgia Institute of Technology, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
213 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: A.
주기사항  
Advisor: Kumar, Srijan.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2025.
초록/해제  
요약Large artificial intelligence (AI) models have attracted widespread attention for theirimpressive, sometimes "superhuman," performance on standardized benchmarks. Yet,a decade of research applying these models in domains such as healthcare, web safety,and education has revealed persistent challenges. For instance, many popular AI modelsare brittle to small input variations, large language models (LLMs) exhibit sensitivity toprompt formatting, and their performance deteriorates in highly specialized settings. Theseshortcomings are often amplified when AI systems struggle to effectively support diverseuser groups, such as users with a lower Need For Cognition. Consequently, the successfuladoption of large AI models in specialized verticals requires strategic efforts.The key contributions made in this thesis - along with the thesis statement, overview, andimpact - are presented in Chapter 1. To address the aforementioned challenges, this thesisstarts by presenting a framework that helps researchers and developers optimize large-modeldevelopment and deployment (Part 1; Chapter 2). The framework articulates four modularlayers: (i) starting with large AI models at the bottom, (ii) vertical-agnostic properties, (iii)vertical-specific applications, and (iv) finally, vertical-user interfacing. Beyond capturingthe modular steps in developing practical systems that deliver real value to end users, theframework also serves as the foundation for further contributions made in this thesis.Large AI models that process, understand, and generate multimodal data-spanningvision and language-form the backbone of human-like AI interactions and richer worldunderstanding than unimodal systems. To this end, this thesis advances the use of largemultimodal models in multiple verticals by addressing three vertical-agnostic properties(Part 2): (a) robustness to realistic data variations (Chapter 3), (b) efficient cross-modalmapping (Chapter 4), and (c) adaptability to new tasks and domains (Chapter 5). First, weassess how current multimodal models respond to cross-modally grounded input variationsand expose their brittleness. Next, we propose a straightforward yet effective methodto capture text "visualness," thereby making text-to-image retrieval and generation moreefficient. Finally, we show how to adapt multimodal agents to custom workflows withminimal human demonstrations. Addressing these issues of robustness, efficiency, andadaptability reduces barriers to integrating multimodal AI across verticals, enabling moreeffective remediation techniques.Figure 1: Overview of the thesis; read from bottom to top. We address vertical-agnosticproperties as well as vertical-specific challenges of applying large multimodal models.Building on these foundational elements, the thesis transitions from vertical-agnosticproperties to vertical-specific applications (Part 3). We show that delivering value inspecific verticals requires tailored data, models, and evaluation methods. Focusing onweb safety and well-being, we collaborate with domain experts to (a) characterize anddetect violence-provoking speech (Chapter 6), and (b) use large language models (LLMs) touncover personal well-being insights that can inform policymaking (Chapter 7). Throughthese focused efforts, we develop a nuanced view of current AI strengths and limitations.Concluding these explorations, we demonstrate how vertical-specific insights can loop backto improve large AI models, by spotlighting inequitable outcomes across languages andshowing that multimodal training mitigates these disparities (Chapter 8).In summary, while large multimodal models have the potential to revolutionize verticalsystems and deliver tangible value to end users, they remain challenging to adopt acrossspecialized domains. This thesis offers a modular framework for addressing both verticalagnostic properties - robustness, efficiency, and adaptability - and vertical-specificchallenges, exemplified by web safety and well-being. Beyond these contributions, weemphasize that effective user interfacing stands as a major open challenge, providingopportunities for future work (Chapter 9).
일반주제명  
Anxiety
일반주제명  
Adaptability
일반주제명  
Humanitarianism
일반주제명  
Adaptation
일반주제명  
Violence
일반주제명  
Community
일반주제명  
Likert scale
일반주제명  
Keywords
일반주제명  
Image retrieval
일반주제명  
False information
일반주제명  
Multilingualism
일반주제명  
Speech
일반주제명  
Large language models
일반주제명  
Bilingual education
일반주제명  
Clinical psychology
일반주제명  
Web studies
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-05A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017360456
■00520260202105529
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263343934
■035    ▼a(MiAaPQ)AAI32309806
■035    ▼a(MiAaPQ)GeorgiaTech77829
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a616.8522
■1001  ▼aVerma,  Gaurav.
■24510▼aRobust,  Efficient,  and  Adaptable  Multimodal  Artificial  Intelligence  for  Vertical  Applications
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a213  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  A.
■500    ▼aAdvisor:  Kumar,  Srijan.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2025.
■520    ▼aLarge  artificial  intelligence  (AI)  models  have  attracted  widespread  attention  for  theirimpressive,  sometimes  "superhuman,"  performance  on  standardized  benchmarks.  Yet,a  decade  of  research  applying  these  models  in  domains  such  as  healthcare,  web  safety,and  education  has  revealed  persistent  challenges.  For  instance,  many  popular  AI  modelsare  brittle  to  small  input  variations,  large  language  models  (LLMs)  exhibit  sensitivity  toprompt  formatting,  and  their  performance  deteriorates  in  highly  specialized  settings.  Theseshortcomings  are  often  amplified  when  AI  systems  struggle  to  effectively  support  diverseuser  groups,  such  as  users  with  a  lower  Need  For  Cognition.  Consequently,  the  successfuladoption  of  large  AI  models  in  specialized  verticals  requires  strategic  efforts.The  key  contributions  made  in  this  thesis  -  along  with  the  thesis  statement,  overview,  andimpact  -  are  presented  in  Chapter  1.  To  address  the  aforementioned  challenges,  this  thesisstarts  by  presenting  a  framework  that  helps  researchers  and  developers  optimize  large-modeldevelopment  and  deployment  (Part  1;  Chapter  2).  The  framework  articulates  four  modularlayers:  (i)  starting  with  large  AI  models  at  the  bottom,  (ii)  vertical-agnostic  properties,  (iii)vertical-specific  applications,  and  (iv)  finally,  vertical-user  interfacing.  Beyond  capturingthe  modular  steps  in  developing  practical  systems  that  deliver  real  value  to  end  users,  theframework  also  serves  as  the  foundation  for  further  contributions  made  in  this  thesis.Large  AI  models  that  process,  understand,  and  generate  multimodal  data-spanningvision  and  language-form  the  backbone  of  human-like  AI  interactions  and  richer  worldunderstanding  than  unimodal  systems.  To  this  end,  this  thesis  advances  the  use  of  largemultimodal  models  in  multiple  verticals  by  addressing  three  vertical-agnostic  properties(Part  2):  (a)  robustness  to  realistic  data  variations  (Chapter  3),  (b)  efficient  cross-modalmapping  (Chapter  4),  and  (c)  adaptability  to  new  tasks  and  domains  (Chapter  5).  First,  weassess  how  current  multimodal  models  respond  to  cross-modally  grounded  input  variationsand  expose  their  brittleness.  Next,  we  propose  a  straightforward  yet  effective  methodto  capture  text  "visualness,"  thereby  making  text-to-image  retrieval  and  generation  moreefficient.  Finally,  we  show  how  to  adapt  multimodal  agents  to  custom  workflows  withminimal  human  demonstrations.  Addressing  these  issues  of  robustness,  efficiency,  andadaptability  reduces  barriers  to  integrating  multimodal  AI  across  verticals,  enabling  moreeffective  remediation  techniques.Figure  1:  Overview  of  the  thesis;  read  from  bottom  to  top.  We  address  vertical-agnosticproperties  as  well  as  vertical-specific  challenges  of  applying  large  multimodal  models.Building  on  these  foundational  elements,  the  thesis  transitions  from  vertical-agnosticproperties  to  vertical-specific  applications  (Part  3).  We  show  that  delivering  value  inspecific  verticals  requires  tailored  data,  models,  and  evaluation  methods.  Focusing  onweb  safety  and  well-being,  we  collaborate  with  domain  experts  to  (a)  characterize  anddetect  violence-provoking  speech  (Chapter  6),  and  (b)  use  large  language  models  (LLMs)  touncover  personal  well-being  insights  that  can  inform  policymaking  (Chapter  7).  Throughthese  focused  efforts,  we  develop  a  nuanced  view  of  current  AI  strengths  and  limitations.Concluding  these  explorations,  we  demonstrate  how  vertical-specific  insights  can  loop  backto  improve  large  AI  models,  by  spotlighting  inequitable  outcomes  across  languages  andshowing  that  multimodal  training  mitigates  these  disparities  (Chapter  8).In  summary,  while  large  multimodal  models  have  the  potential  to  revolutionize  verticalsystems  and  deliver  tangible  value  to  end  users,  they  remain  challenging  to  adopt  acrossspecialized  domains.  This  thesis  offers  a  modular  framework  for  addressing  both  verticalagnostic  properties  -  robustness,  efficiency,  and  adaptability  -  and  vertical-specificchallenges,  exemplified  by  web  safety  and  well-being.  Beyond  these  contributions,  weemphasize  that  effective  user  interfacing  stands  as  a  major  open  challenge,  providingopportunities  for  future  work  (Chapter  9).
■590    ▼aSchool  code:  0078.
■650  4▼aAnxiety
■650  4▼aAdaptability
■650  4▼aHumanitarianism
■650  4▼aAdaptation
■650  4▼aViolence
■650  4▼aCommunity
■650  4▼aLikert  scale
■650  4▼aKeywords
■650  4▼aImage  retrieval
■650  4▼aFalse  information
■650  4▼aMultilingualism
■650  4▼aSpeech
■650  4▼aLarge  language  models
■650  4▼aBilingual  education
■650  4▼aClinical  psychology
■650  4▼aWeb  studies
■690    ▼a0800
■690    ▼a0282
■690    ▼a0622
■690    ▼a0646
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-05A.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360456▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17299 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.