본문

서브메뉴

Uncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models
Uncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models
Uncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260202105532
ISBN  
9798263345709
DDC  
519.2
저자명  
Li, Yinghao.
서명/저자  
Uncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models
발행사항  
[Sl] : Georgia Institute of Technology, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
196 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
주기사항  
Advisor: Zhang, Chao.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2025.
초록/해제  
요약Pre-trained foundation models have become indispensable in modern Natural Language Processing (NLP) and scientific domains, evolving into Large Language Models (LLMs) with impressive zero- and few-shot capabilities. Despite their widespread success, challenges persist in real-world applications due to distribution shifts between training and inference data, opaque inference processes, and limited in-domain manually labeled examples. These issues complicate reliable confidence estimation and application of foundation models in downstream tasks. To address these concerns, this thesis explores two primary research directions: 1) reliable Uncertainty Quantification (UQ) and 2) data-efficient model learning.In addressing reliability, we investigate different UQ methods and develop novel techniques to enhance model calibration. Molecular Uncertainty Benchmark (MUBen) establishes a best-practice benchmark for UQ in molecular representation models, thoroughly evaluating uncertainty calibration and predictive accuracy in large-scale discriminative tasks. Expanding uncertainty estimation techniques to autoregressive LLMs, we introduce Uncertainty Quantification with Attention Chain (UQAC), an approach that employs iterative attention-chain backtracking to approximate an otherwise intractable marginalization over Chain of Thought (CoT) reasoning paths, thus enhancing confidence estimation robustness for LLMs.Regarding data efficiency, our work targets scenarios characterized by limited or noisy labeled data. In zero-shot Named Entity Recognition (NER), we develop Conditional Hidden Markov Model (CHMM) and Sparse Conditional Hidden Markov Model (SparseCHMM), which effectively exploit weak supervision signals through contextual embeddings from autoencoding foundation models, employing sparsity regularization to improve robustness. Additionally, we propose Generate and Organize (G&O), a zero-shot Information Extraction (IE) framework leveraging the powerful reasoning abilities of autore gressive LLMs. Lastly, we introduce Ensembles of Low-Rank Expert Adapters (ELREA), designed for date-efficient multi-task fine-tuning, which clusters training instructions based on gradient directions and applies task-specific Low-Rank Adaptation (LoRA) experts through ensemble techniques. ELREA mitigates task interference, promoting better generalization and parameter efficiency.Together, our proposed methods enhance the trustworthiness and adaptability of pretrained models in critical domains by addressing uncertainty concerns and reducing dependency on extensive labeled data. The thesis underscores the importance of calibration, interpretability, and scalable fine-tuning strategies in developing robust, data-efficient solutions suitable for high-stakes real-world applications.
일반주제명  
Probability
일반주제명  
Correlation analysis
일반주제명  
Computer engineering
키워드  
Natural Language Processing
키워드  
Large Language Models
키워드  
Uncertainty Quantification
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-06B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017360472
■00520260202105532
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263345709
■035    ▼a(MiAaPQ)AAI32309875
■035    ▼a(MiAaPQ)GeorgiaTech77890
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a519.2
■1001  ▼aLi,  Yinghao.
■24510▼aUncertainty-Aware  and  Data-Efficient  Fine-Tuning  and  Application  of  Foundation  Models
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a196  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-06,  Section:  B.
■500    ▼aAdvisor:  Zhang,  Chao.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2025.
■520    ▼aPre-trained  foundation  models  have  become  indispensable  in  modern  Natural  Language  Processing  (NLP)  and  scientific  domains,  evolving  into  Large  Language  Models  (LLMs)  with  impressive  zero-  and  few-shot  capabilities.  Despite  their  widespread  success,  challenges  persist  in  real-world  applications  due  to  distribution  shifts  between  training  and  inference  data,  opaque  inference  processes,  and  limited  in-domain  manually  labeled  examples.  These  issues  complicate  reliable  confidence  estimation  and  application  of  foundation  models  in  downstream  tasks.  To  address  these  concerns,  this  thesis  explores  two  primary  research  directions:  1)  reliable  Uncertainty  Quantification  (UQ)  and  2)  data-efficient  model  learning.In  addressing  reliability,  we  investigate  different  UQ  methods  and  develop  novel  techniques  to  enhance  model  calibration.  Molecular  Uncertainty  Benchmark  (MUBen)  establishes  a  best-practice  benchmark  for  UQ  in  molecular  representation  models,  thoroughly  evaluating  uncertainty  calibration  and  predictive  accuracy  in  large-scale  discriminative  tasks.  Expanding  uncertainty  estimation  techniques  to  autoregressive  LLMs,  we  introduce  Uncertainty  Quantification  with  Attention  Chain  (UQAC),  an  approach  that  employs  iterative  attention-chain  backtracking  to  approximate  an  otherwise  intractable  marginalization  over  Chain  of  Thought  (CoT)  reasoning  paths,  thus  enhancing  confidence  estimation  robustness  for  LLMs.Regarding  data  efficiency,  our  work  targets  scenarios  characterized  by  limited  or  noisy  labeled  data.  In  zero-shot  Named  Entity  Recognition  (NER),  we  develop  Conditional  Hidden  Markov  Model  (CHMM)  and  Sparse  Conditional  Hidden  Markov  Model  (SparseCHMM),  which  effectively  exploit  weak  supervision  signals  through  contextual  embeddings  from  autoencoding  foundation  models,  employing  sparsity  regularization  to  improve  robustness.  Additionally,  we  propose  Generate  and  Organize  (G&O),  a  zero-shot  Information  Extraction  (IE)  framework  leveraging  the  powerful  reasoning  abilities  of  autore  gressive  LLMs.  Lastly,  we  introduce  Ensembles  of  Low-Rank  Expert  Adapters  (ELREA),  designed  for  date-efficient  multi-task  fine-tuning,  which  clusters  training  instructions  based  on  gradient  directions  and  applies  task-specific  Low-Rank  Adaptation  (LoRA)  experts  through  ensemble  techniques.  ELREA  mitigates  task  interference,  promoting  better  generalization  and  parameter  efficiency.Together,  our  proposed  methods  enhance  the  trustworthiness  and  adaptability  of  pretrained  models  in  critical  domains  by  addressing  uncertainty  concerns  and  reducing  dependency  on  extensive  labeled  data.  The  thesis  underscores  the  importance  of  calibration,  interpretability,  and  scalable  fine-tuning  strategies  in  developing  robust,  data-efficient  solutions  suitable  for  high-stakes  real-world  applications.
■590    ▼aSchool  code:  0078.
■650  4▼aProbability
■650  4▼aCorrelation  analysis
■650  4▼aComputer  engineering
■653    ▼aNatural  Language  Processing
■653    ▼aLarge  Language  Models
■653    ▼aUncertainty  Quantification
■690    ▼a0464
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-06B.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360472▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Buch Status

    • Reservierung
    • frei buchen
    • Meine Mappe
    • Erste Aufräumarbeiten Anfrage
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    Sammlungen
    Registrierungsnummer callnumber Standort Verkehr Status Verkehr Info
    TF17315 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Kredite nur für Ihre Daten gebucht werden. Wenn Sie buchen möchten Reservierungen, klicken Sie auf den Button.

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.