서브메뉴
검색
Uncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models
Uncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105532
- ISBN
- 9798263345709
- DDC
- 519.2
- 저자명
- Li, Yinghao.
- 서명/저자
- Uncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models
- 발행사항
- [Sl] : Georgia Institute of Technology, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 196 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
- 주기사항
- Advisor: Zhang, Chao.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2025.
- 초록/해제
- 요약Pre-trained foundation models have become indispensable in modern Natural Language Processing (NLP) and scientific domains, evolving into Large Language Models (LLMs) with impressive zero- and few-shot capabilities. Despite their widespread success, challenges persist in real-world applications due to distribution shifts between training and inference data, opaque inference processes, and limited in-domain manually labeled examples. These issues complicate reliable confidence estimation and application of foundation models in downstream tasks. To address these concerns, this thesis explores two primary research directions: 1) reliable Uncertainty Quantification (UQ) and 2) data-efficient model learning.In addressing reliability, we investigate different UQ methods and develop novel techniques to enhance model calibration. Molecular Uncertainty Benchmark (MUBen) establishes a best-practice benchmark for UQ in molecular representation models, thoroughly evaluating uncertainty calibration and predictive accuracy in large-scale discriminative tasks. Expanding uncertainty estimation techniques to autoregressive LLMs, we introduce Uncertainty Quantification with Attention Chain (UQAC), an approach that employs iterative attention-chain backtracking to approximate an otherwise intractable marginalization over Chain of Thought (CoT) reasoning paths, thus enhancing confidence estimation robustness for LLMs.Regarding data efficiency, our work targets scenarios characterized by limited or noisy labeled data. In zero-shot Named Entity Recognition (NER), we develop Conditional Hidden Markov Model (CHMM) and Sparse Conditional Hidden Markov Model (SparseCHMM), which effectively exploit weak supervision signals through contextual embeddings from autoencoding foundation models, employing sparsity regularization to improve robustness. Additionally, we propose Generate and Organize (G&O), a zero-shot Information Extraction (IE) framework leveraging the powerful reasoning abilities of autore gressive LLMs. Lastly, we introduce Ensembles of Low-Rank Expert Adapters (ELREA), designed for date-efficient multi-task fine-tuning, which clusters training instructions based on gradient directions and applies task-specific Low-Rank Adaptation (LoRA) experts through ensemble techniques. ELREA mitigates task interference, promoting better generalization and parameter efficiency.Together, our proposed methods enhance the trustworthiness and adaptability of pretrained models in critical domains by addressing uncertainty concerns and reducing dependency on extensive labeled data. The thesis underscores the importance of calibration, interpretability, and scalable fine-tuning strategies in developing robust, data-efficient solutions suitable for high-stakes real-world applications.
- 일반주제명
- Probability
- 일반주제명
- Correlation analysis
- 일반주제명
- Computer engineering
- 기본자료저록
- Dissertations Abstracts International. 87-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017360472
■00520260202105532
■006m o d
■007cr#unu||||||||
■020 ▼a9798263345709
■035 ▼a(MiAaPQ)AAI32309875
■035 ▼a(MiAaPQ)GeorgiaTech77890
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a519.2
■1001 ▼aLi, Yinghao.
■24510▼aUncertainty-Aware and Data-Efficient Fine-Tuning and Application of Foundation Models
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a196 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: B.
■500 ▼aAdvisor: Zhang, Chao.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2025.
■520 ▼aPre-trained foundation models have become indispensable in modern Natural Language Processing (NLP) and scientific domains, evolving into Large Language Models (LLMs) with impressive zero- and few-shot capabilities. Despite their widespread success, challenges persist in real-world applications due to distribution shifts between training and inference data, opaque inference processes, and limited in-domain manually labeled examples. These issues complicate reliable confidence estimation and application of foundation models in downstream tasks. To address these concerns, this thesis explores two primary research directions: 1) reliable Uncertainty Quantification (UQ) and 2) data-efficient model learning.In addressing reliability, we investigate different UQ methods and develop novel techniques to enhance model calibration. Molecular Uncertainty Benchmark (MUBen) establishes a best-practice benchmark for UQ in molecular representation models, thoroughly evaluating uncertainty calibration and predictive accuracy in large-scale discriminative tasks. Expanding uncertainty estimation techniques to autoregressive LLMs, we introduce Uncertainty Quantification with Attention Chain (UQAC), an approach that employs iterative attention-chain backtracking to approximate an otherwise intractable marginalization over Chain of Thought (CoT) reasoning paths, thus enhancing confidence estimation robustness for LLMs.Regarding data efficiency, our work targets scenarios characterized by limited or noisy labeled data. In zero-shot Named Entity Recognition (NER), we develop Conditional Hidden Markov Model (CHMM) and Sparse Conditional Hidden Markov Model (SparseCHMM), which effectively exploit weak supervision signals through contextual embeddings from autoencoding foundation models, employing sparsity regularization to improve robustness. Additionally, we propose Generate and Organize (G&O), a zero-shot Information Extraction (IE) framework leveraging the powerful reasoning abilities of autore gressive LLMs. Lastly, we introduce Ensembles of Low-Rank Expert Adapters (ELREA), designed for date-efficient multi-task fine-tuning, which clusters training instructions based on gradient directions and applies task-specific Low-Rank Adaptation (LoRA) experts through ensemble techniques. ELREA mitigates task interference, promoting better generalization and parameter efficiency.Together, our proposed methods enhance the trustworthiness and adaptability of pretrained models in critical domains by addressing uncertainty concerns and reducing dependency on extensive labeled data. The thesis underscores the importance of calibration, interpretability, and scalable fine-tuning strategies in developing robust, data-efficient solutions suitable for high-stakes real-world applications.
■590 ▼aSchool code: 0078.
■650 4▼aProbability
■650 4▼aCorrelation analysis
■650 4▼aComputer engineering
■653 ▼aNatural Language Processing
■653 ▼aLarge Language Models
■653 ▼aUncertainty Quantification
■690 ▼a0464
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-06B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360472▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


