서브메뉴
검색
Understanding Generalization in Deep Learning Through Occam's Razor
Understanding Generalization in Deep Learning Through Occam's Razor
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103053
- ISBN
- 9798286424351
- DDC
- 519
- 저자명
- Lotfi, Sanae.
- 서명/저자
- Understanding Generalization in Deep Learning Through Occams Razor
- 발행사항
- [Sl] : New York University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 260 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
- 주기사항
- Advisor: Gordon Wilson, Andrew.
- 학위논문주기
- Thesis (Ph.D.)--New York University, 2025.
- 초록/해제
- 요약Gaining insight into the mechanisms behind the generalization of deep learning models is crucial to build on their strengths, address their limitations, and deploy them in safety-critical applications. As state-of-the-art models for various data modalities become increasingly large and are trained on internet-scale data, the notion of generalization becomes more challenging to define. In this thesis, I study generalization through the lens of Occam's razor: among models that can fit the training data, the simplest is most likely to perform well on unseen data. Compression bounds provide a principled way to capture this intuition through a trade-off between the model's training performance and its compressed size.First, I present our work on deriving state-of-the-art generalization bounds for image classification models, providing key insights into why these models generalize effectively in practice. I then explore the challenges of extending these bounds to pretrained large language models (LLMs), establishing the first non-vacuous bounds for LLMs. Our findings reveal that larger LLMs not only yield better bounds but also find simpler representations of the data. Furthermore, we demonstrate that LLMs retain their understanding of patterns but forget highly unstructured data more rapidly as we compress them more aggressively. Finally, I discuss the connection between generalization bounds and the marginal likelihood, a Bayesian tool used for model selection and hyperparameter tuning. Specifically, I demonstrate how generalization bounds can predict practical issues like overfitting and underfitting when using the marginal likelihood for model selection, and propose a remedy that is more aligned with generalization.
- 일반주제명
- Applied mathematics
- 일반주제명
- Statistics
- 일반주제명
- Computer science
- 일반주제명
- Information science
- 키워드
- Occam's razor
- 키워드
- Generalization
- 키워드
- Deep learning
- 기타저자
- New York University Center for Data Science
- 기본자료저록
- Dissertations Abstracts International. 86-12A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356871
■00520260202103053
■006m o d
■007cr#unu||||||||
■020 ▼a9798286424351
■035 ▼a(MiAaPQ)AAI31932110
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a519
■1001 ▼aLotfi, Sanae.
■24510▼aUnderstanding Generalization in Deep Learning Through Occam's Razor
■260 ▼a[Sl]▼bNew York University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a260 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: A.
■500 ▼aAdvisor: Gordon Wilson, Andrew.
■5021 ▼aThesis (Ph.D.)--New York University, 2025.
■520 ▼aGaining insight into the mechanisms behind the generalization of deep learning models is crucial to build on their strengths, address their limitations, and deploy them in safety-critical applications. As state-of-the-art models for various data modalities become increasingly large and are trained on internet-scale data, the notion of generalization becomes more challenging to define. In this thesis, I study generalization through the lens of Occam's razor: among models that can fit the training data, the simplest is most likely to perform well on unseen data. Compression bounds provide a principled way to capture this intuition through a trade-off between the model's training performance and its compressed size.First, I present our work on deriving state-of-the-art generalization bounds for image classification models, providing key insights into why these models generalize effectively in practice. I then explore the challenges of extending these bounds to pretrained large language models (LLMs), establishing the first non-vacuous bounds for LLMs. Our findings reveal that larger LLMs not only yield better bounds but also find simpler representations of the data. Furthermore, we demonstrate that LLMs retain their understanding of patterns but forget highly unstructured data more rapidly as we compress them more aggressively. Finally, I discuss the connection between generalization bounds and the marginal likelihood, a Bayesian tool used for model selection and hyperparameter tuning. Specifically, I demonstrate how generalization bounds can predict practical issues like overfitting and underfitting when using the marginal likelihood for model selection, and propose a remedy that is more aligned with generalization.
■590 ▼aSchool code: 0146.
■650 4▼aApplied mathematics
■650 4▼aStatistics
■650 4▼aComputer science
■650 4▼aInformation science
■653 ▼aOccam's razor
■653 ▼aGeneralization
■653 ▼aDeep learning
■653 ▼aLarge language models
■653 ▼aCompression bounds
■690 ▼a0364
■690 ▼a0984
■690 ▼a0723
■690 ▼a0800
■690 ▼a0463
■71020▼aNew York University▼bCenter for Data Science.
■7730 ▼tDissertations Abstracts International▼g86-12A.
■790 ▼a0146
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356871▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


