본문

서브메뉴

Contextures: The Mechanism of Representation Learning
Contextures: The Mechanism of Representation Learning
Contextures: The Mechanism of Representation Learning

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260202103130
ISBN  
9798286448548
DDC  
004
저자명  
Zhai, Runtian.
서명/저자  
Contextures: The Mechanism of Representation Learning
발행사항  
[Sl] : Carnegie Mellon University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
166 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
주기사항  
Advisor: Ravikumar, Pradeep;Kolter, Zico.
학위논문주기  
Thesis (Ph.D.)--Carnegie Mellon University, 2025.
초록/해제  
요약This dissertation establishes the contexture theory to mathematically characterize the mechanism of representation learning, also known as pretraining. Despite the remarkable empirical success of foundation models, it is not very clear what representations they learn, and why these representations are useful for various disparate downstream tasks. A scientific understanding of representation learning is critical, especially at this point when scaling up the model size is producing diminishing returns, and designing new pretraining methods is imperative for further progress.Prior work treated different representation learning methods quite differently, whereas the contexture theory provides a unified framework for delineating the representations these methods learn. The central argument is that a representation is learned from the association between the input X and a context variable A. We prove that if an encoder captures the maximum information of this association, in which case we say that the encoder learns the contexture, then it will be optimal on the class of tasks that are compatible with the context. We also show that a context is the most useful when the association between X and A is neither too strong nor too weak. The important implication of the contexture theory is that increasing the model size alone will achieve diminishing returns, and further advancements require better contexts.We demonstrate that lots of existing pretraining objectives can learn the contexture, including supervised learning, self-supervised learning, generative models, etc. Based on that, we introduce two general objectives-SVME and KISE, for learning the contexture. We also show how to mix multiple contexts together, which is an effortless way to create better contexts from existing ones. Then, we prove statistical learning bounds for representation learning, and extend the framework to spectrally transformed kernel regression for semi-supervised learning. Finally, we discuss the effect of the data distribution shift from pretraining to the downstream task.
일반주제명  
Computer science
일반주제명  
Computer engineering
키워드  
Foundation models
키워드  
Learning theory
키워드  
Machine learning
키워드  
Representation learning
키워드  
Downstream task
기타저자  
Carnegie Mellon University Computer Science
기본자료저록  
Dissertations Abstracts International. 87-01B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357097
■00520260202103130
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798286448548
■035    ▼a(MiAaPQ)AAI31939492
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aZhai,  Runtian.▼0(orcid)0000-0003-3332-3466
■24510▼aContextures:  The  Mechanism  of  Representation  Learning
■260    ▼a[Sl]▼bCarnegie  Mellon  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a166  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  B.
■500    ▼aAdvisor:  Ravikumar,  Pradeep;Kolter,  Zico.
■5021  ▼aThesis  (Ph.D.)--Carnegie  Mellon  University,  2025.
■520    ▼aThis  dissertation  establishes  the  contexture  theory  to  mathematically  characterize  the  mechanism  of  representation  learning,  also  known  as  pretraining.  Despite  the  remarkable  empirical  success  of  foundation  models,  it  is  not  very  clear  what  representations  they  learn,  and  why  these  representations  are  useful  for  various  disparate  downstream  tasks.  A  scientific  understanding  of  representation  learning  is  critical,  especially  at  this  point  when  scaling  up  the  model  size  is  producing  diminishing  returns,  and  designing  new  pretraining  methods  is  imperative  for  further  progress.Prior  work  treated  different  representation  learning  methods  quite  differently,  whereas  the  contexture  theory  provides  a  unified  framework  for  delineating  the  representations  these  methods  learn.  The  central  argument  is  that  a  representation  is  learned  from  the  association  between  the  input  X  and  a  context  variable  A.  We  prove  that  if  an  encoder  captures  the  maximum  information  of  this  association,  in  which  case  we  say  that  the  encoder  learns  the  contexture,  then  it  will  be  optimal  on  the  class  of  tasks  that  are  compatible  with  the  context.  We  also  show  that  a  context  is  the  most  useful  when  the  association  between  X  and  A  is  neither  too  strong  nor  too  weak.  The  important  implication  of  the  contexture  theory  is  that  increasing  the  model  size  alone  will  achieve  diminishing  returns,  and  further  advancements  require  better  contexts.We  demonstrate  that  lots  of  existing  pretraining  objectives  can  learn  the  contexture,  including  supervised  learning,  self-supervised  learning,  generative  models,  etc.  Based  on  that,  we  introduce  two  general  objectives-SVME  and  KISE,  for  learning  the  contexture.  We  also  show  how  to  mix  multiple  contexts  together,  which  is  an  effortless  way  to  create  better  contexts  from  existing  ones.  Then,  we  prove  statistical  learning  bounds  for  representation  learning,  and  extend  the  framework  to  spectrally  transformed  kernel  regression  for  semi-supervised  learning.  Finally,  we  discuss  the  effect  of  the  data  distribution  shift  from  pretraining  to  the  downstream  task.
■590    ▼aSchool  code:  0041.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■653    ▼aFoundation  models
■653    ▼aLearning  theory
■653    ▼aMachine  learning
■653    ▼aRepresentation  learning
■653    ▼aDownstream  task
■690    ▼a0984
■690    ▼a0800
■690    ▼a0464
■71020▼aCarnegie  Mellon  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g87-01B.
■790    ▼a0041
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357097▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    פרט מידע

    • הזמנה
    • לא קיים
    • התיקיה שלי
    • צפה הראשון בקשה
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    גשמי
    Reg No. Call No. מיקום מצב להשאיל מידע
    TF16588 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * הזמנות זמינים בספר ההשאלה. כדי להזמין, נא לחץ על כפתור ההזמנה

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.