본문

서브메뉴

Understanding the Mechanism of Pretraining Stabilization Heuristics: A Variance-Oriented Perspective
Understanding the Mechanism of Pretraining Stabilization Heuristics: A Variance-Oriented P...
Understanding the Mechanism of Pretraining Stabilization Heuristics: A Variance-Oriented Perspective

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105703
ISBN  
9798263308032
DDC  
004
저자명  
Liu, Liyuan.
서명/저자  
Understanding the Mechanism of Pretraining Stabilization Heuristics: A Variance-Oriented Perspective
발행사항  
[Sl] : University of Illinois at Urbana-Champaign, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
112 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Han, Jiawei.
학위논문주기  
Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
초록/해제  
요약Language model pretraining has been breaking the glass ceiling for various natural language processing tasks and has been viewed as one of the most significant successes of deep learning, continuously challenging our understanding of learning and cognition. Recently models, including GPT-4 and BART, fueled by an unprecedented scale of computing and data, exhibit unprecedented intelligence, that some even refer to as "sparks of artificial general intelligence". The success of large-scale pretraining hinges on intricate engineering heuristics. While the empirical benefits of these heuristics are evident, their underlying mechanisms remain elusive. This dissertation endeavors to demystify the mathematical principles underlying these pretraining heuristics, aiming to illuminate their mechanisms and potentially guide future algorithm developments. Adopting a variance-oriented perspective, my research rigorously inspects the heuristics that are pivotal to the stability of current pretraining practices, emphasizing learning rate warmup, model initialization, and gradient approximation. In this dissertation, I show that these pretraining stabilization heuristics can be coherently elucidated with a unified framework anchored in variance, a classical metric for stability. First, I analyze the variance of adaptive learning rate and model outputs, revealing that both learning rate warmup and model initialization function as variance modulators. Then, I move to explore the variance-bias tradeoff in the discrete variable gradient approximation, i.e., employing a numerical ODE framework, I unveil the underlying dynamics of the approximation bias, achieving second order precision with minimal computational overhead. Besides theoretical results, empirical verifications are conducted to verify the assumptions and applicability of the recognized principles. Building upon these insights, this dissertation introduces novel techniques designed to advance the pretraining practices, including RAdam for learning rate warmup, Admin for Transformer model initialization, ReinMax and SparseMixer for gradient approximation. Under the guidance of the recognized principles, all proposed methods require minimal trial-and-error configurations, thereby emerging as robust and high-perform tools for pretraining practices for adaptations.
일반주제명  
Computer science
일반주제명  
Applied mathematics
키워드  
Training stability
키워드  
Variance
키워드  
Model initialization
기타저자  
University of Illinois at Urbana-Champaign Computer Science
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017361083
■00520260202105703
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263308032
■035    ▼a(MiAaPQ)AAI32409894
■035    ▼a(MiAaPQ)124230
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aLiu,  Liyuan.
■24510▼aUnderstanding  the  Mechanism  of  Pretraining  Stabilization  Heuristics:  A  Variance-Oriented  Perspective
■260    ▼a[Sl]▼bUniversity  of  Illinois  at  Urbana-Champaign▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a112  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Han,  Jiawei.
■5021  ▼aThesis  (Ph.D.)--University  of  Illinois  at  Urbana-Champaign,  2024.
■520    ▼aLanguage  model  pretraining  has  been  breaking  the  glass  ceiling  for  various  natural  language  processing  tasks  and  has  been  viewed  as  one  of  the  most  significant  successes  of  deep  learning,  continuously  challenging  our  understanding  of  learning  and  cognition.  Recently  models,  including  GPT-4  and  BART,  fueled  by  an  unprecedented  scale  of  computing  and  data,  exhibit  unprecedented  intelligence,  that  some  even  refer  to  as  "sparks  of  artificial  general  intelligence".                          The  success  of  large-scale  pretraining  hinges  on  intricate  engineering  heuristics.  While  the  empirical  benefits  of  these  heuristics  are  evident,  their  underlying  mechanisms  remain  elusive.  This  dissertation  endeavors  to  demystify  the  mathematical  principles  underlying  these  pretraining  heuristics,  aiming  to  illuminate  their  mechanisms  and  potentially  guide  future  algorithm  developments.  Adopting  a  variance-oriented  perspective,  my  research  rigorously  inspects  the  heuristics  that  are  pivotal  to  the  stability  of  current  pretraining  practices,  emphasizing  learning  rate  warmup,  model  initialization,  and  gradient  approximation.                        In  this  dissertation,  I  show  that  these  pretraining  stabilization  heuristics  can  be  coherently  elucidated  with  a  unified  framework  anchored  in  variance,  a  classical  metric  for  stability.  First,  I  analyze  the  variance  of  adaptive  learning  rate  and  model  outputs,  revealing  that  both  learning  rate  warmup  and  model  initialization  function  as  variance  modulators.  Then,  I  move  to  explore  the  variance-bias  tradeoff  in  the  discrete  variable  gradient  approximation,  i.e.,  employing  a  numerical  ODE  framework,  I  unveil  the  underlying  dynamics  of  the  approximation  bias,  achieving  second  order  precision  with  minimal  computational  overhead.  Besides  theoretical  results,  empirical  verifications  are  conducted  to  verify  the  assumptions  and  applicability  of  the  recognized  principles.                          Building  upon  these  insights,  this  dissertation  introduces  novel  techniques  designed  to  advance  the  pretraining  practices,  including  RAdam  for  learning  rate  warmup,  Admin  for  Transformer  model  initialization,  ReinMax  and  SparseMixer  for  gradient  approximation.  Under  the  guidance  of  the  recognized  principles,  all  proposed  methods  require  minimal  trial-and-error  configurations,  thereby  emerging  as  robust  and  high-perform  tools  for  pretraining  practices  for  adaptations.
■590    ▼aSchool  code:  0090.
■650  4▼aComputer  science
■650  4▼aApplied  mathematics
■653    ▼aTraining  stability
■653    ▼aVariance
■653    ▼aModel  initialization
■690    ▼a0984
■690    ▼a0364
■71020▼aUniversity  of  Illinois  at  Urbana-Champaign▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0090
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361083▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF16760 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.