서브메뉴
검색
PDE Methods for Deep Learning Analysis and Optimization
PDE Methods for Deep Learning Analysis and Optimization
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105602
- ISBN
- 9798263395636
- DDC
- 541.34
- 저자명
- Sun, Yuxin.
- 서명/저자
- PDE Methods for Deep Learning Analysis and Optimization
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 122 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Yezzi, Anthony J.;Sundaramoorthi, Ganesh.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약The objective of this dissertation is to use tools from partial differential equations (PDEs) to understand and construct deep learning algorithms. This research includes the theory-inspired design of optimization algorithms for deep learning, theoretical analysis of deep network training, and applications in computer vision. We introduce a recently developed framework (PDE acceleration), which is a variational approach to accelerated optimization with PDEs, in the context of optimization of deep networks, which leads to a novel and simple extension of stochastic gradient descent (SGD) with momentum. We empirically validate the theory and evaluate our new algorithm on image classification showing empirical improvement over SGD. To further enhance the performance of deep learning algorithms, we need a better understanding of the stability and convergence properties. We discovered restrained numerical instabilities in current training practices of deep networks. To explain this phenomenon, we present a theoretical framework using numerical analysis of PDE and analyzing the gradient descent PDE of a simplified convolutional neural network (CNN). We also link restrained instabilities to the recently discovered Edge of Stability (EoS) phenomena and provide new insights and predictions about the EoS. Further, the special potential of "geometric" PDEs in particular to advance deep learning applications is explored in this dissertation. Under the geometric PDE's framework, we provide a theoretical analysis to understand the instability caused by the Eikonal loss and explain how some existing approaches can unknowingly mitigate this instability. Furthermore, those regularization enables the use of new neural networks with higher representation power that can capture finer scale details of shape. In summary, we believe the tools we've introduced could improve deep learning practice.
- 일반주제명
- Diffusion
- 일반주제명
- Deep learning
- 일반주제명
- Computer vision
- 일반주제명
- Neural networks
- 일반주제명
- Computer science
- 일반주제명
- Mathematics
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360658
■00520260202105602
■006m o d
■007cr#unu||||||||
■020 ▼a9798263395636
■035 ▼a(MiAaPQ)AAI32315993
■035 ▼a(MiAaPQ)GeorgiaTech76924
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a541.34
■1001 ▼aSun, Yuxin.
■24510▼aPDE Methods for Deep Learning Analysis and Optimization
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a122 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Yezzi, Anthony J.;Sundaramoorthi, Ganesh.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aThe objective of this dissertation is to use tools from partial differential equations (PDEs) to understand and construct deep learning algorithms. This research includes the theory-inspired design of optimization algorithms for deep learning, theoretical analysis of deep network training, and applications in computer vision. We introduce a recently developed framework (PDE acceleration), which is a variational approach to accelerated optimization with PDEs, in the context of optimization of deep networks, which leads to a novel and simple extension of stochastic gradient descent (SGD) with momentum. We empirically validate the theory and evaluate our new algorithm on image classification showing empirical improvement over SGD. To further enhance the performance of deep learning algorithms, we need a better understanding of the stability and convergence properties. We discovered restrained numerical instabilities in current training practices of deep networks. To explain this phenomenon, we present a theoretical framework using numerical analysis of PDE and analyzing the gradient descent PDE of a simplified convolutional neural network (CNN). We also link restrained instabilities to the recently discovered Edge of Stability (EoS) phenomena and provide new insights and predictions about the EoS. Further, the special potential of "geometric" PDEs in particular to advance deep learning applications is explored in this dissertation. Under the geometric PDE's framework, we provide a theoretical analysis to understand the instability caused by the Eikonal loss and explain how some existing approaches can unknowingly mitigate this instability. Furthermore, those regularization enables the use of new neural networks with higher representation power that can capture finer scale details of shape. In summary, we believe the tools we've introduced could improve deep learning practice.
■590 ▼aSchool code: 0078.
■650 4▼aDiffusion
■650 4▼aPartial differential equations
■650 4▼aDeep learning
■650 4▼aComputer vision
■650 4▼aOrdinary differential equations
■650 4▼aNeural networks
■650 4▼aComputer science
■650 4▼aMathematics
■690 ▼a0984
■690 ▼a0800
■690 ▼a0405
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360658▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


