서브메뉴
검색
Energy-Efficient On-Chip Deep Neural Network (DNN) Inference and Training with Emerging Non-Volatile Memory Technologies
Energy-Efficient On-Chip Deep Neural Network (DNN) Inference and Training with Emerging Non-Volatile Memory Technologies
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105551
- ISBN
- 9798263395377
- DDC
- 741
- 저자명
- Luo, Yandong.
- 서명/저자
- Energy-Efficient On-Chip Deep Neural Network (DNN) Inference and Training with Emerging Non-Volatile Memory Technologies
- 발행사항
- [Sl] : Georgia Institute of Technology, 2023
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2023
- 형태사항
- 159 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: A.
- 주기사항
- Advisor: Yu, Shimeng.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2023.
- 초록/해제
- 요약Artificial intelligence (AI) based applications are becoming pervasive in our daily life. The emerging non-volatile memory (eNVM) technologies are regarded as promising technological candidates for building energy-efficient AI hardware for edge devices. This thesis identifies and resolves several challenges related to AI hardware design for deep neural network (DNN) training and inference using eNVM-based technologies. For DNN inference, the existing compute-in-memory (CIM)-based AI accelerator suffers from area scalability and lacks reconfigurability for different DNN models. Besides, for large DNN models such as transformers, it is challenging to store the model parameters on-chip and execute various types of matrix multiplication workloads efficiently. For DNN training, the existing memory technologies, such as static random access memory (SRAM) and embedded dynamic random access memory (eDRAM), are not satisfactory in buffering a large amount of training data. Besides, the performance of eNVM-based in-memory training is limited by the high write energy and low write endurance of eNVM devices.In this thesis, I explored and presented cross-layer solutions to the abovementioned challenges involving device technology selection, circuit/architecture design, and neural network model optimization. For DNN inference, I proposed a monolithic 3D integration scheme based on the back-end-of-line (BEOL) compatible semiconducting oxide transistors. Together with a reconfigurable interconnect design using the ferroelectric field-effect transistor (FeFET) based routing switch, the scalability and reconfigurability of the CIM-based DNN accelerator are improved. A 3D heterogeneous computing platform using stacked FeFET CIM dies, and digital logic die is demonstrated for transformer models. It leverages the high memory density to store model parameters and the heterogeneity of computing paradigms to optimize the hardware performance for different matrix multiplication workloads.For DNN training, I designed an innovative dual-mode memory architecture based on ferroelectric random access memory (FeRAM). It demonstrates excellent dynamic and static performance for the DNN training accelerators by optimally switching between the volatile and non-volatile operation modes. To further improve the energy efficiency of DNN training, I proposed an in-memory training architecture using a novel hybrid weight cell design. It overcomes the high write energy and low write endurance of eNVM devices.The works presented in this thesis can be a significant milestone toward energy-efficient AI hardware design with eNVM technologies, which potentially allow AI applications to deploy on edge devices with power and area budget constraints.
- 일반주제명
- Design
- 일반주제명
- Energy efficiency
- 일반주제명
- Energy consumption
- 일반주제명
- Sustainability
- 기본자료저록
- Dissertations Abstracts International. 87-05A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2023 us c eng d■001000017360590
■00520260202105551
■006m o d
■007cr#unu||||||||
■020 ▼a9798263395377
■035 ▼a(MiAaPQ)AAI32315737
■035 ▼a(MiAaPQ)GeorgiaTech75066
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a741
■1001 ▼aLuo, Yandong.
■24510▼aEnergy-Efficient On-Chip Deep Neural Network (DNN) Inference and Training with Emerging Non-Volatile Memory Technologies
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2023
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2023
■300 ▼a159 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: A.
■500 ▼aAdvisor: Yu, Shimeng.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2023.
■520 ▼aArtificial intelligence (AI) based applications are becoming pervasive in our daily life. The emerging non-volatile memory (eNVM) technologies are regarded as promising technological candidates for building energy-efficient AI hardware for edge devices. This thesis identifies and resolves several challenges related to AI hardware design for deep neural network (DNN) training and inference using eNVM-based technologies. For DNN inference, the existing compute-in-memory (CIM)-based AI accelerator suffers from area scalability and lacks reconfigurability for different DNN models. Besides, for large DNN models such as transformers, it is challenging to store the model parameters on-chip and execute various types of matrix multiplication workloads efficiently. For DNN training, the existing memory technologies, such as static random access memory (SRAM) and embedded dynamic random access memory (eDRAM), are not satisfactory in buffering a large amount of training data. Besides, the performance of eNVM-based in-memory training is limited by the high write energy and low write endurance of eNVM devices.In this thesis, I explored and presented cross-layer solutions to the abovementioned challenges involving device technology selection, circuit/architecture design, and neural network model optimization. For DNN inference, I proposed a monolithic 3D integration scheme based on the back-end-of-line (BEOL) compatible semiconducting oxide transistors. Together with a reconfigurable interconnect design using the ferroelectric field-effect transistor (FeFET) based routing switch, the scalability and reconfigurability of the CIM-based DNN accelerator are improved. A 3D heterogeneous computing platform using stacked FeFET CIM dies, and digital logic die is demonstrated for transformer models. It leverages the high memory density to store model parameters and the heterogeneity of computing paradigms to optimize the hardware performance for different matrix multiplication workloads.For DNN training, I designed an innovative dual-mode memory architecture based on ferroelectric random access memory (FeRAM). It demonstrates excellent dynamic and static performance for the DNN training accelerators by optimally switching between the volatile and non-volatile operation modes. To further improve the energy efficiency of DNN training, I proposed an in-memory training architecture using a novel hybrid weight cell design. It overcomes the high write energy and low write endurance of eNVM devices.The works presented in this thesis can be a significant milestone toward energy-efficient AI hardware design with eNVM technologies, which potentially allow AI applications to deploy on edge devices with power and area budget constraints.
■590 ▼aSchool code: 0078.
■650 4▼aDesign
■650 4▼aEnergy efficiency
■650 4▼aNatural language processing
■650 4▼aEnergy consumption
■650 4▼aSustainability
■690 ▼a0389
■690 ▼a0800
■690 ▼a0640
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05A.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360590▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


