서브메뉴
검색
Compiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Optimization
Compiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Optimization
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153025
- ISBN
- 9798346759195
- DDC
- 621.3
- 저자명
- Li, Yi.
- 서명/저자
- Compiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Optimization
- 발행사항
- [Sl] : Princeton University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 157 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-06, Section: B.
- 주기사항
- Advisor: Malik, Sharad.
- 학위논문주기
- Thesis (Ph.D.)--Princeton University, 2024.
- 초록/해제
- 요약Specialized hardware accelerators have been developed to enhance power-performance efficiency for Deep Neural Network (DNN) applications. A primary challenge in DNN accelerator development is the early-stage evaluation of design prototypes on real-world applications. Such evaluations are crucial: modern DNN accelerators are equipped with several techniques to boost power-performance, but these techniques can introduce numerical discrepancies, such as data quantization with customized numerical representation or reformulated operators. Given the deeply-connected layered nature of DNN applications, these numerical errors can accumulate and result in significant deviations from reference results. Additionally, the energy and performance costs of data movement between host machine and the accelerator's on-chip memory are substantial, making the reduction of the data transfer a critical optimization focus for mapping DNN applications to accelerators.To address these challenges, this thesis proposes several innovative solutions. First, we introduce "3LA" - an end-to-end compiler pipeline that facilitates application-level testing of hardware accelerator prototypes on unmodified DNN applications. Built upon a recently proposed formal hardware specification named Instruction-Level Abstraction (ILA), 3LA allows for automated application-level simulation, providing crucial development feedback with much reduced manual engineering effort.Second, we proposed Shoehorn, an optimized scheduler designed for mapping DNN operators to hardware accelerators that co-optimizes loop tiling, loop ordering and on-chip memory partitioning decisions. This scheduler creates an optimal mapping schedule for single application-level operators to a specific accelerator, minimizing off-chip memory access.Lastly, this thesis introduces "COSMA," an optimization framework that aims to minimize total off-chip data access when deploying entire or segments of DNN applications to the target accelerator. COSMA collectively optimizes operator scheduling, memory allocation and tensor replacement strategies, presenting a comprehensive solution to data movement minimization.These contributions are expected to significantly streamline the process of DNN accelerator development, from early-stage design to final application deployment, enhancing both efficiency and effectiveness in the field.
- 일반주제명
- Electrical engineering
- 일반주제명
- Computer engineering
- 일반주제명
- Computer science
- 키워드
- Deep learning
- 키워드
- On-chip memory
- 기타저자
- Princeton University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164634
■00520250211153025
■006m o d
■007cr#unu||||||||
■020 ▼a9798346759195
■035 ▼a(MiAaPQ)AAI31633675
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aLi, Yi.▼0(orcid)0009-0000-4837-2282
■24510▼aCompiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Optimization
■260 ▼a[Sl]▼bPrinceton University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a157 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-06, Section: B.
■500 ▼aAdvisor: Malik, Sharad.
■5021 ▼aThesis (Ph.D.)--Princeton University, 2024.
■520 ▼aSpecialized hardware accelerators have been developed to enhance power-performance efficiency for Deep Neural Network (DNN) applications. A primary challenge in DNN accelerator development is the early-stage evaluation of design prototypes on real-world applications. Such evaluations are crucial: modern DNN accelerators are equipped with several techniques to boost power-performance, but these techniques can introduce numerical discrepancies, such as data quantization with customized numerical representation or reformulated operators. Given the deeply-connected layered nature of DNN applications, these numerical errors can accumulate and result in significant deviations from reference results. Additionally, the energy and performance costs of data movement between host machine and the accelerator's on-chip memory are substantial, making the reduction of the data transfer a critical optimization focus for mapping DNN applications to accelerators.To address these challenges, this thesis proposes several innovative solutions. First, we introduce "3LA" - an end-to-end compiler pipeline that facilitates application-level testing of hardware accelerator prototypes on unmodified DNN applications. Built upon a recently proposed formal hardware specification named Instruction-Level Abstraction (ILA), 3LA allows for automated application-level simulation, providing crucial development feedback with much reduced manual engineering effort.Second, we proposed Shoehorn, an optimized scheduler designed for mapping DNN operators to hardware accelerators that co-optimizes loop tiling, loop ordering and on-chip memory partitioning decisions. This scheduler creates an optimal mapping schedule for single application-level operators to a specific accelerator, minimizing off-chip memory access.Lastly, this thesis introduces "COSMA," an optimization framework that aims to minimize total off-chip data access when deploying entire or segments of DNN applications to the target accelerator. COSMA collectively optimizes operator scheduling, memory allocation and tensor replacement strategies, presenting a comprehensive solution to data movement minimization.These contributions are expected to significantly streamline the process of DNN accelerator development, from early-stage design to final application deployment, enhancing both efficiency and effectiveness in the field.
■590 ▼aSchool code: 0181.
■650 4▼aElectrical engineering
■650 4▼aComputer engineering
■650 4▼aComputer science
■653 ▼aDeep learning
■653 ▼aDeep Neural Network
■653 ▼aInstruction-Level Abstraction
■653 ▼aOn-chip memory
■653 ▼aHardware accelerators
■690 ▼a0544
■690 ▼a0464
■690 ▼a0984
■71020▼aPrinceton University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g86-06B.
■790 ▼a0181
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164634▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


