본문

서브메뉴

Compiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Optimization
Compiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Opt...
Compiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Optimization

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211153025
ISBN  
9798346759195
DDC  
621.3
저자명  
Li, Yi.
서명/저자  
Compiler Support for Deep Learning Accelerators: End-to-End Evaluation and Data Access Optimization
발행사항  
[Sl] : Princeton University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
157 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-06, Section: B.
주기사항  
Advisor: Malik, Sharad.
학위논문주기  
Thesis (Ph.D.)--Princeton University, 2024.
초록/해제  
요약Specialized hardware accelerators have been developed to enhance power-performance efficiency for Deep Neural Network (DNN) applications. A primary challenge in DNN accelerator development is the early-stage evaluation of design prototypes on real-world applications. Such evaluations are crucial: modern DNN accelerators are equipped with several techniques to boost power-performance, but these techniques can introduce numerical discrepancies, such as data quantization with customized numerical representation or reformulated operators. Given the deeply-connected layered nature of DNN applications, these numerical errors can accumulate and result in significant deviations from reference results. Additionally, the energy and performance costs of data movement between host machine and the accelerator's on-chip memory are substantial, making the reduction of the data transfer a critical optimization focus for mapping DNN applications to accelerators.To address these challenges, this thesis proposes several innovative solutions. First, we introduce "3LA" - an end-to-end compiler pipeline that facilitates application-level testing of hardware accelerator prototypes on unmodified DNN applications. Built upon a recently proposed formal hardware specification named Instruction-Level Abstraction (ILA), 3LA allows for automated application-level simulation, providing crucial development feedback with much reduced manual engineering effort.Second, we proposed Shoehorn, an optimized scheduler designed for mapping DNN operators to hardware accelerators that co-optimizes loop tiling, loop ordering and on-chip memory partitioning decisions. This scheduler creates an optimal mapping schedule for single application-level operators to a specific accelerator, minimizing off-chip memory access.Lastly, this thesis introduces "COSMA," an optimization framework that aims to minimize total off-chip data access when deploying entire or segments of DNN applications to the target accelerator. COSMA collectively optimizes operator scheduling, memory allocation and tensor replacement strategies, presenting a comprehensive solution to data movement minimization.These contributions are expected to significantly streamline the process of DNN accelerator development, from early-stage design to final application deployment, enhancing both efficiency and effectiveness in the field.
일반주제명  
Electrical engineering
일반주제명  
Computer engineering
일반주제명  
Computer science
키워드  
Deep learning
키워드  
Deep Neural Network
키워드  
Instruction-Level Abstraction
키워드  
On-chip memory
키워드  
Hardware accelerators
기타저자  
Princeton University Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 86-06B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164634
■00520250211153025
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798346759195
■035    ▼a(MiAaPQ)AAI31633675
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aLi,  Yi.▼0(orcid)0009-0000-4837-2282
■24510▼aCompiler  Support  for  Deep  Learning  Accelerators:  End-to-End  Evaluation  and  Data  Access  Optimization
■260    ▼a[Sl]▼bPrinceton  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a157  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-06,  Section:  B.
■500    ▼aAdvisor:  Malik,  Sharad.
■5021  ▼aThesis  (Ph.D.)--Princeton  University,  2024.
■520    ▼aSpecialized  hardware  accelerators  have  been  developed  to  enhance  power-performance  efficiency  for  Deep  Neural  Network  (DNN)  applications.  A  primary  challenge  in  DNN  accelerator  development  is  the  early-stage  evaluation  of  design  prototypes  on  real-world  applications.  Such  evaluations  are  crucial:  modern  DNN  accelerators  are  equipped  with  several  techniques  to  boost  power-performance,  but  these  techniques  can  introduce  numerical  discrepancies,  such  as  data  quantization  with  customized  numerical  representation  or  reformulated  operators.  Given  the  deeply-connected  layered  nature  of  DNN  applications,  these  numerical  errors  can  accumulate  and  result  in  significant  deviations  from  reference  results.  Additionally,  the  energy  and  performance  costs  of  data  movement  between  host  machine  and  the  accelerator's  on-chip  memory  are  substantial,  making  the  reduction  of  the  data  transfer  a  critical  optimization  focus  for  mapping  DNN  applications  to  accelerators.To  address  these  challenges,  this  thesis  proposes  several  innovative  solutions.  First,  we  introduce  "3LA"  -  an  end-to-end  compiler  pipeline  that  facilitates  application-level  testing  of  hardware  accelerator  prototypes  on  unmodified  DNN  applications.  Built  upon  a  recently  proposed  formal  hardware  specification  named  Instruction-Level  Abstraction  (ILA),  3LA  allows  for  automated  application-level  simulation,  providing  crucial  development  feedback  with  much  reduced  manual  engineering  effort.Second,  we  proposed  Shoehorn,  an  optimized  scheduler  designed  for  mapping  DNN  operators  to  hardware  accelerators  that  co-optimizes  loop  tiling,  loop  ordering  and  on-chip  memory  partitioning  decisions.  This  scheduler  creates  an  optimal  mapping  schedule  for  single  application-level  operators  to  a  specific  accelerator,  minimizing  off-chip  memory  access.Lastly,  this  thesis  introduces  "COSMA,"  an  optimization  framework  that  aims  to  minimize  total  off-chip  data  access  when  deploying  entire  or  segments  of  DNN  applications  to  the  target  accelerator.  COSMA  collectively  optimizes  operator  scheduling,  memory  allocation  and  tensor  replacement  strategies,  presenting  a  comprehensive  solution  to  data  movement  minimization.These  contributions  are  expected  to  significantly  streamline  the  process  of  DNN  accelerator  development,  from  early-stage  design  to  final  application  deployment,  enhancing  both  efficiency  and  effectiveness  in  the  field.
■590    ▼aSchool  code:  0181.
■650  4▼aElectrical  engineering
■650  4▼aComputer  engineering
■650  4▼aComputer  science
■653    ▼aDeep  learning
■653    ▼aDeep  Neural  Network
■653    ▼aInstruction-Level  Abstraction
■653    ▼aOn-chip  memory
■653    ▼aHardware  accelerators
■690    ▼a0544
■690    ▼a0464
■690    ▼a0984
■71020▼aPrinceton  University▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-06B.
■790    ▼a0181
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164634▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11761 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.