서브메뉴
검색
Building Efficient Tensor Accelerators for Sparse and Irregular Workloads
Building Efficient Tensor Accelerators for Sparse and Irregular Workloads
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105536
- ISBN
- 9798263392710
- DDC
- 330
- 저자명
- Qin, Eric.
- 서명/저자
- Building Efficient Tensor Accelerators for Sparse and Irregular Workloads
- 발행사항
- [Sl] : Georgia Institute of Technology, 2022
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2022
- 형태사항
- 144 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Krishna, Tushar.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2022.
- 초록/해제
- 요약Popular Machine Learning (ML) and High Performance Computing (HPC) workloads contribute to a significant portion of runtime on data centers. Applications include image classification, speech recognition, recommendation systems, social network analysis, robotic problems, chemical process simulations, and so on. Recently due to large computational demands from emerging workloads, there is a surge of custom hardware accelerator development for computing tensor kernels with high performance and energy efficiency. For example, the Google Tensor Processing Unit (TPU) is a custom hardware accelerator targeting efficient matrix multiplications for Deep Neural Networks (DNNs). However, there are limitations with state-of-the-art accelerators, stemming from (1) a vast spectrum of sparsity across various workloads and (2) irregularity of tensor dimensions (e.g. tallskinny matrices). This thesis explores novel methodologies and architectures for building efficient accelerators for sparse tensor algebra.The first major contribution of this thesis is the proposal of using specialized on-chip interconnects to provide flexible computational mappings of sparse and irregular matrices onto processing elements (PEs). This enables close to full PE utilization and significantly improves the performance over TPU, which has a rigid on-chip interconnect. With the proposed specialized interconnects, this thesis presents a new sparse DNN accelerator targeting workloads with ∼ 30% to 100% density (percentage of nonzeros) named SIGMA.Unlike popular DNNs, HPC workloads utilize tensors spanning from ∼ 10−6% dense to fully dense. The second major contribution of this thesis explores the system impact of utilizing various compression formats across all sparsity regions. The key insights gathered is that different workloads prefer different compression formats, and the best compression format used for memory storage may not be the same as the best compression format used for computation. This thesis proposes a predictor to determine the the best compression format combination and a custom hardware compression format converter named MINT.Together, they provide significant energy-delay product (EDP) improvement over state-ofthe-art accelerators.The third major contribution of this thesis analyzes popular state-of-the-art sparse accelerators using a new tool named Hard TACO. This tool utilizes the open source Tensor Algebra Compiler (TACO) and High Level Synthesis (HLS) to generate functional sparse accelerator of different dataflows, e.g. inner product vs output product SpGEMM. The impact of Hard TACO is that it allows realistic architectural exploration of homogeneous and heterogeneous accelerators.
- 일반주제명
- Sparsity
- 일반주제명
- Scheduling
- 일반주제명
- Space exploration
- 일반주제명
- Energy consumption
- 일반주제명
- Conversion
- 일반주제명
- Flexibility
- 일반주제명
- Aerospace engineering
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2022 us c eng d■001000017360494
■00520260202105536
■006m o d
■007cr#unu||||||||
■020 ▼a9798263392710
■035 ▼a(MiAaPQ)AAI32314798
■035 ▼a(MiAaPQ)GeorgiaTech66540
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a330
■1001 ▼aQin, Eric.
■24510▼aBuilding Efficient Tensor Accelerators for Sparse and Irregular Workloads
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2022
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2022
■300 ▼a144 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Krishna, Tushar.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2022.
■520 ▼aPopular Machine Learning (ML) and High Performance Computing (HPC) workloads contribute to a significant portion of runtime on data centers. Applications include image classification, speech recognition, recommendation systems, social network analysis, robotic problems, chemical process simulations, and so on. Recently due to large computational demands from emerging workloads, there is a surge of custom hardware accelerator development for computing tensor kernels with high performance and energy efficiency. For example, the Google Tensor Processing Unit (TPU) is a custom hardware accelerator targeting efficient matrix multiplications for Deep Neural Networks (DNNs). However, there are limitations with state-of-the-art accelerators, stemming from (1) a vast spectrum of sparsity across various workloads and (2) irregularity of tensor dimensions (e.g. tallskinny matrices). This thesis explores novel methodologies and architectures for building efficient accelerators for sparse tensor algebra.The first major contribution of this thesis is the proposal of using specialized on-chip interconnects to provide flexible computational mappings of sparse and irregular matrices onto processing elements (PEs). This enables close to full PE utilization and significantly improves the performance over TPU, which has a rigid on-chip interconnect. With the proposed specialized interconnects, this thesis presents a new sparse DNN accelerator targeting workloads with ∼ 30% to 100% density (percentage of nonzeros) named SIGMA.Unlike popular DNNs, HPC workloads utilize tensors spanning from ∼ 10−6% dense to fully dense. The second major contribution of this thesis explores the system impact of utilizing various compression formats across all sparsity regions. The key insights gathered is that different workloads prefer different compression formats, and the best compression format used for memory storage may not be the same as the best compression format used for computation. This thesis proposes a predictor to determine the the best compression format combination and a custom hardware compression format converter named MINT.Together, they provide significant energy-delay product (EDP) improvement over state-ofthe-art accelerators.The third major contribution of this thesis analyzes popular state-of-the-art sparse accelerators using a new tool named Hard TACO. This tool utilizes the open source Tensor Algebra Compiler (TACO) and High Level Synthesis (HLS) to generate functional sparse accelerator of different dataflows, e.g. inner product vs output product SpGEMM. The impact of Hard TACO is that it allows realistic architectural exploration of homogeneous and heterogeneous accelerators.
■590 ▼aSchool code: 0078.
■650 4▼aSparsity
■650 4▼aScheduling
■650 4▼aSpace exploration
■650 4▼aEnergy consumption
■650 4▼aConversion
■650 4▼aFlexibility
■650 4▼aAerospace engineering
■690 ▼a0538
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2022
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360494▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


