본문

서브메뉴

Tuning Sparse Matrix Kernel Performance Via Lightweight Signatures
Tuning Sparse Matrix Kernel Performance Via Lightweight Signatures
Tuning Sparse Matrix Kernel Performance Via Lightweight Signatures

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260202105524
ISBN  
9798263343385
DDC  
513.2
저자명  
Jain, Anirudh.
서명/저자  
Tuning Sparse Matrix Kernel Performance Via Lightweight Signatures
발행사항  
[Sl] : Georgia Institute of Technology, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
158 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Conte, Thomas M.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2025.
초록/해제  
요약This dissertation introduces lightweight sparse-matrix pattern and occupancy signatures for automatic input-dependent tiling of sparse kernels such as sparse-dense matrix multiplication (SpMM), sampled dense-dense matrix multiplication (SDDMM), etc. Sparse matrix kernels are typically bandwidth limited and can benefit from optimizations such as tiling to improve cache effectiveness. However, tiling sparse kernels is challenging. Irregular sparse matrices often present intra-matrix variations in the distribution and structure of non-zeros and a one-size-fits all approach to tiling these kernels can result in sub-optimal performance. I present lightweight signatures, Residues, that use down-sampling techniques and bit-vectors to capture this irregularity. Residues capture non-zero occupancy and structure of rectangular regions of the sparse-matrix plane in such a manner that combinations of these allow the evaluation of arbitrarily larger regions of the sparse matrix. I demonstrate how Residues can be used for making intelligent tiling decisions that are both data reuseand data-movement-aware, tiling, specifically for single sparse matrix kernels like SpMM and SDDMM. These tiling techniques, ResGeMM and RASSM, greedily combine residue entries to analyze different tile shapes and generate tiles with a high cache volume footprint and data-reuse potential to improve performance. The maximum cache resident volume (temporal volume) of sparse kernels varies during execution, and statically determining this is not straightforward. I make the observation that the temporal volume problem for single-sparse-matrix-kernels is input dependent and can be mapped to the maximum overlapping interval analysis problem. I augment RASSM with the ability to leverage this analysis for improved tiling. This results in higher performance over static-spatial techniques and other state-of-the-art sparse tiling methods. Finally, this dissertation demonstrates the use of signatures in tiling sparse-sparse matrix multiplication (SpGeMM) when hardware accelerators are used for the partial product reduction phase of the algorithm.
일반주제명  
Multiplication & division
일반주제명  
Sparsity
일반주제명  
Spatial analysis
일반주제명  
Mathematics
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017360431
■00520260202105524
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263343385
■035    ▼a(MiAaPQ)AAI32309741
■035    ▼a(MiAaPQ)GeorgiaTech77772
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a513.2
■1001  ▼aJain,  Anirudh.
■24510▼aTuning  Sparse  Matrix  Kernel  Performance  Via  Lightweight  Signatures
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a158  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Conte,  Thomas  M.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2025.
■520    ▼aThis  dissertation  introduces  lightweight  sparse-matrix  pattern  and  occupancy  signatures  for  automatic  input-dependent  tiling  of  sparse  kernels  such  as  sparse-dense  matrix  multiplication  (SpMM),  sampled  dense-dense  matrix  multiplication  (SDDMM),  etc.  Sparse  matrix  kernels  are  typically  bandwidth  limited  and  can  benefit  from  optimizations  such  as  tiling  to  improve  cache  effectiveness.  However,  tiling  sparse  kernels  is  challenging.  Irregular  sparse  matrices  often  present  intra-matrix  variations  in  the  distribution  and  structure  of  non-zeros  and  a  one-size-fits  all  approach  to  tiling  these  kernels  can  result  in  sub-optimal  performance.  I  present  lightweight  signatures,  Residues,  that  use  down-sampling  techniques  and  bit-vectors  to  capture  this  irregularity.  Residues  capture  non-zero  occupancy  and  structure  of  rectangular  regions  of  the  sparse-matrix  plane  in  such  a  manner  that  combinations  of  these  allow  the  evaluation  of  arbitrarily  larger  regions  of  the  sparse  matrix.  I  demonstrate  how  Residues  can  be  used  for  making  intelligent  tiling  decisions  that  are  both  data  reuseand  data-movement-aware,  tiling,  specifically  for  single  sparse  matrix  kernels  like  SpMM  and  SDDMM.  These  tiling  techniques,  ResGeMM  and  RASSM,  greedily  combine  residue  entries  to  analyze  different  tile  shapes  and  generate  tiles  with  a  high  cache  volume  footprint  and  data-reuse  potential  to  improve  performance.  The  maximum  cache  resident  volume  (temporal  volume)  of  sparse  kernels  varies  during  execution,  and  statically  determining  this  is  not  straightforward.  I  make  the  observation  that  the  temporal  volume  problem  for  single-sparse-matrix-kernels  is  input  dependent  and  can  be  mapped  to  the  maximum  overlapping  interval  analysis  problem.  I  augment  RASSM  with  the  ability  to  leverage  this  analysis  for  improved  tiling.  This  results  in  higher  performance  over  static-spatial  techniques  and  other  state-of-the-art  sparse  tiling  methods.  Finally,  this  dissertation  demonstrates  the  use  of  signatures  in  tiling  sparse-sparse  matrix  multiplication  (SpGeMM)  when  hardware  accelerators  are  used  for  the  partial  product  reduction  phase  of  the  algorithm.
■590    ▼aSchool  code:  0078.
■650  4▼aMultiplication  &  division
■650  4▼aSparsity
■650  4▼aSpatial  analysis
■650  4▼aMathematics
■690    ▼a0800
■690    ▼a0405
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360431▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF17219 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.