서브메뉴
검색
Accelerating Multilinear Maps and Structured Sparse Tensor Kernels
Accelerating Multilinear Maps and Structured Sparse Tensor Kernels
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104837
- ISBN
- 9798297600379
- DDC
- 004
- 서명/저자
- Accelerating Multilinear Maps and Structured Sparse Tensor Kernels
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 184 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
- 주기사항
- Advisor: Demmel, James;Buluc, Aydin.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Linear maps dominate machine learning and scientific computing workloads today. What about multilinear maps? Just as a linear map with one argument can be represented by a 2D matrix, a D-dimensional multilinear map is represented by a (D + 1)-dimensional tensor. We apply the map by flattening the tensor into a matrix and multiplying it by the Kronecker product of the inputs. When a batch of inputs is provided, this primitive is known as the Matricized-Tensor-Times-Khatri-Rao Product (MTTKRP). Efficient multilinear maps are essential in computational chemistry, multi-way data analysis, and signal processing. Unfortunately, they receive comparatively less interest from theorists and high-performance kernel designers.We optimize the multilinear map in two applications, making contributions that span theory and practical implementation. We first examine equivariant graph neural networks, which use a structured sparse tensor to interact node features with edge features. In response, we design a GPU kernel generator that matches or exceeds the best closed-source implementations for the problem. Our package, OpenEquivariance, provides 5-6x end-to-end speedup for training quantum chemical foundation models. Our focus then shifts to Candecomp / PARAFAC decomposition, a higher-dimensional analogue of the matrix singular value decomposition. Here, we use randomized linear algebra to accelerate the MTTKRP in tall, overdetermined linear least-squares problems, scaling our work to thousands of CPU cores. The remaining chapters detour by adapting this randomized algorithm to sketch tensor trains, structures that originated in quantum mechanical computations. We also design communication-avoiding algorithms for a pair of kernels used in matrix completion and graph attention networks. Our work demonstrates that sustained attention to the multilinear map yields fruit across the computational stack.
- 일반주제명
- Computer science
- 일반주제명
- Applied mathematics
- 키워드
- Machine learning
- 키워드
- Multilinearity
- 키워드
- Tensors
- 기타저자
- University of California, Berkeley Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359120
■00520260202104837
■006m o d
■007cr#unu||||||||
■020 ▼a9798297600379
■035 ▼a(MiAaPQ)AAI32171981
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aBharadwaj, Vivek.
■24510▼aAccelerating Multilinear Maps and Structured Sparse Tensor Kernels
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a184 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-04, Section: B.
■500 ▼aAdvisor: Demmel, James;Buluc, Aydin.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aLinear maps dominate machine learning and scientific computing workloads today. What about multilinear maps? Just as a linear map with one argument can be represented by a 2D matrix, a D-dimensional multilinear map is represented by a (D + 1)-dimensional tensor. We apply the map by flattening the tensor into a matrix and multiplying it by the Kronecker product of the inputs. When a batch of inputs is provided, this primitive is known as the Matricized-Tensor-Times-Khatri-Rao Product (MTTKRP). Efficient multilinear maps are essential in computational chemistry, multi-way data analysis, and signal processing. Unfortunately, they receive comparatively less interest from theorists and high-performance kernel designers.We optimize the multilinear map in two applications, making contributions that span theory and practical implementation. We first examine equivariant graph neural networks, which use a structured sparse tensor to interact node features with edge features. In response, we design a GPU kernel generator that matches or exceeds the best closed-source implementations for the problem. Our package, OpenEquivariance, provides 5-6x end-to-end speedup for training quantum chemical foundation models. Our focus then shifts to Candecomp / PARAFAC decomposition, a higher-dimensional analogue of the matrix singular value decomposition. Here, we use randomized linear algebra to accelerate the MTTKRP in tall, overdetermined linear least-squares problems, scaling our work to thousands of CPU cores. The remaining chapters detour by adapting this randomized algorithm to sketch tensor trains, structures that originated in quantum mechanical computations. We also design communication-avoiding algorithms for a pair of kernels used in matrix completion and graph attention networks. Our work demonstrates that sustained attention to the multilinear map yields fruit across the computational stack.
■590 ▼aSchool code: 0028.
■650 4▼aComputer science
■650 4▼aApplied mathematics
■653 ▼aGraph neural networks
■653 ▼aMachine learning
■653 ▼aMultilinearity
■653 ▼aTensors
■690 ▼a0984
■690 ▼a0364
■690 ▼a0800
■71020▼aUniversity of California, Berkeley▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g87-04B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359120▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


