서브메뉴
검색
Energy-Performance Tradeoffs in Data Centers and Machine Learning
Energy-Performance Tradeoffs in Data Centers and Machine Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260209102915
- ISBN
- 9798265427021
- DDC
- 001
- 서명/저자
- Energy-Performance Tradeoffs in Data Centers and Machine Learning
- 발행사항
- [Sl] : Stanford University, 2023
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2023
- 형태사항
- 218 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: A.
- 주기사항
- Advisor: Bambos, Nicholas.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2023.
- 초록/해제
- 요약The ability to scale computing over the coming decade is limited by power and energy: whether that be the power limits of individual processor chips due to heating effects, the ability of the power grid to supply a certain amount of power relative to the massive demand of hyperscale data centers, or the growing emphasis on reducing the carbon impact of computing. All of these problems benefit not only from revolutionary technological improvements, but also from fine-grained power consumption control.In this thesis, a general three tiered approach to control the performance-energy tradeoff in computing is presented, which combines tools from stochastic modeling, dynamic programming, and operating systems. A concrete instance of this approach is developed for the processor speed control setting, where a higher speed improves a program's slowdown performance at the convex cost of unit energy or power. The first tier is the low level control algorithm. For the modern slowdown performance metric, I show that the processor-queue's state must be modeled with its underlying multilevel structure otherwise the control policy will be sub-optimal even in expectation. To avoid the complexity of the complete multi-level policy solution, I develop an approximate control policy that accounts for the multi-level state and functions under any scheduling policy. The second tier is a meta-algorithm that leverages the first tier's control policy to achieve a particular target. It functions for a wide range of particular performance or energy targets, which itself enables increased robustness to parameter misestimation. Of particular interest are the 5-minute and 1-hour average power targets of data centers' power purchase agreements. The final tier adjusts the reconfiguration frequency of the computations required for the lower tiers and the estimation of system parameters, in order to tradeoff between the algorithms' computational overhead and the control accuracy.While in some settings like processor speed control a clear performance-energy tradeoff is possible, in modern machine learning (and deep neural networks in particular) large gains can be achieved by actually redesigning the training algorithms themselves for energy efficiency. Distributed Distillation (D-Dist) is one example of this in the small, power-limited devices setting. Presented in this thesis, D-Dist achieves a 10,000x reduction in the power-hungry communication required for distributed on-device training compared to the vanilla distributed stochastic gradient decent algorithm. Both of these approaches (improved algorithms and tradeoff control policies) can be combined to produce even further improvement. Today the current use of large-language-models is limited by inference time and expense; to begin to address this and demonstrate how these two approaches can be combined, in the final chapter, the application of the performance-energy tradeoff management stack is outlined for deep neural network inference-stage batch-size control.
- 일반주제명
- Software
- 일반주제명
- Dynamic programming
- 일반주제명
- Cooling
- 일반주제명
- Neural networks
- 일반주제명
- Energy management
- 일반주제명
- Hard disks
- 일반주제명
- Energy efficiency
- 일반주제명
- Energy resources
- 일반주제명
- Transistors
- 일반주제명
- Cloud computing
- 일반주제명
- Energy consumption
- 일반주제명
- Alternative energy
- 일반주제명
- Electrical engineering
- 일반주제명
- Sustainability
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-05A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260203s2023 us c eng d■001000017366021
■00520260209102915
■006m o d
■007cr#unu||||||||
■020 ▼a9798265427021
■035 ▼a(MiAaPQ)AAI32316407
■035 ▼a(MiAaPQ)Stanfordqz590zt0478
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a001
■1001 ▼aMann, Ariana Joy.
■24510▼aEnergy-Performance Tradeoffs in Data Centers and Machine Learning
■260 ▼a[Sl]▼bStanford University▼c2023
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2023
■300 ▼a218 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: A.
■500 ▼aAdvisor: Bambos, Nicholas.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2023.
■520 ▼aThe ability to scale computing over the coming decade is limited by power and energy: whether that be the power limits of individual processor chips due to heating effects, the ability of the power grid to supply a certain amount of power relative to the massive demand of hyperscale data centers, or the growing emphasis on reducing the carbon impact of computing. All of these problems benefit not only from revolutionary technological improvements, but also from fine-grained power consumption control.In this thesis, a general three tiered approach to control the performance-energy tradeoff in computing is presented, which combines tools from stochastic modeling, dynamic programming, and operating systems. A concrete instance of this approach is developed for the processor speed control setting, where a higher speed improves a program's slowdown performance at the convex cost of unit energy or power. The first tier is the low level control algorithm. For the modern slowdown performance metric, I show that the processor-queue's state must be modeled with its underlying multilevel structure otherwise the control policy will be sub-optimal even in expectation. To avoid the complexity of the complete multi-level policy solution, I develop an approximate control policy that accounts for the multi-level state and functions under any scheduling policy. The second tier is a meta-algorithm that leverages the first tier's control policy to achieve a particular target. It functions for a wide range of particular performance or energy targets, which itself enables increased robustness to parameter misestimation. Of particular interest are the 5-minute and 1-hour average power targets of data centers' power purchase agreements. The final tier adjusts the reconfiguration frequency of the computations required for the lower tiers and the estimation of system parameters, in order to tradeoff between the algorithms' computational overhead and the control accuracy.While in some settings like processor speed control a clear performance-energy tradeoff is possible, in modern machine learning (and deep neural networks in particular) large gains can be achieved by actually redesigning the training algorithms themselves for energy efficiency. Distributed Distillation (D-Dist) is one example of this in the small, power-limited devices setting. Presented in this thesis, D-Dist achieves a 10,000x reduction in the power-hungry communication required for distributed on-device training compared to the vanilla distributed stochastic gradient decent algorithm. Both of these approaches (improved algorithms and tradeoff control policies) can be combined to produce even further improvement. Today the current use of large-language-models is limited by inference time and expense; to begin to address this and demonstrate how these two approaches can be combined, in the final chapter, the application of the performance-energy tradeoff management stack is outlined for deep neural network inference-stage batch-size control.
■590 ▼aSchool code: 0212.
■650 4▼aSoftware
■650 4▼aDynamic programming
■650 4▼aCooling
■650 4▼aNeural networks
■650 4▼aEnergy management
■650 4▼aHard disks
■650 4▼aEnergy efficiency
■650 4▼aAlternative energy sources
■650 4▼aEnergy resources
■650 4▼aTransistors
■650 4▼aCloud computing
■650 4▼aEnergy consumption
■650 4▼aAlternative energy
■650 4▼aElectrical engineering
■650 4▼aSustainability
■690 ▼a0363
■690 ▼a0800
■690 ▼a0544
■690 ▼a0640
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-05A.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17366021▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


