서브메뉴
검색
Designing and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Application Level
Designing and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Application Level
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105529
- ISBN
- 9798263341794
- DDC
- 741
- 서명/저자
- Designing and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Application Level
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 153 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Sarkar, Vivek.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약Moore's law is dead or dying, but demands for compute continue to grow faster each year. The hardware scaling trends that have driven the growth of HPC are tapering off, which is forcing the industry to explore new approaches to continue scaling. Further, there is a growing public concern about the environmental impact of extreme-scale computing. Consequently, researchers in the cloud computing, machine learning, and embedded computing areas are exploring reduced-reliability computing as a means to improve both performance and efficiency. In HPC, however, the current Global Checkpoint/Recovery (GCR) approach to dealing with reduced hardware reliability is fundamentally unscalable. The costs of GCR are rising faster than the performance of leading supercomputers. It is critical for application resilience to scale with, rather than against, increasing hardware fault rates if HPC is to continue scaling while reigning in its environmental footprint. To avoid the exponential scaling of GCR, applications must localize the cost of hardware faults, which requires several changes in the traditional approach to fault tolerance. First, fault tolerance must be flexible to application-specific refinements while managing application developers' reticence to implement complex resilience code. We describe a layer-based resilience taxonomy and approach that exposes the imperative configurability mechanisms to make fault-tolerance tools that can flexibly combine to utilize general application- and platformtailored fault recovery. We demonstrate this by extending contemporary resilience tools to enable flexible and simple online recovery into applications with a multi-layered approach. Next, we define the key properties of localized recovery by creating a general analytical model. We demonstrate a pseudo-local approach using modern ULFM MPI features and show high scalability despite ULFM's collective nature. Finally, we design a task-based recovery model capable of extending pseudo-locality to applications which do not fit the ideal model. These works build the path for HPC to maintain environmental accountability, meet growing compute demands, and benefit from novel upcoming hardware trends.
- 일반주제명
- Design
- 일반주제명
- Fault tolerance
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 일반주제명
- Information technology
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360457
■00520260202105529
■006m o d
■007cr#unu||||||||
■020 ▼a9798263341794
■035 ▼a(MiAaPQ)AAI32309808
■035 ▼a(MiAaPQ)GeorgiaTech77831
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a741
■1001 ▼aWhitlock, Matthew.
■24510▼aDesigning and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Application Level
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a153 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Sarkar, Vivek.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aMoore's law is dead or dying, but demands for compute continue to grow faster each year. The hardware scaling trends that have driven the growth of HPC are tapering off, which is forcing the industry to explore new approaches to continue scaling. Further, there is a growing public concern about the environmental impact of extreme-scale computing. Consequently, researchers in the cloud computing, machine learning, and embedded computing areas are exploring reduced-reliability computing as a means to improve both performance and efficiency. In HPC, however, the current Global Checkpoint/Recovery (GCR) approach to dealing with reduced hardware reliability is fundamentally unscalable. The costs of GCR are rising faster than the performance of leading supercomputers. It is critical for application resilience to scale with, rather than against, increasing hardware fault rates if HPC is to continue scaling while reigning in its environmental footprint. To avoid the exponential scaling of GCR, applications must localize the cost of hardware faults, which requires several changes in the traditional approach to fault tolerance. First, fault tolerance must be flexible to application-specific refinements while managing application developers' reticence to implement complex resilience code. We describe a layer-based resilience taxonomy and approach that exposes the imperative configurability mechanisms to make fault-tolerance tools that can flexibly combine to utilize general application- and platformtailored fault recovery. We demonstrate this by extending contemporary resilience tools to enable flexible and simple online recovery into applications with a multi-layered approach. Next, we define the key properties of localized recovery by creating a general analytical model. We demonstrate a pseudo-local approach using modern ULFM MPI features and show high scalability despite ULFM's collective nature. Finally, we design a task-based recovery model capable of extending pseudo-locality to applications which do not fit the ideal model. These works build the path for HPC to maintain environmental accountability, meet growing compute demands, and benefit from novel upcoming hardware trends.
■590 ▼aSchool code: 0078.
■650 4▼aDesign
■650 4▼aFault tolerance
■650 4▼aHigh performance computing
■650 4▼aField programmable gate arrays
■650 4▼aComputer science
■650 4▼aElectrical engineering
■650 4▼aInformation technology
■690 ▼a0389
■690 ▼a0984
■690 ▼a0544
■690 ▼a0489
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360457▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


