본문

서브메뉴

Designing and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Application Level
Designing and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Appli...
Designing and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Application Level

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105529
ISBN  
9798263341794
DDC  
741
저자명  
Whitlock, Matthew.
서명/저자  
Designing and Automating Asynchronous, Localized, Multi-Level Fault-Tolerance at the Application Level
발행사항  
[Sl] : Georgia Institute of Technology, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
153 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Sarkar, Vivek.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
초록/해제  
요약Moore's law is dead or dying, but demands for compute continue to grow faster each year. The hardware scaling trends that have driven the growth of HPC are tapering off, which is forcing the industry to explore new approaches to continue scaling. Further, there is a growing public concern about the environmental impact of extreme-scale computing. Consequently, researchers in the cloud computing, machine learning, and embedded computing areas are exploring reduced-reliability computing as a means to improve both performance and efficiency. In HPC, however, the current Global Checkpoint/Recovery (GCR) approach to dealing with reduced hardware reliability is fundamentally unscalable. The costs of GCR are rising faster than the performance of leading supercomputers. It is critical for application resilience to scale with, rather than against, increasing hardware fault rates if HPC is to continue scaling while reigning in its environmental footprint. To avoid the exponential scaling of GCR, applications must localize the cost of hardware faults, which requires several changes in the traditional approach to fault tolerance. First, fault tolerance must be flexible to application-specific refinements while managing application developers' reticence to implement complex resilience code. We describe a layer-based resilience taxonomy and approach that exposes the imperative configurability mechanisms to make fault-tolerance tools that can flexibly combine to utilize general application- and platformtailored fault recovery. We demonstrate this by extending contemporary resilience tools to enable flexible and simple online recovery into applications with a multi-layered approach. Next, we define the key properties of localized recovery by creating a general analytical model. We demonstrate a pseudo-local approach using modern ULFM MPI features and show high scalability despite ULFM's collective nature. Finally, we design a task-based recovery model capable of extending pseudo-locality to applications which do not fit the ideal model. These works build the path for HPC to maintain environmental accountability, meet growing compute demands, and benefit from novel upcoming hardware trends.
일반주제명  
Design
일반주제명  
Fault tolerance
일반주제명  
High performance computing
일반주제명  
Field programmable gate arrays
일반주제명  
Computer science
일반주제명  
Electrical engineering
일반주제명  
Information technology
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017360457
■00520260202105529
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263341794
■035    ▼a(MiAaPQ)AAI32309808
■035    ▼a(MiAaPQ)GeorgiaTech77831
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a741
■1001  ▼aWhitlock,  Matthew.
■24510▼aDesigning  and  Automating  Asynchronous,  Localized,  Multi-Level  Fault-Tolerance  at  the  Application  Level
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a153  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Sarkar,  Vivek.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2024.
■520    ▼aMoore's  law  is  dead  or  dying,  but  demands  for  compute  continue  to  grow  faster  each  year.  The  hardware  scaling  trends  that  have  driven  the  growth  of  HPC  are  tapering  off,  which  is  forcing  the  industry  to  explore  new  approaches  to  continue  scaling.  Further,  there  is  a  growing  public  concern  about  the  environmental  impact  of  extreme-scale  computing.  Consequently,  researchers  in  the  cloud  computing,  machine  learning,  and  embedded  computing  areas  are  exploring  reduced-reliability  computing  as  a  means  to  improve  both  performance  and  efficiency.  In  HPC,  however,  the  current  Global  Checkpoint/Recovery  (GCR)  approach  to  dealing  with  reduced  hardware  reliability  is  fundamentally  unscalable.  The  costs  of  GCR  are  rising  faster  than  the  performance  of  leading  supercomputers.  It  is  critical  for  application  resilience  to  scale  with,  rather  than  against,  increasing  hardware  fault  rates  if  HPC  is  to  continue  scaling  while  reigning  in  its  environmental  footprint.  To  avoid  the  exponential  scaling  of  GCR,  applications  must  localize  the  cost  of  hardware  faults,  which  requires  several  changes  in  the  traditional  approach  to  fault  tolerance.  First,  fault  tolerance  must  be  flexible  to  application-specific  refinements  while  managing  application  developers'  reticence  to  implement  complex  resilience  code.  We  describe  a  layer-based  resilience  taxonomy  and  approach  that  exposes  the  imperative  configurability  mechanisms  to  make  fault-tolerance  tools  that  can  flexibly  combine  to  utilize  general  application-  and  platformtailored  fault  recovery.  We  demonstrate  this  by  extending  contemporary  resilience  tools  to  enable  flexible  and  simple  online  recovery  into  applications  with  a  multi-layered  approach.  Next,  we  define  the  key  properties  of  localized  recovery  by  creating  a  general  analytical  model.  We  demonstrate  a  pseudo-local  approach  using  modern  ULFM  MPI  features  and  show  high  scalability  despite  ULFM's  collective  nature.  Finally,  we  design  a  task-based  recovery  model  capable  of  extending  pseudo-locality  to  applications  which  do  not  fit  the  ideal  model.  These  works  build  the  path  for  HPC  to  maintain  environmental  accountability,  meet  growing  compute  demands,  and  benefit  from  novel  upcoming  hardware  trends.
■590    ▼aSchool  code:  0078.
■650  4▼aDesign
■650  4▼aFault  tolerance
■650  4▼aHigh  performance  computing
■650  4▼aField  programmable  gate  arrays
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■650  4▼aInformation  technology
■690    ▼a0389
■690    ▼a0984
■690    ▼a0544
■690    ▼a0489
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360457▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17300 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.