본문

서브메뉴

Towards Cloud-Scale Debugging
Towards Cloud-Scale Debugging
Towards Cloud-Scale Debugging

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151942
ISBN  
9798382778426
DDC  
004
저자명  
Dogga, Pradeep.
서명/저자  
Towards Cloud-Scale Debugging
발행사항  
[Sl] : University of California, Los Angeles, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
178 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Netravali, Ravi Arun;Varghese, George.
학위논문주기  
Thesis (Ph.D.)--University of California, Los Angeles, 2024.
초록/해제  
요약Cloud computing is an integral part of today's world: it primarily enables individuals and enterprises to provision and manage resources such as compute, storage, etc., for their needs with the click of a button. Modular approach to software development enabled cloud providers to rapidly evolve and deliver increasing number of services to users rendering clouds mission-critical. To insure prompt serviceability of this Achilles' Heel from facing incidents, cloud providers employ significant human resources. However, with the ever increasing number of services offered by clouds and growing types of workloads such as the proliferation of Machine Learning workloads in recent times, it is no longer viable for cloud providers to scale their human resources at this pace to insure prompt serviceability of their clouds.In this dissertation, I present my work towards improving the serviceability of clouds by leveraging insights from my experience with real debugging workflows employed at the three largest clouds today. I present techniques from Machine Learning and Natural Language Processing to leverage the vast amount of historical debugging data in clouds to develop tools that provide assistance to their engineers. I present a 'Coarsening' framework that enables transition towards a centralized debugging plane and discuss practical evaluations of tools built using this framework.I present Revelio, a tool that can generate debugging queries for engineers to execute over system-wide logged data, whose results can likely hint them of the root cause of an incident. To enable benchmarking many techniques, I also built a distributed systems debugging testbed that can inject faults into services, interface with human users and collect execution logs across the system. I present AutoARTS, a tool that can tag a lengthy postmortem report of an incident in the cloud with all root causes from an extensive taxonomy and can also highlight key pieces of information from a postmortem for ease of analysis. I present PerfRCA, a tool that can scale causal discovery to production-scale telemetry to reason performance degradations. I conclude with my vision for a centralized approach to automatically extract generalizable debugging assistance to engineers across a cloud.
일반주제명  
Computer science
일반주제명  
Computer engineering
키워드  
Cloud computing
키워드  
Computer networks
키워드  
Debugging
키워드  
Distributed systems
키워드  
Machine Learning
키워드  
Natural Language Processing
기타저자  
University of California, Los Angeles Computer Science 0201
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017162178
■00520250211151942
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382778426
■035    ▼a(MiAaPQ)AAI31302284
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aDogga,  Pradeep.
■24510▼aTowards  Cloud-Scale  Debugging
■260    ▼a[Sl]▼bUniversity  of  California,  Los  Angeles▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a178  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Netravali,  Ravi  Arun;Varghese,  George.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Los  Angeles,  2024.
■520    ▼aCloud  computing  is  an  integral  part  of  today's  world:  it  primarily  enables  individuals  and  enterprises  to  provision  and  manage  resources  such  as  compute,  storage,  etc.,  for  their  needs  with  the  click  of  a  button.  Modular  approach  to  software  development  enabled  cloud  providers  to  rapidly  evolve  and  deliver  increasing  number  of  services  to  users  rendering  clouds  mission-critical.  To  insure  prompt  serviceability  of  this  Achilles'  Heel  from  facing  incidents,  cloud  providers  employ  significant  human  resources.  However,  with  the  ever  increasing  number  of  services  offered  by  clouds  and  growing  types  of  workloads  such  as  the  proliferation  of  Machine  Learning  workloads  in  recent  times,  it  is  no  longer  viable  for  cloud  providers  to  scale  their  human  resources  at  this  pace  to  insure  prompt  serviceability  of  their  clouds.In  this  dissertation,  I  present  my  work  towards  improving  the  serviceability  of  clouds  by  leveraging  insights  from  my  experience  with  real  debugging  workflows  employed  at  the  three  largest  clouds  today.  I  present  techniques  from  Machine  Learning  and  Natural  Language  Processing  to  leverage  the  vast  amount  of  historical  debugging  data  in  clouds  to  develop  tools  that  provide  assistance  to  their  engineers.  I  present  a  'Coarsening'  framework  that  enables  transition  towards  a  centralized  debugging  plane  and  discuss  practical  evaluations  of  tools  built  using  this  framework.I  present  Revelio,  a  tool  that  can  generate  debugging  queries  for  engineers  to  execute  over  system-wide  logged  data,  whose  results  can  likely  hint  them  of  the  root  cause  of  an  incident.  To  enable  benchmarking  many  techniques,  I  also  built  a  distributed  systems  debugging  testbed  that  can  inject  faults  into  services,  interface  with  human  users  and  collect  execution  logs  across  the  system.  I  present  AutoARTS,  a  tool  that  can  tag  a  lengthy  postmortem  report  of  an  incident  in  the  cloud  with  all  root  causes  from  an  extensive  taxonomy  and  can  also  highlight  key  pieces  of  information  from  a  postmortem  for  ease  of  analysis.  I  present  PerfRCA,  a  tool  that  can  scale  causal  discovery  to  production-scale  telemetry  to  reason  performance  degradations.  I  conclude  with  my  vision  for  a  centralized  approach  to  automatically  extract  generalizable  debugging  assistance  to  engineers  across  a  cloud.
■590    ▼aSchool  code:  0031.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■653    ▼aCloud  computing
■653    ▼aComputer  networks
■653    ▼aDebugging
■653    ▼aDistributed  systems
■653    ▼aMachine  Learning
■653    ▼aNatural  Language  Processing
■690    ▼a0984
■690    ▼a0800
■690    ▼a0464
■71020▼aUniversity  of  California,  Los  Angeles▼bComputer  Science  0201.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0031
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162178▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF10211 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.