본문

서브메뉴

Efficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural Networks
Efficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural...
Efficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural Networks

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105601
ISBN  
9798263393892
DDC  
658.404
저자명  
Khare, Alind.
서명/저자  
Efficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural Networks
발행사항  
[Sl] : Georgia Institute of Technology, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
171 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Tumanov, Alexey.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
초록/해제  
요약The increasing deployment of Deep Neural Networks (DNNs) on the critical path of production applications in both datacenter and the edge require production systems to serve these DNNs under unpredictable and bursty request arrival rates. Serving models under such conditions requires these systems to strike a careful balance between the latency (R1) and accuracy (R2) requirements of the application and the overall efficiency of utilization of scarce resources (R3). To efficiently balance trade-offs in R1-R3, production systems need to navigate between models, choices of hardware, and application contexts. This thesis proposes an efficient AI stack to solve this tension in the R1-R3 trade-off space. The key idea in the efficient AI stack is to produce and consume Pareto-Optimal (w.r.t latency/accuracy) DNNs. On the production side, the thesis proposes several neural architecture search algorithms namely CompOFA, DES, and SuperFedNAS that automatically specialize DNNs to produce the highest accuracy under different hardware and latency targets in centralized and federated data environments. On the consumption side, the thesis proposes a) an inference serving system SuperServe that consumes these DNNs and schedules them under bursty workloads with resource efficiency, and b) DSched that schedules data pipelines to DNNs in a timely and cost-efficient manner. Overall, the proposed efficient stack co-optimizes R1-R2 under dynamic workloads with resource efficiency R3.
일반주제명  
Scheduling
일반주제명  
Neural networks
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017360650
■00520260202105601
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263393892
■035    ▼a(MiAaPQ)AAI32315976
■035    ▼a(MiAaPQ)GeorgiaTech76909
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a658.404
■1001  ▼aKhare,  Alind.
■24510▼aEfficient  AI  Stack:  Deployment-Aware  Neural  Architecture  Search  and  Serving  of  Deep  Neural  Networks
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a171  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Tumanov,  Alexey.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2024.
■520    ▼aThe  increasing  deployment  of  Deep  Neural  Networks  (DNNs)  on  the  critical  path  of  production  applications  in  both  datacenter  and  the  edge  require  production  systems  to  serve  these  DNNs  under  unpredictable  and  bursty  request  arrival  rates.  Serving  models  under  such  conditions  requires  these  systems  to  strike  a  careful  balance  between  the  latency  (R1)  and  accuracy  (R2)  requirements  of  the  application  and  the  overall  efficiency  of  utilization  of  scarce  resources  (R3).  To  efficiently  balance  trade-offs  in  R1-R3,  production  systems  need  to  navigate  between  models,  choices  of  hardware,  and  application  contexts.  This  thesis  proposes  an  efficient  AI  stack  to  solve  this  tension  in  the  R1-R3  trade-off  space.  The  key  idea  in  the  efficient  AI  stack  is  to  produce  and  consume  Pareto-Optimal  (w.r.t  latency/accuracy)  DNNs.  On  the  production  side,  the  thesis  proposes  several  neural  architecture  search  algorithms  namely  CompOFA,  DES,  and  SuperFedNAS  that  automatically  specialize  DNNs  to  produce  the  highest  accuracy  under  different  hardware  and  latency  targets  in  centralized  and  federated  data  environments.  On  the  consumption  side,  the  thesis  proposes  a)  an  inference  serving  system  SuperServe  that  consumes  these  DNNs  and  schedules  them  under  bursty  workloads  with  resource  efficiency,  and  b)  DSched  that  schedules  data  pipelines  to  DNNs  in  a  timely  and  cost-efficient  manner.  Overall,  the  proposed  efficient  stack  co-optimizes  R1-R2  under  dynamic  workloads  with  resource  efficiency  R3.
■590    ▼aSchool  code:  0078.
■650  4▼aScheduling
■650  4▼aNeural  networks
■690    ▼a0800
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360650▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17457 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.