서브메뉴
검색
Efficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural Networks
Efficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural Networks
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105601
- ISBN
- 9798263393892
- DDC
- 658.404
- 저자명
- Khare, Alind.
- 서명/저자
- Efficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural Networks
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 171 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Tumanov, Alexey.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약The increasing deployment of Deep Neural Networks (DNNs) on the critical path of production applications in both datacenter and the edge require production systems to serve these DNNs under unpredictable and bursty request arrival rates. Serving models under such conditions requires these systems to strike a careful balance between the latency (R1) and accuracy (R2) requirements of the application and the overall efficiency of utilization of scarce resources (R3). To efficiently balance trade-offs in R1-R3, production systems need to navigate between models, choices of hardware, and application contexts. This thesis proposes an efficient AI stack to solve this tension in the R1-R3 trade-off space. The key idea in the efficient AI stack is to produce and consume Pareto-Optimal (w.r.t latency/accuracy) DNNs. On the production side, the thesis proposes several neural architecture search algorithms namely CompOFA, DES, and SuperFedNAS that automatically specialize DNNs to produce the highest accuracy under different hardware and latency targets in centralized and federated data environments. On the consumption side, the thesis proposes a) an inference serving system SuperServe that consumes these DNNs and schedules them under bursty workloads with resource efficiency, and b) DSched that schedules data pipelines to DNNs in a timely and cost-efficient manner. Overall, the proposed efficient stack co-optimizes R1-R2 under dynamic workloads with resource efficiency R3.
- 일반주제명
- Scheduling
- 일반주제명
- Neural networks
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360650
■00520260202105601
■006m o d
■007cr#unu||||||||
■020 ▼a9798263393892
■035 ▼a(MiAaPQ)AAI32315976
■035 ▼a(MiAaPQ)GeorgiaTech76909
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a658.404
■1001 ▼aKhare, Alind.
■24510▼aEfficient AI Stack: Deployment-Aware Neural Architecture Search and Serving of Deep Neural Networks
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a171 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Tumanov, Alexey.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aThe increasing deployment of Deep Neural Networks (DNNs) on the critical path of production applications in both datacenter and the edge require production systems to serve these DNNs under unpredictable and bursty request arrival rates. Serving models under such conditions requires these systems to strike a careful balance between the latency (R1) and accuracy (R2) requirements of the application and the overall efficiency of utilization of scarce resources (R3). To efficiently balance trade-offs in R1-R3, production systems need to navigate between models, choices of hardware, and application contexts. This thesis proposes an efficient AI stack to solve this tension in the R1-R3 trade-off space. The key idea in the efficient AI stack is to produce and consume Pareto-Optimal (w.r.t latency/accuracy) DNNs. On the production side, the thesis proposes several neural architecture search algorithms namely CompOFA, DES, and SuperFedNAS that automatically specialize DNNs to produce the highest accuracy under different hardware and latency targets in centralized and federated data environments. On the consumption side, the thesis proposes a) an inference serving system SuperServe that consumes these DNNs and schedules them under bursty workloads with resource efficiency, and b) DSched that schedules data pipelines to DNNs in a timely and cost-efficient manner. Overall, the proposed efficient stack co-optimizes R1-R2 under dynamic workloads with resource efficiency R3.
■590 ▼aSchool code: 0078.
■650 4▼aScheduling
■650 4▼aNeural networks
■690 ▼a0800
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360650▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


