본문

서브메뉴

Hierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits to Modules to Models
Hierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits ...
Hierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits to Modules to Models

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105534
ISBN  
9798263348441
DDC  
658.404
저자명  
Fu, Yonggan.
서명/저자  
Hierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits to Modules to Models
발행사항  
[Sl] : Georgia Institute of Technology, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
150 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Lin, Yingyan Celine.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2025.
초록/해제  
요약Despite the remarkable advancements of AI foundation models, such as large language models (LLMs), in numerous tasks and applications, deploying these powerful models on everyday devices remains challenging due to their growing computational and memory demands. This challenge hinders the realization of immersive and interactive user experiences that require real-time AI processing on resource-constrained devices.This PhD thesis aims to bridge this gap by performing hierarchical and hardware-aware optimization of AI models, maximizing accuracy-efficiency trade-offs to enable ubiquitous edge intelligence. Specifically, this thesis addresses redundancy at the bit, module, and model levels and leverages hardware characteristics to achieve real-device speed-ups. The proposed techniques include cyclic precision training (CPT) for efficient and accurate bitlevel quantization, DepthShrinker and AmoebaLLM for delivering real-hardware-efficient LLMs through module-level optimization, and a new language model architecture, Hymba, for efficient language processing, as well as Omni-Recon for efficient 3D understanding at the model level. These techniques collectively enable real-time execution of complex AI models on everyday devices, advancing the development of efficient AI solutions for ubiquitous edge intelligence.
일반주제명  
Schedules
일반주제명  
Embedded systems
일반주제명  
Human performance
일반주제명  
Large language models
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017360485
■00520260202105534
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263348441
■035    ▼a(MiAaPQ)AAI32310115
■035    ▼a(MiAaPQ)GeorgiaTech77951
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a658.404
■1001  ▼aFu,  Yonggan.
■24510▼aHierarchical  and  Hardware-Aware  Optimization  for  Enhancing  AI  Model  Efficiency:  From  Bits  to  Modules  to  Models
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a150  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Lin,  Yingyan  Celine.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2025.
■520    ▼aDespite  the  remarkable  advancements  of  AI  foundation  models,  such  as  large  language  models  (LLMs),  in  numerous  tasks  and  applications,  deploying  these  powerful  models  on  everyday  devices  remains  challenging  due  to  their  growing  computational  and  memory  demands.  This  challenge  hinders  the  realization  of  immersive  and  interactive  user  experiences  that  require  real-time  AI  processing  on  resource-constrained  devices.This  PhD  thesis  aims  to  bridge  this  gap  by  performing  hierarchical  and  hardware-aware  optimization  of  AI  models,  maximizing  accuracy-efficiency  trade-offs  to  enable  ubiquitous  edge  intelligence.  Specifically,  this  thesis  addresses  redundancy  at  the  bit,  module,  and  model  levels  and  leverages  hardware  characteristics  to  achieve  real-device  speed-ups.  The  proposed  techniques  include  cyclic  precision  training  (CPT)  for  efficient  and  accurate  bitlevel  quantization,  DepthShrinker  and  AmoebaLLM  for  delivering  real-hardware-efficient  LLMs  through  module-level  optimization,  and  a  new  language  model  architecture,  Hymba,  for  efficient  language  processing,  as  well  as  Omni-Recon  for  efficient  3D  understanding  at  the  model  level.  These  techniques  collectively  enable  real-time  execution  of  complex  AI  models  on  everyday  devices,  advancing  the  development  of  efficient  AI  solutions  for  ubiquitous  edge  intelligence.
■590    ▼aSchool  code:  0078.
■650  4▼aSchedules
■650  4▼aEmbedded  systems
■650  4▼aHuman  performance
■650  4▼aLarge  language  models
■690    ▼a0800
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360485▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17451 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.