서브메뉴
검색
Hierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits to Modules to Models
Hierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits to Modules to Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105534
- ISBN
- 9798263348441
- DDC
- 658.404
- 저자명
- Fu, Yonggan.
- 서명/저자
- Hierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits to Modules to Models
- 발행사항
- [Sl] : Georgia Institute of Technology, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 150 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Lin, Yingyan Celine.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2025.
- 초록/해제
- 요약Despite the remarkable advancements of AI foundation models, such as large language models (LLMs), in numerous tasks and applications, deploying these powerful models on everyday devices remains challenging due to their growing computational and memory demands. This challenge hinders the realization of immersive and interactive user experiences that require real-time AI processing on resource-constrained devices.This PhD thesis aims to bridge this gap by performing hierarchical and hardware-aware optimization of AI models, maximizing accuracy-efficiency trade-offs to enable ubiquitous edge intelligence. Specifically, this thesis addresses redundancy at the bit, module, and model levels and leverages hardware characteristics to achieve real-device speed-ups. The proposed techniques include cyclic precision training (CPT) for efficient and accurate bitlevel quantization, DepthShrinker and AmoebaLLM for delivering real-hardware-efficient LLMs through module-level optimization, and a new language model architecture, Hymba, for efficient language processing, as well as Omni-Recon for efficient 3D understanding at the model level. These techniques collectively enable real-time execution of complex AI models on everyday devices, advancing the development of efficient AI solutions for ubiquitous edge intelligence.
- 일반주제명
- Schedules
- 일반주제명
- Embedded systems
- 일반주제명
- Human performance
- 일반주제명
- Large language models
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017360485
■00520260202105534
■006m o d
■007cr#unu||||||||
■020 ▼a9798263348441
■035 ▼a(MiAaPQ)AAI32310115
■035 ▼a(MiAaPQ)GeorgiaTech77951
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a658.404
■1001 ▼aFu, Yonggan.
■24510▼aHierarchical and Hardware-Aware Optimization for Enhancing AI Model Efficiency: From Bits to Modules to Models
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a150 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Lin, Yingyan Celine.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2025.
■520 ▼aDespite the remarkable advancements of AI foundation models, such as large language models (LLMs), in numerous tasks and applications, deploying these powerful models on everyday devices remains challenging due to their growing computational and memory demands. This challenge hinders the realization of immersive and interactive user experiences that require real-time AI processing on resource-constrained devices.This PhD thesis aims to bridge this gap by performing hierarchical and hardware-aware optimization of AI models, maximizing accuracy-efficiency trade-offs to enable ubiquitous edge intelligence. Specifically, this thesis addresses redundancy at the bit, module, and model levels and leverages hardware characteristics to achieve real-device speed-ups. The proposed techniques include cyclic precision training (CPT) for efficient and accurate bitlevel quantization, DepthShrinker and AmoebaLLM for delivering real-hardware-efficient LLMs through module-level optimization, and a new language model architecture, Hymba, for efficient language processing, as well as Omni-Recon for efficient 3D understanding at the model level. These techniques collectively enable real-time execution of complex AI models on everyday devices, advancing the development of efficient AI solutions for ubiquitous edge intelligence.
■590 ▼aSchool code: 0078.
■650 4▼aSchedules
■650 4▼aEmbedded systems
■650 4▼aHuman performance
■650 4▼aLarge language models
■690 ▼a0800
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360485▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


