서브메뉴
검색
Error Resilient and Adaptive Deep Learning Systems
Error Resilient and Adaptive Deep Learning Systems
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105604
- ISBN
- 9798265405890
- DDC
- 006
- 저자명
- Ma, Kwondo.
- 서명/저자
- Error Resilient and Adaptive Deep Learning Systems
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 171 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: A.
- 주기사항
- Advisor: Chatterjee, Abhijit.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약As deep learning systems become integral to a wide array of applications, including autonomous systems, healthcare, and finance, their complexity and deployment in hardware bring new challenges. In particular, the susceptibility of deep learning systems to hardware-induced errors, manufacturing process variability, and resource constraints presents critical obstacles to their reliable and efficient operation. This dissertation addresses these issues by introducing methodologies that enhance error resilience, adaptability, and energy efficiency in deep learning systems. The motivation for this work stems from the increasing integration of deep learning systems into real-world applications where reliability and robustness are paramount. The inherent variability in hardware-such as resistive RAM (RRAM)-and the need for efficient testing and tuning processes highlight the need for adaptive systems that can mitigate the impact of these variabilities. Additionally, the demand for low-power, high-performance hardware accelerators in edge computing environments presents further challenges in balancing computational efficiency and energy consumption. In response to these challenges, this research proposes a signature-based predictive testing framework for detecting performance degradation caused by process variability in hardware implementations of deep neural networks (DNNs). This framework introduces a compact, efficient testing mechanism that significantly improves the ability to identify defective devices during manufacturing, while also adapting to evolving manufacturing conditions through continuous retraining. Furthermore, a learning-assisted postmanufacture tuning framework is developed to optimize the performance of DNN accelerators, ensuring higher yields and greater reliability in fault-sensitive environments. This framework allows the system to adapt its tuning strategies over time, reducing the need for exhaustive retraining while maintaining operational efficiency. The dissertation also addresses the resilience of Transformer architectures to soft errors, a growing concern in high-performance applications such as natural language processing, and vision and image processing. The proposed approach combines error detection and suppression techniques to restore model performance under various error conditions, demonstrating the robustness of Transformer networks when deployed in real-world, error-prone environments. Finally, the work presents a novel energy-efficient DNN accelerator design that replaces traditional multiplication operations with shift-add computations, substantially reducing power consumption and latency. This architecture is particularly suited for low-power applications in edge and Internet of Things (IoT) devices, offering a practical solution for the deployment of deep learning models in energy-constrained settings. Overall, this research makes significant contributions toward improving the reliability and adaptability of deep learning systems, addressing key limitations in error resilience, manufacturing yield, and energy efficiency. These methodologies pave the way for the development of robust, efficient AI technologies capable of thriving in diverse and challenging environments.
- 일반주제명
- Deep learning
- 일반주제명
- Voice recognition
- 일반주제명
- Neural networks
- 일반주제명
- Adaptation
- 일반주제명
- Energy efficiency
- 일반주제명
- Machine translation
- 일반주제명
- Fault tolerance
- 일반주제명
- Industrial engineering
- 일반주제명
- Sustainability
- 기본자료저록
- Dissertations Abstracts International. 87-05A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360675
■00520260202105604
■006m o d
■007cr#unu||||||||
■020 ▼a9798265405890
■035 ▼a(MiAaPQ)AAI32316059
■035 ▼a(MiAaPQ)GeorgiaTech76975
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a006
■1001 ▼aMa, Kwondo.
■24510▼aError Resilient and Adaptive Deep Learning Systems
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a171 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: A.
■500 ▼aAdvisor: Chatterjee, Abhijit.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aAs deep learning systems become integral to a wide array of applications, including autonomous systems, healthcare, and finance, their complexity and deployment in hardware bring new challenges. In particular, the susceptibility of deep learning systems to hardware-induced errors, manufacturing process variability, and resource constraints presents critical obstacles to their reliable and efficient operation. This dissertation addresses these issues by introducing methodologies that enhance error resilience, adaptability, and energy efficiency in deep learning systems. The motivation for this work stems from the increasing integration of deep learning systems into real-world applications where reliability and robustness are paramount. The inherent variability in hardware-such as resistive RAM (RRAM)-and the need for efficient testing and tuning processes highlight the need for adaptive systems that can mitigate the impact of these variabilities. Additionally, the demand for low-power, high-performance hardware accelerators in edge computing environments presents further challenges in balancing computational efficiency and energy consumption. In response to these challenges, this research proposes a signature-based predictive testing framework for detecting performance degradation caused by process variability in hardware implementations of deep neural networks (DNNs). This framework introduces a compact, efficient testing mechanism that significantly improves the ability to identify defective devices during manufacturing, while also adapting to evolving manufacturing conditions through continuous retraining. Furthermore, a learning-assisted postmanufacture tuning framework is developed to optimize the performance of DNN accelerators, ensuring higher yields and greater reliability in fault-sensitive environments. This framework allows the system to adapt its tuning strategies over time, reducing the need for exhaustive retraining while maintaining operational efficiency. The dissertation also addresses the resilience of Transformer architectures to soft errors, a growing concern in high-performance applications such as natural language processing, and vision and image processing. The proposed approach combines error detection and suppression techniques to restore model performance under various error conditions, demonstrating the robustness of Transformer networks when deployed in real-world, error-prone environments. Finally, the work presents a novel energy-efficient DNN accelerator design that replaces traditional multiplication operations with shift-add computations, substantially reducing power consumption and latency. This architecture is particularly suited for low-power applications in edge and Internet of Things (IoT) devices, offering a practical solution for the deployment of deep learning models in energy-constrained settings. Overall, this research makes significant contributions toward improving the reliability and adaptability of deep learning systems, addressing key limitations in error resilience, manufacturing yield, and energy efficiency. These methodologies pave the way for the development of robust, efficient AI technologies capable of thriving in diverse and challenging environments.
■590 ▼aSchool code: 0078.
■650 4▼aDeep learning
■650 4▼aError correction & detection
■650 4▼aVoice recognition
■650 4▼aNeural networks
■650 4▼aAdaptation
■650 4▼aEnergy efficiency
■650 4▼aMachine translation
■650 4▼aNatural language processing
■650 4▼aFault tolerance
■650 4▼aIndustrial engineering
■650 4▼aSustainability
■690 ▼a0800
■690 ▼a0546
■690 ▼a0640
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05A.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360675▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


