서브메뉴
검색
Beyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and Nuclear Applications
Beyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and Nuclear Applications
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260311091517.5
- ISBN
- 9798270232672
- DDC
- 006
- 서명/저자
- Beyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and Nuclear Applications / Selma Liliane Peterson Wanna
- 발행사항
- [Sl] : The University of Texas at Austin, 2025
- 형태사항
- 1 electronic resource (216 pages)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
- 주기사항
- Advisors: Landsberger, Sheldon; Pryor, Mitch Committee members: Clarno, Kevin; Martín-Martín, Roberto; Moore, Juston.
- 학위논문주기
- - Ph.D. : The University of Texas at Austin, 2025.
- 초록/해제
- 요약Artificial intelligence (AI) is increasingly deployed in high-stakes environments, from autonomous systems in hazardous industrial settings to AI-driven decision-making in safety-critical applications. AI failures are rarely simple or predictable; they often occur in opaque, uninterpretable, and potentially catastrophic ways. This work explores the challenges that arise when AI is applied to complex, real-world applications, focusing on two key areas: perception in active safety systems and large language model (LLM)-driven task planning for Embodied AI (EAI).The first half of this dissertation examines uncertainty quantification (UQ) in real-time semantic segmentation, particularly within the context of glovebox safety systems used in nuclear and industrial environments. The research develops novel uncertainty-aware models, including Laplacian Segmentation Networks (LSNs), and introduces the Hand and Glove Segmentation (HAGS) dataset, a benchmark for evaluating segmentation robustness in human-robot collaboration tasks. Through a series of empirical studies, despite their strong theoretical foundations, Bayesian deep learning techniques are at best only comparable to less theoretically grounded methods for OOD detection.The second half of this dissertation investigates robustness in LLM-driven task planning for EAI, particularly in high-risk, unstructured domains such as nuclear robotics. While LLMs have shown promise in structured, household environments, they struggle significantly in out-of-distribution (OOD) scenarios, e.g, hierarchical and cyclical task planning. Through a series of prompt robustness experiments and dataset evaluations, this research identifies critical failure modes in current AI task planning methodologies. Beyond this, the work moves towards a data-centric approach, analyzing the biases present in existing EAI instruction-following datasets and proposing strategies to improve dataset diversity and mitigate spurious correlations.A unifying theme across both research areas is the necessity of trustworthy AI systems that go beyond achieving high benchmark accuracy and instead prioritize robustness, transparency, and safety in real-world deployment. This dissertation argues that the path forward requires a paradigm shift toward adhering to engineering-driven application requirements, improved dataset design, and AI models that incorporate meaningful uncertainty estimation and human oversight. The key findings presented below reflect this shift:1. Dataset characteristics, model architecture, and uncertainty estimation techniques interact in complex ways that often defy expectations, resulting in performance outcomes that challenge theoretically motivated uncertainty quantification methods, grounded in Information Theory and Bayesian uncertainty estimation frameworks, for OOD detection performance. While a new model (LSN) for uncertainty-aware segmentation was contributed; our experiments demonstrate that ensembles generalize more robustly across dataset domains.2. LLM-based task planners in EAI perform well on linear, sequential instruction chains. However, they struggle with conditional, cyclical, or parallel tasks, reflecting a deeper bias in current EAI datasets, which rarely include examples of complex control flow or decision-making logic.3. These same planners also underperform in domains with specialized vocabulary, e.g., terms found in nuclear and industrial environments, indicating that domain adaptation and terminology grounding remain open challenges in language-conditioned robotics, despite the breadth of LLM pretraining datasets.4. Instruction-following datasets in EAI are often linguistically shallow, relying heavily on templated commands with limited syntactic or semantic variation. This undermines the generalization ability of LLM-based models to natural modes of human communication.Together, these findings underscore the urgent need for interdisciplinary collaboration to build systems that are not only capable but also trustworthy, interpretable, and safe in high-stakes environments.
- 언어주기
- English
- 일반주제명
- Computer engineering
- 일반주제명
- Information technology
- 일반주제명
- Robotics
- 기타저자
- The University of Texas at Austin Mechanical Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260311s2025 us eng d■001000017361246
■00520260311091517.5
■006m o d
■007cr|nu||||||||
■020 ▼a9798270232672
■040 ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082 ▼a006
■1001 ▼aWanna, Selma Liliane Peterson▼eauthor.
■24510▼aBeyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and Nuclear Applications ▼cSelma Liliane Peterson Wanna
■260 ▼a[Sl]▼bThe University of Texas at Austin▼c2025
■264 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a1 electronic resource (216 pages)
■336 ▼atext▼btxt▼2rdacontent
■337 ▼acomputer▼bc▼2rdamedia
■338 ▼aonline resource▼bcr▼2rdacarrier
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: B.
■500 ▼aAdvisors: Landsberger, Sheldon; Pryor, Mitch Committee members: Clarno, Kevin; Martín-Martín, Roberto; Moore, Juston.
■5021 ▼bPh.D.▼cThe University of Texas at Austin▼d2025.
■520 ▼aArtificial intelligence (AI) is increasingly deployed in high-stakes environments, from autonomous systems in hazardous industrial settings to AI-driven decision-making in safety-critical applications. AI failures are rarely simple or predictable; they often occur in opaque, uninterpretable, and potentially catastrophic ways. This work explores the challenges that arise when AI is applied to complex, real-world applications, focusing on two key areas: perception in active safety systems and large language model (LLM)-driven task planning for Embodied AI (EAI).The first half of this dissertation examines uncertainty quantification (UQ) in real-time semantic segmentation, particularly within the context of glovebox safety systems used in nuclear and industrial environments. The research develops novel uncertainty-aware models, including Laplacian Segmentation Networks (LSNs), and introduces the Hand and Glove Segmentation (HAGS) dataset, a benchmark for evaluating segmentation robustness in human-robot collaboration tasks. Through a series of empirical studies, despite their strong theoretical foundations, Bayesian deep learning techniques are at best only comparable to less theoretically grounded methods for OOD detection.The second half of this dissertation investigates robustness in LLM-driven task planning for EAI, particularly in high-risk, unstructured domains such as nuclear robotics. While LLMs have shown promise in structured, household environments, they struggle significantly in out-of-distribution (OOD) scenarios, e.g, hierarchical and cyclical task planning. Through a series of prompt robustness experiments and dataset evaluations, this research identifies critical failure modes in current AI task planning methodologies. Beyond this, the work moves towards a data-centric approach, analyzing the biases present in existing EAI instruction-following datasets and proposing strategies to improve dataset diversity and mitigate spurious correlations.A unifying theme across both research areas is the necessity of trustworthy AI systems that go beyond achieving high benchmark accuracy and instead prioritize robustness, transparency, and safety in real-world deployment. This dissertation argues that the path forward requires a paradigm shift toward adhering to engineering-driven application requirements, improved dataset design, and AI models that incorporate meaningful uncertainty estimation and human oversight. The key findings presented below reflect this shift:1. Dataset characteristics, model architecture, and uncertainty estimation techniques interact in complex ways that often defy expectations, resulting in performance outcomes that challenge theoretically motivated uncertainty quantification methods, grounded in Information Theory and Bayesian uncertainty estimation frameworks, for OOD detection performance. While a new model (LSN) for uncertainty-aware segmentation was contributed; our experiments demonstrate that ensembles generalize more robustly across dataset domains.2. LLM-based task planners in EAI perform well on linear, sequential instruction chains. However, they struggle with conditional, cyclical, or parallel tasks, reflecting a deeper bias in current EAI datasets, which rarely include examples of complex control flow or decision-making logic.3. These same planners also underperform in domains with specialized vocabulary, e.g., terms found in nuclear and industrial environments, indicating that domain adaptation and terminology grounding remain open challenges in language-conditioned robotics, despite the breadth of LLM pretraining datasets.4. Instruction-following datasets in EAI are often linguistically shallow, relying heavily on templated commands with limited syntactic or semantic variation. This undermines the generalization ability of LLM-based models to natural modes of human communication.Together, these findings underscore the urgent need for interdisciplinary collaboration to build systems that are not only capable but also trustworthy, interpretable, and safe in high-stakes environments.
■546 ▼aEnglish
■590 ▼aSchool code: 0227
■650 4▼aComputer engineering
■650 4▼aInformation technology
■650 4▼aRobotics
■653 ▼aLarge language model
■653 ▼aLaplacian Segmentation Networks
■653 ▼aHand and Glove Segmentation
■653 ▼aOut-of-distribution
■7102 ▼aThe University of Texas at Austin▼bMechanical Engineering.▼edegree granting institution.
■7201 ▼aLandsberger, Sheldon▼edegree supervisor.
■7201 ▼aPryor, Mitch▼edegree supervisor.
■7730 ▼tDissertations Abstracts International▼g87-06B.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361246▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


