본문

서브메뉴

Beyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and Nuclear Applications
Beyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and N...
Beyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and Nuclear Applications

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260311091517.5
ISBN  
9798270232672
DDC  
006
저자명  
Wanna, Selma Liliane Peterson
서명/저자  
Beyond Standard Benchmarking: Towards Robust and Trustworthy Robotics for Industrial and Nuclear Applications / Selma Liliane Peterson Wanna
발행사항  
[Sl] : The University of Texas at Austin, 2025
형태사항  
1 electronic resource (216 pages)
주기사항  
Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
주기사항  
Advisors: Landsberger, Sheldon; Pryor, Mitch Committee members: Clarno, Kevin; Martín-Martín, Roberto; Moore, Juston.
학위논문주기  
- Ph.D. : The University of Texas at Austin, 2025.
초록/해제  
요약Artificial intelligence (AI) is increasingly deployed in high-stakes environments, from autonomous systems in hazardous industrial settings to AI-driven decision-making in safety-critical applications. AI failures are rarely simple or predictable; they often occur in opaque, uninterpretable, and potentially catastrophic ways. This work explores the challenges that arise when AI is applied to complex, real-world applications, focusing on two key areas: perception in active safety systems and large language model (LLM)-driven task planning for Embodied AI (EAI).The first half of this dissertation examines uncertainty quantification (UQ) in real-time semantic segmentation, particularly within the context of glovebox safety systems used in nuclear and industrial environments. The research develops novel uncertainty-aware models, including Laplacian Segmentation Networks (LSNs), and introduces the Hand and Glove Segmentation (HAGS) dataset, a benchmark for evaluating segmentation robustness in human-robot collaboration tasks. Through a series of empirical studies, despite their strong theoretical foundations, Bayesian deep learning techniques are at best only comparable to less theoretically grounded methods for OOD detection.The second half of this dissertation investigates robustness in LLM-driven task planning for EAI, particularly in high-risk, unstructured domains such as nuclear robotics. While LLMs have shown promise in structured, household environments, they struggle significantly in out-of-distribution (OOD) scenarios, e.g, hierarchical and cyclical task planning. Through a series of prompt robustness experiments and dataset evaluations, this research identifies critical failure modes in current AI task planning methodologies. Beyond this, the work moves towards a data-centric approach, analyzing the biases present in existing EAI instruction-following datasets and proposing strategies to improve dataset diversity and mitigate spurious correlations.A unifying theme across both research areas is the necessity of trustworthy AI systems that go beyond achieving high benchmark accuracy and instead prioritize robustness, transparency, and safety in real-world deployment. This dissertation argues that the path forward requires a paradigm shift toward adhering to engineering-driven application requirements, improved dataset design, and AI models that incorporate meaningful uncertainty estimation and human oversight. The key findings presented below reflect this shift:1. Dataset characteristics, model architecture, and uncertainty estimation techniques interact in complex ways that often defy expectations, resulting in performance outcomes that challenge theoretically motivated uncertainty quantification methods, grounded in Information Theory and Bayesian uncertainty estimation frameworks, for OOD detection performance. While a new model (LSN) for uncertainty-aware segmentation was contributed; our experiments demonstrate that ensembles generalize more robustly across dataset domains.2. LLM-based task planners in EAI perform well on linear, sequential instruction chains. However, they struggle with conditional, cyclical, or parallel tasks, reflecting a deeper bias in current EAI datasets, which rarely include examples of complex control flow or decision-making logic.3. These same planners also underperform in domains with specialized vocabulary, e.g., terms found in nuclear and industrial environments, indicating that domain adaptation and terminology grounding remain open challenges in language-conditioned robotics, despite the breadth of LLM pretraining datasets.4. Instruction-following datasets in EAI are often linguistically shallow, relying heavily on templated commands with limited syntactic or semantic variation. This undermines the generalization ability of LLM-based models to natural modes of human communication.Together, these findings underscore the urgent need for interdisciplinary collaboration to build systems that are not only capable but also trustworthy, interpretable, and safe in high-stakes environments.
언어주기  
English
일반주제명  
Computer engineering
일반주제명  
Information technology
일반주제명  
Robotics
키워드  
Large language model
키워드  
Laplacian Segmentation Networks
키워드  
Hand and Glove Segmentation
키워드  
Out-of-distribution
기타저자  
The University of Texas at Austin Mechanical Engineering
기본자료저록  
Dissertations Abstracts International. 87-06B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260311s2025        us                                    eng  d
■001000017361246
■00520260311091517.5
■006m          o    d                
■007cr|nu||||||||
■020    ▼a9798270232672
■040    ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082    ▼a006
■1001  ▼aWanna,  Selma  Liliane  Peterson▼eauthor.
■24510▼aBeyond  Standard  Benchmarking:  Towards  Robust  and  Trustworthy  Robotics  for  Industrial  and  Nuclear  Applications  ▼cSelma  Liliane  Peterson  Wanna
■260    ▼a[Sl]▼bThe  University  of  Texas  at  Austin▼c2025
■264  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a1  electronic  resource  (216  pages)
■336    ▼atext▼btxt▼2rdacontent
■337    ▼acomputer▼bc▼2rdamedia
■338    ▼aonline  resource▼bcr▼2rdacarrier
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-06,  Section:  B.
■500    ▼aAdvisors:  Landsberger,  Sheldon;  Pryor,  Mitch    Committee  members:  Clarno,  Kevin;  Martín-Martín,  Roberto;  Moore,  Juston.
■5021  ▼bPh.D.▼cThe  University  of  Texas  at  Austin▼d2025.
■520    ▼aArtificial  intelligence  (AI)  is  increasingly  deployed  in  high-stakes  environments,  from  autonomous  systems  in  hazardous  industrial  settings  to  AI-driven  decision-making  in  safety-critical  applications.  AI  failures  are  rarely  simple  or  predictable;  they  often  occur  in  opaque,  uninterpretable,  and  potentially  catastrophic  ways.  This  work  explores  the  challenges  that  arise  when  AI  is  applied  to  complex,  real-world  applications,  focusing  on  two  key  areas:  perception  in  active  safety  systems  and  large  language  model  (LLM)-driven  task  planning  for  Embodied  AI  (EAI).The  first  half  of  this  dissertation  examines  uncertainty  quantification  (UQ)  in  real-time  semantic  segmentation,  particularly  within  the  context  of  glovebox  safety  systems  used  in  nuclear  and  industrial  environments.  The  research  develops  novel  uncertainty-aware  models,  including  Laplacian  Segmentation  Networks  (LSNs),  and  introduces  the  Hand  and  Glove  Segmentation  (HAGS)  dataset,  a  benchmark  for  evaluating  segmentation  robustness  in  human-robot  collaboration  tasks.  Through  a  series  of  empirical  studies,  despite  their  strong  theoretical  foundations,  Bayesian  deep  learning  techniques  are  at  best  only  comparable  to  less  theoretically  grounded  methods  for  OOD  detection.The  second  half  of  this  dissertation  investigates  robustness  in  LLM-driven  task  planning  for  EAI,  particularly  in  high-risk,  unstructured  domains  such  as  nuclear  robotics.  While  LLMs  have  shown  promise  in  structured,  household  environments,  they  struggle  significantly  in  out-of-distribution  (OOD)  scenarios,  e.g,  hierarchical  and  cyclical  task  planning.  Through  a  series  of  prompt  robustness  experiments  and  dataset  evaluations,  this  research  identifies  critical  failure  modes  in  current  AI  task  planning  methodologies.  Beyond  this,  the  work  moves  towards  a  data-centric  approach,  analyzing  the  biases  present  in  existing  EAI  instruction-following  datasets  and  proposing  strategies  to  improve  dataset  diversity  and  mitigate  spurious  correlations.A  unifying  theme  across  both  research  areas  is  the  necessity  of  trustworthy  AI  systems  that  go  beyond  achieving  high  benchmark  accuracy  and  instead  prioritize  robustness,  transparency,  and  safety  in  real-world  deployment.  This  dissertation  argues  that  the  path  forward  requires  a  paradigm  shift  toward  adhering  to  engineering-driven  application  requirements,  improved  dataset  design,  and  AI  models  that  incorporate  meaningful  uncertainty  estimation  and  human  oversight.  The  key  findings  presented  below  reflect  this  shift:1.  Dataset  characteristics,  model  architecture,  and  uncertainty  estimation  techniques  interact  in  complex  ways  that  often  defy  expectations,  resulting  in  performance  outcomes  that  challenge  theoretically  motivated  uncertainty  quantification  methods,  grounded  in  Information  Theory  and  Bayesian  uncertainty  estimation  frameworks,  for  OOD  detection  performance.  While  a  new  model  (LSN)  for  uncertainty-aware  segmentation  was  contributed;  our  experiments  demonstrate  that  ensembles  generalize  more  robustly  across  dataset  domains.2.  LLM-based  task  planners  in  EAI  perform  well  on  linear,  sequential  instruction  chains.  However,  they  struggle  with  conditional,  cyclical,  or  parallel  tasks,  reflecting  a  deeper  bias  in  current  EAI  datasets,  which  rarely  include  examples  of  complex  control  flow  or  decision-making  logic.3.  These  same  planners  also  underperform  in  domains  with  specialized  vocabulary,  e.g.,  terms  found  in  nuclear  and  industrial  environments,  indicating  that  domain  adaptation  and  terminology  grounding  remain  open  challenges  in  language-conditioned  robotics,  despite  the  breadth  of  LLM  pretraining  datasets.4.  Instruction-following  datasets  in  EAI  are  often  linguistically  shallow,  relying  heavily  on  templated  commands  with  limited  syntactic  or  semantic  variation.  This  undermines  the  generalization  ability  of  LLM-based  models  to  natural  modes  of  human  communication.Together,  these  findings  underscore  the  urgent  need  for  interdisciplinary  collaboration  to  build  systems  that  are  not  only  capable  but  also  trustworthy,  interpretable,  and  safe  in  high-stakes  environments. 
■546    ▼aEnglish
■590    ▼aSchool  code:  0227
■650  4▼aComputer  engineering
■650  4▼aInformation  technology
■650  4▼aRobotics
■653    ▼aLarge  language  model
■653    ▼aLaplacian  Segmentation  Networks
■653    ▼aHand  and  Glove  Segmentation
■653    ▼aOut-of-distribution
■7102  ▼aThe  University  of  Texas  at  Austin▼bMechanical  Engineering.▼edegree  granting  institution.
■7201  ▼aLandsberger,  Sheldon▼edegree  supervisor.
■7201  ▼aPryor,  Mitch▼edegree  supervisor.
■7730  ▼tDissertations  Abstracts  International▼g87-06B.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361246▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    ค้นหาข้อมูลรายละเอียด

    • จองห้องพัก
    • ไม่อยู่
    • โฟลเดอร์ของฉัน
    • ขอดูแรก
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    วัสดุ
    Reg No. Call No. ตำแหน่งที่ตั้ง สถานะ ยืมข้อมูล
    TF18080 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * จองมีอยู่ในหนังสือยืม เพื่อให้การสำรองที่นั่งคลิกที่ปุ่มจองห้องพัก

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.