서브메뉴
검색
Integration of Learning-Based and Model-Based Autonomy: From System Modeling to Control Policy
Integration of Learning-Based and Model-Based Autonomy: From System Modeling to Control Policy
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103556
- ISBN
- 9798288862083
- DDC
- 621
- 저자명
- Li, Chenran.
- 서명/저자
- Integration of Learning-Based and Model-Based Autonomy: From System Modeling to Control Policy
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 164 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Tomizuka, Masayoshi.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Autonomous systems operating in dynamic, uncertain, and socially interactive environments must effectively integrate learning-based and model-based methods, while reconciling data-driven and reward-driven objectives. Learning-based approaches offer flexibility and scalability but often suffer from brittleness under distribution shifts. Model-based methods provide robustness and predictive foresight, yet face challenges in modeling complex, high-dimensional dynamics. Similarly, data-driven policies trained from demonstrations may lack adaptability to new objectives, while reward-driven optimization can be inefficient and unstable without strong priors.This dissertation addresses these challenges by proposing a set of approaches that combine predictive modeling with residual Q-learning frameworks to enhance the robustness and adaptability of autonomous systems. At a high level, the work builds methods that enable autonomous agents to extract structured models from data, predict and reason about dynamic environments, and flexibly adapt their behavior to changing objectives without discarding prior knowledge.The first part of the dissertation focuses on building models from data to support reliable estimation and planning. We develop a dual estimation framework that jointly estimates latent system states and time-varying dynamic parameters in real time, enabling adaptive model-based control under changing conditions. Additionally, we propose a game-theoretic planning framework that incorporates predictive heuristics into Monte Carlo Tree Search, allowing the agent to reason about socially compliant interactions while maintaining computational efficiency in multi-agent environments.The second part of the dissertation introduces residual Q-learning as a principled mechanism for integrating data-driven behaviors with reward-driven adaptation. We formulate the policy customization problem and propose Residual Q-learning, a method that enables policies to adapt to new task objectives without requiring access to the original reward signals. We extend residual Q-learning to policy gradient methods, developing a unified structure that connects data-driven and reward-driven objectives. Finally, we integrate residual Q-learning into model-predictive path integral control, enabling fast, adaptive continuous control by combining learned priors with real-time model-based optimization.Through these, the dissertation advances scalable, adaptable, and robust decision-making frameworks for autonomous systems, bridging the gap between offline learning and real-world deployment.
- 일반주제명
- Mechanical engineering
- 일반주제명
- Computer science
- 일반주제명
- Robotics
- 키워드
- Planning
- 기타저자
- University of California, Berkeley Mechanical Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357756
■00520260202103556
■006m o d
■007cr#unu||||||||
■020 ▼a9798288862083
■035 ▼a(MiAaPQ)AAI32042064
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621
■1001 ▼aLi, Chenran.
■24510▼aIntegration of Learning-Based and Model-Based Autonomy: From System Modeling to Control Policy
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a164 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Tomizuka, Masayoshi.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aAutonomous systems operating in dynamic, uncertain, and socially interactive environments must effectively integrate learning-based and model-based methods, while reconciling data-driven and reward-driven objectives. Learning-based approaches offer flexibility and scalability but often suffer from brittleness under distribution shifts. Model-based methods provide robustness and predictive foresight, yet face challenges in modeling complex, high-dimensional dynamics. Similarly, data-driven policies trained from demonstrations may lack adaptability to new objectives, while reward-driven optimization can be inefficient and unstable without strong priors.This dissertation addresses these challenges by proposing a set of approaches that combine predictive modeling with residual Q-learning frameworks to enhance the robustness and adaptability of autonomous systems. At a high level, the work builds methods that enable autonomous agents to extract structured models from data, predict and reason about dynamic environments, and flexibly adapt their behavior to changing objectives without discarding prior knowledge.The first part of the dissertation focuses on building models from data to support reliable estimation and planning. We develop a dual estimation framework that jointly estimates latent system states and time-varying dynamic parameters in real time, enabling adaptive model-based control under changing conditions. Additionally, we propose a game-theoretic planning framework that incorporates predictive heuristics into Monte Carlo Tree Search, allowing the agent to reason about socially compliant interactions while maintaining computational efficiency in multi-agent environments.The second part of the dissertation introduces residual Q-learning as a principled mechanism for integrating data-driven behaviors with reward-driven adaptation. We formulate the policy customization problem and propose Residual Q-learning, a method that enables policies to adapt to new task objectives without requiring access to the original reward signals. We extend residual Q-learning to policy gradient methods, developing a unified structure that connects data-driven and reward-driven objectives. Finally, we integrate residual Q-learning into model-predictive path integral control, enabling fast, adaptive continuous control by combining learned priors with real-time model-based optimization.Through these, the dissertation advances scalable, adaptable, and robust decision-making frameworks for autonomous systems, bridging the gap between offline learning and real-world deployment.
■590 ▼aSchool code: 0028.
■650 4▼aMechanical engineering
■650 4▼aComputer science
■650 4▼aRobotics
■653 ▼aBehavior modeling
■653 ▼aImitation learning
■653 ▼aPlanning
■653 ▼aReinforcement learning
■690 ▼a0548
■690 ▼a0984
■690 ▼a0771
■690 ▼a0800
■71020▼aUniversity of California, Berkeley▼bMechanical Engineering.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357756▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


