서브메뉴
검색
Data-Efficient Decision-Making- [electronic resource]
Data-Efficient Decision-Making- [electronic resource]
Detailed Information
- 자료유형
- 학위논문파일 국외
- 최종처리일시
- 20240214100126
- ISBN
- 9798379711665
- DDC
- 310
- 저자명
- Hu, Yichun.
- 서명/저자
- Data-Efficient Decision-Making - [electronic resource]
- 발행사항
- [S.l.]: : Cornell University., 2023
- 발행사항
- Ann Arbor : : ProQuest Dissertations & Theses,, 2023
- 형태사항
- 1 online resource(322 p.)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 84-12, Section: B.
- 주기사항
- Advisor: Kallus, Nathan.
- 학위논문주기
- Thesis (Ph.D.)--Cornell University, 2023.
- 사용제한주기
- This item must not be sold to any third party vendors.
- 초록/해제
- 요약This thesis is focused on the development of sample-efficient algorithms for personalized data-driven decision-making. In particular, the dissertation aims to address the following questions in both online (sequential) and offline (batch) settings: (i) What problem structures allow for achieving instance-specific fast regret rates? (ii) How can these problem structures be leveraged to design practical algorithms that achieve fast theoretical rates?Part I of this thesis investigates the above questions from an online perspective. Chapter 2 studies the smooth contextual bandit problem, where we use the smoothness property of the function class to design contextual bandit algorithms that interpolate between two extremes previously studied in isolation: nondifferentiable bandits and parametric-response bandits. Chapter 3 examines the DTR bandit problem, where we develop the first online algorithm with logarithmic regret for dynamic treatment regimes that involve personalized, adaptive, multi-stage treatment plans.Part II of this work delves into fast regret rates for offline problems by leveraging a probabilistic condition that measures the distribution of the reward gap between the optimal and second-optimal decisions, which we term the margin condition. In the case of contextual linear optimization, Chapter 4 shows that the naive plug-in approach actually achieves regret convergence rates that are significantly faster than methods that directly optimize downstream decision performance. In the case of offline reinforcement learning, Chapter 5 presents a finer regret analysis that characterizes the faster-than-square-root regret convergence rate we observe in practice.
- 일반주제명
- Statistics.
- 일반주제명
- Computer science.
- 키워드
- Data
- 키워드
- Decision-making
- 키워드
- Algorithms
- 기타저자
- Cornell University Operations Research and Information Engineering
- 기본자료저록
- Dissertations Abstracts International. 84-12B.
- 기본자료저록
- Dissertation Abstract International
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008240612s2023 us |||||||||||||||c||eng d■001000016931842
■00520240214100126
■006m o d
■007cr#unu||||||||
■020 ▼a9798379711665
■035 ▼a(MiAaPQ)AAI30425306
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aHu, Yichun.▼0(orcid)0000-0002-5826-9665
■24510▼aData-Efficient Decision-Making▼h[electronic resource]
■260 ▼a[S.l.]:▼bCornell University. ▼c2023
■260 1▼aAnn Arbor :▼bProQuest Dissertations & Theses, ▼c2023
■300 ▼a1 online resource(322 p.)
■500 ▼aSource: Dissertations Abstracts International, Volume: 84-12, Section: B.
■500 ▼aAdvisor: Kallus, Nathan.
■5021 ▼aThesis (Ph.D.)--Cornell University, 2023.
■506 ▼aThis item must not be sold to any third party vendors.
■520 ▼aThis thesis is focused on the development of sample-efficient algorithms for personalized data-driven decision-making. In particular, the dissertation aims to address the following questions in both online (sequential) and offline (batch) settings: (i) What problem structures allow for achieving instance-specific fast regret rates? (ii) How can these problem structures be leveraged to design practical algorithms that achieve fast theoretical rates?Part I of this thesis investigates the above questions from an online perspective. Chapter 2 studies the smooth contextual bandit problem, where we use the smoothness property of the function class to design contextual bandit algorithms that interpolate between two extremes previously studied in isolation: nondifferentiable bandits and parametric-response bandits. Chapter 3 examines the DTR bandit problem, where we develop the first online algorithm with logarithmic regret for dynamic treatment regimes that involve personalized, adaptive, multi-stage treatment plans.Part II of this work delves into fast regret rates for offline problems by leveraging a probabilistic condition that measures the distribution of the reward gap between the optimal and second-optimal decisions, which we term the margin condition. In the case of contextual linear optimization, Chapter 4 shows that the naive plug-in approach actually achieves regret convergence rates that are significantly faster than methods that directly optimize downstream decision performance. In the case of offline reinforcement learning, Chapter 5 presents a finer regret analysis that characterizes the faster-than-square-root regret convergence rate we observe in practice.
■590 ▼aSchool code: 0058.
■650 4▼aStatistics.
■650 4▼aComputer science.
■653 ▼aData
■653 ▼aDecision-making
■653 ▼aAlgorithms
■653 ▼aFast regret rates
■653 ▼aContextual bandit algorithms
■690 ▼a0796
■690 ▼a0463
■690 ▼a0984
■71020▼aCornell University▼bOperations Research and Information Engineering.
■7730 ▼tDissertations Abstracts International▼g84-12B.
■773 ▼tDissertation Abstract International
■790 ▼a0058
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T16931842▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
■980 ▼a202402▼f2024
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
detalle info
- Reserva
- No existe
- Mi carpeta
- Primera solicitud
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.
![Data-Efficient Decision-Making - [electronic resource]](/Users/Baul/Images/book.png)

