서브메뉴
검색
Fidelity, Fairness and Responsibility Through the Lens of Sequential Decision Making
Fidelity, Fairness and Responsibility Through the Lens of Sequential Decision Making
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211150927
- ISBN
- 9798381947687
- DDC
- 004
- 저자명
- Sun, He.
- 서명/저자
- Fidelity, Fairness and Responsibility Through the Lens of Sequential Decision Making
- 발행사항
- [Sl] : Harvard University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 168 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-09, Section: B.
- 주기사항
- Advisor: Parkes, David.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2024.
- 초록/해제
- 요약As methods of artificial intelligence continue to become increasingly important to support robust decision making in regard to deciding how to act on the basis of the right data, learning to act over time while supporting fairness to participants, and helping individuals make better sequential decisions.This thesis expands in these directions, developing algorithms for enhancing decision-making processes, ensuring fairness in automated decisions, and optimizing user engagement. Motivating settings come from financial time series generation and portfolio optimization, the study of reinforcement learning with fairness constraints in the context of making loans, and the formulation of user engagement optimization in online platforms.First, I introduce the decision-aware time-series conditional generative adversarial network (DAT- CGAN), which is a new method for time-series generation that is aware of the way in which data will be used. In particular, the framework adopts a multi-Wasserstein loss on decision-related quantities and is designed to support decision-making. DAT-CGAN uses an overlapped block-sampling approach for sample efficiency. The main results characterize the generalization properties of DAT-CGAN, and apply to financial time series and a multi-period portfolio choice problem. The proposed method demonstrates better training stability and generative quality in regard to both raw data and decision-related quantities than GAN-based baselines.Second, I introduce the study of reinforcement learning (RL) with stepwise fairness constraints, which requires group fairness at each time step. This problem is motivated by the increasing use of AI methods in societally important settings, ranging from credit to employment to housing, and where it is crucial to provide fairness in regard to automated decision making. Moreover, many such settings are dynamic, with populations responding to sequential decision policies. In the case of tabular episodic RL, I provide a learning algorithm with a strong theoretical guarantee in regard to policy optimality and fairness violations. The experimental results also show that the proposed algorithm outperforms strong learning-based baselines.Third, I formulate and solve a learning problem to handle content recommendation while also learning when to recommend users take a break during a user session. User engagement optimization plays a crucial role in online platforms, with platform designers putting great efforts into recommending interesting content to attract users. At the same time, blindly pushing users to extend a session can lead to burn out and regret, which is harmful to users' long-term well-being. In response, many platforms now provide a service that reminds users to take a break. However, this timing is typically set manually, which motivates an interest in algorithms to automatically pop-out a reminder. Technically, I formulate the problem as an optimal stopping problem for a Markov decision process, and give an offline Q-learning based algorithm with a rigorous theoretical guarantee. I demonstrate the effectiveness of the algorithm on online click-stream data in an online shopping setting.
- 일반주제명
- Computer science
- 키워드
- Fairness
- 키워드
- Optimal stopping
- 키워드
- Time series
- 기타저자
- Harvard University Engineering and Applied Sciences - Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-09B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160180
■00520250211150927
■006m o d
■007cr#unu||||||||
■020 ▼a9798381947687
■035 ▼a(MiAaPQ)AAI30989722
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aSun, He.▼0(orcid)0009-0000-9146-1563
■24510▼aFidelity, Fairness and Responsibility Through the Lens of Sequential Decision Making
■260 ▼a[Sl]▼bHarvard University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a168 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-09, Section: B.
■500 ▼aAdvisor: Parkes, David.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2024.
■520 ▼aAs methods of artificial intelligence continue to become increasingly important to support robust decision making in regard to deciding how to act on the basis of the right data, learning to act over time while supporting fairness to participants, and helping individuals make better sequential decisions.This thesis expands in these directions, developing algorithms for enhancing decision-making processes, ensuring fairness in automated decisions, and optimizing user engagement. Motivating settings come from financial time series generation and portfolio optimization, the study of reinforcement learning with fairness constraints in the context of making loans, and the formulation of user engagement optimization in online platforms.First, I introduce the decision-aware time-series conditional generative adversarial network (DAT- CGAN), which is a new method for time-series generation that is aware of the way in which data will be used. In particular, the framework adopts a multi-Wasserstein loss on decision-related quantities and is designed to support decision-making. DAT-CGAN uses an overlapped block-sampling approach for sample efficiency. The main results characterize the generalization properties of DAT-CGAN, and apply to financial time series and a multi-period portfolio choice problem. The proposed method demonstrates better training stability and generative quality in regard to both raw data and decision-related quantities than GAN-based baselines.Second, I introduce the study of reinforcement learning (RL) with stepwise fairness constraints, which requires group fairness at each time step. This problem is motivated by the increasing use of AI methods in societally important settings, ranging from credit to employment to housing, and where it is crucial to provide fairness in regard to automated decision making. Moreover, many such settings are dynamic, with populations responding to sequential decision policies. In the case of tabular episodic RL, I provide a learning algorithm with a strong theoretical guarantee in regard to policy optimality and fairness violations. The experimental results also show that the proposed algorithm outperforms strong learning-based baselines.Third, I formulate and solve a learning problem to handle content recommendation while also learning when to recommend users take a break during a user session. User engagement optimization plays a crucial role in online platforms, with platform designers putting great efforts into recommending interesting content to attract users. At the same time, blindly pushing users to extend a session can lead to burn out and regret, which is harmful to users' long-term well-being. In response, many platforms now provide a service that reminds users to take a break. However, this timing is typically set manually, which motivates an interest in algorithms to automatically pop-out a reminder. Technically, I formulate the problem as an optimal stopping problem for a Markov decision process, and give an offline Q-learning based algorithm with a rigorous theoretical guarantee. I demonstrate the effectiveness of the algorithm on online click-stream data in an online shopping setting.
■590 ▼aSchool code: 0084.
■650 4▼aComputer science
■653 ▼aFairness
■653 ▼aGenerative adversarial network
■653 ▼aOptimal stopping
■653 ▼aReinforcement learning
■653 ▼aTime series
■690 ▼a0984
■690 ▼a0800
■690 ▼a0796
■71020▼aHarvard University▼bEngineering and Applied Sciences - Computer Science.
■7730 ▼tDissertations Abstracts International▼g85-09B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160180▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


