RT DW A1 Hong, Kihyuk. T1 Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings