Fetching the paper…
Reading the bibliography…
We consider reinforcement learning (RL) in episodic MDPs with adversarial full-information reward feedback and unknown fixed transition kernels.
Online convex optimization in adversarial Markov decision processes
Aviv Rosenberg and Yishay Mansour · 1905
Earlier work this paper cites.
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 1906
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 1907
Earlier work this paper cites.
Bandit convex optimization in non-stationary environments
Peng Zhao, Guanghui Wang, Lijun Zhang, and Zhi-Hua Zhou · 1907
Earlier work this paper cites.
Understand dynamic regret with switching cost for online decision making
Yawei Zhao, Qian Zhao, Xingxing Zhang, En Zhu, Xinwang Liu, and Jianping Yin · 1911
Earlier work this paper cites.
Learning adversarial mdps with bandit feedback and unknown transition
Chi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra, and Tiancheng Yu · 1912
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Prox-method with rate of convergence o (1/t) for variational inequalities with lipschitz continuous monotone operators and smooth convex-concave saddle point problems
Arkadi Nemirovski · 2004
Earlier work this paper cites.
Experts in a Markov decision process
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2005
Earlier work this paper cites.
Online learning in Markov decision processes with arbitrarily changing rewards and transitions
Jia Yuan Yu and Shie Mannor · 2009
Earlier work this paper cites.
Markov decision processes with arbitrary reward processes
Jia Yuan Yu, Shie Mannor, and Nahum Shimkin · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
The online loop-free stochastic shortest-path problem
Gergely Neu, András György, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Deterministic mdps with adversarial rewards and bandit feedback
Raman Arora, Ofer Dekel, and Ambuj Tewari · 2012
Earlier work this paper cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Gergely Neu, Andras Gyorgy, and Csaba Szepesvári · 2012
Earlier work this paper cites.
Online learning with predictable sequences
Alexander Rakhlin and Karthik Sridharan · 2012
Earlier work this paper cites.
Online learning in Markov decision processes with adversarially chosen transition probability distributions
Yasin Abbasi-Yadkori, Peter L. Bartlett, Varun Kanade, Yevgeny Seldin, and Csaba Szepesvári · 2013
Earlier work this paper cites.
Dynamical models and tracking regret in online convex programming
Eric C. Hall and Rebecca M. Willett · 2013
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
Stochastic multi-armed bandit with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Using response functions to measure strategy strength
Trevor Davis, Neil Burch, and Michael Bowling · 2014
Cited alongside, same era.
Online learning in Markov decision processes with changing cost sequences
Travis Dick, Andras Gyorgy, and Csaba Szepesvari · 2014
Cited alongside, same era.
Non-stationary stochastic optimization
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2015
Cited alongside, same era.
Online convex optimization in dynamic environments
Eric C. Hall and Rebecca M. Willett · 2015
Cited alongside, same era.
Online optimization: Competing with dynamic comparators
Ali Jadbabaie, Alexander Rakhlin, Shahin Shahrampour, and Karthik Sridharan · 2015
Cited alongside, same era.
Continual reinforcement learning with complex synapses
Christos Kaplanis, Murray Shanahan, and Claudia Clopath · 2018
Later among the works it cites.
Efficient contextual bandits in non-stationary worlds
Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Proximal online gradient is optimum for dynamic regret
Yawei Zhao, Shuang Qiu, and Ji Liu · 2018
Later among the works it cites.
Adaptively tracking the best bandit arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2019
Later among the works it cites.
Lilian Besson and Emilie Kaufmann · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E. Schapire · 2015
Cited alongside, same era.
Rl: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J. Russell, Pieter Abbeel, and Anca Dragan · 2016
Cited alongside, same era.
Multi-armed bandits: Competing with optimal sequences
Zohar S. Karnin and Oren Anava · 2016
Cited alongside, same era.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z. Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Cited alongside, same era.
Tracking slowly moving clairvoyant: Optimal dynamic regret of online learning with true and noisy gradient
Tianbao Yang, Lijun Zhang, Rong Jin, and Jinfeng Yi · 2016
Cited alongside, same era.
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
Large scale Markov decision processes with changing rewards
Adrian Rivera Cardoso, He Wang, and Huan Xu · 2019
Later among the works it cites.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Later among the works it cites.
Chelsea Finn, Aravind Rajeswaran, Sham Kakade, and Sergey Levine · 2019
Later among the works it cites.
Corruption robust exploration in episodic reinforcement learning
Thodoris Lykouris, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2019
Later among the works it cites.
Variational regret bounds for reinforcement learning
Ronald Ortner, Pratik Gajane, and Peter Auer · 2019
Later among the works it cites.
Learning to collaborate in Markov decision processes
Goran Radanovic, Rati Devidze, David Parkes, and Adish Singla · 2019
Later among the works it cites.
Prediction in online convex optimization for parametrizable objective functions
Robert Ravier and Vahid Tarokh · 2019
Later among the works it cites.
Multi-point bandit algorithms for nonstationary online nonconvex optimization
Abhishek Roy, Krishnakumar Balasubramanian, Saeed Ghadimi, and Prasant Mohapatra · 2019
Later among the works it cites.
Weighted linear bandits for non-stationary environments
Yoan Russac, Claire Vernade, and Olivier Cappé · 2019
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Closest in time.
Optimistic policy optimization with bandit feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg, and Shie Mannor · 2020
Closest in time.
A survey of reinforcement learning algorithms for dynamically varying environments
Sindhu Padakandla · 2020
Closest in time.
Minimizing dynamic regret and adaptive regret simultaneously
Lijun Zhang, Shiyin Lu, and Tianbao Yang · 2020
Closest in time.