Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) provides a theoretical framework for continuously improving an agent's behavior via trial and error.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Complexity analysis of real-time reinforcement learning
Koenig, S. and Simmons, R. G · 1993
Earlier work this paper cites.
Learning from demonstration
Schaal, S. et al · 1997
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Effective reinforcement learning for mobile robots
Smart, W. and Pack Kaelbling, L · 2002
Earlier work this paper cites.
Policy search by dynamic programming
Bagnell, J., Kakade, S. M., Schneider, J., and Ng, A · 2003
Earlier work this paper cites.
Learning decisions: Robustness, uncertainty, and approximation
Bagnell, J. A · 2004
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
Langford, J. and Zhang, T · 2007
Earlier work this paper cites.
Variational policy gradient method for reinforcement learning with general utilities
Zhang, J., Koppel, A., Bedi, A. S., Szepesvari, C., and Wang, M · 2007
Earlier work this paper cites.
Imitation and reinforcement learning for motor primitives with perceptual coupling
Kober, J., Mohler, B., and Peters, J · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R · 2011
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Osband, I. and Van Roy, B · 2014
Earlier work this paper cites.
Approximate policy iteration schemes: A comparison
Scherrer, B · 2014
Earlier work this paper cites.
Playing atari games with deep reinforcement learning and human checkpoint replay
Hosu, I.-A. and Rebedea, T · 2016
Earlier work this paper cites.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Earlier work this paper cites.
Learning unknown markov decision processes: A thompson sampling approach
Ouyang, Y., Gagrani, M., Nayyar, A., and Jain, R · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2017
Cited alongside, same era.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Agarwal, A., Henaff, M., Kakade, S., and Sun, W · 2020
Later among the works it cites.
D4RL: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Batch policy learning in average reward Markov decision processes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M · 2018
Cited alongside, same era.
Backplay:” man muss immer umkehren”
Resnick, C., Raileanu, R., Kapoor, S., Peysakhovich, A., Cho, K., and Bruna, J · 2018
Cited alongside, same era.
Learning montezuma’s revenge from a single demonstration
Salimans, T. and Chen, R · 2018
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards, 2018
Vecerik, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N · 2019
Cited alongside, same era.
Barc: Backward reachability curriculum for robotic reinforcement learning
Ivanovic, B., Harrison, J., Sharma, A., Chen, M., and Pavone, M · 2019
Cited alongside, same era.
On value functions and the agent-environment boundary
Jiang, N · 2019
Cited alongside, same era.
Liao, P., Qi, Z., and Murphy, S · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
Simchi-Levi, D. and Xu, Y · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Zanette, A., Lazaric, A., Kochenderfer, M., and Brunskill, E · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
First return, then explore
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Later among the works it cites.
Aw-opt: Learning robotic skills with imitation andreinforcement at scale
Lu, Y., Hausman, K., Chebotar, Y., Yan, M., Jang, E., Herzog, A., Xiao, T., Irpan, A., Khansari, M., Kalashnikov, D., and Levine, S · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y · 2021
Later among the works it cites.
Zheng, Q., Zhang, A., and Grover, A · 2022
Closest in time.