Fetching the paper…
Reading the bibliography…
Reinforcement learning is a learning paradigm for solving sequential decision-making problems.
R. Bellman, “A markovian decision process,” Journal of mathematics and mechanics , 1957
1957
Earlier work this paper cites.
R. Bellman, “Dynamic programming,” Science , 1966
1966
Earlier work this paper cites.
R. S. Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” in Machine learning proceedings 1990 . Elsevier, 1990, pp. 216–224
1990
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning , 1992
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , 1992
1992
Earlier work this paper cites.
P. Dayan and G. E. Hinton, “Feudal reinforcement learning,” NeurIPS , 1993
1993
Earlier work this paper cites.
R. J. Williams and L. C. Baird, “Tight performance bounds on greedy policies based on imperfect value functions,” Tech. Rep., 1993
1993
Earlier work this paper cites.
P. Dayan, “Improving generalization for temporal difference learning: The successor representation,” Neural Computation , 1993
1993
Earlier work this paper cites.
G. A. Rummery and M. Niranjan, On-line Q-learning using connectionist systems . University of Cambridge, Department of Engineering Cambridge, England, 1994
1994
Earlier work this paper cites.
S. Schaal, “Learning from demonstration,” NeurIPS , 1997
1997
Earlier work this paper cites.
R. Parr and S. J. Russell, “Reinforcement learning with hierarchies of machines,” NeurIPS , 1998
1998
Earlier work this paper cites.
J. Moody, L. Wu, Y. Liao, and M. Saffell, “Performance functions and reinforcement learning for trading systems and portfolios,” Journal of Forecasting , 1998
1998
Earlier work this paper cites.
R. Neuneier, “Enhancing q-learning for optimal asset allocation,” NeurIPS , 1998
1998
Earlier work this paper cites.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence , 1999
1999
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” ICML , 1999
1999
Earlier work this paper cites.
V. Konda and J. Tsitsiklis, “Actor-critic algorithms,” NeurIPS , 2000
2000
Earlier work this paper cites.
T. G. Dietterich, “Hierarchical reinforcement learning with the maxq value function decomposition,” Journal of artificial intelligence research , 2000
2000
Earlier work this paper cites.
S. P. Singh, M. J. Kearns, D. J. Litman, and M. A. Walker, “Reinforcement learning for spoken dialogue systems,” NeurIPS , 2000
2000
Earlier work this paper cites.
E. Wiewiora, G. W. Cottrell, and C. Elkan, “Principled methods for advising reinforcement learning agents,” ICML , 2003
2003
Earlier work this paper cites.
B. Zadrozny, “Learning and evaluating classifiers under sample selection bias,” ICML , 2004
2004
Earlier work this paper cites.
L. Torrey, T. Walker, J. Shavlik, and R. Maclin, “Using advice to transfer knowledge acquired in one reinforcement learning task to another,” European Conference on Machine Learning , 2005
2005
Earlier work this paper cites.
A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Incorporating prior knowledge into q-learning for drug delivery individualization,” Fourth International Conference on Machine Learning and Applications , 2005
2005
Earlier work this paper cites.
F. Fernández and M. Veloso, “Probabilistic policy reuse in a reinforcement learning agent,” Proceedings of the fifth international joint conference on Autonomous agents and multiagent systems , 2006
2006
Earlier work this paper cites.
G. Konidaris and A. Barto, “Autonomous shaping: Knowledge transfer in reinforcement learning,” ICML , 2006
2006
Earlier work this paper cites.
M. E. Taylor, P. Stone, and Y. Liu, “Transfer learning via inter-task mappings for temporal difference learning,” Journal of Machine Learning Research , 2007
2007
Earlier work this paper cites.
A. Lazaric, M. Restelli, and A. Bonarini, “Transfer of samples in batch reinforcement learning,” ICML , 2008
2008
Earlier work this paper cites.
N. Mehta, S. Natarajan, P. Tadepalli, and A. Fern, “Transfer in variable-reward hierarchical reinforcement learning,” Machine Learning , 2008
2008
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering , 2009
2009
Earlier work this paper cites.
M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research , 2009
2009
Earlier work this paper cites.
H. Van Seijen, H. Van Hasselt, S. Whiteson, and M. Wiering, “A theoretical and empirical analysis of expected sarsa,” IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning , 2009
2009
Earlier work this paper cites.
C. H. Lampert, H. Nickisch, and S. Harmeling, “Learning to detect unseen object classes by between-class attribute transfer,” IEEE Conference on Computer Vision and Pattern Recognition , 2009
2009
Earlier work this paper cites.
M. Grzes and D. Kudenko, “Learning shaping rewards in model-based reinforcement learning,” Proc. AAMAS Workshop on Adaptive Learning Agents , 2009
2009
Earlier work this paper cites.
C. Wang and S. Mahadevan, “Manifold alignment without correspondence,” International Joint Conference on Artificial Intelligence , 2009
2009
Earlier work this paper cites.
I. Harvey, “The microbial genetic algorithm,” European Conference on Artificial Life , 2009
2009
Earlier work this paper cites.
B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and autonomous systems , 2009
2009
Earlier work this paper cites.
A. Lazaric and M. Ghavamzadeh, “Bayesian multi-task reinforcement learning,” in ICML-27th international conference on machine learning . Omnipress, 2010, pp. 599–606
2010
Earlier work this paper cites.
A. C. Tenorio-Gonzalez, E. F. Morales, and L. Villaseñor-Pineda, “Dynamic reward shaping: Training a robot by voice,” Advances in Artificial Intelligence – IBERAMIA , 2010
2010
Earlier work this paper cites.
M. Deisenroth and C. E. Rasmussen, “Pilco: A model-based and data-efficient approach to policy search,” in Proceedings of the 28th International Conference on machine learning (ICML-11) , 2011, pp. 465–472
2011
Earlier work this paper cites.
D. P. Bertsekas, “Approximate policy iteration: A survey and some new methods,” Journal of Control Theory and Applications , 2011
2011
Earlier work this paper cites.
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” AISTATS , 2011
2011
Earlier work this paper cites.
A. Lazaric, “Transfer in reinforcement learning: a framework and a survey.” Springer, 2012
2012
Earlier work this paper cites.
S. M. Devlin and D. Kudenko, “Dynamic potential-based reward shaping,” ICAAMAS , 2012
2012
Earlier work this paper cites.
H. B. Ammar and M. E. Taylor, “Reinforcement learning transfer via common subspaces,” Proceedings of the 11th International Conference on Adaptive and Learning Agents , 2012
2012
Earlier work this paper cites.
H. B. Ammar, K. Tuyls, M. E. Taylor, K. Driessens, and G. Weiss, “Reinforcement learning transfer via sparse coding,” ICAAMS , 2012
2012
Earlier work this paper cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” Journal of Artificial Intelligence Research , 2013
2013
Earlier work this paper cites.
S. El-Tantawy, B. Abdulhai, and H. Abdelgawad, “Multiagent reinforcement learning for integrated network of adaptive traffic signal controllers (marlin-atsc): methodology and large-scale application on downtown toronto,” IEEE Transactions on Intelligent Transportation Systems , 2013
2013
Earlier work this paper cites.
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer, “The cross-entropy method for optimization,” in Handbook of statistics . Elsevier, 2013, vol. 31, pp. 35–59
2013
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in International conference on machine learning . PMLR, 2013, pp. 1–9
2013
Earlier work this paper cites.
B. Kim, A.-m. Farahmand, J. Pineau, and D. Precup, “Learning from limited demonstrations,” NeurIPS , 2013
2013
Earlier work this paper cites.
B. Bocsi, L. Csató, and J. Peters, “Alignment-based transfer learning for robot models,” The 2013 International Joint Conference on Neural Networks (IJCNN) , 2013
2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” NeurIPS , pp. 2672–2680, 2014
2014
Earlier work this paper cites.
S. Devlin, L. Yliniemi, D. Kudenko, and K. Tumer, “Potential-based difference rewards for multiagent reinforcement learning,” ICAAMS , 2014
2014
Earlier work this paper cites.
B. Piot, M. Geist, and O. Pietquin, “Boosted bellman residual minimization handling expert demonstrations,” Joint European Conference on Machine Learning and Knowledge Discovery in Databases , 2014
2014
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” Deep Learning and Representation Learning Workshop, NeurIPS , 2014
2014
Earlier work this paper cites.
M. R. Kosorok and E. E. Moodie, Adaptive TreatmentStrategies in Practice: Planning Trials and Analyzing Data for Personalized Medicine . SIAM, 2015
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” ICML , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” ICML , 2015
2015
Earlier work this paper cites.
A. Harutyunyan, S. Devlin, P. Vrancx, and A. Nowé, “Expressing arbitrary reward functions as potential-based advice,” AAAI , 2015
2015
Cited alongside, same era.
T. Brys, A. Harutyunyan, M. E. Taylor, and A. Nowé, “Policy transfer using reward shaping,” ICAAMS , 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
J. Chemali and A. Lazaric, “Direct policy iteration with demonstrations,” International Joint Conference on Artificial Intelligence , 2015
2015
Cited alongside, same era.
T. Brys, A. Harutyunyan, H. B. Suay, S. Chernova, M. E. Taylor, and A. Nowé, “Reinforcement learning from demonstration through shaping,” International Joint Conference on Artificial Intelligence , 2015
Y. Li, J. Song, and S. Ermon, “Infogail: Interpretable imitation learning from visual demonstrations,” NeurIPS , 2017
2017
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Later among the works it cites.
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen, “Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection,” The International Journal of Robotics Research , 2018
2018
Later among the works it cites.
H. Wei, G. Zheng, H. Yao, and Z. Li, “Intellilight: A reinforcement learning approach for intelligent traffic light control,” ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , 2018
2018
Later among the works it cites.
C. Florensa, D. Held, X. Geng, and P. Abbeel, “Automatic goal generation for reinforcement learning agents,” in International conference on machine learning . PMLR, 2018, pp. 1515–1528
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2015
Cited alongside, same era.
H. B. Ammar, E. Eaton, P. Ruvolo, and M. E. Taylor, “Unsupervised cross-domain transfer in policy gradient reinforcement learning via manifold alignment,” AAAI , 2015
2015
Cited alongside, same era.
B. Kehoe, S. Patil, P. Abbeel, and K. Goldberg, “A survey of research on cloud robotics and automation,” IEEE Transactions on automation science and engineering , 2015
2015
Cited alongside, same era.
K.-W. Chang, A. Krishnamurthy, A. Agarwal, J. Langford, and H. Daumé III, “Learning to search better than your teacher,” 2015
2015
Cited alongside, same era.
Z. Wen, D. O’Neill, and H. Maei, “Optimal demand response using device-based reinforcement learning,” IEEE Transactions on Smart Grid , 2015
2015
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , 2016
2016
Cited alongside, same era.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” ICML , 2016
2016
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” International Conference on Machine Learning , 2018
2018
Later among the works it cites.
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver, “Rainbow: Combining improvements in deep reinforcement learning,” AAAI , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 7559–7566
2018
Later among the works it cites.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” NeurIPS , vol. 31, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Barreto, D. Borsa, J. Quan, T. Schaul, D. Silver, M. Hessel, D. Mankowitz, A. Žídek, and R. Munos, “Transfer in deep reinforcement learning using successor features and generalised policy improvement,” ICML , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
O. Marom and B. Rosman, “Belief reward shaping in reinforcement learning,” AAAI , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband et al. , “Deep q-learning from demonstrations,” AAAI , 2018
2018
Later among the works it cites.
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Overcoming exploration in reinforcement learning with demonstrations,” IEEE International Conference on Robotics and Automation (ICRA) , 2018
2018
Later among the works it cites.
B. Kang, Z. Jie, and J. Feng, “Policy optimization with demonstrations,” ICML , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
K. Bousmalis, A. Irpan, P. Wohlhart, Y. Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige et al. , “Using simulation and domain adaptation to improve efficiency of deep robotic grasping,” IEEE International Conference on Robotics and Automation (ICRA) , 2018
2018
Later among the works it cites.
A. Serrano, B. Imbernón, H. Pérez-Sánchez, J. M. Cecilia, A. Bueno-Crespo, and J. L. Abellán, “Accelerating drugs discovery with deep reinforcement learning: An early approach,” International Conference on Parallel Processing Companion , 2018
2018
Later among the works it cites.
M. Popova, O. Isayev, and A. Tropsha, “Deep reinforcement learning for de novo drug design,” Science advances , 2018
2018
Later among the works it cites.
K. Lin, R. Zhao, Z. Xu, and J. Zhou, “Efficient large-scale fleet management via multi-agent deep reinforcement learning,” ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2018
2018
Later among the works it cites.
W. Czarnecki, R. Pascanu, S. Osindero, S. Jayakumar, G. Swirszcz, and M. Jaderberg, “Distilling policy distillation,” The 22nd International Conference on Artificial Intelligence and Statistics , 2019
2019
Later among the works it cites.
C. Finn and S. Levine, “Meta-learning: from few-shot learning to rapid reinforcement learning,” ICML , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Ma, Y.-X. Wang, and B. Narayanaswamy, “Imitation-regularized offline learning,” International Conference on Artificial Intelligence and Statistics , 2019
2019
Later among the works it cites.
K. Brantley, W. Sun, and M. Henaff, “Disagreement-regularized imitation learning,” ICLR , 2019
2019
Later among the works it cites.
D. Borsa, A. Barreto, J. Quan, D. Mankowitz, R. Munos, H. van Hasselt, D. Silver, and T. Schaul, “Universal successor features approximators,” ICLR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Bharadhwaj, Z. Wang, Y. Bengio, and L. Paull, “A data-efficient framework for training and sim-to-real transfer of navigation policies,” International Conference on Robotics and Automation (ICRA) , 2019
2019
Later among the works it cites.
OpenAI. (2019) Dotal2 blog. [Online]. Available: https://openai.com/blog/openai-five/
2019
Later among the works it cites.
N. Justesen, P. Bontrager, J. Togelius, and S. Risi, “Deep learning for video game playing,” IEEE Transactions on Games , 2019
2019
Later among the works it cites.
F. Godin, A. Kumar, and A. Mittal, “Learning when not to answer: a ternary reward structure for reinforcement learning based question answering,” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Alansary, O. Oktay, Y. Li, L. Le Folgoc, B. Hou, G. Vaillant, K. Kamnitsas, A. Vlontzos, B. Glocker, B. Kainz et al. , “Evaluating reinforcement learning agents for anatomical landmark detection,” 2019
2019
Later among the works it cites.
Z. Xu and A. Tewari, “Reinforcement learning in factored mdps: Oracle-efficient algorithms and tighter regret bounds for the non-episodic setting,” NeurIPS , vol. 33, pp. 18 226–18 236, 2020
2020
Closest in time.
H. Bharadhwaj, K. Xie, and F. Shkurti, “Model-predictive control via cross-entropy and gradient-based optimization,” in Learning for Dynamics and Control . PMLR, 2020, pp. 277–286
2020
Closest in time.
R. Yang, H. Xu, Y. Wu, and X. Wang, “Multi-task reinforcement learning with soft modularization,” NeurIPS , vol. 33, pp. 4767–4777, 2020
2020
Closest in time.
2020
Closest in time.
Z. Zhu, K. Lin, B. Dai, and J. Zhou, “Off-policy imitation learning from observations,” NeurIPS , 2020
2020
Closest in time.
W. Zhao, J. P. Queralta, and T. Westerlund, “Sim-to-real transfer in deep reinforcement learning for robotics: a survey,” in 2020 IEEE symposium series on computational intelligence (SSCI) . IEEE, 2020, pp. 737–744
2020
Closest in time.
N. Vithayathil Varghese and Q. H. Mahmoud, “A survey of multi-task deep reinforcement learning,” Electronics , vol. 9, no. 9, p. 1363, 2020
2020
Closest in time.
K. Kim, Y. Gu, J. Song, S. Zhao, and S. Ermon, “Domain adaptive imitation learning,” ICML , 2020
2020
Closest in time.
M. Jing, X. Ma, W. Huang, F. Sun, C. Yang, B. Fang, and H. Liu, “Reinforcement learning from imperfect demonstrations under soft expert guidance.” AAAI , 2020
2020
Closest in time.
Y. Zhang and Q. Yang, “A survey on multi-task learning,” IEEE Transactions on Knowledge and Data Engineering , vol. 34, no. 12, pp. 5586–5609, 2021
2021
Closest in time.
T. Hospedales, A. Antoniou, P. Micaelli, and A. Storkey, “Meta-learning in neural networks: A survey,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 9, pp. 5149–5169, 2021
2021
Closest in time.
M. Muller-Brockhausen, M. Preuss, and A. Plaat, “Procedural content generation: Better benchmarks for transfer reinforcement learning,” in 2021 IEEE Conference on games (CoG) . IEEE, 2021, pp. 01–08
2021
Closest in time.
2021
Closest in time.
2022
Closest in time.
C. Yu, A. Velu, E. Vinitsky, J. Gao, Y. Wang, A. Bayen, and Y. Wu, “The surprising effectiveness of ppo in cooperative multi-agent games,” NeurIPS , vol. 35, pp. 24 611–24 624, 2022
2022
Closest in time.
Z. Jia, X. Li, Z. Ling, S. Liu, Y. Wu, and H. Su, “Improving policy optimization with generalist-specialist learning,” in International Conference on Machine Learning . PMLR, 2022, pp. 10 104–10 119
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
R. Kirk, A. Zhang, E. Grefenstette, and T. Rocktäschel, “A survey of zero-shot generalisation in deep reinforcement learning,” Journal of Artificial Intelligence Research , vol. 76, pp. 201–264, 2023
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” arXiv , 2023
2023
Closest in time.