Fetching the paper…
Reading the bibliography…
We study Policy-extended Value Function Approximator (PeVFA) in Reinforcement Learning (RL), which extends conventional value function approximator (VFA) to take as input not only the state (and action) but also an explicit policy representation.
Neuro-dynamic programming , volume 3 of Optimization and neural computation series
Bertsekas, D. P.; and Tsitsiklis, J. N. 1996 · 1996
Earlier work this paper cites.
Reinforcement learning - an introduction
Sutton, R. S.; and Barto, A. G. 1998 · 1998
Earlier work this paper cites.
A Simple Framework for Contrastive Learning of Visual Representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. E. 2020 · 2002
Earlier work this paper cites.
Harb, J.; Schaul, T.; Precup, D.; and Bacon, P. 2020 · 2002
Earlier work this paper cites.
Predicting Neural Network Accuracy from Weights
Unterthiner, T.; Keysers, D.; Gelly, S.; Bousquet, O.; and Tolstikhin, I. O. 2020 · 2002
Earlier work this paper cites.
Least-Squares Policy Iteration
Lagoudakis, M. G.; and Parr, R. 2003 · 2003
Earlier work this paper cites.
Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
Kostrikov, I.; Yarats, D.; and Fergus, R. 2020 · 2004
Earlier work this paper cites.
Reinforcement Learning with Augmented Data
Laskin, M.; Lee, K.; Stooke, A.; Pinto, L.; Abbeel, P.; and Srinivas, A. 2020 · 2004
Earlier work this paper cites.
CURL: Contrastive Unsupervised Representations for Reinforcement Learning
Srinivas, A.; Laskin, M.; and Abbeel, P. 2020 · 2004
Earlier work this paper cites.
Context-aware Dynamics Model for Generalization in Model-Based Reinforcement Learning
Lee, K.; Seo, Y.; Lee, S.; Lee, H.; and Shin, J. 2020 · 2005
Earlier work this paper cites.
What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Andrychowicz, M.; Raichuk, A.; Stanczyk, P.; Orsini, M.; Girgin, S.; Marinier, R.; Hussenot, L.; Geist, M.; Pietquin, O.; Michalski, M.; Gelly, S.; and Bachem, O. 2020 · 2006
Earlier work this paper cites.
Parameter-based Value Functions
Faccio, F.; and Schmidhuber, J. 2020 · 2006
Earlier work this paper cites.
The Impact of Non-stationarity on Generalisation in Deep Reinforcement Learning
Igl, M.; Farquhar, G.; Luketina, J.; Boehmer, W.; and Whiteson, S. 2020 · 2006
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Nesterov, Y. E.; and Polyak, B. T. 2006 · 2006
Earlier work this paper cites.
Demystifying Self-Supervised Learning: An Information-Theoretical Framework
Tsai, Y. H.; Wu, Y.; Salakhutdinov, R.; and Morency, L. 2020 · 2006
Earlier work this paper cites.
Fast Adaptation via Policy-Dynamics Value Functions
Raileanu, R.; Goldstein, M.; Szlam, A.; and Fergus, R. 2020 · 2007
Earlier work this paper cites.
Data-Efficient Reinforcement Learning with Momentum Predictive Representations
Schwarzer, M.; Anand, A.; Goel, R.; Hjelm, R. D.; Courville, A. C.; and Bachman, P. 2020 · 2007
Earlier work this paper cites.
Visualizing Data using t-SNE
Maaten, L. V. D.; and Hinton, G. E. 2008 · 2008
Earlier work this paper cites.
Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive Learning
Fu, H.; Tang, H.; Hao, J.; Chen, C.; Feng, X.; Li, D.; and Liu, W. 2020 · 2009
Earlier work this paper cites.
Double Q-learning
v. Hasselt, H. 2010 · 2010
Earlier work this paper cites.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S.; Modayil, J.; Delp, M.; Degris, T.; Pilarski, P. M.; White, A.; and Precup, D. 2011 · 2011
Earlier work this paper cites.
Degris, T.; White, M.; and Sutton, R. S. 2012 · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P.; and Welling, M. 2014 · 2014
Earlier work this paper cites.
Deterministic Policy Gradient Algorithms
Silver, D.; Lever, G.; Heess, N.; Degris, T.; Wierstra, D.; and Riedmiller, M. A. 2014 · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M. A.; Fidjeland, A.; Ostrovski, G.; Petersen, S.; Beattie, C.; Sadik, A.; Antonoglou, I.; King, H.; Kumaran, D.; Wierstra, D.; Legg, S.; and Hassabis, D. 2015 · 2015
Cited alongside, same era.
Universal Value Function Approximators
Schaul, T.; Horgan, D.; Gregor, K.; and Silver, D. 2015 · 2015
Cited alongside, same era.
Approximate modified policy iteration and its application to the game of Tetris
Scherrer, B.; Ghavamzadeh, M.; Gabillon, V.; Lesner, B.; and Geist, M. 2015 · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Identifying Generalization Properties in Neural Networks
Wang, H.; Keskar, N. S.; Xiong, C.; and Socher, R. 2018 · 2018
Later among the works it cites.
Graph Convolutional Policy Network for Goal-Directed Molecular Graph Generation
You, J.; Liu, B.; Ying, Z.; Pande, V. S.; and Leskovec, J. 2018 · 2018
Later among the works it cites.
VPE: Variational Policy Embedding for Transfer Reinforcement Learning
Arnekvist, I.; Kragic, D.; and Stork, J. A. 2019 · 2019
Later among the works it cites.
Learning Multi-Level Hierarchies with Hindsight
Levy, A.; Konidaris, G. D.; Jr., R. P.; and Saenko, K. 2019 · 2019
Later among the works it cites.
DARTS: Differentiable Architecture Search
Liu, H.; Simonyan, K.; and Yang, Y. 2019 · 2019
Later among the works it cites.
Near-Optimal Representation Learning for Hierarchical Reinforcement Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M. I.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Opponent Modeling in Deep Reinforcement Learning
He, H.; and Boyd-Graber, J. L. 2016 · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T. P.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J.; Moritz, P.; Levine, S.; Jordan, M. I.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; van den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; Dieleman, S.; Grewe, D.; Nham, J.; Kalchbrenner, N.; Sutskever, I.; Lillicrap, T. P.; Leach, M.; Kavukcuoglu, K.; Graepel, T.; and Hassabis, D. 2016 · 2016
Cited alongside, same era.
Hindsight Experience Replay
Andrychowicz, M.; Crow, D.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Abbeel, P.; and Zaremba, W. 2017 · 2017
Cited alongside, same era.
Nachum, O.; Gu, S.; Lee, H.; and Levine, S. 2019 · 2019
Later among the works it cites.
Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
Rakelly, K.; Zhou, A.; Finn, C.; Levine, S.; and Quillen, D. 2019 · 2019
Later among the works it cites.
Learning retrosynthetic planning through simulated experience
Schreck, J. S.; Coley, C. W.; and Bishop, K. J. 2019 · 2019
Later among the works it cites.
Relational Forward Models for Multi-Agent Learning
Tacchetti, A.; Song, H. F.; Mediano, P. A. M.; Zambaldi, V. F.; Kramár, J.; Rabinowitz, N. C.; Graepel, T.; Botvinick, M.; and Battaglia, P. W. 2019 · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O.; Babuschkin, I.; Czarnecki, W. M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D. H.; Powell, R.; Ewalds, T.; Georgiev, P.; Oh, J.; Horgan, D.; Kroiss, M.; Danihelka, I.; Huang, A.; Sifre, L.; Cai, T.; Agapiou, J. P.; Jaderberg, M.; Vezhnevets, A. S.; Leblond, R.; Pohlen, T.; Dalibard, V.; Budden, D.; Sulsky, Y.; Molloy, J.; Paine, T. L.; Gulcehre, C.; Wang, Z.; Pfaff, T.; Wu, Y.; Ring, R.; Yogatama, D.; Wünsch, D.; McKinney, K.; Smith, O.; Schaul, T.; Lillicrap, T.; Kavukcuoglu, K.; Hassabis, D.; Apps, C.; and Silver, D. 2019 · 2019
Later among the works it cites.
How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy Optimization
D’Oro, P.; and Jaskowski, W. 2020 · 2020
Closest in time.
Implementation Matters in Deep RL: A Case Study on PPO and TRPO
Engstrom, L.; Ilyas, A.; Santurkar, S.; Tsipras, D.; Janoos, F.; Rudolph, L.; and Madry, A. 2020 · 2020
Closest in time.
Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
Eysenbach, B.; Geng, X.; Levine, S.; and Salakhutdinov, R. R. 2020 · 2020
Closest in time.
Discovering representations for black-box optimization
Gaier, A.; Asteroth, A.; and Mouret, J. 2020 · 2020
Closest in time.
Dream to Control: Learning Behaviors by Latent Imagination
Hafner, D.; Lillicrap, T. P.; Ba, J.; and Norouzi, M. 2020 · 2020
Closest in time.
Momentum Contrast for Unsupervised Visual Representation Learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. B. 2020 · 2020
Closest in time.
Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics
Kuznetsov, A.; Shvechikov, P.; Grishin, A.; and Vetrov, D. P. 2020 · 2020
Closest in time.
Maxmin Q-learning: Controlling the Estimation Bias of Q-learning
Lan, Q.; Pan, Y.; Fyshe, A.; and White, M. 2020 · 2020
Closest in time.
I 2 HRL: Interactive Influence-based Hierarchical Reinforcement Learning
Wang, R.; Yu, R.; An, B.; and Rabinovich, Z. 2020 · 2020
Closest in time.
High-Dimensional Bayesian Optimisation with Variational Autoencoders and Deep Metric Learning
Grosnit, A.; Tutunov, R.; Maraval, A.; Griffiths, R.; Cowen-Rivers, A.; Yang, L.; Zhu, L.; Lyu, W.; Chen, Z.; Wang, J.; Peters, J.; and Bou-Ammar, H. 2021 · 2021
Closest in time.
Improving black-box optimization in VAE latent space using decoder uncertainty
Notin, P.; Hernández-Lobato, J.; and Gal, Y. 2021 · 2021
Closest in time.
Policy Manifold Search: Exploring the Manifold Hypothesis for Diversity-based Neuroevolution
Rakicevic, N.; Cully, A.; and Kormushev, P. 2021 · 2021
Closest in time.
Phasic Policy Gradient
Cobbe, K.; Hilton, J.; Klimov, O.; and Schulman, J. 2021 · 2027
Closest in time.