Fetching the paper…
Reading the bibliography…
Sampled environment transitions are a critical input to deep reinforcement learning (DRL) algorithms.
Replab: A reproducible low-cost arm benchmark platform for robotic learning
Brian Yang et al · 1905
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner et al · 1912
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng et al · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare et al · 2013
Earlier work this paper cites.
Algorithmic progress in six domains
Katja Grace · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih et al · 2013
Earlier work this paper cites.
The ingredients of real-world robotic reinforcement learning
Henry Zhu et al · 2013
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess et al · 2015
Earlier work this paper cites.
State of the art control of atari games using shallow reinforcement learning
Yitao Liang et al · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap et al · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih et al · 2015
Earlier work this paper cites.
Tom Schaul et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman et al · 2015
Earlier work this paper cites.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Earlier work this paper cites.
Greg Brockman et al · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan et al · 2016
Earlier work this paper cites.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu et al · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih et al · 2016
Earlier work this paper cites.
Combining policy gradient and q-learning
Brendan O’Donoghue et al · 2016
Earlier work this paper cites.
End-to-end deep reinforcement learning for lane keeping assist
Ahmad El Sallab et al · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado P. van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
EFF AI progress measurement project
Peter Eckersley and Yomna Nasser · 2017
Earlier work this paper cites.
Noisy networks for exploration
Meire Fortunato et al · 2017
Earlier work this paper cites.
The future of employment: How susceptible are jobs to computerisation?
Carl Benedikt Frey and Michael A. Osborne · 2017
Earlier work this paper cites.
The reactor: A sample-efficient actor-critic architecture
Audrunas Gruslys et al · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja et al · 2017
Earlier work this paper cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Riashat Islam et al · 2017
Earlier work this paper cites.
Trust-pcl: An off-policy trust region method for continuous control
Ofir Nachum et al · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman et al · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver et al · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani et al · 2017
Earlier work this paper cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu et al · 2017
Earlier work this paper cites.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron et al · 2018
Cited alongside, same era.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Miles Brundage et al · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Jacob Buckman et al · 2018
Cited alongside, same era.
Expected policy gradients for reinforcement learning
Kamil Ciosek and Shimon Whiteson · 2018
Cited alongside, same era.
Model-based reinforcement learning via meta-policy optimization
Ignasi Clavera et al · 2018
Fixing the train-test resolution discrepancy
Hugo Touvron et al · 2019
Later among the works it cites.
Sim-to-real transfer learning using robustified controllers in robotic tasks involving complex dynamics
Jeroen van Baar et al · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
Hado P. van Hasselt, Matteo Hessel, and John Aslanides · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals et al · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Tingwu Wang and Jimmy Ba · 2019
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt et al · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Clipped action policy gradient
Yasuhiro Fujita and Shin-ichi Maeda · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja et al · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson et al · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel et al · 2018
Cited alongside, same era.
Distributed prioritized experience replay
Dan Horgan et al · 2018
Cited alongside, same era.
Denis Yarats et al · 2019
Later among the works it cites.
Dac: The double actor-critic architecture for learning options
Shangtong Zhang and Shimon Whiteson · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Later among the works it cites.
Robel: Robotics benchmarks for learning with low-cost robots
Michael Ahn et al · 2020
Later among the works it cites.
AI and compute
Dario Amodei and Danny Hernandez · 2020
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
Adrià Puigdomènech Badia et al · 2020
Later among the works it cites.
Regularizing model-based planning with energy-based models
Rinu Boney, Juho Kannala, and Alexander Ilin · 2020
Later among the works it cites.
Dawnbench: An end-to-end deep learning benchmark and competition
Cody Coleman et al · 2020
Later among the works it cites.
Convergence rates of accelerated markov gradient descent with applications in reinforcement learning
Thinh T. Doan et al · 2020
Later among the works it cites.
A time leap challenge for sat solving
Johannes K. Fichte, Markus Hecher, and Stefan Szeider · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu et al · 2020
Later among the works it cites.
Monte-carlo tree search as regularized policy optimization
Jean-Bastien Grill et al · 2020
Later among the works it cites.
Measuring the algorithmic efficiency of neural networks
Danny Hernandez and Tom B. Brown · 2020
Later among the works it cites.
AI evaluation: On broken yardsticks and measurement scales
Jose Hernandez-Orallo · 2020
Later among the works it cites.
Let’s discuss OpenAI’s rubik’s cube result
Alex Irpan · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan et al · 2020
Later among the works it cites.
Do recent advancements in model-based deep reinforcement learning really improve data efficiency?
Kacper Kielak · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Later among the works it cites.
Reinforcement learning with augmented data
Michael Laskin et al · 2020
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee et al · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine et al · 2020
Later among the works it cites.
Leveraging exploration in off-policy algorithms via normalizing flows
Bogdan Mazoure et al · 2020
Later among the works it cites.
Dreaming: Model-based reinforcement learning by latent imagination without reconstruction
Masashi Okada and Tadahiro Taniguchi · 2020
Later among the works it cites.
Learning dexterous in-hand manipulation
OpenAI et al · 2020
Later among the works it cites.
Spot mini mini
Maurice Rahme · 2020
Later among the works it cites.
Adaptive trade-offs in off-policy learning
Mark Rowland, Will Dabney, and Rémi Munos · 2020
Later among the works it cites.
Data-efficient reinforcement learning with momentum predictive representations
Max Schwarzer et al · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar et al · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Aravind Srinivas, Michael Laskin, and Pieter Abbeel · 2020
Later among the works it cites.
Evolve to control: Evolution-based soft actor-critic for scalable reinforcement learning
Karush Suri et al · 2020
Later among the works it cites.
Social and governance implications of improved data efficiency
Aaron D. Tucker, Markus Anderljung, and Allan Dafoe · 2020
Later among the works it cites.
An AI just beat a human F-16 pilot in a dogfight - again
Patrick Tucker · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Zhongwen Xu et al · 2020
Later among the works it cites.
Policy search by target distribution learning for continuous control
Chuheng Zhang, Yuanqi Li, and Jian Li · 2020
Later among the works it cites.