Fetching the paper…
Reading the bibliography…
Despite the recent success of reinforcement learning in various domains, these approaches remain, for the most part, deterringly sensitive to hyper-parameters and are often riddled with essential engineering feats allowing their success.
On the Theory of the Brownian Motion
G E Uhlenbeck and L S Ornstein · 1930
Earlier work this paper cites.
Some Aspects of the Sequential Design of Experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Sample Selection Bias as a Specification Error
James J Heckman · 1979
Earlier work this paper cites.
Incremental Learning from Noisy Data
Jeffrey C Schlimmer and Richard H Granger, Jr · 1986
Earlier work this paper cites.
A massively parallel architecture for a self-organizing neural pattern recognition machine
Gail A Carpenter and Stephen Grossberg · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
ALVINN: An Autonomous Land Vehicle in a Neural Network
Dean Pomerleau · 1989
Earlier work this paper cites.
Learning from Delayed Rewards
Christopher J C H Watkins · 1989
Earlier work this paper cites.
Rapidly Adapting Artificial Neural Networks for Autonomous Navigation
Dean Pomerleau · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Technical Note: Q-Learning
Christopher J C H Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Learning to Achieve Goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
Functional approximation by feed-forward networks: a least-squares approach to generalization
A R Webb · 1994
Earlier work this paper cites.
The non-stochastic multi-armed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 1995
Earlier work this paper cites.
Training with Noise is Equivalent to Tikhonov Regularization
Chris M Bishop · 1995
Earlier work this paper cites.
Incremental Multi-Step Q-Learning
Jing Peng and Ronald J Williams · 1996
Earlier work this paper cites.
Robot learning from demonstration
Christopher G Atkeson and Stefan Schaal · 1997
Earlier work this paper cites.
Reinforcement Learning with Time
Daishi Harada · 1997
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning from Demonstration
Stefan Schaal · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Apprenticeship Learning via Inverse Reinforcement Learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Experts in a Markov Decision Process
Eyal Even-dar, Sham M Kakade, and Yishay Mansour · 2005
Earlier work this paper cites.
Robust Control of Markov Decision Processes with Uncertain Transition Matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Dealing with Non-Stationary Environments using Context Detection
Bruno C Da Silva, Eduardo W Basso, Ana L C Bazzan, and Paulo M Engel · 2006
Earlier work this paper cites.
Imitation learning for locomotion and manipulation
Nathan Ratliff, J Andrew Bagnell, and Siddhartha S Srinivasa · 2007
Earlier work this paper cites.
The Robustness-Performance Tradeoff in Markov Decision Processes
Huan Xu and Shie Mannor · 2007
Earlier work this paper cites.
Robot Programming by Demonstration
Aude Billard, Sylvain Calinon, Rüdiger Dillmann, and Stefan Schaal · 2008
Earlier work this paper cites.
Apprenticeship Learning Using Linear Programming
Umar Syed, Michael Bowling, and Robert E Schapire · 2008
Earlier work this paper cites.
A Game-Theoretic Approach to Apprenticeship Learning
Umar Syed and Robert E Schapire · 2008
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Where Do Rewards Come From?
Satinder Singh, Richard L Lewis, and Andrew G Barto · 2009
Earlier work this paper cites.
Arbitrarily modulated Markov decision processes
J Y Yu and S Mannor · 2009
Earlier work this paper cites.
Online learning in Markov decision processes with arbitrarily changing rewards and transitions
J Y Yu and S Mannor · 2009
Earlier work this paper cites.
Model-Free Monte Carlo–like Policy Evaluation
Raphael Fonteneau, Susan A Murphy, Louis Wehenkel, and Damien Ernst · 2010
Earlier work this paper cites.
Near-optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Efficient Reductions for Imitation Learning
Stéphane Ross and J Andrew Bagnell · 2010
Earlier work this paper cites.
Double Q-learning
Hado van Hasselt · 2010
Earlier work this paper cites.
On Upper-Confidence Bound Policies for Switching Bandit Problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
Reinforcement learning in feedback control
Roland Hafner and Martin Riedmiller · 2011
Earlier work this paper cites.
Curriculum learning for motor skills
Andrej Karpathy and Michiel Van De Panne · 2012
Earlier work this paper cites.
Batch Reinforcement Learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Online Learning in Markov Decision Processes with Adversarially Chosen Transition Probability Distributions
Yasin Abbasi-Yadkori, Peter L Bartlett, and Csaba Szepesvari · 2013
Earlier work this paper cites.
Batch Mode Reinforcement Learning based on the Synthesis of Artificial Trajectories
Raphael Fonteneau, Susan A Murphy, Louis Wehenkel, and Damien Ernst · 2013
Earlier work this paper cites.
Reinforcement Learning in Robust Markov Decision Processes
Shiau Hong Lim, Huan Xu, and Shie Mannor · 2013
Earlier work this paper cites.
Rectifier Nonlinearities Improve Neural Network Acoustic Models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Stochastic Multi-Armed-Bandit Problem with Non-stationary Rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Online Learning in Markov Decision Processes with Changing Cost Sequences
Travis Dick, Andras Gyorgy, and Csaba Szepesvari · 2014
Cited alongside, same era.
A Survey on Concept Drift Adaptation
João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia · 2014
Cited alongside, same era.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Klimov Oleg · 2017
Later among the works it cites.
Robust Imitation of Diverse Behaviors
Ziyu Wang, Josh Merel, Scott Reed, Greg Wayne, Nando de Freitas, and Nicolas Heess · 2017
Later among the works it cites.
mixup: Beyond Empirical Risk Minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Later among the works it cites.
Energy-based Generative Adversarial Network
Junbo Zhao, Michael Mathieu, and Yann LeCun · 2017
Later among the works it cites.
Exploration by Random Network Distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Adapting Auxiliary Losses Using Gradient Similarity
Yunshu Du, Wojciech M Czarnecki, Siddhant M Jayakumar, Razvan Pascanu, and Balaji Lakshminarayanan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Deterministic Policy Gradient Algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
An invitation to imitation
J Andrew Bagnell · 2015
Cited alongside, same era.
Unsupervised Visual Representation Learning by Context Prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
Lipschitz regularized Deep Neural Networks generalize and are adversarially robust
Chris Finlay, Jeff Calder, Bilal Abbasi, and Adam Oberman · 2018
Later among the works it cites.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Later among the works it cites.
A Sliding-Window Algorithm for Markov Decision Processes with Arbitrarily Changing Rewards and Transitions
Pratik Gajane, Ronald Ortner, and Peter Auer · 2018
Later among the works it cites.
The GAN Landscape: Losses, Architectures, Regularization, and Normalization
Karol Kurach, Mario Lucic, Xiaohua Zhai, Marcin Michalski, and Sylvain Gelly · 2018
Later among the works it cites.
Visualizing the Loss Landscape of Neural Nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
Efficient Contextual Bandits in Non-stationary Worlds
Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Which Training Methods for GANs do actually Converge?
Lars Mescheder, Andreas Geiger, and Sebastian Nowozin · 2018
Later among the works it cites.
Spectral Normalization for Generative Adversarial Networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Later among the works it cites.
Time Limits in Reinforcement Learning
Fabio Pardo, Arash Tavakoli, Vitaly Levdik, and Petar Kormushev · 2018
Later among the works it cites.
Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow
Xue Bin Peng, Angjoo Kanazawa, Sam Toyer, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Parameter Space Noise for Exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2018
Later among the works it cites.
Visual Imitation With a Minimal Adversary
Scott Reed, Yusuf Aytar, Ziyu Wang, Tom Paine, Aäron van den Oord, Tobias Pfaff, Sergio Gomez, Alexander Novikov, David Budden, and Oriol Vinyals · 2018
Later among the works it cites.
Reward Estimation for Variance Reduction in Deep Reinforcement Learning
Joshua Romoff, Peter Henderson, Alexandre Piché, Vincent Francois-Lavet, and Joelle Pineau · 2018
Later among the works it cites.
Deep Reinforcement Learning and the Deadly Triad
Hado van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil · 2018
Later among the works it cites.
Towards Characterizing Divergence in Deep Q-Learning
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 2019
Later among the works it cites.
Adaptively Tracking the Best Bandit Arm with an Unknown Number of Distribution Changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2019
Later among the works it cites.
Sample-Efficient Imitation Learning via Generative Adversarial Nets
Lionel Blondé and Alexandros Kalousis · 2019
Later among the works it cites.
A New Algorithm for Non-stationary Contextual Bandits: Efficient, Optimal, and Parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Later among the works it cites.
Learning to Optimize under Non-Stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Later among the works it cites.
Reinforcement Learning under Drift
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Later among the works it cites.
Understanding Multi-Step Deep Reinforcement Learning: A Systematic Study of the DQN Target
J Fernando Hernandez-Garcia and Richard S Sutton · 2019
Later among the works it cites.
Diagnosing Bottlenecks in Deep Q-learning Algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Later among the works it cites.
Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2019
Later among the works it cites.
Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning
Erwan Lecarpentier and Emmanuel Rachelson · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Later among the works it cites.
Solving Rubik’s Cube with a Robot Hand
OpenAI · 2019
Later among the works it cites.
Variational Regret Bounds for Reinforcement Learning
Ronald Ortner, Pratik Gajane, and Peter Auer · 2019
Later among the works it cites.
Reinforcement Learning in Non-Stationary Environments
Sindhu Padakandla, K J Prabuchandran, and Shalabh Bhatnagar · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Later among the works it cites.
Weighted Linear Bandits for Non-Stationary Environments
Yoan Russac, Claire Vernade, and Olivier Cappé · 2019
Later among the works it cites.
Scalability in Perception for Autonomous Driving: Waymo Open Dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, Sheng Zhao, Shuyang Cheng, Yu Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P Agapiou, Max Jaderberg, Alexander S Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom L Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver · 2019
Later among the works it cites.
Random Expert Distillation: Imitation Learning via Expert Policy Support Estimation
Ruohan Wang, Carlo Ciliberto, Pierluigi Amadori, and Yiannis Demiris · 2019
Later among the works it cites.
Positive-Unlabeled Reward Learning
Danfei Xu and Misha Denil · 2019
Later among the works it cites.
Efficient Policy Learning for Non-Stationary MDPs under Adversarial Manipulation
Tiancheng Yu and Suvrit Sra · 2019
Later among the works it cites.
Task-Relevant Adversarial Imitation Learning
Konrad Zolna, Scott Reed, Alexander Novikov, Sergio Gomez Colmenarej, David Budden, Serkan Cabi, Misha Denil, Nando de Freitas, and Ziyu Wang · 2019
Later among the works it cites.
Experiment Tracking with Weights and Biases, 2020
Lukas Biewald · 2020
Closest in time.
An LSTM-Based Autonomous Driving Model Using Waymo Open Dataset
Zhicheng Gu, Zhihao Li, Xuan Di, and Rongye Shi · 2020
Closest in time.
Learning to Walk in the Real World with Minimal Human Effort
Sehoon Ha, Peng Xu, Zhenyu Tan, Sergey Levine, and Jie Tan · 2020
Closest in time.
Measuring the Algorithmic Efficiency of Neural Networks
Danny Hernandez and Tom B Brown · 2020
Closest in time.
A case for new neural network smoothness constraints
Mihaela Rosca, Theophane Weber, Arthur Gretton, and Shakir Mohamed · 2020
Closest in time.
Adversarial Robustness Through Local Lipschitzness
Yao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov, and Kamalika Chaudhuri · 2020
Closest in time.
Regularisation of neural networks by enforcing Lipschitz continuity
Henry Gouk, Eibe Frank, Bernhard Pfahringer, and Michael J Cree · 2021
Closest in time.
What Matters for Adversarial Imitation Learning?
Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent, Robert Dadashi, Sertan Girgin, Matthieu Geist, Olivier Bachem, Olivier Pietquin, and Marcin Andrychowicz · 2021
Closest in time.