Fetching the paper…
Reading the bibliography…
One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions.
A note on the pure theory of consumer’s behaviour
Paul A Samuelson · 1938
Earlier work this paper cites.
Runaround
Isaac Asimov · 1942
Earlier work this paper cites.
Some moral and technical consequences of automation
Norbert Wiener · 1960
Earlier work this paper cites.
Reducibility among combinatorial problems
Richard Karp · 1972
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
The Expected-Outcome Model of Two-Player Games
Bruce Abramson · 1987
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Yann LeCun, Bernhard E Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne E Hubbard, and Lawrence D Jackel · 1990
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
Q-learning
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald Williams · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey Hinton · 1993
Earlier work this paper cites.
The first law of robotics (a call to arms)
Oren Etzioni and Daniel Weld · 1994
Earlier work this paper cites.
Learning agents for uncertain environments
Stuart Russell · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Ethical issues in advanced artificial intelligence
Nick Bostrom · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Ng · 2004
Earlier work this paper cites.
Coherent extrapolated volition
Eliezer Yudkowsky · 2004
Earlier work this paper cites.
Universal artificial intelligence
Marcus Hutter · 2005
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Shane Legg and Marcus Hutter · 2007
Earlier work this paper cites.
Gödel machines: Self-referential universal problem solvers making provably optimal self-improvements
Jürgen Schmidhuber · 2007
Earlier work this paper cites.
The basic AI drives
Stephen Omohundro · 2008
Earlier work this paper cites.
Teachable robots: Understanding human teaching behavior to build more effective robot learners
Andrea L Thomaz and Cynthia Breazeal · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Computational Complexity: A Modern Approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
Dataset shift in machine learning, 2009
J Quiñonero Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence · 2009
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
William Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Learning what to value
Daniel Dewey · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Delusion, survival, and intelligent agents
Mark Ring and Laurent Orseau · 2011
Earlier work this paper cites.
APRIL: Active preference learning-based reinforcement learning
Riad Akrour, Marc Schoenauer, and Michèle Sebag · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park · 2012
Earlier work this paper cites.
Learning from human-generated reward
William Bradley Knox · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton · 2012
Earlier work this paper cites.
Red teams
Brendan Mulvaney · 2012
Earlier work this paper cites.
Space-time embedded intelligence
Laurent Orseau and Mark Ring · 2012
Earlier work this paper cites.
Active learning
Burr Settles · 2012
Earlier work this paper cites.
Leakproofing the singularity: Artificial intelligence confinement problem
Roman Yampolskiy · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Maxout networks
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz · 2013
Earlier work this paper cites.
Asymptotic non-learnability of universal agents with computable horizon functions
Laurent Orseau · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Programming by feedback
Riad Akrour, Marc Schoenauer, Michèle Sebag, and Jean-Christophe Souplet · 2014
Earlier work this paper cites.
Robust cooperation in the prisoner’s dilemma: Program equilibrium via provability logic
Mihaly Barasz, Paul Christiano, Benja Fallenstein, Marcello Herreshoff, Patrick LaVictoire, and Eliezer Yudkowsky · 2014
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Earlier work this paper cites.
SNES Super Mario World (USA) “arbitrary code execution” in 02:25.19
Masterjun · 2014
Earlier work this paper cites.
Watching all the movies ever made
Philip Peter · 2014
Earlier work this paper cites.
Motivated value selection for artificial agents
Stuart Armstrong · 2015
Earlier work this paper cites.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier García and Fernando Fernández · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
José Miguel Hernández-Lobato and Ryan Adams · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Diederik P Kingma, Tim Salimans, and Max Welling · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Bad universal priors and notions of optimality
Jan Leike and Marcus Hutter · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
On a formal model of safe and scalable self-driving cars
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Graepel Thore, and Demis Hassabis · 2017
Later among the works it cites.
Agent foundations for aligning machine intelligence with human interests: a technical research agenda
Nate Soares and Benya Fallenstein · 2017
Later among the works it cites.
Third-person imitation learning
Bradly C Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
The value learning problem
Nate Soares · 2015
Cited alongside, same era.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Optimization daemons
Arbital · 2016
Cited alongside, same era.
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Online learning with gated linear networks
Joel Veness, Tor Lattimore, Avishkar Bhoopchand, Agnieszka Grabska-Barwinska, Christopher Mattern, and Peter Toth · 2017
Later among the works it cites.
FeUdal Networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2017
Later among the works it cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Learning with latent language
Jacob Andreas, Dan Klein, and Sergey Levine · 2018
Closest in time.
Occam’s razor is insufficient to infer the preferences of irrational agents
Stuart Armstrong and Sören Mindermann · 2018
Closest in time.
Deep reinforcement learning from policy-dependent human feedback
Dilip Arumugam, Jun Ki Lee, Sophie Saskin, and Michael L Littman · 2018
Closest in time.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Closest in time.
Playing hard exploration games by watching YouTube
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Closest in time.
Learning to follow language instructions with adversarial reward induction
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Pushmeet Kohli, and Edward Grefenstette · 2018
Closest in time.
Shane Barratt and Rishi Sharma · 2018
Closest in time.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva TB, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Closest in time.
Verifiable reinforcement learning via policy extraction
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama · 2018
Closest in time.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu · 2018
Closest in time.
Unrestricted adversarial examples
Tom B Brown, Nicholas Carlini, Chiyuan Zhang, Catherine Olsson, Paul Christiano, and Ian Goodfellow · 2018
Closest in time.
Supervising strong learners by amplifying weak experts
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Closest in time.
Embedded agency
Abram Demski and Scott Garrabrant · 2018
Closest in time.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Closest in time.
Predicting human deliberative judgments with machine learning
Owain Evans, Andreas Stuhlmüller, Chris Cundy, Ryan Carey, Zachary Kenton, Thomas McGrath, and Andrew Schreiber · 2018
Closest in time.
Towards Safe Artificial General Intelligence
Tom Everitt · 2018
Closest in time.
The alignment problem for Bayesian history-based reinforcement learners
Tom Everitt and Marcus Hutter · 2018
Closest in time.
Tom Everitt, Gary Lea, and Marcus Hutter · 2018
Closest in time.
Sources of intuitions and data on AGI
Scott Garrabrant · 2018
Closest in time.
Motivating the rules of the game for adversarial example research
Justin Gilmer, Ryan P Adams, Ian Goodfellow, David Andersen, and George E Dahl · 2018
Closest in time.
Visualizing and understanding Atari agents
Sam Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern · 2018
Closest in time.
Deep Q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Gabriel Dulac-Arnold, John Agapiou, Joel Z Leibo, and Audrunas Gruslys · 2018
Closest in time.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Closest in time.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Closest in time.
Deep reinforcement learning doesn’t work yet
Alex Irpan · 2018
Closest in time.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Closest in time.
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel · 2018
Closest in time.
Measuring and avoiding side effects using relative reachability
Victoria Krakovna, Laurent Orseau, Miljan Martic, and Shane Legg · 2018
Closest in time.
Julia Kreutzer, Joshua Uyheng, and Stefan Riezler · 2018
Closest in time.
Accurate uncertainties for deep learning using calibrated regression
Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon · 2018
Closest in time.
Human-in-the-loop interpretability prior
Isaac Lage, Andrew Slavin Ross, Been Kim, Samuel J Gershman, and Finale Doshi-Velez · 2018
Closest in time.
The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation and artificial life research communities
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Julie Beaulieu, Peter J Bentley, Samuel Bernard, Guillaume Belson, David M Bryson, Nick Cheney, Antoine Cully, Stephane Doncieux, Fred C Dyer, Kai Olav Ellefsen, Robert Feldt, Stephan Fischer, Stephanie Forrest, Antoine Frénoy, Christian Gagné, Leni Le Goff, Laura M Grabowski, Babak Hodjat, Frank Hutter, Laurent Keller, Carole Knibbe, Peter Krcah, Richard E Lenski, Hod Lipson, Robert MacCurdy, Carlos Maestre, Risto Miikkulainen, Sara Mitri, David E Moriarty, Jean-Baptiste Mouret, Anh Nguyen, Charles Ofria, Marc Parizeau, David Parsons, Robert T Pennock, William F Punch, Thomas S Ray, Marc Schoenauer, Eric Shulte, Karl Sims, Kenneth O Stanley, François Taddei, Danesh Tarapore, Simon Thibault, Westley Weimer, Richard Watson, and Jason Yosinski · 2018
Closest in time.
Deep learning: A critical appraisal
Gary Marcus · 2018
Closest in time.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Closest in time.
Building safe artificial intelligence: specification, robustness, and assurance
Pedro A Ortega, Vishal Maini, and the DeepMind safety team · 2018
Closest in time.
Social choice and the value alignment problem
Mahendra Prasad · 2018
Closest in time.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2018
Closest in time.
Active reinforcement learning with monte-carlo tree search
Sebastian Schulze and Owain Evans · 2018
Closest in time.
Does your model know the digit 6 is not a cat? a less biased evaluation of “outlier” detectors
Alireza Shafaei, Mark Schmidt, and James J Little · 2018
Closest in time.
Thilo Stadelmann, Mohammadreza Amirian, Ismail Arabaci, Marek Arnold, Gilbert François Duivesteijn, Ismail Elezi, Melanie Geiger, Stefan Lörwald, Benjamin Bruno Meier, Katharina Rombach, and Lukas Tuggener · 2018
Closest in time.
Factored cognition
Andreas Stuhlmüller · 2018
Closest in time.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 2018
Closest in time.
Neural arithmetic logic units
Andrew Trask, Felix Hill, Scott Reed, Jack Rae, Chris Dyer, and Phil Blunsom · 2018
Closest in time.
Reward learning from narrated demonstrations
Hsiao-Yu Fish Tung, Adam W Harley, Liang-Kang Huang, and Katerina Fragkiadaki · 2018
Closest in time.
Adversarial risk and the dangers of evaluating against weak attacks
Jonathan Uesato, Brendan O’Donoghue, Aaron van den Oord, and Pushmeet Kohli · 2018
Closest in time.
Scaling provable adversarial defenses
Eric Wong, Frank Schmidt, Jan Hendrik Metzen, and J Zico Kolter · 2018
Closest in time.
Bridging the gap: Converting human advice into imagined examples
Eric Yeh, Melinda Gervasio, Daniel Sanchez, Matthew Crossley, and Karen Myers · 2018
Closest in time.