Fetching the paper…
Reading the bibliography…
Artificial reinforcement learning (RL) is a widely used technique in artificial intelligence that provides a general method for training agents to perform a wide variety of behaviours.
Animal Intelligence: Experimental Studies
Edward Lee Thorndike · 1911
Earlier work this paper cites.
The Principles of Humane Experimental Technique
William Moy Stratton Russell and Rex Leonard Burch · 1959
Earlier work this paper cites.
Anarchy, State, and Utopia
Robert Nozick · 1974
Earlier work this paper cites.
An opponent-process theory of motivation: I. temporal dynamics of affect
Richard L. Solomon and John D. Corbit · 1974
Earlier work this paper cites.
Psychologism and behaviorism
Ned Block · 1981
Earlier work this paper cites.
Reward learning in normal and mutant Drosophila
Bruce L. Tempel, Nancy Bonini, Douglas R. Dawson, and William G. Quinn · 1983
Earlier work this paper cites.
Xenopsychology
Robert A. Freitas, Jr · 1984
Earlier work this paper cites.
A unified biosocial theory of personality and its role in the development of anxiety states
C. Robert Cloninger · 1986
Earlier work this paper cites.
Artificial intelligence and ethics: An exercise in the moral imagination
Michael R. LaChat · 1986
Earlier work this paper cites.
Behavioural investigation of pain in animals
Manfred Zimmermann · 1986
Earlier work this paper cites.
The moral standing of insects and the ethics of extinction
Jeffrey A. Lockwood · 1987
Earlier work this paper cites.
Uneven pattern of dopamine loss in the striatum of patients with idiopathic Parkinson’s disease
Stephen J. Kish, Kathleen Shannak, and Oleh Hornykiewicz · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G. Barto, Richard S. Sutton, and Charles W. Anderson · 1990
Earlier work this paper cites.
Integrated architecture for learning, planning, and reacting based on approximating dynamic programming
Richard S. Sutton · 1990
Earlier work this paper cites.
Consciousness Explained
Daniel C. Dennett · 1991
Earlier work this paper cites.
Overcoming incomplete perception with utile distinction memory
Andrew McCallum · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Bugs in the System: Insects and Their Impact on Human Affairs
May R. Berenbaum · 1995
Earlier work this paper cites.
Towards welfare biology: Evolutionary economics of animal consciousness and suffering
Yew-Kwang Ng · 1995
Earlier work this paper cites.
Improving elevator performance using reinforcement learning
Robert H. Crites and Andrew G. Barto · 1996
Earlier work this paper cites.
Consciousness: More like fame than television
Daniel C. Dennett · 1996
Earlier work this paper cites.
Full House: The Spread of Excellence from Plato to Darwin
Stephen Jay Gould · 1996
Earlier work this paper cites.
A neostriatal habit learning system in humans
Barbara J. Knowlton, Jennifer A. Mangels, and Larry R. Squire · 1996
Earlier work this paper cites.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Sridhar Mahadevan · 1996
Earlier work this paper cites.
Reinforcement learning and animat emotions
Ian Wright · 1996
Earlier work this paper cites.
Babies and Beasts: The Argument from Marginal Cases
Daniel A. Dombrowski · 1997
Earlier work this paper cites.
A neural substrate of prediction and reward
Wolfram Schultz, Peter Dayan, and P. Read Montague · 1997
Earlier work this paper cites.
Creatures: Entertainment software agents with artificial life
Stephen Grand and Dave Cliff · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Harm avoidance and serotonin
Michel Hansenne and Marc Ansseau · 1999
Earlier work this paper cites.
Using reinforcement learning to spider the web efficiently
Jason Rennie and Andrew McCallum · 1999
Earlier work this paper cites.
Dynamic job-shop scheduling using reinforcement learning agents
M. Emin Aydin and Ercan Öztemel · 2000
Earlier work this paper cites.
Will robots rise up and demand their rights?
Rodney Brooks · 2000
Earlier work this paper cites.
Evolutionary computation versus reinforcement learning
Jürgen Schmidhuber · 2000
Earlier work this paper cites.
Towards an ethics for epersons
Steve Torrance · 2000
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bram Bakker · 2001
Earlier work this paper cites.
Are we explaining consciousness yet?
Daniel C. Dennett · 2001
Earlier work this paper cites.
Animal suffering: An invertebrate perspective
Jennifer A. Mather · 2001
Earlier work this paper cites.
Opponent interactions between serotonin and dopamine
Nathaniel D. Daw, Sham Kakade, and Peter Dayan · 2002
Earlier work this paper cites.
The mirror test
Gordon G. Gallup, Jr., James R. Anderson, and Daniel J. Shillito · 2002
Earlier work this paper cites.
Actor–critic models of the basal ganglia: New anatomical and computational perspectives
Daphna Joel, Yael Niv, and Eytan Ruppin · 2002
Earlier work this paper cites.
Efficient reinforcement learning through evolving neural network topologies
Kenneth O. Stanley and Risto Miikkulainen · 2002
Earlier work this paper cites.
A robot that reinforcement-learns to identify and memorize important previous observations
Bram Bakker, Viktor Zhumatiy, Gabriel Gruener, and Jürgen Schmidhuber · 2003
Earlier work this paper cites.
Astronomical waste: The opportunity cost of delayed technological development
Nick Bostrom · 2003
Earlier work this paper cites.
Energy Landscapes: Applications to Clusters, Biomolecules and Glasses
David Wales · 2003
Earlier work this paper cites.
The Power of Reinforcement
Stephen Ray Flora · 2004
Cited alongside, same era.
By carrot or by stick: Cognitive reinforcement learning in Parkinsonism
Michael J. Frank, Lauren C. Seeberger, and Randall C. O’Reilly · 2004
Cited alongside, same era.
Dog training methods: Their use, effectiveness and interaction with behaviour and welfare
E. F. Hiby, N. J. Rooney, and J. W. S. Bradshaw · 2004
Cited alongside, same era.
Temporal difference models describe higher-order learning in humans
Ben Seymour, John P. O’Doherty, Peter Dayan, Martin Koltzenburg, Anthony K. Jones, Raymond J. Dolan, Karl J. Friston, and Richard S. Frackowiak · 2004
Cited alongside, same era.
Global workspace theory of consciousness: Toward a cognitive neuroscience of human experience
Bernard J. Baars · 2005
Cited alongside, same era.
Android science and the animal rights movement: Are there analogies?
Artificial Intelligence: A Modern Approach
Stuart Russell and Peter Norvig · 2009
Later among the works it cites.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Yoav Shoham and Kevin Leyton-Brown · 2009
Later among the works it cites.
Animal Liberation: The Definitive Classic of the Animal Movement
Peter Singer · 2009
Later among the works it cites.
Neural coding of pleasure: “rose-tinted glasses” of the ventral pallidum
J. Wayne Aldridge and Kent C. Berridge · 2010
Later among the works it cites.
High-level reinforcement learning in strategy games
Christopher Amato and Guy Shani · 2010
Later among the works it cites.
Robot rights? towards a social-relational justification of moral consideration
Mark Coeckelbergh · 2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David J. Calverley · 2005
Cited alongside, same era.
Aspects of the biology and welfare of animals used for experimental and other scientific purposes
EFSA · 2005
Cited alongside, same era.
Actor–critic models of reinforcement learning in the basal ganglia: From natural to artificial rats
Mehdi Khamassi, Loïc Lachèze, Benoît Girard, Alain Berthoz, and Agnès Guillot · 2005
Cited alongside, same era.
Dopamine cells respond to predicted events during classical conditioning: Evidence for eligibility traces in the reward-learning network
Wei-Xing Pan, Robert Schmidt, Jeffery R. Wickens, and Brian I. Hyland · 2005
Cited alongside, same era.
The remote roots of consciousness in fruit-fly selective attention?
Bruno van Swinderen · 2005
Cited alongside, same era.
The metacognitive loop I: Enhancing reinforcement learning with metacognitive monitoring and control for improved perturbation tolerance
Michael L. Anderson, Tim Oates, Waiyian Chong, and Don Perlis · 2006
Cited alongside, same era.
Quantity of experience: Brain-duplication and degrees of consciousness
Nick Bostrom · 2006
Cited alongside, same era.
The neuronal replicator hypothesis
Chrisantha Fernando, Richard Goldstein, and Eörs Szathmáry · 2010
Later among the works it cites.
Debunking the idyllic view of natural processes: Population dynamics and suffering in the wild
Oscar Horta · 2010
Later among the works it cites.
Combining self-motivation with logical planning and inference in a reward-seeking agent
Daphne Liu and Lenhart K. Schubert · 2010
Later among the works it cites.
Learning to overtake in TORCS using simple reinforcement learning
Daniele Loiacono, Alessandro Prete, Pier Luca Lanzi, and Luigi Cardamone · 2010
Later among the works it cites.
Two-factor theory, the actor-critic model, and conditioned avoidance
Tiago V. Maia · 2010
Later among the works it cites.
Reinforcement learning in first person shooter games
Michelle McPartland and Marcus Gallagher · 2010
Later among the works it cites.
Phenomenal and access consciousness and the “hard” problem: A view from the designer stance
Aaron Sloman · 2010
Later among the works it cites.
Machine Ethics
Michael Anderson and Susan Leigh Anderson, editors · 2011
Later among the works it cites.
Assessing the impact of planned social change
Donald T. Campbell · 2011
Later among the works it cites.
Anesthesia, analgesia, and euthanasia of invertebrates
John E. Cooper · 2011
Later among the works it cites.
Understanding dopamine and reinforcement learning: The dopamine reward prediction error hypothesis
Paul W. Glimcher · 2011
Later among the works it cites.
Neuroevolutionary reinforcement learning for generalized control of simulated helicopters
Rogier Koppejan and Shimon Whiteson · 2011
Later among the works it cites.
Empirical support for higher-order theories of conscious awareness
Hakwan Lau and David Rosenthal · 2011
Later among the works it cites.
Robot Ethics: The Ethical and Social Implications of Robotics
Patrick Lin, Keith Abney, and George A. Bekey, editors · 2011
Later among the works it cites.
Optimizing drug therapy with reinforcement learning: The case of anemia management
Jordan M. Malof and Adam E. Gaweda · 2011
Later among the works it cites.
Invertebrate welfare: Where is the real evidence for conscious affective states?
Georgia J. Mason · 2011
Later among the works it cites.
Tuning computer gaming agents using Q-learning
Purvag G. Patel, Norman Carver, and Shahram Rahimi · 2011
Later among the works it cites.
A neural signature of hierarchical reinforcement learning
José J. F. Ribas-Fernandes, Alec Solway, Carlos Diuk, Joseph T. McGuire, Andrew G. Barto, Yael Niv, and Matthew M. Botvinick · 2011
Later among the works it cites.
The ubiquity of model-based reinforcement learning
Bradley B. Doll, Dylan A. Simon, and Nathaniel D. Daw · 2012
Later among the works it cites.
Selectionist and evolutionary approaches to brain function: A critical appraisal
Chrisantha Fernando, Eörs Szathmáry, and Phil Husbands · 2012
Later among the works it cites.
Global Workspace Theory, its LIDA model and the underlying neuroscience
Stan Franklin, Steve Strain, Javier Snaider, Ryan McCall, and Usef Faghihi · 2012
Later among the works it cites.
The Machine Question: Critical Perspectives on AI, Robots, and Ethics
David J. Gunkel · 2012
Later among the works it cites.
The phenomenal stance revisited
Anthony I. Jack and Philip Robbins · 2012
Later among the works it cites.
The Mario AI benchmark and competitions
Sergey Karakovskiy and Julian Togelius · 2012
Later among the works it cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Later among the works it cites.
Evaluating the TD model of classical conditioning
Elliot A. Ludvig, Richard S. Sutton, and E. James Kehoe · 2012
Later among the works it cites.
Philosophers & futurists, catch up! response to the singularity
Jürgen Schmidhuber · 2012
Later among the works it cites.
Evolutionary computation for reinforcement learning
Shimon Whiteson · 2012
Later among the works it cites.
Applying reinforcement learning to an insurgency agent-based simulation
Andrew J. Collins, John Sokolowski, and Catherine Banks · 2013
Later among the works it cites.
Hierarchical learning induces two simultaneous, but separable, prediction errors in human basal ganglia
Carlos Diuk, Karin Tsai, Jonathan Wallis, Matthew Botvinick, and Yael Niv · 2013
Later among the works it cites.
Reinforcement learning in robotics: A survey
Jens Kober, J. Andrew Bagnell, and Jan Peters · 2013
Later among the works it cites.
Evolving large-scale neural networks for vision-based reinforcement learning
Jan Koutník, Giuseppe Cuccu, Jürgen Schmidhuber, and Faustino Gomez · 2013
Later among the works it cites.
Behavior selection using utility-based reinforcement learning in irregular warfare simulation models
Sotiris Papadopoulos, Francisco Baez, Jonathan Alt, and Christian Darken · 2013
Later among the works it cites.
Reinforcement learning and human behavior
Hanan Shteingart and Yonatan Loewenstein · 2013
Later among the works it cites.
Suffering subroutines: On the humanity of making a computer that feels pain
Meghan Winsby · 2013
Later among the works it cites.
The not-craving brain
Rick Hanson · 2014
Closest in time.
The hedonistic imperative, 2007
David Pearce · 2014
Closest in time.
If materialism is true, the United States is probably conscious, 2012
Eric Schwitzgebel · 2014
Closest in time.
When robots have feelings
Peter Singer and Agata Sagan · 2014
Closest in time.
Are wireheads happy?
Scott Siskind · 2014
Closest in time.