Fetching the paper…
Reading the bibliography…
To widen their accessibility and increase their utility, intelligent agents must be able to learn complex behaviors as specified by (non-expert) human users.
Function optimization using connectionist reinforcement learning algorithms
Williams, Ronald J. and Peng, Jing · 1902
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, Andrew G., Sutton, Richard S., and Anderson, Charles W · 1983
Earlier work this paper cites.
Neurocomputing: Foundations of research
Rumelhart, David E., Hinton, Geoffrey E., and Williams, Ronald J · 1988
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Puterman, Martin L · 1994
Earlier work this paper cites.
Rewarding behaviors
Bacchus, Fahiem, Boutilier, Craig, and Grove, Adam · 1996
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, Robert M · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S., McAllester, David A., Singh, Satinder P., and Mansour, Yishay · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, Pieter and Ng, Andrew Y · 2004
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G E and Salakhutdinov, R R · 2006
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Thomaz, Andrea Lockerd and Breazeal, Cynthia · 2006
Earlier work this paper cites.
TAMER: Training an agent manually via evaluative reinforcement
Knox, W. Bradley and Stone, Peter · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, Brenna, Chernova, Sonia, Veloso, Manuela M., and Browning, Brett · 2009
Earlier work this paper cites.
Combining manual feedback with subsequent MDP reward signals for reinforcement learning
Knox, W. Bradley and Stone, Peter · 2010
Earlier work this paper cites.
Stacked convolutional auto-encoders for hierarchical feature extraction
Masci, Jonathan, Meier, Ueli, Ciresan, Dan C., and Schmidhüber, Jurgen · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, Stéphane, Gordon, Geoffrey J., and Bagnell, J. Andrew · 2011
Cited alongside, same era.
Degris, Thomas, White, Martha, and Sutton, Richard S · 2012
Cited alongside, same era.
Reinforcement learning from simultaneous human and MDP reward
Knox, W. Bradley and Stone, Peter · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, Martín, Barham, Paul, Chen, Jianmin, Chen, Zhifeng, Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Irving, Geoffrey, Isard, Michael, Kudlur, Manjunath, Levenberg, Josh, Monga, Rajat, Moore, Sherry, Murray, Derek Gordon, Steiner, Benoit, Tucker, Paul A., Vasudevan, Vijay, Warden, Pete, Wicke, Martin, Yu, Yuan, and Zhang, Xiaoqiang · 2016
Later among the works it cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, Chelsea, Levine, Sergey, and Abbeel, Pieter · 2016
Later among the works it cites.
The Malmo platform for artificial intelligence experimentation
Johnson, Matthew, Hofmann, Katja, Hutton, Tim, and Bignell, David · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adrià Puigdomènech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P., Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Control of memory, active perception, and action in minecraft
Oh, Junhyuk, Chockalingam, Valliappa, Singh, Satinder P., and Lee, Honglak · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vien, Ngo Anh and Ertel, Wolfgang · 2012
Cited alongside, same era.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, Shane, Subramanian, Kaushik, Scholz, Jonathan, Isbell, Charles Lee, and Thomaz, Andrea Lockerd · 2013
Cited alongside, same era.
Power to the people: The role of humans in interactive machine learning
Amershi, Saleema, Cakmak, Maya, Knox, W. Bradley, and Kulesza, Todd · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2014
Cited alongside, same era.
Minecraft as an experimental world for AI in robotics
Aluru, Krishna, Tellex, Stefanie, Oberlin, John, and Macglashan, James · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P., Hunt, Jonathan J., Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Cited alongside, same era.
Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
Loftin, Robert Tyler, Peng, Bei, MacGlashan, James, Littman, Michael L., Taylor, Matthew E., Huang, Jeff, and Roberts, David L · 2015
Cited alongside, same era.
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J., Guez, Arthur, Sifre, Laurent, van den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Vedavyas, Lanctot, Marc, Dieleman, Sander, Grewe, Dominik, Nham, John, Kalchbrenner, Nal, Sutskever, Ilya, Lillicrap, Timothy P., Leach, Madeleine, Kavukcuoglu, Koray, Graepel, Thore, and Hassabis, Demis · 2016
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, Chen, Givony, Shahar, Zahavy, Tom, Mankowitz, Daniel J., and Mannor, Shie · 2016
Later among the works it cites.
Deep reinforcement learning from human preferences
Christiano, Paul F., Leike, Jan, Brown, Tom B., Martic, Miljan, Legg, Shane, and Amodei, Dario · 2017
Later among the works it cites.
Environment-independent task specifications via gltl
Littman, Michael L., Topcu, Ufuk, Fu, Jie, Isbell, Charles Lee, Wen, Min, and MacGlashan, James · 2017
Later among the works it cites.
Interactive learning from policy-dependent human feedback
MacGlashan, James, Ho, Mark K., Loftin, Robert Tyler, Peng, Bei, Roberts, David L., Taylor, Matthew E., and Littman, Michael L · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Later among the works it cites.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Warnell, Garrett, Waytowich, Nicholas, Lawhern, Vernon, and Stone, Peter · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, Matteo, Modayil, Joseph, van Hasselt, Hado, Schaul, Tom, Ostrovski, Georg, Dabney, Will, Horgan, Daniel, Piot, Bilal, Azar, Mohammad Gheshlaghi, and Silver, David · 2018
Later among the works it cites.