Fetching the paper…
Reading the bibliography…
Robots will eventually be part of every household.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Incorporating advice into agents that learn from reinforcements
Richard Maclin and Jude W. Shavlik · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Guiding a reinforcement learner with natural language advice: Initial results in robocup soccer
G. Kuhlmann, P. Stone, R. Mooney, and J. Shavlik · 2004
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance
A. Thomaz and C. Breazeal · 2006
Earlier work this paper cites.
Reinforcement learning via practice and critique advice
K. Judah, S. Roy, A. Fern, and T. Dietterich · 2010
Earlier work this paper cites.
Reinforcement learning from simultaneous human and mdp reward
W. Bradley Knox and Peter Stone · 2012
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L. Isbell, and Andrea Lockerd Thomaz · 2013
Earlier work this paper cites.
Training a robot via human feedback: A case study
W. Bradley Knox, Cynthia Breazeal, , and Peter Stone · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky · 2014
Earlier work this paper cites.
How to construct deep recurrent neural networks
Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Remi Lebret, Pedro O. Pinheiro, and Ronan Collobert · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2015
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Michel Galley, Jianfeng Gao, and Dan Jurafsky · 2016
Later among the works it cites.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2016
Later among the works it cites.
Self-critical sequence training for image captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neuroaesthetics in fashion: Modeling the perception of beauty
Edgar Simo-Serra, Sanja Fidler, Francesc Moreno-Noguer, and Raquel Urtasun · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Oriol Vinyals and Quoc Le · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Tega: A social robot
Jacqueline Kory Westlund, Jin Joo Lee, Luke Plummer, Fardad Faridi, Jesse Gray, Matt Berlin, Harald Quintus-Bosz, Robert Hartmann, Mike Hess, Stacy Dyer, Kristopher dos Santos, Sigurdhur Örn Adhalgeirsson, Goren Gordon, Samuel Spaulding, Marayna Martinez, Madhurima Das, Maryam Archie, Sooyeon Jeong, and Cynthia Breazeal · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
phi-lstm: A phrase-based hierarchical lstm model for image captioning
Ying Hua Tan and Chee Seng Chan · 2016
Later among the works it cites.
Dialog-based language learning
Jason Weston · 2016
Later among the works it cites.
Training an adaptive dialogue policy for interactive learning of visually grounded word meanings
Yanchao Yu, Arash Eshghi, and Oliver Lemon · 2016
Later among the works it cites.
Towards diverse and natural image descriptions via a conditional gan
Bo Dai, Dahua Lin, Raquel Urtasun, and Sanja Fidler · 2017
Closest in time.
Beating atari with natural language guided reinforcement learning
Russell Kaplan, Christopher Sauer, and Alexander Sosa · 2017
Closest in time.
A hierarchical approach for generating descriptive image paragraphs
Jonathan Krause, Justin Johnson, Ranjay Krishna, and Li Fei-Fei · 2017
Closest in time.
Socially assistive robotics: Human augmentation vs. automation
Maja J. Matarič · 2017
Closest in time.
The burchak corpus: a challenge data set for interactive learning of visually grounded word meanings
Yanchao Yu, Arash Eshghi, Gregory Mills, and Oliver Lemon · 2017
Closest in time.