Fetching the paper…
Reading the bibliography…
Reinforcement learning problems are often described through rewards that indicate if an agent has completed some task.
The exponentially weighted moving average
J Stuart Hunter · 1986
Earlier work this paper cites.
A tutorial on visual servo control
Seth Hutchinson, Gregory D Hager, Peter Corke, et al · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Recognizing movement using motion histograms
James Davis · 1999
Earlier work this paper cites.
Comprehensive database for facial expression analysis
Takeo Kanade, Jeffrey F Cohn, and Yingli Tian · 2000
Earlier work this paper cites.
The recognition of human movement using temporal templates
Aaron F Bobick and James W Davis · 2001
Earlier work this paper cites.
Hand-drawn maps for robot navigation
Marjorie Skubic, Sam Blisard, Andy Carle, and Pascal Matsakis · 2002
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Cited alongside, same era.
Learning OpenCV: Computer vision with the OpenCV library
Gary Bradski and Adrian Kaehler · 2008
Cited alongside, same era.
The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression
Patrick Lucey, Jeffrey F Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews · 2010
Cited alongside, same era.
Web-enabled robots
Moritz Tenorth, Ulrich Klank, Dejan Pangercic, and Michael Beetz · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Cited alongside, same era.
Development of expressive robotic head for bipedal humanoid robot
Tatsuhiro Kishi, Takuya Otani, Nobutsuna Endo, Przemyslaw Kryczka, Kenji Hashimoto, K Nakata, and Atsuo Takanishi · 2012
A sketch interface for robust and natural robot control
Danelle Shah, Joseph Schneider, and Mark Campbell · 2012
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Later among the works it cites.
Deep spatial autoencoders for visuomotor learning
Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel · 2015
Later among the works it cites.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robot learning manipulation action plans by “Watching” unconstrained videos from world wide web
Yezhou Yang, Yi Li, Cornelia Fermuller, and Yiannis Aloimonos · 2015
Later among the works it cites.