Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (RL) has proven a powerful technique in many sequential decision making domains.
R. E. Kalman et al. , “A new approach to linear filtering and prediction problems,” Journal of basic Engineering , 1960
1960
Earlier work this paper cites.
Z. Balorda, “Reducing uncertainty of objects by robot pushing,” in ICRA 1990
1990
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization , 1992
1992
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research , 1996
1996
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Advances in neural information processing systems , 2000, pp. 1008–1014
2000
Earlier work this paper cites.
A. Bicchi and V. Kumar, “Robotic grasping and contact: a review,” in ICRA 2000
2000
Earlier work this paper cites.
D. Precup, R. S. Sutton, and S. Dasgupta, “Off-policy temporal-difference learning with function approximation,” in ICML , 2001
2001
Earlier work this paper cites.
T. Schlegl, M. Buss, T. Omata, and G. Schmidt, “Fast dextrous re-grasping with optimal contact forces and contact sensor-based impedance control,” in ICRA 2001
2001
Earlier work this paper cites.
J. Z. Kolter and A. Y. Ng, “Learning omnidirectional path following using dimensionality reduction.” in RSS , 2007
2007
Earlier work this paper cites.
K.-J. Oh, S. Yea, and Y.-S. Ho, “Hole filling method using depth based in-painting for view synthesis in free viewpoint television and 3-d video,” in PCS 2009
2009
Earlier work this paper cites.
B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and autonomous systems , 2009
2009
Earlier work this paper cites.
S. Ross, G. J. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in International Conference on Artificial Intelligence and Statistics , 2011
2011
Earlier work this paper cites.
M. Dogar and S. Srinivasa, “A framework for push-grasping in clutter,” Robotics: Science and Systems (RSS) , 2011
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in IROS , 2012
2012
Earlier work this paper cites.
N.-E. Yang, Y.-G. Kim, and R.-H. Park, “Depth hole filling using the depth distribution of neighboring regions of depth holes in the kinect sensor,” in ICSPCC 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NIPS 2012
2012
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR 2014
2014
Cited alongside, same era.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in ICML 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
V. Mnih et al. , “Human-level control through deep reinforcement learning,” Nature , 2015
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in ICML 2015
2015
Cited alongside, same era.
S. R. Richter, V. Vineet, S. Roth, and V. Koltun, “Playing for data: Ground truth from computer games,” in ECCV 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
L. Pinto and A. Gupta, “Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours,” ICRA 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
M. Cutler and J. P. How, “Efficient reinforcement learning for robots using informative simulated priors,” in ICRA 2015
2015
Cited alongside, same era.
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” in ICML 2015
2015
Cited alongside, same era.
D. Silver et al. , “Mastering the game of go with deep neural networks and tree search,” Nature , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Closest in time.
D. Held, Z. McCarthy, M. Zhang, F. Shentu, and P. Abbeel, “Probabilistically safe policy transfer,” ICRA 2017
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, “Robust adversarial reinforcement learning,” ICML , 2017
2017
Closest in time.
2017
Closest in time.
L. Pinto and A. Gupta, “Learning to push by grasping: Using multiple tasks for effective learning,” in ICRA 2017
2017
Closest in time.