Fetching the paper…
Reading the bibliography…
We consider real-world reinforcement learning (RL) of robotic manipulation tasks that involve both visuomotor skills and contact-rich skills.
R. P. Rao and D. H. Ballard, “Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects,” Nature neuroscience , vol. 2, no. 1, pp. 79–87, 1999
1999
Earlier work this paper cites.
E. N. Eskandar and J. A. Assad, “Dissociation of visual, motor and predictive signals in parietal cortex during visual guidance,” Nature neuroscience , vol. 2, no. 1, pp. 88–93, 1999
1999
Earlier work this paper cites.
E. P. Pednault, “Representation is everything,” Communications of the ACM , vol. 43, no. 8, pp. 80–83, 2000
2000
Earlier work this paper cites.
T. S. Lee and D. Mumford, “Hierarchical bayesian inference in the visual cortex,” JOSA A , vol. 20, no. 7, pp. 1434–1448, 2003
2003
Earlier work this paper cites.
R. S. Sutton and B. Tanner, “Temporal-difference networks,” in Advances in neural information processing systems , 2005, pp. 1377–1384
2005
Earlier work this paper cites.
F. Chaumette and S. Hutchinson, “Visual servo control. ii. advanced approaches [tutorial],” IEEE Robotics & Automation Magazine , vol. 14, no. 1, pp. 109–118, 2007
2007
Earlier work this paper cites.
M. Ponsen, M. E. Taylor, and K. Tuyls, “Abstraction and generalization in reinforcement learning: A summary and framework,” in International Workshop on Adaptive and Learning Agents . Springer, 2009, pp. 1–32
2009
Earlier work this paper cites.
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup, “Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction,” in The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2 , 2011, pp. 761–768
2011
Earlier work this paper cites.
——, “Whatever next? predictive brains, situated agents, and the future of cognitive science,” Behavioral and brain sciences , vol. 36, no. 3, pp. 181–204, 2013
2013
Earlier work this paper cites.
T. Schaul and M. Ring, “Better generalization with forecasts,” in Twenty-Third International Joint Conference on Artificial Intelligence , 2013
2013
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning , 2013, pp. 1–9
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. Clark, “Perceiving as predicting,” Perception and its modalities , pp. 23–43, 2014
2014
Earlier work this paper cites.
H.-C. Song, Y.-L. Kim, and J.-B. Song, “Automated guidance of peg-in-hole assembly tasks for complex-shaped parts,” in 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2014, pp. 4517–4522
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
T. Tang, H.-C. Lin, Y. Zhao, Y. Fan, W. Chen, and M. Tomizuka, “Teach industrial robots peg-hole-insertion by human demonstration,” in 2016 IEEE International Conference on Advanced Intelligent Mechatronics (AIM) . IEEE, 2016, pp. 488–494
2016
Earlier work this paper cites.
A. Patterson and M. Schlegel, “A comparison of general value functions and temporal-difference networks,” in International conference on autonomous agents and multiagent systems (AAMAS) , 2016
2016
Earlier work this paper cites.
S. Gu, E. Holly, T. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,” in 2017 IEEE international conference on robotics and automation (ICRA) . IEEE, 2017, pp. 3389–3396
2017
Cited alongside, same era.
2017
Cited alongside, same era.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba, “Hindsight experience replay,” in Advances in neural information processing systems , 2017, pp. 5048–5058
2017
Cited alongside, same era.
T. Inoue, G. De Magistris, A. Munawar, T. Yokoya, and R. Tachibana, “Deep reinforcement learning for high precision assembly tasks,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2017, pp. 819–825
M. Vecerik, O. Sushkov, D. Barker, T. Rothörl, T. Hester, and J. Scholz, “A practical approach to insertion with variable socket position using deep reinforcement learning,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 754–760
2019
Later among the works it cites.
X. Wu, D. Zhang, F. Qin, and D. Xu, “Deep reinforcement learning of robotic precision insertion skill accelerated by demonstrations,” in 2019 IEEE 15th International Conference on Automation Science and Engineering (CASE) . IEEE, 2019, pp. 1651–1656
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Jin, L. Petrich, M. Dehghan, Z. Zhang, and M. Jagersand, “Robot eye-hand coordination learning by watching human demonstrations: a task function approximation approach,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 6624–6630
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
A. R. Mahmood, D. Korenkevych, B. J. Komer, and J. Bergstra, “Setting up a reinforcement learning task with a real-world robot,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 4635–4640
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 1–8
2018
Cited alongside, same era.
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine, “Variational inverse control with events: A general framework for data-driven reward definition,” in Advances in Neural Information Processing Systems , 2018, pp. 8538–8547
2018
Cited alongside, same era.
Z. Hou, M. Philipp, K. Zhang, Y. Guan, K. Chen, and J. Xu, “The learning-based optimization algorithm for robotic dual peg-in-hole assembly,” Assembly Automation , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in International conference on machine learning . PMLR, 2019, pp. 2555–2565
2019
Later among the works it cites.
2019
Later among the works it cites.
V. Veeriah, M. Hessel, Z. Xu, J. Rajendran, R. L. Lewis, J. Oh, H. P. van Hasselt, D. Silver, and S. Singh, “Discovery of useful questions as auxiliary tasks,” in Advances in Neural Information Processing Systems , 2019, pp. 9310–9321
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
J. Jin, N. M. Nguyen, N. Sakib, D. Graves, H. Yao, and M. Jagersand, “Mapless navigation among dynamics with social-safety-awareness: a reinforcement learning approach from 2d laser scans,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 6979–6985
2020
Closest in time.
M. Okada and T. Taniguchi, “Dreaming: Model-based reinforcement learning by latent imagination without reconstruction,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 4209–4215
2021
Closest in time.
D. Graves, N. M. Nguyen, K. Hassanzadeh, J. Jin, and J. Luo, “Learning robust driving policies without online exploration,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 13 186–13 193
2021
Closest in time.
J. Luo, E. Solowjow, C. Wen, J. A. Ojea, and A. M. Agogino, “Deep reinforcement learning for robotic assembly of mixed deformable and rigid objects,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 2062–2069
2069
Closest in time.