Fetching the paper…
Reading the bibliography…
Learning from diverse offline datasets is a promising path towards learning general purpose robotic agents.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR09 , 2009
2009
Earlier work this paper cites.
J. Schmidhuber, “Formal theory of creativity, fun, and intrinsic motivation (1990–2010),” IEEE Transactions on Autonomous Mental Development , vol. 2, no. 3, 2010
2010
Earlier work this paper cites.
M. Lopes, T. Lang, M. Toussaint, and P. yves Oudeyer, “Exploration in model-based reinforcement learning by empirically estimating learning progress,” in Advances in Neural Information Processing Systems 25 , F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2012
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012
2012
Earlier work this paper cites.
R. Y. Rubinstein and D. P. Kroese, The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning . Springer Science & Business Media, 2013
2013
Earlier work this paper cites.
B. Piot, M. Geist, and O. Pietquin, “Boosted bellman residual minimization handling expert demonstrations,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2014, pp. 549–564
2014
Earlier work this paper cites.
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” in Advances in neural information processing systems , 2016
2016
Earlier work this paper cites.
L. Pinto and A. Gupta, “Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours,” in IEEE international conference on robotics and automation (ICRA) , 2016
2016
Earlier work this paper cites.
K.-T. Yu, M. Bauza, N. Fazeli, and A. Rodriguez, “More than a million ways to be pushed. a high-fidelity experimental dataset of planar pushing,” 2016
2016
Earlier work this paper cites.
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine, “Learning to poke by poking: Experiential learning of intuitive physics,” in Advances in neural information processing systems , 2016
2016
Earlier work this paper cites.
Y. Chebotar, K. Hausman, Z. Su, A. Molchanov, O. Kroemer, G. S. Sukhatme, and S. Schaal, “Bigs: Biotac grasp stability dataset,” 2016
2016
Earlier work this paper cites.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Vime: Variational information maximizing exploration,” in Advances in Neural Information Processing Systems , 2016
2016
Earlier work this paper cites.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in ICML , 2017
2017
Earlier work this paper cites.
C. Finn and S. Levine, “Deep visual foresight for planning robot motion,” in IEEE International Conference on Robotics and Automation (ICRA) , 2017
2017
Earlier work this paper cites.
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. X. Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel, “# exploration: A study of count-based exploration for deep reinforcement learning,” in Advances in neural information processing systems , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, S. Savarese, and L. Fei-Fei, “Roboturk: A crowdsourcing platform for robotic skill learning through imitation,” in Conference on Robot Learning , 2018
2018
Cited alongside, same era.
A. Zeng, S. Song, S. Welker, J. Lee, A. Rodriguez, and T. Funkhouser, “Learning synergies between pushing and grasping with self-supervised deep reinforcement learning,” Proceedings of the IEEE International Conference on Intelligent Robots and Systems (IROS) , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Gupta, A. Murali, D. P. Gandhi, and L. Pinto, “Robot learning in homes: Improving generalization and reducing dataset bias,” in Advances in Neural Information Processing Systems , 2018
2018
Y. Burda, H. Edwards, A. Storkey, and O. Klimov, “Exploration by random network distillation,” in International Conference on Learning Representations , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
I. Osband, B. Van Roy, D. J. Russo, and Z. Wen, “Deep exploration via randomized value functions.” Journal of Machine Learning Research , vol. 20, no. 124, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine, “Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation,” 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen, “Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection,” The International Journal of Robotics Research , vol. 37, no. 4-5, 2018
2018
Cited alongside, same era.
P.-Y. Oudeyer, “Computational theories of curiosity-driven learning,” arXiv:1802.10546 , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine, “Variational inverse control with events: A general framework for data-driven reward definition,” in Advances in Neural Information Processing Systems , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=r1Ddp1-Rb
2018
Cited alongside, same era.
2019
Later among the works it cites.
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak, “Planning to explore via self-supervised world models,” in ICML , 2020
2020
Closest in time.
A. Mandlekar, F. Ramos, B. Boots, S. Savarese, L. Fei-Fei, A. Garg, and D. Fox, “Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data,” 2020 IEEE International Conference on Robotics and Automation (ICRA) , May 2020. [Online]. Available: http://dx.doi.org/10.1109/ICRA40945.2020.9196935
2020
Closest in time.
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine, “Skew-fit: State-covering self-supervised reinforcement learning,” in ICML , 2020
2020
Closest in time.
X. Chen, Y. Gao, A. Ghadirzadeh, M. Bjorkman, G. Castellano, and P. Jensfelt, “Skew-explore: Learn faster in continuous spaces with sparse rewards,” 2020
2020
Closest in time.
2020
Closest in time.
R. Simmons-Edler, B. Eisner, D. Yang, A. Bisulco, E. Mitchell, S. Seung, and D. Lee, “{QX}plore: Q-learning exploration by maximizing temporal difference error,” 2020
2020
Closest in time.
2020
Closest in time.
L. Lee, B. Eysenbach, R. Salakhutdinov, C. Finn, et al. , “Weakly-supervised reinforcement learning for controllable behavior,” in Advances in Neural Information Processing Systems , 2020
2020
Closest in time.
2020
Closest in time.
K. Koreitem, F. Shkurti, T. Manderson, W.-D. Chang, J. C. G. Higuera, and G. Dudek, “One-shot informed robotic visual search in the wild,” 2020
2020
Closest in time.
2020
Closest in time.
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in Conference on Robot Learning , 2020
2020
Closest in time.