Fetching the paper…
Reading the bibliography…
We present a self-supervised sensorimotor pre-training approach for robotics.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Self-supervised correspondence in visuomotor policy learning
P. Florence, L. Manuelli, and R. Tedrake · 2019
Earlier work this paper cites.
On-policy dataset synthesis for learning robot grasping policies using fully convolutional deep networks
V. Satish, J. Mahler, and K. Goldberg · 2019
Earlier work this paper cites.
Generative pretraining from pixels
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever · 2020
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Understanding human hands in contact at internet scale
D. Shan, J. Geng, M. Shu, and D. F. Fouhey · 2020
Cited alongside, same era.
Exploratory grasping: Asymptotically optimal algorithms for grasping challenging polyhedral objects
M. Danielczuk, A. Balakrishna, D. S. Brown, S. Devgon, and K. Goldberg · 2020
Cited alongside, same era.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2022
Later among the works it cites.
Learning visual robotic control efficiently with contrastive pre-training and data augmentation
A. Zhan, R. Zhao, L. Pinto, P. Abbeel, and M. Laskin · 2022
Later among the works it cites.
From play to policy: Conditional behavior generation from uncurated robot data
Z. J. Cui, Y. Wang, N. Muhammad, L. Pinto, et al · 2022
Later among the works it cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2021
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2021
Cited alongside, same era.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
D. Damen, H. Doughty, G. M. Farinella, , A. Furnari, J. Ma, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2021
Cited alongside, same era.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnkov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Cited alongside, same era.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Cited alongside, same era.
Pre-training for robots: Offline rl enables learning new tasks from a handful of trials
A. Kumar, A. Singh, F. Ebert, Y. Yang, C. Finn, and S. Levine · 2022
Later among the works it cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Later among the works it cites.
Multi-view masked world models for visual robotic manipulation
Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Closest in time.