Fetching the paper…
Reading the bibliography…
In this work, we explore self-supervised visual pre-training on images from diverse, in-the-wild videos for real-world robotic tasks.
Rrt-connect: An efficient approach to single-query path planning
J. J. Kuffner and S. M. LaValle · 2000
Earlier work this paper cites.
Automatic grasp planning using shape primitives
A. T. Miller, S. Knoop, H. I. Christensen, and P. K. Allen · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Wavenet: A generative model for raw audio
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Earlier work this paper cites.
The curious robot: Learning visual representations via physical interactions
L. Pinto, D. Gandhi, Y. Han, Y.-L. Park, and A. Gupta · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Transferring end-to-end visuomotor control from simulation to real world for a multi-stage task
S. James, A. J. Davison, and E. Johns · 2017
Earlier work this paper cites.
Unsupervised perceptual rewards for imitation learning
P. Sermanet, K. Xu, and S. Levine · 2017
Earlier work this paper cites.
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Learning synergies between pushing and grasping with self-supervised deep reinforcement learning
A. Zeng, S. Song, S. Welker, J. Lee, A. Rodriguez, and T. Funkhouser · 2018
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen · 2018
Cited alongside, same era.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
The surprising effectiveness of representation learning for visual imitation
J. Pari, N. Muhammad, S. P. Arunachalam, L. Pinto, et al · 2021
Later among the works it cites.
What should not be contrastive in contrastive learning
T. Xiao, X. Wang, A. A. Efros, and T. Darrell · 2021
Later among the works it cites.
Benchmarking detection transfer learning with vision transformers
Y. Li, S. Xie, X. Chen, P. Dollar, K. He, and R. Girshick · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Self-supervised correspondence in visuomotor policy learning
P. Florence, L. Manuelli, and R. Tedrake · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Jukebox: A generative model for music
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Understanding human hands in contact at internet scale
D. Shan, J. Geng, M. Shu, and D. F. Fouhey · 2020
Cited alongside, same era.
Learning dexterous in-hand manipulation
OpenAI, M. Andrychowicz, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba · 2020
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Early convolutions help transformers see better
T. Xiao, P. Dollar, M. Singh, E. Mintun, T. Darrell, and R. Girshick · 2021
Later among the works it cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al · 2021
Later among the works it cites.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Closest in time.
Q-attention: Enabling Efficient Learning for Vision-based Robotic Manipulation
S. James and A. J. Davison · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Closest in time.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Z. Tong, Y. Song, J. Wang, and L. Wang · 2022
Closest in time.
Masked autoencoders as spatiotemporal learners
C. Feichtenhofer, H. Fan, Y. Li, and K. He · 2022
Closest in time.
Multimae: Multi-modal multi-task masked autoencoders
R. Bachmann, D. Mizrahi, A. Atanov, and A. Zamir · 2022
Closest in time.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Closest in time.
Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via Discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Closest in time.
Dexterous imitation made easy: A learning-based framework for efficient dexterous manipulation
S. P. Arunachalam, S. Silwal, B. Evans, and L. Pinto · 2022
Closest in time.