Fetching the paper…
Reading the bibliography…
Visual representation learning hold great promise for robotics, but is severely hampered by the scarcity and homogeneity of robotics datasets.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
S. Schaal · 1999
Earlier work this paper cites.
Survey: Robot programming by demonstration
A. Billard, S. Calinon, R. Dillmann, and S. Schaal · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Robot learning in homes: Improving generalization and reducing dataset bias
A. Gupta, A. Murali, D. P. Gandhi, and L. Pinto · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen · 2018
Earlier work this paper cites.
Visual reinforcement learning with imagined goals
A. V. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Earlier work this paper cites.
Policy optimization with demonstrations
B. Kang, Z. Jie, and J. Feng · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Earlier work this paper cites.
From virtual demonstration to real-world manipulation using lstm and mdn
R. Rahmatizadeh, P. Abolghasemi, A. Behal, and L. Bölöni · 2018
Cited alongside, same era.
W. Whitney, R. Agarwal, K. Cho, and A. Gupta · 2019
Cited alongside, same era.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Cited alongside, same era.
Relay policy learning: Solving long horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Cited alongside, same era.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2022
Later among the works it cites.
Simple but effective: Clip embeddings for embodied ai
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi · 2022
Later among the works it cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On network design spaces for visual recognition
I. Radosavovic, J. Johnson, S. X. W.-Y. Lo, and P. Dollár · 2019
Cited alongside, same era.
Learning by cheating
D. Chen, B. Zhou, V. Koltun, and P. Krähenbühl · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
A short note on the kinetics-700-2020 human action dataset
L. Smaira, J. Carreira, E. Noland, E. Clancy, A. Wu, and A. Zisserman · 2020
Cited alongside, same era.
Understanding human hands in contact at internet scale
D. Shan, J. Geng, M. Shu, and D. Fouhey · 2020
Cited alongside, same era.
Designing network design spaces
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár · 2020
Cited alongside, same era.
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Later among the works it cites.
BEit: BERT pre-training of image transformers
H. Bao, L. Dong, S. Piao, and F. Wei · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Z. Tong, Y. Song, J. Wang, and L. Wang · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
C. Feichtenhofer, Y. Li, K. He, et al · 2022
Later among the works it cites.
Mavil: Masked audio-video learners
P.-Y. Huang, V. Sharma, H. Xu, C. Ryali, H. Fan, Y. Li, S.-W. Li, G. Ghosh, J. Malik, and C. Feichtenhofer · 2022
Later among the works it cites.
Multimae: Multi-modal multi-task masked autoencoders
R. Bachmann, D. Mizrahi, A. Atanov, and A. Zamir · 2022
Later among the works it cites.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Later among the works it cites.
Teach a robot to fish: Versatile imitation from one minute of demonstrations
S. Haldar, J. Pari, A. Rai, and L. Pinto · 2023
Closest in time.
Where are we in the search for an artificial visual cortex for embodied intelligence?
A. Majumdar, K. Yadav, S. Arnaud, Y. J. Ma, C. Chen, S. Silwal, A. Jain, V.-P. Berges, P. Abbeel, J. Malik, et al · 2023
Closest in time.
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song · 2023
Closest in time.
Videomae v2: Scaling video masked autoencoders with dual masking
L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao · 2023
Closest in time.
The effectiveness of mae pre-pretraining for billion-scale pretraining
M. Singh, Q. Duval, K. V. Alwala, H. Fan, V. Aggarwal, A. Adcock, A. Joulin, P. Dollár, C. Feichtenhofer, R. Girshick, et al · 2023
Closest in time.
R-mae: Regions meet masked autoencoders
D.-K. Nguyen, V. Aggarwal, Y. Li, M. R. Oswald, A. Kirillov, C. G. Snoek, and X. Chen · 2023
Closest in time.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
N. Hansen, Z. Yuan, Y. Ze, T. Mu, A. Rajeswaran, H. Su, H. Xu, and X. Wang · 2023
Closest in time.
Train offline, test online: A real robot learning benchmark
G. Zhou, V. Dean, M. K. Srirama, A. Rajeswaran, J. Pari, K. Hatch, A. Jain, T. Yu, P. Abbeel, L. Pinto, et al · 2023
Closest in time.