Fetching the paper…
Reading the bibliography…
Vision-based robot policy learning, which maps visual inputs to actions, necessitates a holistic understanding of diverse visual tasks beyond single-task needs like classification or segmentation.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Manipulators and manipulation in high dimensional spaces
V. Kumar · 2016
Earlier work this paper cites.
The ”something something” video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Habitat: A platform for embodied ai research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, et al · 2019
Earlier work this paper cites.
Object-centric learning with slot attention
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin · 2020
Earlier work this paper cites.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2020
Earlier work this paper cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
D. Yarats, I. Kostrikov, and R. Fergus · 2020
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Earlier work this paper cites.
Understanding human hands in contact at internet scale
D. Shan, J. Geng, M. Shu, and D. Fouhey · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
Rrl: Resnet as representation for reinforcement learning
R. M. Shah and V. Kumar · 2021
Cited alongside, same era.
Self-supervised disentangled representation learning for third-person imitation learning
J. Shang and M. S. Ryoo · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Cited alongside, same era.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Cited alongside, same era.
Equidiff: A conditional equivariant diffusion model for trajectory prediction
K. Chen, X. Chen, Z. Yu, M. Zhu, and H. Yang · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Minigpt-v2: Large language model as a unified interface for vision-language multi-task learning
J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P. Zhang, R. Krishnamoorthi, V. Chandra, Y. Xiong, and M. Elhoseiny · 2023
Later among the works it cites.
u-llava: Unifying multi-modal tasks via large language model
J. Xu, L. Xu, Y. Yang, X. Li, Y. Xie, Y.-J. Huang, and Y. Li · 2023
Later among the works it cites.
0.1% data makes segment anything slim
Z. Chen, G. Fang, X. Ma, and X. Wang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Habitat 2.0: Training home assistants to rearrange their habitat
A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y. Zhao, J. Turner, N. Maestre, M. Mukadam, D. S. Chaplot, O. Maksymets, et al · 2021
Cited alongside, same era.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Cited alongside, same era.
Image BERT pre-training with online tokenizer
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next-generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Cited alongside, same era.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Cited alongside, same era.
Later among the works it cites.
Efficientsam: Leveraged masked image pretraining for efficient segment anything
Y. Xiong, B. Varadarajan, L. Wu, X. Xiang, F. Xiao, C. Zhu, X. Dai, D. Wang, F. Sun, F. Iandola, et al · 2023
Later among the works it cites.
Sam-clip: Merging vision foundation models towards semantic and spatial understanding
H. Wang, P. K. A. Vasu, F. Faghri, R. Vemulapalli, M. Farajtabar, S. Mehta, M. Rastegari, O. Tuzel, and H. Pouransari · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Later among the works it cites.
Interactive language: Talking to robots in real time
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al · 2023
Later among the works it cites.
Habitat-matterport 3d semantics dataset
K. Yadav, R. Ramrakhya, S. K. Ramakrishnan, T. Gervet, J. Turner, A. Gokaslan, N. Maestre, A. X. Chang, D. Batra, M. Savva, et al · 2023
Later among the works it cites.
Ovrl-v2: A simple state-of-art baseline for imagenav and objectnav
K. Yadav, A. Majumdar, R. Ramrakhya, N. Yokoyama, A. Baevski, Z. Kira, O. Maksymets, and D. Batra · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Decomposing the generalization gap in imitation learning for visual robotic manipulation, 2023
A. Xie, L. Lee, T. Xiao, and C. Finn · 2023
Later among the works it cites.
Depth anything: Unleashing the power of large-scale unlabeled data
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao · 2024
Closest in time.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Where are we in the search for an artificial visual cortex for embodied intelligence?
A. Majumdar, K. Yadav, S. Arnaud, J. Ma, C. Chen, S. Silwal, A. Jain, V.-P. Berges, T. Wu, J. Vakil, et al · 2024
Closest in time.
Vision transformers need registers
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski · 2024
Closest in time.
Llara: Supercharging robot learning data for vision-language policy
X. Li, C. Mata, J. Park, K. Kahatapitiya, Y. S. Jang, J. Shang, K. Ranasinghe, R. Burgert, M. Cai, Y. J. Lee, and M. S. Ryoo · 2024
Closest in time.
Openvla: An open-source vision-language-action model
M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Closest in time.
J. Shi, J. Qian, Y. J. Ma, and D. Jayaraman · 2024
Closest in time.
Efficientvit-sam: Accelerated segment anything model without performance loss
Z. Zhang, H. Cai, and S. Han · 2024
Closest in time.
Am-radio: Agglomerative vision foundation model reduce all domains into one
M. Ranzinger, G. Heinrich, J. Kautz, and P. Molchanov · 2024
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2024
Closest in time.
Probing the 3D Awareness of Visual Foundation Models
M. El Banani, A. Raj, K.-K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani · 2024
Closest in time.