Fetching the paper…
Reading the bibliography…
Good pre-trained visual representations could enable robots to learn visuomotor policy efficiently.
Object recognition from local scale-invariant features
D. G. Lowe · 1999
Earlier work this paper cites.
Histograms of oriented gradients for human detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
Efficient visual search of videos cast as text retrieval
J. Sivic and A. Zisserman · 2008
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie · 2016
Earlier work this paper cites.
Discovering objects and their relations from entangled scene representations
D. Raposo, A. Santoro, D. Barrett, R. Pascanu, T. Lillicrap, and P. Battaglia · 2017
Earlier work this paper cites.
See and think: Disentangling semantic scene completion
S. Liu, Y. Hu, Y. Zeng, Q. Tang, B. Jin, Y. Han, and X. Li · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Earlier work this paper cites.
Flexible neural representation for physics prediction
D. Mrowca, C. Zhuang, E. Wang, N. Haber, L. Fei-Fei, J. B. Tenenbaum, and D. Yamins · 2018
Earlier work this paper cites.
Learning particle dynamics for manipulating rigid bodies, deformable objects, and fluids
Y. Li, J. Wu, R. Tedrake, J. B. Tenenbaum, and A. Torralba · 2018
Earlier work this paper cites.
Theory and evaluation metrics for learning disentangled representations
K. Do and T. Tran · 2019
Earlier work this paper cites.
Are disentangled representations helpful for abstract visual reasoning?
S. Van Steenkiste, F. Locatello, J. Schmidhuber, and O. Bachem · 2019
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2019
Earlier work this paper cites.
Contrastive learning of structured world models
T. Kipf, E. van der Pol, and M. Welling · 2019
Earlier work this paper cites.
Multi-object representation learning with iterative variational inference
K. Greff, R. L. Kaufman, R. Kabra, N. Watters, C. P. Burgess, D. Zoran, L. Matthey, M. M. Botvinick, and A. Lerchner · 2019
Earlier work this paper cites.
Genesis: Generative scene inference and sampling with object-centric latent representations
M. Engelcke, A. R. Kosiorek, O. P. Jones, and I. Posner · 2019
Earlier work this paper cites.
Unsupervised learning of object keypoints for perception and control
T. D. Kulkarni, A. Gupta, C. Ionescu, S. Borgeaud, M. Reynolds, A. Zisserman, and V. Mnih · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. E. Hinton · 2020
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
X. Chen, H. Fan, R. Girshick, and K. He · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin · 2020
Cited alongside, same era.
Object-centric learning with slot attention
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
Sornet: Spatial object-centric representations for sequential manipulation
W. Yuan, C. Paxton, K. Desingh, and D. Fox · 2021
Cited alongside, same era.
Hierarchical object map estimation for efficient and robust navigation
K. Ok, K. Liu, and N. Roy · 2021
What makes pre-trained visual representations successful for robust manipulation?
K. Burns, Z. Witzel, J. I. Hamid, T. Yu, C. Finn, and K. Hausman · 2023
Later among the works it cites.
Distilled feature fields enable few-shot language-guided manipulation
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola · 2023
Later among the works it cites.
What makes pre-trained visual representations successful for robust manipulation?
K. Burns, Z. Witzel, J. I. Hamid, T. Yu, C. Finn, and K. Hausman · 2023
Later among the works it cites.
For pre-trained vision models in motor control, not all policy learning methods are created equal
Y. Hu, R. Wang, L. E. Li, and Y. Gao · 2023
Later among the works it cites.
Unsupervised open-vocabulary object localization in videos
K. Fan, Z. Bai, T. Xiao, D. Zietlow, M. Horn, Z. Zhao, C.-J. Simon-Gabriel, M. Z. Shou, F. Locatello, B. Schiele, T. Brox, Z. Zhang, Y. Fu, and T. He · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-task reinforcement learning with context-based representations
S. Sodhani, A. Zhang, and J. Pineau · 2021
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Cited alongside, same era.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
N. Hansen, Z. Yuan, Y. Ze, T. Mu, A. Rajeswaran, H. Su, H. Xu, and X. Wang · 2022
Cited alongside, same era.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
N. Hansen, Z. Yuan, Y. Ze, T. Mu, A. Rajeswaran, H. Su, H. Xu, and X. Wang · 2022
Cited alongside, same era.
Illiterate dall-e learns to compose
G. Singh, F. Deng, and S. Ahn · 2022
Cited alongside, same era.
Discovering deformable keypoint pyramids
J. Qian, A. Panagopoulos, and D. Jayaraman · 2022
Cited alongside, same era.
Later among the works it cites.
Learning generalizable manipulation policies with object-centric 3d representations
Y. Zhu, Z. Jiang, P. Stone, and Y. Zhu · 2023
Later among the works it cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. B. Girshick · 2023
Later among the works it cites.
Dynamic-resolution model learning for object pile manipulation
Y. Wang, Y. Li, K. R. Driggs-Campbell, L. Fei-Fei, and J. Wu · 2023
Later among the works it cites.
Partmanip: Learning cross-category generalizable part manipulation policy from point cloud observations
H. Geng, Z. Li, Y. Geng, J. Chen, H. Dong, and H. Wang · 2023
Later among the works it cites.
Kite: Keypoint-conditioned policies for semantic manipulation
P. Sundaresan, S. Belkhale, D. Sadigh, and J. Bohg · 2023
Later among the works it cites.
Slap: Spatial-language attention policies
P. Parashar, V. Jain, X. Zhang, J. Vakil, S. Powers, Y. Bisk, and C. Paxton · 2023
Later among the works it cites.
Selective visual representations improve convergence and generalization for embodied ai
A. Eftekhar, K.-H. Zeng, J. Duan, A. Farhadi, A. Kembhavi, and R. Krishna · 2023
Later among the works it cites.
Volumetric disentanglement for 3d scene manipulation
S. Benaim, F. Warburg, P. E. Christensen, and S. Belongie · 2024
Closest in time.
Plug-and-play object-centric representations from “what” and “where” foundation models
J. Shi*, J. Qian*, Y. J. Ma, and D. Jayaraman · 2024
Closest in time.
Recasting generic pretrained vision transformers as object-centric scene encoders for manipulation policies
J. Qian, A. Panagopoulos, and D. Jayaraman · 2024
Closest in time.
Mrest: Multi-resolution sensing for real-time control with vision-language models
S. Saxena, M. Sharma, and O. Kroemer · 2024
Closest in time.
Composable part-based manipulation
W. Liu, J. Mao, J. Hsu, T. Hermans, A. Garg, and J. Wu · 2024
Closest in time.
Task-conditioned adaptation of visual features in multi-task policy learning
P. Marza, L. Matignon, O. Simonin, and C. Wolf · 2024
Closest in time.
Vision-language models provide promptable representations for reinforcement learning
W. Chen, O. Mees, A. Kumar, and S. Levine · 2024
Closest in time.
Explorllm: Guiding exploration in reinforcement learning with large language models
R. Ma, J. Luijkx, Z. Ajanovic, and J. Kober · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang · 2024
Closest in time.