Fetching the paper…
Reading the bibliography…
We present R+X, a framework which enables robots to learn skills from long, unlabelled, first-person videos of humans performing everyday tasks.
1908
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2015
2015
Earlier work this paper cites.
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos, “Orb-slam: A versatile and accurate monocular slam system,” IEEE Transactions on Robotics , vol. 31, no. 5, p. 1147–1163, Oct. 2015. [Online]. Available: http://dx.doi.org/10.1109/TRO.2015.2463671
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” 2015
2015
Earlier work this paper cites.
J. Romero, D. Tzionas, and M. J. Black, “Embodied hands: modeling and capturing hands and bodies together,” ACM Transactions on Graphics , vol. 36, no. 6, p. 1–17, Nov. 2017. [Online]. Available: http://dx.doi.org/10.1145/3130800.3130883
2017
Earlier work this paper cites.
T. Brown et al. , “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , H. Larochelle et al. , Eds., vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
2020
Earlier work this paper cites.
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg, “Concept2robot: Learning manipulation concepts from instructions and human demonstrations,” The International Journal of Robotics Research , vol. 40, pp. 1419 – 1434, 2020. [Online]. Available: https://api.semanticscholar.org/CorpusID:220069237
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Amir, Y. Gandelsman, S. Bagon, and T. Dekel, “Deep vit features as dense visual descriptors,” 2022
2022
Earlier work this paper cites.
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta, “R3m: A universal visual representation for robot manipulation,” 2022
2022
Earlier work this paper cites.
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. van den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. de Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. W. Rae, E. Elsen, and L. Sifre, “Improving language models by retrieving from trillions of tokens,” 2022
2022
Cited alongside, same era.
T. Lüddecke and A. S. Ecker, “Image segmentation using text and image prompts,” 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
D. Driess et al. , “PaLM-e: An embodied multimodal language model,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, A. Krause et al. , Eds., vol. 202. PMLR, 23–29 Jul 2023, pp. 8469–8488. [Online]. Available: https://proceedings.mlr.press/v202/driess23a.html
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Gemini-Team, “Gemini: A family of highly capable multimodal models,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
OpenAI, “GPT-4 Technical Report,” arXiv e-prints , p. arXiv:2303.08774, Mar. 2023
2023
Cited alongside, same era.
G. Pavlakos, D. Shan, I. Radosavovic, A. Kanazawa, D. Fouhey, and J. Malik, “Reconstructing hands in 3d with transformers,” 2023
2023
Cited alongside, same era.
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
G. Team, “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” 2024
2024
Closest in time.
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” 2024
2024
Closest in time.
L. Wang, X. Zhang, H. Su, and J. Zhu, “A comprehensive survey of continual learning: Theory, method and application,” 2024
2024
Closest in time.
Y. Zhu, Z. Ou, X. Mou, and J. Tang, “Retrieval-augmented embodied agents,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Matsuki, R. Murai, P. H. J. Kelly, and A. J. Davison, “Gaussian splatting slam,” 2024
2024
Closest in time.
G. Papagiannis and E. Johns, “Miles: Making imitation learning easy with self-supervision,” in Proceedings of the Conference on Robot Learning (CoRL) , 2024
2024
Closest in time.
J. Luo, Z. Hu, C. Xu, Y. L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine, “Serl: A software suite for sample-efficient robotic reinforcement learning,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 16 961–16 969
2024
Closest in time.