Fetching the paper…
Reading the bibliography…
Vision-language-action (VLA) models trained on large-scale internet data and robot demonstrations have the potential to serve as generalist robot policies.
G. Shafer and V. Vovk, “A tutorial on conformal prediction.” Journal of Machine Learning Research (JMLR) , 2008
2008
Earlier work this paper cites.
D. Bahdanau, “Neural machine translation by jointly learning to align and translate,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
M. T. Ribeiro, S. Singh, and C. Guestrin, “" why should i trust you?" explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining , 2016
2016
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-Cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Proceedings of Advances in Neural Information Processing Systems (NIPS) , 2017
2017
Earlier work this paper cites.
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proceedings of Advances in Neural Information Processing Systems (NIPS) , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Igl, K. Ciosek, Y. Li, S. Tschiatschek, C. Zhang, S. Devlin, and K. Hofmann, “Generalization in reinforcement learning with selective noise injection and information bottleneck,” Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2019
2019
Earlier work this paper cites.
A. Ghorbani, A. Abid, and J. Zou, “Interpretation of neural networks is fragile,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2019
2019
Earlier work this paper cites.
P.-J. Kindermans, S. Hooker, J. Adebayo, M. Alber, K. T. Schütt, S. Dähne, D. Erhan, and B. Kim, “The (un) reliability of saliency methods,” Explainable AI: Interpreting, Explaining and Visualizing Deep Learning , 2019
2019
Earlier work this paper cites.
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas, “Reinforcement learning with augmented data,” Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2020
2020
Earlier work this paper cites.
Y. Tang, D. Nguyen, and D. Ha, “Neuroevolution of self-interpretable agents,” in Proceedings of the Genetic and Evolutionary Computation Conference , 2020
2020
Earlier work this paper cites.
V. Pacelli and A. Majumdar, “Learning task-driven control policies via information bottlenecks,” in Proceedings of Robotics: Science and Systems (RSS) , 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine, “Learning invariant representations for reinforcement learning without reconstruction,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
R. Agarwal, M. C. Machado, P. S. Castro, and M. G. Bellemare, “Contrastive behavioral similarity embeddings for generalization in reinforcement learning,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2021
2021
Cited alongside, same era.
H. Chefer, S. Gur, and L. Wolf, “Transformer interpretability beyond attention visualization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Cited alongside, same era.
2022
Cited alongside, same era.
S. James and A. J. Davison, “Q-attention: Enabling efficient learning for vision-based robotic manipulation,” IEEE Robotics and Automation Letters , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al. , “Bridgedata v2: A dataset for robot learning at scale,” in Proceedings of Conference on Robot Learning (CoRL) , 2023
2023
Later among the works it cites.
A. N. Angelopoulos and S. Bates, “A gentle introduction to conformal prediction and distribution-free uncertainty quantification,” Foundations and Trends in Machine Learning , 2023
2023
Later among the works it cites.
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” in Proceedings of Robotics: Science and Systems (RSS) , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Elhage, T. Hume, C. Olsson, N. Nanda, T. Henighan, S. Johnston, S. ElShowk, N. Joseph, N. DasSarma, B. Mann, D. Hernandez, A. Askell, K. Ndousse, A. Jones, D. Drain, A. Chen, Y. Bai, D. Ganguli, L. Lovitt, Z. Hatfield-Dodds, J. Kernion, T. Conerly, S. Kravec, S. Fort, S. Kadavath, J. Jacobson, E. Tran-Johnson, J. Kaplan, J. Clark, T. Brown, S. McCandlish, D. Amodei, and C. Olah, “Softmax linear units,” Transformer Circuits Thread , 2022, https://transformer-circuits.pub/2022/solu/index.html
2022
Cited alongside, same era.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al. , “RT-1: Robotics transformer for real-world control at scale,” in Proceedings of Robotics: Science and Systems (RSS) , 2023
2023
Cited alongside, same era.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al. , “RT-2: Vision-language-action models transfer web knowledge to robotic control,” in Proceedings of the Conference on Robot Learning (CoRL) , 2023
2023
Cited alongside, same era.
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al. , “Do as i can, not as i say: Grounding language in robotic affordances,” in Proceedings of Conference on robot learning (CoRL) , 2023
2023
Cited alongside, same era.
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al. , “Scaling robot learning with semantically imagined experience,” in Proceedings of Robotics: Science and Systems (RSS) , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, S. Kirmani, B. Zitkovich, F. Xia, et al. , “Open-world object manipulation using pre-trained vision-language models,” in Proceedings of the Conference on Robot Learning (CoRL) , 2023
2023
Cited alongside, same era.
Y. Zhu, Z. Jiang, P. Stone, and Y. Zhu, “Learning generalizable manipulation policies with object-centric 3d representations,” in Proceedings of the Conference on Robot Learning (CoRL) , 2023
2023
Cited alongside, same era.
2024
Closest in time.
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al. , “Open x-embodiment: Robotic learning datasets and rt-x models,” in Proceedings of IEEE International Conference on Robotics and Automation (ICRA) , 2024
2024
Closest in time.
2024
Closest in time.
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al. , “Droid: A large-scale in-the-wild robot manipulation dataset,” in Proceedings of Robotics: Science and Systems (RSS) , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
B. Grooten, T. Tomilin, G. Vasan, M. E. Taylor, A. R. Mahmood, M. Fang, M. Pechenizkiy, and D. C. Mocanu, “Madi: Learning to mask distractions for generalization in visual deep reinforcement learning,” in Proceedings of International Foundation for Autonomous Agents and Multiagent Systems (AAMAS) , 2024
2024
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Closest in time.
2024
Closest in time.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2024
2024
Closest in time.
T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang, “Grounded sam: Assembling open-world models for diverse visual tasks,” 2024
2024
Closest in time.
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan, “Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet,” Transformer Circuits Thread , 2024. [Online]. Available: https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
2024
Closest in time.