Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning.
On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function
Aigner, D. J., Amemiya, T., and Poirier, D. J · 1976
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Sutton, R. S · 1984
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S · 2007
Earlier work this paper cites.
Orb: An efficient alternative to sift or surf
Rublee, E., Rabaud, V., Konolige, K., and Bradski, G · 2011
Earlier work this paper cites.
Geoadditive expectile regression
Sobotka, F. and Kneib, T · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Sohn, K., Lee, H., and Yan, X · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Reinformer: Max-return sequence modeling for offline rl
Zhuang, Z., Peng, D., Liu, J., Zhang, Z., and Wang, D · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Earlier work this paper cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, F., Rae, J., Pascanu, R., Gulcehre, C., Jayakumar, S., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Cited alongside, same era.
Catformer: Designing stable transformers via sensitivity analysis
Davis, J. Q., Gu, A., Choromanski, K., Dao, T., Re, C., Finn, C., and Liang, P · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Vuong, Q., Levine, S., Walke, H. R., Pertsch, K., Singh, A., Doshi, R., Xu, C., Luo, J., Tan, L., Shah, D., et al · 2023
Later among the works it cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., et al · 2023
Later among the works it cites.
Quar-vla: Vision-language-action model for quadruped robots
Ding, P., Zhao, H., Zhang, W., Song, W., Zhang, M., Huang, S., Yang, N., and Wang, D · 2024
Later among the works it cites.
On transforming reinforcement learning with transformers: The development trajectory
Hu, S., Shen, L., Zhang, Y., Chen, Y., and Tao, D · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Cited alongside, same era.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
James, S. and Davison, A. J · 2022
Cited alongside, same era.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Mees, O., Hermann, L., Rosete-Beas, E., and Burgard, W · 2022
Cited alongside, same era.
Behavior transformers: Cloning k k modes with one stone
Shafiullah, N. M., Cui, Z., Altanzaya, A. A., and Pinto, L · 2022
Cited alongside, same era.
Mark, M. S., Gao, T., Sampaio, G. G., Srirama, M. K., Sharma, A., Finn, C., and Kumar, A · 2024
Later among the works it cites.
Predictive inverse dynamics models are scalable learners for robotic manipulation
Tian, Y., Yang, S., Zeng, J., Wang, P., Lin, D., Dong, H., and Pang, J · 2024
Later among the works it cites.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Zhai, S., Bai, H., Lin, Z., Pan, J., Tong, P., Zhou, Y., Suhr, A., Xie, S., LeCun, Y., Ma, Y., et al · 2024
Later among the works it cites.
Grape: Generalizing robot policy via preference alignment
Zhang, Z., Zheng, K., Chen, Z., Jang, J., Li, Y., Wang, C., Ding, M., Fox, D., and Yao, H · 2024
Later among the works it cites.
Vlmpc: Vision-language model predictive control for robotic manipulation
Zhao, W., Chen, J., Meng, Z., Mao, D., Song, R., and Zhang, W · 2024
Later among the works it cites.
Bai, S., Zhou, W., Ding, P., Zhao, W., Wang, D., and Chen, B · 2025
Closest in time.
Fdpp: Fine-tune diffusion policy with human preference
Chen, Y., Jha, D. K., Tomizuka, M., and Romeres, D · 2025
Closest in time.
Improving vision-language-action model with online reinforcement learning
Guo, Y., Zhang, J., Chen, X., Ji, X., Wang, Y.-J., Hu, Y., and Chen, J · 2025
Closest in time.
Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy
Li, P., Wu, H., Huang, Y., Cheang, C., Wang, L., and Kong, T · 2025
Closest in time.
Gevrm: Goal-expressive video generation model for robust visual manipulation
Zhang, H., Ding, P., Lyu, S., Peng, Y., and Wang, D · 2025
Closest in time.