Fetching the paper…
Reading the bibliography…
We introduce $\textbf{F}$uture $\textbf{LA}$tent $\textbf{RE}$presentation Alignment ($\textbf{FLARE}$), a novel framework that integrates predictive latent world modeling into robot policy learning.
When to trust your model: Model-based policy optimization, 2020
X. Jiang, Q. Chen, S. Han, M. Li, J. Dong, and R. Zhang · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Mastering atari with discrete world models
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2020
Earlier work this paper cites.
Pretraining representations for data-efficient reinforcement learning
M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, R. D. Hjelm, P. Bachman, and A. C. Courville · 2021
Earlier work this paper cites.
Temporal difference learning for model predictive control
N. Hansen, X. Wang, and H. Su · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Earlier work this paper cites.
Dexmv: Imitation learning for dexterous manipulation from human videos
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang · 2022
Earlier work this paper cites.
Interactive language: Talking to robots in real time, 2022
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence · 2022
Earlier work this paper cites.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. B. Tenenbaum, D. Schuurmans, and P. Abbeel · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Earlier work this paper cites.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Earlier work this paper cites.
COPlanner: Plan to roll out conservatively but to explore optimistically for model-based RL
X. Wang, R. Zheng, Y. Sun, R. Jia, W. Wongkamjan, H. Xu, and F. Huang · 2023
Earlier work this paper cites.
Is model ensemble necessary? model-based RL via a single model with lipschitz regularized value function
R. Zheng, X. Wang, H. Xu, and F. Huang · 2023
Earlier work this paper cites.
Taco: Temporal latent action-driven contrastive loss for visual reinforcement learning
R. Zheng, X. Wang, Y. Sun, S. Ma, J. Zhao, H. Xu, H. Daumé III, and F. Huang · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Earlier work this paper cites.
Vision-language foundation models as effective robot imitators
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y. Jing, W. Zhang, H. Liu, et al · 2023
Earlier work this paper cites.
Affordances from human videos as a versatile representation for robotics
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak · 2023
Earlier work this paper cites.
Deft: Dexterous fine-tuning for real-world hand policies
A. Kannan, K. Shaw, S. Bahl, P. Mannam, and D. Pathak · 2023
Earlier work this paper cites.
Videodex: Learning dexterity from internet videos
K. Shaw, S. Bahl, and D. Pathak · 2023
Earlier work this paper cites.
Any-point trajectory modeling for policy learning
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel · 2023
Earlier work this paper cites.
Mimicplay: Long-horizon imitation learning by watching human play
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar · 2023
Earlier work this paper cites.
Zero-shot robot manipulation from passive human videos
H. Bharadhwaj, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Earlier work this paper cites.
Learning continuous grasping function with a dexterous hand from human demonstrations
J. Ye, J. Wang, B. Huang, Y. Qin, and X. Wang · 2023
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine · 2023
Cited alongside, same era.
Mutex: Learning unified policies from multimodal task specifications
R. Shah, R. Martín-Martín, and Y. Zhu · 2023
Cited alongside, same era.
Plex: Making the most of the available data for robotic manipulation pretraining
G. Thomas, C.-A. Cheng, R. Loynd, F. V. Frujeri, V. Vineet, M. Jalobeanu, and A. Kolobov · 2023
Cited alongside, same era.
Unleashing large-scale video generative pre-training for visual robot manipulation
H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong · 2024
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
C.-L. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, H. Zhang, and M. Zhu · 2024
Cited alongside, same era.
Equibot: Sim (3)-equivariant diffusion policy for generalizable and data efficient learning
J. Yang, Z.-a. Cao, C. Deng, R. Antonova, S. Song, and J. Bohg · 2024
Later among the works it cites.
Genie: Generative interactive environments, 2024
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel · 2024
Later among the works it cites.
Learning to act without actions
D. Schmidt and M. Jiang · 2024
Later among the works it cites.
Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
π 0 \pi_{0} : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Open X-Embodiment: Robotic learning datasets and RT-X models
Open X-Embodiment Collaboration et al · 2024
Cited alongside, same era.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu · 2024
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2024
Cited alongside, same era.
Video language planning
Y. Du, S. Yang, P. Florence, F. Xia, A. Wahid, brian ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, L. P. Kaelbling, A. Zeng, and J. Tompson · 2024
Cited alongside, same era.
Dino-wm: World models on pre-trained visual features enable zero-shot planning
G. Zhou, H. Pan, Y. LeCun, and L. Pinto · 2024
Cited alongside, same era.
Premier-taco is a few-shot policy learner: pretraining multitask representation via temporal action-driven contrastive loss
R. Zheng, Y. Liang, X. Wang, S. Ma, H. Daumé III, H. Xu, J. Langford, P. Palanisamy, K. S. Basu, and F. Huang · 2024
Cited alongside, same era.
S. Li, Y. Gao, D. Sadigh, and S. Song · 2025
Closest in time.
Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets
C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, M.-Y. Liu, D. Xiang, G. Wetzstein, and T.-Y. Lin · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots, 2025
NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu · 2025
Closest in time.
Representation alignment for generation: Training diffusion transformers is easier than you think
S. Yu, S. Kwak, H. Jang, J. Jeong, J. Huang, J. Shin, and S. Xie · 2025
Closest in time.
M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa, et al · 2025
Closest in time.
Eagle 2: Building post-training data strategies from scratch for frontier vision-language models
Z. Li, G. Chen, S. Liu, S. Wang, V. VS, Y. Ji, S. Lan, H. Zhang, Y. Zhao, S. Radhakrishnan, et al · 2025
Closest in time.
Rambo: Rl-augmented model-based optimal control for whole-body loco-manipulation, 2025
J. Cheng, D. Kang, G. Fadini, G. Shi, and S. Coros · 2025
Closest in time.
Ardup: Active region video diffusion for universal policies, 2025
S. Huang, M. Levy, Z. Jiang, A. Anandkumar, Y. Zhu, L. Fan, D.-A. Huang, and A. Shrivastava · 2025
Closest in time.
Y. Gao, L. Gong, Q. Guo, X. Hou, Z. Lai, F. Li, L. Li, X. Lian, C. Liao, L. Liu, et al · 2025
Closest in time.
TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
R. Zheng, Y. Liang, S. Huang, J. Gao, H. D. III, A. Kolobov, F. Huang, and J. Yang · 2025
Closest in time.
Latent action pretraining from videos
S. Ye, J. Jang, B. Jeon, S. J. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo · 2025
Closest in time.
Magma: A foundation model for multimodal ai agents, 2025
J. Yang, R. Tan, Q. Wu, R. Zheng, B. Peng, Y. Liang, Y. Gu, M. Cai, S. Ye, J. Jang, Y. Deng, L. Liden, and J. Gao · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models, 2025
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
M. J. Kim, C. Finn, and P. Liang · 2025
Closest in time.
Moto: Latent motion token as the bridging language for learning robot manipulation from videos, 2025
Y. Chen, Y. Ge, W. Tang, Y. Li, Y. Ge, M. Ding, Y. Shan, and X. Liu · 2025
Closest in time.
Videoworld: Exploring knowledge learning from unlabeled videos, 2025
Z. Ren, Y. Wei, X. Guo, Y. Zhao, B. Kang, J. Feng, and X. Jin · 2025
Closest in time.
Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Huang, S. Jiang, et al · 2025
Closest in time.