Fetching the paper…
Reading the bibliography…
This paper introduces ManiFlow, a visuomotor imitation learning policy for general robot manipulation that generates precise, high-dimensional actions conditioned on diverse visual, language and proprioceptive inputs.
Logistic-normal distributions: Some properties and uses
J. Atchison and S. M. Shen · 1980
Earlier work this paper cites.
Manipulators and Manipulation in high dimensional spaces
V. Kumar · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Earlier work this paper cites.
Scalable diffusion models with transformers. 2023 ieee
W. S. Peebles and S. Xie · 2022
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
X. Liu, C. Gong, and Q. Liu · 2022
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Earlier work this paper cites.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2023
Earlier work this paper cites.
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever · 2023
Earlier work this paper cites.
Dexart: Benchmarking generalizable dexterous manipulation with articulated objects
C. Bao, H. Xu, Y. Qin, and X. Wang · 2023
Earlier work this paper cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Earlier work this paper cites.
Rvt: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Earlier work this paper cites.
Gnfactor: Multi-task real robot learning with generalizable neural feature fields
Y. Ze, G. Yan, Y.-H. Wu, A. Macaluso, Y. Ge, J. Ye, N. Hansen, L. E. Li, and X. Wang · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Cited alongside, same era.
Uni3d: Exploring unified 3d representation at scale
J. Zhou, J. Wang, B. Ma, Y.-S. Liu, T. Huang, and X. Wang · 2023
Cited alongside, same era.
Vision-language foundation models as effective robot imitators
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y. Jing, W. Zhang, H. Liu, et al · 2023
Cited alongside, same era.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine · 2023
Cited alongside, same era.
Unleashing large-scale video generative pre-training for visual robot manipulation
Manicm: Real-time 3d diffusion policy via consistency model for robotic manipulation
G. Lu, Z. Gao, T. Chen, W. Dai, Z. Wang, W. Ding, and Y. Tang · 2024
Later among the works it cites.
B. Jia, P. Ding, C. Cui, M. Sun, P. Qian, S. Huang, Z. Fan, and D. Wang · 2024
Later among the works it cites.
Rvt-2: Learning precise manipulation from few demonstrations
A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox · 2024
Later among the works it cites.
Dnact: Diffusion guided multi-task 3d policy learning
G. Yan, Y.-H. Wu, and X. Wang · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong · 2023
Cited alongside, same era.
Learning robotic manipulation policies from point clouds with conditional flow matching
E. Chisari, N. Heppert, M. Argus, T. Welschehold, T. Brox, and A. Valada · 2024
Cited alongside, same era.
π \pi 0: A vision-language-action flow model for general robot control, 2024
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Riemannian flow matching policy for robot motion learning
M. Braun, N. Jaquier, L. Rozo, and T. Asfour · 2024
Cited alongside, same era.
Affordance-based robot manipulation with flow matching
F. Zhang and M. Gienger · 2024
Cited alongside, same era.
Consistency policy: Accelerated visuomotor policies via consistency distillation
A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg · 2024
Cited alongside, same era.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
M. Reuss, Ö. E. Yağmurlu, F. Wenzel, and R. Lioutikov · 2024
Cited alongside, same era.
3d diffusion policy
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Cited alongside, same era.
Later among the works it cites.
The ingredients for robotic diffusion transformers
S. Dasari, O. Mees, S. Zhao, M. K. Srirama, and S. Levine · 2024
Later among the works it cites.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song · 2024
Later among the works it cites.
Consistency flow matching: Defining straight flows with velocity consistency
L. Yang, Z. Zhang, Z. Zhang, X. Liu, M. Xu, W. Zhang, C. Meng, S. Ermon, and B. Cui · 2024
Later among the works it cites.
Bunny-visionpro: Real-time bimanual dexterous teleoperation for imitation learning
R. Ding, Y. Qin, J. Zhu, C. Jia, S. Yang, R. Yang, X. Qi, and X. Wang · 2024
Later among the works it cites.
Open-television: Teleoperation with immersive active visual feedback
X. Cheng, J. Li, S. Yang, G. Yang, and X. Wang · 2024
Later among the works it cites.
One step diffusion via shortcut models
K. Frans, D. Hafner, S. Levine, and P. Abbeel · 2025
Closest in time.
T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Q. Liang, Z. Li, X. Lin, Y. Ge, Z. Gu, et al · 2025
Closest in time.
Integrating lmm planners and 3d skill policies for generalizable manipulation
Y. Li, G. Yan, A. Macaluso, M. Ji, X. Zou, and X. Wang · 2025
Closest in time.
Vggt: Visual geometry grounded transformer
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny · 2025
Closest in time.