Fetching the paper…
Reading the bibliography…
Visuomotor policy learning has witnessed substantial progress in robotic manipulation, with recent approaches predominantly relying on generative models to model the action distribution.
Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex
D. H. Hubel and T. N. Wiesel · 1962
Earlier work this paper cites.
Hierarchical bayesian inference in the visual cortex
T. S. Lee and D. Mumford · 2003
Earlier work this paper cites.
Gaussian mixture models
D. A. Reynolds et al · 2009
Earlier work this paper cites.
Grasping novel objects with depth segmentation
D. Rao, Q. V. Le, T. Phoka, M. Quigley, A. Sudsang, and A. Y. Ng · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
Multimodal transfer: A hierarchical deep convolutional neural network for fast artistic style transfer
X. Wang, G. Oxholm, D. Zhang, and Y.-F. Wang · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Unet++: A nested u-net architecture for medical image segmentation
Z. Zhou, M. M. Rahman Siddiquee, N. Tajbakhsh, and J. Liang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Hierarchical reinforcement learning with advantage-based auxiliary rewards
S. Li, R. Wang, M. Tang, and C. Zhang · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
A. Razavi, A. Van den Oord, and O. Vinyals · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Earlier work this paper cites.
Super-convergence: Very fast training of neural networks using large learning rates
L. N. Smith and N. Topin · 2019
Earlier work this paper cites.
Hierarchical structure is employed by humans during visual motion perception
J. Bill, H. Pailian, S. J. Gershman, and J. Drugowitsch · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. Rovick Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Cited alongside, same era.
Depth-aware object segmentation and grasp detection for robotic picking tasks
S. Ainetter, C. Böhm, R. Dhakate, S. Weiss, and F. Fraundorfer · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Cited alongside, same era.
Hierarchical reinforcement learning: A comprehensive survey
S. Pateria, B. Subagdja, A.-h. Tan, and C. Quek · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Later among the works it cites.
Learning to manipulate anywhere: A visual generalizable framework for reinforcement learning
Z. Yuan, T. Wei, S. Cheng, G. Zhang, Y. Chen, and H. Xu · 2024
Later among the works it cites.
Consistency policy: Accelerated visuomotor policies via consistency distillation
A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg · 2024
Later among the works it cites.
One-step diffusion policy: Fast visuomotor policies via diffusion distillation
Z. Wang, Z. Li, A. Mandlekar, Z. Xu, J. Fan, Y. Narang, L. Fan, Y. Zhu, Y. Balaji, M. Zhou, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Dhariwal and A. Nichol · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Cited alongside, same era.
Behavior transformers: Cloning k k modes with one stone
N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto · 2022
Cited alongside, same era.
Visual motion perception as online hierarchical inference
J. Bill, S. J. Gershman, and J. Drugowitsch · 2022
Cited alongside, same era.
Generative modelling with inverse heat dissipation
S. Rissanen, M. Heinonen, and A. Solin · 2022
Cited alongside, same era.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2022
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow
X. Liu, C. Gong, and Q. Liu · 2022
Cited alongside, same era.
K. Frans, D. Hafner, S. Levine, and P. Abbeel · 2024
Later among the works it cites.
Behavior generation with latent actions
S. Lee, Y. Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto · 2024
Later among the works it cites.
Carp: Visuomotor policy learning via coarse-to-fine autoregressive prediction
Z. Gong, P. Ding, S. Lyu, S. Huang, M. Sun, W. Zhao, Z. Fan, and D. Wang · 2024
Later among the works it cites.
Humanoidbench: Simulated humanoid benchmark for whole-body locomotion and manipulation
C. Sferrazza, D.-M. Huang, X. Lin, Y. Lee, and P. Abbeel · 2024
Later among the works it cites.
Point cloud matters: Rethinking the impact of different observation spaces on robot learning
H. Zhu, Y. Wang, D. Huang, W. Ye, W. Ouyang, and T. He · 2024
Later among the works it cites.
Diffusion is spectral autoregression, 2024
S. Dieleman · 2024
Later among the works it cites.
Generalizable humanoid manipulation with 3d diffusion policies
Y. Ze, Z. Chen, W. Wang, T. Chen, X. He, Y. Yuan, X. B. Peng, and J. Wu · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Later among the works it cites.
Manicm: Real-time 3d diffusion policy via consistency model for robotic manipulation
G. Lu, Z. Gao, T. Chen, W. Dai, Z. Wang, and Y. Tang · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
K. Tian, Y. Jiang, Z. Yuan, B. Peng, and L. Wang · 2024
Later among the works it cites.
Robotwin: Dual-arm robot benchmark with generative digital twins (early version)
Y. Mu, T. Chen, S. Peng, Z. Chen, Z. Gao, Y. Zou, L. Lin, Z. Xie, and P. Luo · 2024
Later among the works it cites.
Dense policy: Bidirectional autoregressive learning of actions
Y. Su, X. Zhan, H. Fang, H. Xue, H.-S. Fang, Y.-L. Li, C. Lu, and L. Yang · 2025
Closest in time.
Roboverse: Towards a unified platform, dataset and benchmark for scalable and generalizable robot learning, April 2025
H. Geng, F. Wang, S. Wei, Y. Li, B. Wang, B. An, C. T. Cheng, H. Lou, P. Li, Y.-J. Wang, Y. Liang, D. Goetting, C. Xu, H. Chen, Y. Qian, Y. Geng, J. Mao, W. Wan, M. Zhang, J. Lyu, S. Zhao, J. Zhang, J. Zhang, C. Zhao, H. Lu, Y. Ding, R. Gong, Y. Wang, Y. Kuang, R. Wu, B. Jia, C. Sferrazza, H. Dong, S. Huang, K. Sreenath, Y. Wang, J. Malik, and P. Abbeel · 2025
Closest in time.
Ddt: Decoupled diffusion transformer
S. Wang, Z. Tian, W. Huang, and L. Wang · 2025
Closest in time.
Demogen: Synthetic demonstration generation for data-efficient visuomotor policy learning
Z. Xue, S. Deng, Z. Chen, Y. Wang, Z. Yuan, and H. Xu · 2025
Closest in time.
Autoregressive action sequence learning for robotic manipulation
X. Zhang, Y. Liu, H. Chang, L. Schramm, and A. Boularias · 2025
Closest in time.
Is noise conditioning necessary for denoising generative models?
Q. Sun, Z. Jiang, H. Zhao, and K. He · 2025
Closest in time.