Fetching the paper…
Reading the bibliography…
Learning real-world robotic manipulation is challenging, particularly when limited demonstrations are available.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Attention is all you need
A. Vaswani · 2017
Earlier work this paper cites.
Deep object pose estimation for semantic robotic grasping of household objects
J. Tremblay, T. To, B. Sundaralingam, Y. Xiang, D. Fox, and S. Birchfield · 2018
Earlier work this paper cites.
Deep object-centric representations for generalizable robot learning
C. Devin, P. Abbeel, T. Darrell, and S. Levine · 2018
Earlier work this paper cites.
Deep object-centric policies for autonomous driving
D. Wang, C. Devin, Q.-Z. Cai, F. Yu, and T. Darrell · 2019
Earlier work this paper cites.
Monet: Unsupervised scene decomposition and representation
C. P. Burgess, L. Matthey, N. Watters, R. Kabra, I. Higgins, M. Botvinick, and A. Lerchner · 2019
Earlier work this paper cites.
Object-centric task and motion planning in dynamic environments
T. Migimatsu and J. Bohg · 2020
Earlier work this paper cites.
Object-centric learning with slot attention
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Generalization through hand-eye coordination: An action space for learning spatially-invariant visuomotor control
C. Wang, R. Wang, A. Mandlekar, L. Fei-Fei, S. Savarese, and D. Xu · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam
C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. Montiel, and J. D. Tardós · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
6-dof pose estimation of household objects for robotic manipulation: An accessible dataset and benchmark
S. Tyree, J. Tremblay, T. To, J. Cheng, T. Mosier, J. Smith, and S. Birchfield · 2022
Earlier work this paper cites.
Object-category aware reinforcement learning
Q. Yi, R. Zhang, J. Guo, X. Hu, Z. Du, Q. Guo, Y. Chen, et al · 2022
Earlier work this paper cites.
Detecting twenty-thousand classes using image-level supervision
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Earlier work this paper cites.
Viola: Imitation learning for vision-based manipulation with object proposal priors
Y. Zhu, A. Joshi, P. Stone, and Y. Zhu · 2023
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Gendexgrasp: Generalizable dexterous grasping
P. Li, T. Liu, Y. Li, Y. Geng, Y. Zhu, Y. Yang, and S. Huang · 2023
Earlier work this paper cites.
Mimicgen: A data generation system for scalable robot learning using human demonstrations
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y. Narang, L. Fan, Y. Zhu, and D. Fox · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2023
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models
A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, et al · 2023
Cited alongside, same era.
Learning generalizable manipulation policies with object-centric 3d representations
Y. Zhu, Z. Jiang, P. Stone, and Y. Zhu · 2023
Cited alongside, same era.
Grasp multiple objects with one hand
Y. Li, B. Liu, Y. Geng, P. Li, Y. Yang, Y. Zhu, T. Liu, and S. Huang · 2024
Later among the works it cites.
Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation
M. Torne, A. Simeonov, Z. Li, A. Chan, T. Chen, A. Gupta, and P. Agrawal · 2024
Later among the works it cites.
Robotwin: Dual-arm robot benchmark with generative digital twins (early version)
Y. Mu, T. Chen, S. Peng, Z. Chen, Z. Gao, Y. Zou, L. Lin, Z. Xie, and P. Luo · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Cited alongside, same era.
Visuomotor control in multi-object scenes using object-aware representations
N. Heravi, A. Wahid, C. Lynch, P. Florence, T. Armstrong, J. Tompson, P. Sermanet, J. Bohg, and D. Dwibedi · 2023
Cited alongside, same era.
An investigation into pre-training object-centric representations for reinforcement learning
J. Yoon, Y.-F. Wu, H. Bae, and S. Ahn · 2023
Cited alongside, same era.
Sa6d: Self-adaptive few-shot 6d pose estimator for novel and occluded objects
N. Gao, V. A. Ngo, H. Ziesche, and G. Neumann · 2023
Cited alongside, same era.
D. Zavadski, J.-F. Feiden, and C. Rother · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Z. Fu, T. Z. Zhao, and C. Finn · 2024
Cited alongside, same era.
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Later among the works it cites.
π \pi 0: A vision-language-action flow model for general robot control, 2024
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Later among the works it cites.
Densematcher: Learning 3d semantic correspondence for category-level manipulation from a single demo
J. Zhu, Y. Ju, J. Zhang, M. Wang, Z. Yuan, K. Hu, and H. Xu · 2024
Later among the works it cites.
Zero-shot object-centric representation learning
A. Didolkar, A. Zadaianchuk, A. Goyal, M. Mozer, Y. Bengio, G. Martius, and M. Seitzer · 2024
Later among the works it cites.
Stable diffusion v1.5 model card
Stability · 2024
Later among the works it cites.
Controlnet++: Improving conditional controls with efficient consistency feedback: Project page: liming-ai. github. io/controlnet_plus_plus
M. Li, T. Yang, H. Kuang, J. Wu, Z. Wang, X. Xiao, and C. Chen · 2024
Later among the works it cites.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Y. Guo, C. Yang, A. Rao, Z. Liang, Y. Wang, Y. Qiao, M. Agrawala, D. Lin, and B. Dai · 2024
Later among the works it cites.
Disco: Disentangled control for realistic human dance generation
T. Wang, L. Li, K. Lin, Y. Zhai, C.-C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang · 2024
Later among the works it cites.
Lumiere: A space-time diffusion model for video generation
O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, G. Liu, A. Raj, et al · 2024
Later among the works it cites.
Motionlcm: Real-time controllable motion generation via latent consistency model
W. Dai, L.-H. Chen, J. Wang, J. Liu, B. Dai, and Y. Tang · 2024
Later among the works it cites.
Omnicontrol: Control any joint at any time for human motion generation
Y. Xie, V. Jampani, L. Zhong, D. Sun, and H. Jiang · 2024
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, et al · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al · 2024
Later among the works it cites.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Later among the works it cites.
Maniptrans: Efficient dexterous bimanual manipulation transfer via residual learning
K. Li, P. Li, T. Liu, Y. Li, and S. Huang · 2025
Closest in time.
Meta Quest: Virtual Reality Headset
Meta Platforms, Inc · 2025
Closest in time.
Astribot S1: AI Robotic Partner
Astribot, Inc · 2025
Closest in time.