Fetching the paper…
Reading the bibliography…
Can we endow visuomotor robots with generalization capabilities to operate in diverse open-world scenarios? In this paper, we propose \textbf{Maniwhere}, a generalizable framework tailored for visual reinforcement learning, enabling the trained robot policies to generalize across a combination of multiple visual disturbance types.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Spatial transformer networks
M. Jaderberg, K. Simonyan, A. Zisserman, et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Earlier work this paper cites.
Temporal cycle-consistency learning
D. Dwibedi, Y. Aytar, J. Tompson, P. Sermanet, and A. Zisserman · 2019
Earlier work this paper cites.
Self-supervised policy adaptation during deployment
N. Hansen, R. Jangir, Y. Sun, G. Alenyà, P. Abbeel, A. A. Efros, L. Pinto, and X. Wang · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al · 2020
Earlier work this paper cites.
Stabilizing deep q-learning with convnets and vision transformers under data augmentation
N. Hansen, H. Su, and X. Wang · 2021
Earlier work this paper cites.
Generalization in reinforcement learning by soft data augmentation
N. Hansen and X. Wang · 2021
Earlier work this paper cites.
An empirical study of training self-supervised vision transformers
X. Chen, S. Xie, and K. He · 2021
Earlier work this paper cites.
Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation
R. Jangir, N. Hansen, S. Ghosal, M. Jain, and X. Wang · 2022
Earlier work this paper cites.
Spectrum random masking for generalization in image-based reinforcement learning
Y. Huang, P. Peng, Y. Zhao, G. Chen, and Y. Tian · 2022
Earlier work this paper cites.
Look where you look! saliency-guided q-networks for generalization in visual reinforcement learning
D. Bertoin, A. Zouitine, M. Zouitine, and E. Rachelson · 2022
Earlier work this paper cites.
A comprehensive survey of data augmentation in visual reinforcement learning
G. Ma, Z. Wang, Z. Yuan, X. Wang, B. Yuan, and D. Tao · 2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Earlier work this paper cites.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
N. Hansen, Z. Yuan, Y. Ze, T. Mu, A. Rajeswaran, H. Su, H. Xu, and X. Wang · 2022
Earlier work this paper cites.
Xirl: Cross-embodiment inverse reinforcement learning
K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Cited alongside, same era.
Visual reinforcement learning with self-supervised 3d representations
Y. Ze, N. Hansen, Y. Chen, M. Jain, and X. Wang · 2023
Cited alongside, same era.
Multi-view masked world models for visual robotic manipulation
Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel · 2023
Cited alongside, same era.
Dynamic handover: Throw and catch with bimanual hands
B. Huang, Y. Chen, T. Wang, Y. Qin, Y. Yang, N. Atanasov, and X. Wang · 2023
Cited alongside, same era.
Transic: Sim-to-real policy transfer by learning from online correction
Y. Jiang, C. Wang, R. Zhang, J. Wu, and L. Fei-Fei · 2024
Closest in time.
Cyberdemo: Augmenting simulated human demonstration for real-world dexterous manipulation
J. Wang, Y. Qin, K. Kuang, Y. Korkmaz, A. Gurumoorthy, H. Su, and X. Wang · 2024
Closest in time.
Spin: Simultaneous perception, interaction and navigation
S. Uppal, A. Agarwal, H. Xiong, K. Shaw, and D. Pathak · 2024
Closest in time.
Y. Ju, K. Hu, G. Zhang, G. Zhang, M. Jiang, and H. Xu · 2024
Closest in time.
A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence
J. Zhang, C. Herrmann, J. Hur, L. Polania Cabrera, V. Jampani, D. Sun, and M.-H. Yang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalizable visual reinforcement learning with segment anything model
Z. Wang, Y. Ze, Y. Sun, Z. Yuan, and H. Xu · 2023
Cited alongside, same era.
Leap hand: Low-cost, efficient, and anthropomorphic hand for robot learning
K. Shaw, A. Agarwal, and D. Pathak · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Cited alongside, same era.
Dextreme: Transfer of agile in-hand manipulation from simulation to reality
A. Handa, A. Allshire, V. Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. Van Wyk, A. Zhurkevich, B. Sundaralingam, et al · 2023
Cited alongside, same era.
Sequential dexterity: Chaining dexterous policies for long-horizon manipulation
Y. Chen, C. Wang, L. Fei-Fei, and C. K. Liu · 2023
Cited alongside, same era.
Rvt: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Cited alongside, same era.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Cited alongside, same era.
Closest in time.
Rl-vigen: A reinforcement learning benchmark for visual generalization
Z. Yuan, S. Yang, P. Hua, C. Chang, K. Hu, and H. Xu · 2024
Closest in time.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Closest in time.
Movie: Visual model-based policy adaptation for view generalization
S. Yang, Y. Ze, and H. Xu · 2024
Closest in time.
Evaluating real-world robot manipulation policies in simulation
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao · 2024
Closest in time.
The colosseum: A benchmark for evaluating generalization for robotic manipulation
W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox · 2024
Closest in time.
Vision-based manipulation from single human video with open-world object graphs
Y. Zhu, A. Lim, P. Stone, and Y. Zhu · 2024
Closest in time.
J. Lyu, L. Wan, X. Li, and Z. Lu · 2024
Closest in time.
Green screen augmentation enables scene generalisation in robotic manipulation
E. Teoh, S. Patidar, X. Ma, and S. James · 2024
Closest in time.
Natural language can help bridge the sim2real gap
A. Yu, A. Foote, R. Mooney, and R. Martín-Martín · 2024
Closest in time.
Peac: Unsupervised pre-training for cross-embodiment reinforcement learning
C. Ying, Z. Hao, X. Zhou, X. Xu, H. Su, X. Zhang, and J. Zhu · 2024
Closest in time.
Premier-TACO is a few-shot policy learner: Pretraining multitask representation via temporal action-driven contrastive loss
R. Zheng, Y. Liang, X. Wang, S. Ma, H. D. III, H. Xu, J. Langford, P. Palanisamy, K. S. Basu, and F. Huang · 2024
Closest in time.
Visual representation learning with stochastic frame prediction
H. Jang, D. Kim, J. Kim, J. Shin, P. Abbeel, and Y. Seo · 2024
Closest in time.
Learning manipulation by predicting interaction
Z. Jia, B. Qingwen, W. Bangjun, X. Wenke, C. Li, D. Hao, S. Haoming, W. Dong, H. Di, L. Ping, C. Heming, Z. Bin, L. Xuelong, Q. Yu, and L. Hongyang · 2024
Closest in time.
H-index: Visual reinforcement learning with hand-informed representations for dexterous manipulation
Y. Ze, Y. Liu, R. Shi, J. Qin, Z. Yuan, J. Wang, and H. Xu · 2024
Closest in time.