Fetching the paper…
Reading the bibliography…
A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
A tutorial on energy-based learning
Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, and F. Huang · 2006
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction, 2016
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2017
Earlier work this paper cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video, 2018
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, and S. Levine · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction, 2018
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Earlier work this paper cites.
Self-supervised correspondence in visuomotor policy learning, 2019
P. Florence, L. Manuelli, and R. Tedrake · 2019
Earlier work this paper cites.
Implicit generation and generalization in energy-based models, 2020
Y. Du and I. Mordatch · 2020
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning, 2020
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Earlier work this paper cites.
Implicit behavioral cloning, 2021
P. Florence, C. Lynch, A. Zeng, O. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2021
Earlier work this paper cites.
Learning the predictability of the future
D. Suris, R. Liu, and C. Vondrick · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Learning generalizable robotic reward functions from ”in-the-wild” human videos, 2021
A. S. Chen, S. Nair, and C. Finn · 2021
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation, 2022
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models, 2022
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans · 2022
Cited alongside, same era.
Megapose: 6d pose estimation of novel objects via render & compare
Y. Labbé, L. Manuelli, A. Mousavian, S. Tyree, S. Birchfield, J. Tremblay, J. Carpentier, M. Aubry, D. Fox, and J. Sivic · 2022
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Video language planning, 2023
Y. Du, M. Yang, P. Florence, F. Xia, A. Wahid, B. Ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, L. Kaelbling, A. Zeng, and J. Tompson · 2023
Later among the works it cites.
Compositional foundation models for hierarchical planning, 2023
A. Ajay, S. Han, Y. Du, S. Li, A. Gupta, T. Jaakkola, J. Tenenbaum, L. Kaelbling, A. Srivastava, and P. Agrawal · 2023
Later among the works it cites.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Lcm-lora: A universal stable-diffusion acceleration module
S. Luo, Y. Tan, S. Patil, D. Gu, P. von Platen, A. Passos, L. Huang, J. Li, and H. Zhao · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards generalizable zero-shot manipulation via translating human interaction plans
H. Bharadhwaj, A. Gupta, V. Kumar, and S. Tulsiani · 2023
Cited alongside, same era.
Learning universal policies via text-guided video generation, 2023
Y. Du, M. Yang, B. Dai, H. Dai, O. Nachum, J. B. Tenenbaum, D. Schuurmans, and P. Abbeel · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models, 2023
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
Masked world models for visual control
Y. Seo, D. Hafner, H. Liu, F. Liu, S. James, K. Lee, and P. Abbeel · 2023
Cited alongside, same era.
Liv: Language-image representations and rewards for robotic control, 2023
Y. J. Ma, W. Liang, V. Som, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman · 2023
Cited alongside, same era.
Video prediction models as rewards for reinforcement learning, 2023
A. Escontrela, A. Adeniji, W. Yan, A. Jain, X. B. Peng, K. Goldberg, Y. Lee, D. Hafner, and P. Abbeel · 2023
Cited alongside, same era.
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models, 2023
S. Zhang, J. Wang, Y. Zhang, K. Zhao, H. Yuan, Z. Qin, X. Wang, D. Zhao, and J. Zhou · 2023
Cited alongside, same era.
Later among the works it cites.
Q-diffusion: Quantizing diffusion models
X. Li, Y. Liu, L. Lian, H. Yang, Z. Dong, D. Kang, S. Zhang, and K. Keutzer · 2023
Later among the works it cites.
Droid: A large-scale in-the-wild robot manipulation dataset, 2024
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y. J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y. Lee, M. Memmel, S. Park, I. Radosavovic, K. Wang, A. Zhan, K. Black, C. Chi, K. B. Hatch, S. Lin, J. Lu, J. Mercat, A. Rehman, P. R. Sanketi, A. Sharma, C. Simpson, Q. Vuong, H. R. Walke, B. Wulfe, T. Xiao, J. H. Yang, A. Yavary, T. Z. Zhao, C. Agia, R. Baijal, M. G. Castro, D. Chen, Q. Chen, T. Chung, J. Drake, E. P. Foster, J. Gao, D. A. Herrera, M. Heo, K. Hsu, J. Hu, D. Jackson, C. Le, Y. Li, K. Lin, R. Lin, Z. Ma, A. Maddukuri, S. Mirchandani, D. Morton, T. Nguyen, A. O’Neill, R. Scalise, D. Seale, V. Son, S. Tian, E. Tran, A. E. Wang, Y. Wu, A. Xie, J. Yang, P. Yin, Y. Zhang, O. Bastani, G. Berseth, J. Bohg, K. Goldberg, A. Gupta, A. Gupta, D. Jayaraman, J. J. Lim, J. Malik, R. Martín-Martín, S. Ramamoorthy, D. Sadigh, S. Song, J. Wu, M. C. Yip, Y. Zhu, T. Kollar, S. Levine, and C. Finn · 2024
Closest in time.
Diffusion policy: Visuomotor policy learning via action diffusion, 2024
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2024
Closest in time.
Behavior generation with latent actions
S. Lee, Y. Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto · 2024
Closest in time.
Video generation models as world simulators
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh · 2024
Closest in time.
Video as the new language for real-world decision making, 2024
S. Yang, J. Walker, J. Parker-Holder, Y. Du, J. Bruce, A. Barreto, P. Abbeel, and D. Schuurmans · 2024
Closest in time.
Generative camera dolly: Extreme monocular dynamic novel view synthesis
B. Van Hoorick, R. Wu, E. Ozguroglu, K. Sargent, R. Liu, P. Tokmakov, A. Dave, C. Zheng, and C. Vondrick · 2024
Closest in time.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song · 2024
Closest in time.