Fetching the paper…
Reading the bibliography…
The main challenge in learning image-conditioned robotic policies is acquiring a visual representation conducive to low-level control.
Using natural language for reward shaping in reinforcement learning
Prasoon Goyal, Scott Niekum, and Raymond Mooney · 1903
Earlier work this paper cites.
A unified approach for motion and force control of robot manipulators: The operational space formulation
Oussama Khatib · 1987
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean Pomerleau · 1988
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal · 1999
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
Corey Lynch and Pierre Sermanet · 2005
Earlier work this paper cites.
Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards
Prasoon Goyal, Scott Niekum, and Raymond Mooney · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance, 2014
Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Unsupervised pixel-level domain adaptation with generative adversarial networks
Konstantinos Bousmalis, Nathan Silberman, David Dohan, Dumitru Erhan, and Dilip Krishnan · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
Preparing for the unknown: Learning a universal policy with online system identification
Wenhao Yu, Jie Tan, C Karen Liu, and Greg Turk · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Earlier work this paper cites.
Sim-to-real reinforcement learning for deformable object manipulation
Jan Matas, Stephen James, and Andrew J Davison · 2018
Earlier work this paper cites.
Driving policy transfer via modularity and abstraction
Matthias Müller, Alexey Dosovitskiy, Bernard Ghanem, and Vladlen Koltun · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2018
Earlier work this paper cites.
Sim-to-real via sim-to-sim: Data-efficient robotic grasping via randomized-to-canonical adaptation networks
Stephen James, Paul Wohlhart, Mrinal Kalakrishnan, Dmitry Kalashnikov, Alex Irpan, Julian Ibarz, Sergey Levine, Raia Hadsell, and Konstantinos Bousmalis · 2019
Earlier work this paper cites.
Solving rubik’s cube with a robot hand, 2019
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 2019
Earlier work this paper cites.
Learning dexterous in-hand manipulation
OpenAI: Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Sim2real transfer for reinforcement learning without dynamics randomization
Manuel Kaspar, Juan D Muñoz Osorio, and Jürgen Bock · 2020
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Cited alongside, same era.
Rl-cyclegan: Reinforcement learning aware simulation-to-real
Kanishka Rao, Chris Harris, Alex Irpan, Sergey Levine, Julian Ibarz, and Mohi Khansari · 2020
Cited alongside, same era.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh · 2020
Cited alongside, same era.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
Lin Shao, Toki Migimatsu, Qiang Zhang, Karen Yang, and Jeannette Bohg · 2020
Cited alongside, same era.
Maskvit: Masked visual pre-training for video prediction
Agrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu, Roberto Martín-Martín, and Li Fei-Fei · 2022
Later among the works it cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan · 2022
Later among the works it cites.
What matters in language conditioned imitation learning
Oier Mees, Lukas Hermann, and Wolfram Burgard · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers, 2020
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2020
Cited alongside, same era.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Martín-Martín, Abhishek Joshi, Soroush Nasiriany, and Yifeng Zhu · 2020
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, et al · 2021
Cited alongside, same era.
Retinagan: An object-aware approach to sim-to-real transfer
Daniel Ho, Kanishka Rao, Zhuo Xu, Eric Jang, Mohi Khansari, and Yunfei Bai · 2021
Cited alongside, same era.
BC-z: Zero-shot task generalization with robotic imitation learning
Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn · 2021
Cited alongside, same era.
Lila: Language-informed latent actions
Siddharth Karamcheti, Megha Srivastava, Percy Liang, and Dorsa Sadigh · 2021
Cited alongside, same era.
Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard · 2021
Cited alongside, same era.
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2022
Later among the works it cites.
Vrl3: A data-driven framework for visual deep reinforcement learning
Che Wang, Xufang Luo, Keith Ross, and Dongsheng Li · 2022
Later among the works it cites.
Using both demonstrations and language instructions to efficiently learn robotic tasks
Albert Yu and Raymond J Mooney · 2022
Later among the works it cites.
Pre-trained image encoder for generalizable visual reinforcement learning
Zhecheng Yuan, Zhengrong Xue, Bo Yuan, Xueqian Wang, Yi Wu, Yang Gao, and Huazhe Xu · 2022
Later among the works it cites.
What makes representation learning from videos hard for control?
Tony Zhao, Siddharth Karamcheti, Thomas Kollar, Chelsea Finn, and Percy Liang · 2022
Later among the works it cites.
Bo Ai, Zhanxin Wu, and David Hsu · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al · 2023
Later among the works it cites.
Exploring visual pre-training for robot manipulation: Datasets, models and methods
Ya Jing, Xuelin Zhu, Xingbin Liu, Qie Sima, Taozheng Yang, Yunhai Feng, and Tao Kong · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2023
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang · 2023
Later among the works it cites.
Liv: Language-image representations and rewards for robotic control
Yecheng Jason Ma, William Liang, Vaidehi Som, Vikash Kumar, Amy Zhang, Osbert Bastani, and Dinesh Jayaraman · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Abhishek Padalkar, Acorn Pooley, Ajinkya Jain, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anikait Singh, Anthony Brohan, et al · 2023
Later among the works it cites.
Real-world robot learning with masked visual pre-training
Ilija Radosavovic, Tete Xiao, Stephen James, Pieter Abbeel, Jitendra Malik, and Trevor Darrell · 2023
Later among the works it cites.
Cape: Corrective actions from precondition errors using large language models
Shreyas Sundara Raman, Vanya Cohen, David Paulius, Ifrah Idrees, Eric Rosen, Ray Mooney, and Stefanie Tellex · 2023
Later among the works it cites.
Masked world models for visual control
Younggyo Seo, Danijar Hafner, Hao Liu, Fangchen Liu, Stephen James, Kimin Lee, and Pieter Abbeel · 2023
Later among the works it cites.
Mutex: Learning unified policies from multimodal task specifications
Rutav Shah, Roberto Martín-Martín, and Yuke Zhu · 2023
Later among the works it cites.
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al · 2023
Later among the works it cites.
Conv-adapter: Exploring parameter efficient transfer learning for convnets, 2024
Hao Chen, Ran Tao, Han Zhang, Yidong Wang, Xiang Li, Wei Ye, Jindong Wang, Guosheng Hu, and Marios Savvides · 2024
Closest in time.