Fetching the paper…
Reading the bibliography…
Scaling up robot learning requires large and diverse datasets, and how to efficiently reuse collected data and transfer policies to new embodiments remains an open question.
Autonomous shaping: Knowledge transfer in reinforcement learning
G. Konidaris and A. Barto · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Transfer learning across heterogeneous robots with action sequence mapping
B. Lakshmanan and R. Balaraman · 2010
Earlier work this paper cites.
Semantic segmentation using regions and parts
P. Arbeláez, B. Hariharan, C. Gu, S. Gupta, L. Bourdev, and J. Malik · 2012
Earlier work this paper cites.
Unsupervised cross-domain transfer in policy gradient reinforcement learning via manifold alignment
H. B. Ammar, E. Eaton, P. Ruvolo, and M. Taylor · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
Transfer from simulation to real world through learning deep inverse dynamics model
P. Christiano, Z. Shah, I. Mordatch, J. Schneider, T. Blackwell, J. Tobin, P. Abbeel, and W. Zaremba · 2016
Earlier work this paper cites.
Learning modular neural network policies for multi-task and multi-robot transfer
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Sim-to-real robot learning from pixels with progressive nets
A. A. Rusu, M. Večerík, T. Rothörl, N. Heess, R. Pascanu, and R. Hadsell · 2017
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2017
Earlier work this paper cites.
Hardware conditioned policies for multi-robot transfer learning
T. Chen, A. Murali, and A. Gupta · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
T. Yu, C. Finn, S. Dasari, A. Xie, T. Zhang, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Y. Liu, A. Gupta, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Nervenet: Learning structured policy with graph neural networks
T. Wang, R. Liao, J. Ba, and S. Fidler · 2018
Earlier work this paper cites.
Graph networks as learnable physics engines for inference and control
A. Sanchez-Gonzalez, N. Heess, J. T. Springenberg, J. Merel, M. Riedmiller, R. Hadsell, and P. Battaglia · 2018
Earlier work this paper cites.
Jacquard: A large scale dataset for robotic grasp detection
A. Depierre, E. Dellandréa, and L. Chen · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen · 2018
Earlier work this paper cites.
Sim2real viewpoint invariant visual servoing by recurrent control
F. Sadeghi, A. Toshev, E. Jang, and S. Levine · 2018
Earlier work this paper cites.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience
Y. Chebotar, A. Handa, V. Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox · 2019
Earlier work this paper cites.
Learning to control self-assembling morphologies: a study of generalization via modularity
D. Pathak, C. Lu, T. Darrell, P. Isola, and A. A. Efros · 2019
Earlier work this paper cites.
Zero-shot generalization using cascaded system-representations
A. Malik · 2019
Earlier work this paper cites.
Sim2real transfer for reinforcement learning without dynamics randomization
M. Kaspar, J. D. M. Osorio, and J. Bock · 2020
Earlier work this paper cites.
Learning one-shot imitation from humans without humans
A. Bonardi, S. James, and A. J. Davison · 2020
Earlier work this paper cites.
Avid: Learning multi-stage tasks via pixel-level translation of human videos
L. Smith, N. Dhawan, M. Zhang, P. Abbeel, and S. Levine · 2020
Earlier work this paper cites.
Learning agile robotic locomotion skills by imitating animals
X. B. Peng, E. Coumans, T. Zhang, T.-W. Lee, J. Tan, and S. Levine · 2020
Earlier work this paper cites.
Domain adaptive imitation learning
K. Kim, Y. Gu, J. Song, S. Zhao, and S. Ermon · 2020
Earlier work this paper cites.
Hierarchically decoupled imitation for morphological transfer
D. Hejna, L. Pinto, and P. Abbeel · 2020
Earlier work this paper cites.
Unigrasp: Learning a unified model to grasp with multifingered robotic hands
L. Shao, F. Ferreira, M. Jorda, V. Nambiar, J. Luo, E. Solowjow, J. A. Ojea, O. Khatib, and J. Bohg · 2020
Earlier work this paper cites.
One policy to control them all: Shared modular policies for agent-agnostic control
W. Huang, I. Mordatch, and D. Pathak · 2020
Earlier work this paper cites.
ACRONYM: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2020
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2020
Earlier work this paper cites.
robosuite: A modular simulation framework and benchmark for robot learning
Y. Zhu, J. Wong, A. Mandlekar, R. Martín-Martín, A. Joshi, S. Nasiriany, and Y. Zhu · 2020
Earlier work this paper cites.
BC-Z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Beyond pick-and-place: Tackling robotic stacking of diverse shapes
A. X. Lee, C. M. Devin, Y. Zhou, T. Lampe, K. Bousmalis, J. T. Springenberg, A. Byravan, A. Abdolmaleki, N. Gileadi, D. Khosid, et al · 2021
Earlier work this paper cites.
Reinforcement learning with videos: Combining offline observations with interaction
K. Schmeckpeper, O. Rybkin, K. Daniilidis, S. Levine, and C. Finn · 2021
Earlier work this paper cites.
Learning by watching: Physical imitation of manipulation skills from human videos
H. Xiong, Q. Li, Y.-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg · 2021
Earlier work this paper cites.
Learning generalizable robotic reward functions from “in-the-wild” human videos
A. S. Chen, S. Nair, and C. Finn · 2021
Earlier work this paper cites.
Manipulator-independent representations for visual imitation
Y. Zhou, Y. Aytar, and K. Bousmalis · 2021
Earlier work this paper cites.
Cross-domain imitation from observations
D. S. Raychaudhuri, S. Paul, J. Vanbaar, and A. K. Roy-Chowdhury · 2021
Earlier work this paper cites.
Bayesian meta-learning for few-shot policy adaptation across robotic platforms
A. Ghadirzadeh, X. Chen, P. Poklukar, C. Finn, M. Björkman, and D. Kragic · 2021
Earlier work this paper cites.
Adagrasp: Learning an adaptive gripper-aware grasping policy
Z. Xu, B. Qi, S. Agrawal, and S. Song · 2021
Earlier work this paper cites.
My body is a cage: the role of morphology in graph- based incompatible control
V. Kurin, M. Igl, T. Rocktaschel, W. Boehmer, and S. Whiteson · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Cited alongside, same era.
Scaling up multi-task robotic reinforcement learning
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2022
Cited alongside, same era.
Know thyself: Transferable visual control policies through robot-awareness
E. S. Hu, K. Huang, O. Rybkin, and D. Jayaraman · 2022
Cited alongside, same era.
Human-to-robot imitation in the wild
S. Bahl, A. Gupta, and D. Pathak · 2022
Cited alongside, same era.
Robotic telekinesis: Learning a robotic hand imitator by watching humans on youtube
A. Sivakumar, K. Shaw, and D. Pathak · 2022
Cited alongside, same era.
Xirl: Cross-embodiment inverse reinforcement learning
Pali-x: On scaling up a multilingual vision and language model, 2023
X. Chen, J. Djolonga, P. Padlewski, B. Mustafa, S. Changpinyo, J. Wu, C. R. Ruiz, S. Goodman, X. Wang, Y. Tay, S. Shakeri, M. Dehghani, D. Salz, M. Lucic, M. Tschannen, A. Nagrani, H. Hu, M. Joshi, B. Pang, C. Montgomery, P. Pietrzyk, M. Ritter, A. Piergiovanni, M. Minderer, F. Pavetic, A. Waters, G. Li, I. Alabdulmohsin, L. Beyer, J. Amelot, K. Lee, A. P. Steiner, Y. Li, D. Keysers, A. Arnab, Y. Xu, K. Rong, A. Kolesnikov, M. Seyedhosseini, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra · 2023
Later among the works it cites.
Dreamfusion: Text-to-3d using 2d diffusion
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2023
Later among the works it cites.
Realfusion: 360deg reconstruction of any object from a single image
L. Melas-Kyriazi, I. Laina, C. Rupprecht, and A. Vedaldi · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi · 2022
Cited alongside, same era.
Transfer rl across observation feature spaces via model-based regularization
Y. Sun, R. Zheng, X. Wang, A. E. Cohen, and F. Huang · 2022
Cited alongside, same era.
Revolver: Continuous evolutionary models for robot-to-robot policy transfer
X. Liu, D. Pathak, and K. Kitani · 2022
Cited alongside, same era.
Learning invariant feature spaces to transfer skills with reinforcement learning
A. Gupta, C. Devin, Y. Liu, P. Abbeel, and S. Levine · 2022
Cited alongside, same era.
Translating robot skills: Learning unsupervised skill correspondences across robots
T. Shankar, Y. Lin, A. Rajeswaran, V. Kumar, S. Anderson, and J. Oh · 2022
Cited alongside, same era.
Learn what matters: cross-domain imitation learning with task-relevant embeddings
T. Franzmeyer, P. Torr, and J. F. Henriques · 2022
Cited alongside, same era.
Cross domain robot imitation with invariant representation
Z.-H. Yin, L. Sun, H. Ma, M. Tomizuka, and W.-J. Li · 2022
Cited alongside, same era.
Zero-1-to-3: Zero-shot one image to 3d object
R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al · 2023
Later among the works it cites.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Later among the works it cites.
Saytap: Language to quadrupedal locomotion
Y. Tang, W. Yu, J. Tan, H. Zen, A. Faust, and T. Harada · 2023
Later among the works it cites.
Prompt a robot to walk with large language models
Y.-J. Wang, B. Zhang, J. Chen, and K. Sreenath · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
W. Yu, N. Gileadi, C. Fu, S. Kirmani, K.-H. Lee, M. G. Arenas, H.-T. L. Chiang, T. Erez, L. Hasenclever, J. Humplik, et al · 2023
Later among the works it cites.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Later among the works it cites.
Genaug: Retargeting behaviors to unseen situations via generative augmentation
Z. Chen, S. Kiami, A. Gupta, and V. Kumar · 2023
Later among the works it cites.
Frame mining: a free lunch for learning robotic manipulation from 3d point clouds
M. Liu, X. Li, Z. Ling, Y. Li, and H. Su · 2023
Later among the works it cites.
Rvt: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Later among the works it cites.
Exaug: Robot-conditioned navigation policies via geometric experience augmentation
N. Hirose, D. Shah, A. Sridhar, and S. Levine · 2023
Later among the works it cites.
Multi-view masked world models for visual robotic manipulation
Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel · 2023
Later among the works it cites.
Nerf in the palm of your hand: Corrective augmentation for robotics via novel-view synthesis
A. Zhou, M. J. Kim, L. Wang, P. Florence, and C. Finn · 2023
Later among the works it cites.
Polybot: Training one policy across robots while embracing variability
J. H. Yang, D. Sadigh, and C. Finn · 2023
Later among the works it cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models, 2023
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, D. Sadigh, C. Finn, and S. Levine · 2023
Later among the works it cites.
Streamdiffusion: A pipeline-level solution for real-time interactive generation
A. Kodaira, C. Xu, T. Hazama, T. Yoshimoto, K. Ohno, S. Mitsuhori, S. Sugano, H. Cho, Z. Liu, and K. Keutzer · 2023
Later among the works it cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Video generation models as world simulators
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh · 2024
Closest in time.
Efficient data collection for robotic manipulation via compositional generalization
J. Gao, A. Xie, T. Xiao, C. Finn, and D. Sadigh · 2024
Closest in time.
Mirage: Cross-embodiment zero-shot policy transfer with cross-painting
L. Y. Chen, K. Hari, K. Dharmarajan, C. Xu, Q. Vuong, and K. Goldberg · 2024
Closest in time.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu · 2024
Closest in time.
Spin: Simultaneous perception, interaction and navigation
S. Uppal, A. Agarwal, H. Xiong, K. Shaw, and D. Pathak · 2024
Closest in time.
Meta-evolve: Continuous robot evolution for one-to-many policy transfer
X. Liu, D. Pathak, and D. Zhao · 2024
Closest in time.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel · 2024
Closest in time.
Learning interactive real-world simulators
S. Yang, Y. Du, S. K. S. Ghasemipour, J. Tompson, L. P. Kaelbling, D. Schuurmans, and P. Abbeel · 2024
Closest in time.
Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar · 2024
Closest in time.
MOSAIC: A modular system for assistive and interactive cooking
H. Wang, K. Kedia, J. Ren, R. Abdullah, A. Bhardwaj, A. Chao, K. Y. Chen, N. Chin, P. Dan, X. Fan, G. Gonzalez-Pumariega, A. Kompella, M. A. Pace, Y. Sharma, X. Sun, N. Sunkara, and S. Choudhury · 2024
Closest in time.
Eureka: Human-level reward design via coding large language models
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar · 2024
Closest in time.
Gen2sim: Scaling up robot learning in simulation with generative models
P. Katara, Z. Xian, and K. Fragkiadaki · 2024
Closest in time.
Robogen: Towards unleashing infinite data for automated robot learning via generative simulation
Y. Wang, Z. Xian, F. Chen, T.-H. Wang, Y. Wang, K. Fragkiadaki, Z. Erickson, D. Held, and C. Gan · 2024
Closest in time.
Zero-shot robotic manipulation with pre-trained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. R. Walke, C. Finn, A. Kumar, and S. Levine · 2024
Closest in time.
Diffusion meets dagger: Supercharging eye-in-hand imitation learning
X. Zhang, M. Chang, P. Kumar, and S. Gupta · 2024
Closest in time.
Pushing the limits of cross-embodiment learning for manipulation and navigation
J. Yang, C. Glossop, A. Bhorkar, D. Shah, Q. Vuong, C. Finn, D. Sadigh, and S. Levine · 2024
Closest in time.
Zeronvs: Zero-shot 360-degree view synthesis from a single image
K. Sargent, Z. Li, T. Shah, C. Herrmann, H.-X. Yu, Y. Zhang, E. R. Chan, D. Lagun, L. Fei-Fei, D. Sun, et al · 2024
Closest in time.
3difftection: 3d object detection with geometry-aware diffusion features
C. Xu, H. Ling, S. Fidler, and O. Litany · 2024
Closest in time.
Ag2manip: Learning novel manipulation skills with agent-agnostic visual and action representations
P. Li, T. Liu, Y. Li, M. Han, H. Geng, S. Wang, Y. Zhu, S.-C. Zhu, and S. Huang · 2024
Closest in time.
Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors
G. Qian, J. Mai, A. Hamdi, J. Ren, A. Siarohin, B. Li, H.-Y. Lee, I. Skorokhodov, P. Wonka, S. Tulyakov, et al · 2024
Closest in time.
Generative camera dolly: Extreme monocular dynamic novel view synthesis
B. Van Hoorick, R. Wu, E. Ozguroglu, K. Sargent, R. Liu, P. Tokmakov, A. Dave, C. Zheng, and C. Vondrick · 2024
Closest in time.