Fetching the paper…
Reading the bibliography…
Generalization capabilities, or rather a lack thereof, is one of the most important unsolved problems in the field of robot learning, and while several large scale efforts have set out to tackle this problem, unsolved it remains.
Learning latent plans from play
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet · 1903
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, H. Zhou, A. Mohamed, and M. Auli · 2006
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
K. Sohn, H. Lee, and X. Yan · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies, 2016
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
One-shot imitation learning, 2017
Y. Duan, M. Andrychowicz, B. C. Stadie, J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning, 2017
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Simple nearest neighbor policy method for continuous control tasks, 2018
E. Mansimov and K. Cho · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Jukebox: A generative model for music, 2020
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever · 2020
Earlier work this paper cites.
Strictly batch imitation learning by energy-based distribution matching
D. Jarrett, I. Bica, and M. van der Schaar · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Earlier work this paper cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning, 2021
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2021
Earlier work this paper cites.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, S. Piao, and F. Wei · 2021
Earlier work this paper cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets, 2021
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Earlier work this paper cites.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. Lim · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Parrot: Data-driven behavioral priors for reinforcement learning
A. Singh, H. Liu, G. Zhou, A. Yu, N. Rhinehart, and S. Levine · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances, 2022
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Later among the works it cites.
Scaling up and distilling down: Language-guided robot skill acquisition
H. Ha, P. Florence, and S. Song · 2023
Later among the works it cites.
H-gap: Humanoid control with a generalist planner
Z. Jiang, Y. Xu, N. Wagener, Y. Luo, M. Janner, E. Grefenstette, T. Rocktäschel, and Y. Tian · 2023
Later among the works it cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Later among the works it cites.
Z. Liang, Y. Mu, H. Ma, M. Tomizuka, M. Ding, and P. Luo · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Cited alongside, same era.
From play to policy: Conditional behavior generation from uncurated robot data
Z. J. Cui, Y. Wang, N. M. M. Shafiullah, and L. Pinto · 2022
Cited alongside, same era.
Implicit behavioral cloning
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization, 2022
D. Lee, C. Kim, S. Kim, M. Cho, and W.-S. Han · 2022
Cited alongside, same era.
Towards more generalizable one-shot visual imitation learning
Z. Mandi, F. Liu, K. Lee, and P. Abbeel · 2022
Cited alongside, same era.
A generalist agent, 2022
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Later among the works it cites.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Later among the works it cites.
Action-quantized offline reinforcement learning for robotic skill learning
J. Luo, P. Dong, J. Wu, A. Kumar, X. Geng, and S. Levine · 2023
Later among the works it cites.
Finite scalar quantization: Vq-vae made simple
F. Mentzer, D. Minnen, E. Agustsson, and M. Tschannen · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Lotus: Continual imitation learning for robot manipulation through unsupervised skill discovery
W. Wan, Y. Zhu, R. Shah, and Y. Zhu · 2023
Later among the works it cites.
Chaineddiffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation
Z. Xian, N. Gkanatsios, T. Gervet, T.-W. Ke, and K. Fragkiadaki · 2023
Later among the works it cites.
Instruction tuning for large language models: A survey
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu, et al · 2023
Later among the works it cites.
Learning fine-grained bimanual manipulation with low-cost hardware, 2023
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Later among the works it cites.
Autort: Embodied foundation models for large scale orchestration of robotic agents, 2024
M. Ahn, D. Dwibedi, C. Finn, M. G. Arenas, K. Gopalakrishnan, K. Hausman, B. Ichter, A. Irpan, N. Joshi, R. Julian, S. Kirmani, I. Leal, E. Lee, S. Levine, Y. Lu, I. Leal, S. Maddineni, K. Rao, D. Sadigh, P. Sanketi, P. Sermanet, Q. Vuong, S. Welker, F. Xia, T. Xiao, P. Xu, S. Xu, and Z. Xu · 2024
Closest in time.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Z. Fu, T. Z. Zhao, and C. Finn · 2024
Closest in time.
Baku: An efficient transformer for multi-task policy learning, 2024
S. Haldar, Z. Peng, and L. Pinto · 2024
Closest in time.
Behavior generation with latent actions
S. Lee, Y. Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto · 2024
Closest in time.
Libero: Benchmarking knowledge transfer for lifelong robot learning
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone · 2024
Closest in time.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Closest in time.
Dinobot: Robot manipulation via retrieval and alignment with vision foundation models
N. D. Palo and E. Johns · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
G. Team · 2024
Closest in time.
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Closest in time.
Segment everything everywhere all at once
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y. J. Lee · 2024
Closest in time.