Fetching the paper…
Reading the bibliography…
We propose BOSS, an approach that automatically learns to solve new long-horizon, complex, and meaningful tasks by growing a learned skill library with minimal supervision.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Policyblocks: An algorithm for creating useful macro-actions in reinforcement learning
M. Pickett and A. G. Barto · 2002
Earlier work this paper cites.
Dynamic movement primitives–a framework for motor control in humans and humanoid robotics
S. Schaal · 2006
Earlier work this paper cites.
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Layer normalization
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
Variational option discovery algorithms
J. Achiam, H. Edwards, D. Amodei, and P. Abbeel · 2018
Earlier work this paper cites.
Relay policy learning: Solving long horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
MCP: Learning composable hierarchical control with multiplicative compositional policies
X. B. Peng, M. Chang, G. Zhang, P. Abbeel, and S. Levine · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning, 2019
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
Composing complex skills by learning transition policies with proximity reward induction
Y. Lee, S.-H. Sun, S. Somasundaram, E. Hu, and J. J. Lim · 2019
Earlier work this paper cites.
Dynamics-aware unsupervised discovery of skills
A. Sharma, S. Gu, S. Levine, V. Kumar, and K. Hausman · 2019
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
Earlier work this paper cites.
Unsupervised control through non-parametric discriminative rewards
D. Warde-Farley, T. V. de Wiele, T. Kulkarni, C. Ionescu, S. Hansen, and V. Mnih · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. J. Lim · 2020
Earlier work this paper cites.
Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data
A. Mandlekar, F. Ramos, B. Boots, S. Savarese, L. Fei-Fei, A. Garg, and D. Fox · 2020
Earlier work this paper cites.
Planning to explore via self-supervised world models
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak · 2020
Cited alongside, same era.
Program guided agent
S.-H. Sun, T.-L. Wu, and J. J. Lim · 2020
Cited alongside, same era.
Language as a cognitive tool to imagine goals in curiosity driven exploration
C. Colas, T. Karch, N. Lair, J.-M. Dussoux, C. Moulin-Frier, F. P. Dominey, and P.-Y. Oudeyer · 2020
Cited alongside, same era.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Cited alongside, same era.
Dynamics-aware unsupervised discovery of skills
A. Sharma, S. Gu, S. Levine, V. Kumar, and K. Hausman · 2020
Cited alongside, same era.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnkov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter · 2022
Later among the works it cites.
ProgPrompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2022
Later among the works it cites.
Skill-based meta-reinforcement learning
T. Nam, S.-H. Sun, K. Pertsch, S. J. Hwang, and J. J. Lim · 2022
Later among the works it cites.
Skill-based model-based reinforcement learning
L. X. Shi, J. J. Lim, and Y. Lee · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale, 2022
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Beyond pick-and-place: Tackling robotic stacking of diverse shapes
A. X. Lee, C. Devin, Y. Zhou, T. Lampe, K. Bousmalis, J. T. Springenberg, A. Byravan, A. Abdolmaleki, N. Gileadi, D. Khosid, C. Fantacci, J. E. Chen, A. Raju, R. Jeong, M. Neunert, A. Laurens, S. Saliceti, F. Casarini, M. Riedmiller, R. Hadsell, and F. Nori · 2021
Cited alongside, same era.
Demonstration-guided reinforcement learning with learned skills
K. Pertsch, Y. Lee, Y. Wu, and J. J. Lim · 2021
Cited alongside, same era.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets, 2021
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2021
Cited alongside, same era.
Learning to synthesize programs as interpretable and generalizable policies
D. Trivedi, J. Zhang, S.-H. Sun, and J. J. Lim · 2021
Cited alongside, same era.
Hierarchical reinforcement learning by discovering intrinsic options
J. Zhang, H. Yu, and W. Xu · 2021
Cited alongside, same era.
Later among the works it cites.
Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
D. Shah, B. Osinski, B. Ichter, and S. Levine · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2022
Later among the works it cites.
CIC: Contrastive intrinsic control for unsupervised skill discovery, 2022
M. Laskin, H. Liu, X. B. Peng, D. Yarats, A. Rajeswaran, and P. Abbeel · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. V. Le, and J. Wei · 2022
Later among the works it cites.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Later among the works it cites.
Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation
M. Heo, Y. Lee, D. Lee, and J. J. Lim · 2023
Closest in time.
Hierarchical programmatic reinforcement learning via learning to compose programs
G.-T. Liu, E.-P. Hu, P.-J. Cheng, H.-Y. Lee, and S.-H. Sun · 2023
Closest in time.
Controllability-aware unsupervised skill discovery
S. Park, K. Lee, Y. Lee, and P. Abbeel · 2023
Closest in time.
Sprint: Scalable policy pre-training via language instruction relabeling, 2023
J. Zhang, K. Pertsch, J. Zhang, and J. J. Lim · 2023
Closest in time.
Tail: Task-specific adapters for imitation learning with large pretrained models, 2023
Z. Liu, J. Zhang, K. Asadi, Y. Liu, D. Zhao, S. Sabach, and R. Fakoor · 2023
Closest in time.
Guiding pretraining in reinforcement learning with large language models
Y. Du, O. Watkins, Z. Wang, C. Colas, T. Darrell, P. Abbeel, A. Gupta, and J. Andreas · 2023
Closest in time.
Augmenting autotelic agents with large language models
C. Colas, L. Teodorescu, P.-Y. Oudeyer, X. Yuan, and M.-A. Côté · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
Efficient online reinforcement learning with offline data
P. J. Ball, L. Smith, I. Kostrikov, and S. Levine · 2023
Closest in time.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Closest in time.