Fetching the paper…
Reading the bibliography…
We study generalizable policy learning from demonstrations for complex low-level control (e.g., contact-rich object manipulations).
Dynamic programming and markov processes
Howard, R. A · 1960
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1988
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1992
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Policyblocks: An algorithm for creating useful macro-actions in reinforcement learning
Pickett, M. and Barto, A. G · 2002
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
Gaussian process change point models
Saatçi, Y., Turner, R. D., and Rasmussen, C. E · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Long short-term memory
Graves, A. and Graves, A · 2012
Earlier work this paper cites.
Optimal detection of changepoints with a linear computational cost
Killick, R., Fearnhead, P., and Eckley, I. A · 2012
Earlier work this paper cites.
Robot learning from demonstration by constructing skill trees
Konidaris, G., Kuindersma, S., Grupen, R., and Barto, A · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Mujoco haptix: A virtual reality system for hand manipulation
Kumar, V. and Todorov, E · 2015
Earlier work this paper cites.
Neural programmer-interpreters
Reed, S. and De Freitas, N · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Dart: Noise injection for robust imitation learning
Laskey, M., Lee, J., Fox, R., Dragan, A., and Goldberg, K · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Qi, C. R., Su, H., Mo, K., and Guibas, L. J · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2017
Earlier work this paper cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al · 2018
Earlier work this paper cites.
Policy optimization with demonstrations
Kang, B., Jie, Z., and Feng, J · 2018
Earlier work this paper cites.
Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration
Rahmatizadeh, R., Abolghasemi, P., Bölöni, L., and Levine, S · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Neural task programming: Learning to generalize across hierarchical tasks
Xu, D., Nair, S., Zhu, Y., Gao, J., Garg, A., Fei-Fei, L., and Savarese, S · 2018
Earlier work this paper cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
Zhang, T., McCarthy, Z., Jow, O., Lee, D., Chen, X., Goldberg, K., and Abbeel, P · 2018
Earlier work this paper cites.
Disagreement-regularized imitation learning
Brantley, K., Sun, W., and Henaff, M · 2019
Earlier work this paper cites.
Self-supervised correspondence in visuomotor policy learning
Florence, P., Manuelli, L., and Tedrake, R · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K · 2019
Earlier work this paper cites.
Compile: Compositional imitation learning and execution
Kipf, T., Li, Y., Dai, H., Zambaldi, V., Sanchez-Gonzalez, A., Grefenstette, E., Kohli, P., and Battaglia, P · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Pvnet: Pixel-wise voting network for 6dof pose estimation
Peng, S., Liu, Y., Huang, Q., Zhou, X., and Bao, H · 2019
Cited alongside, same era.
Motion planning networks
Qureshi, A. H., Simeonov, A., Bency, M. J., and Yip, M. C · 2019
Cited alongside, same era.
Explain yourself! leveraging language models for commonsense reasoning
Rajani, N. F., McCann, B., Xiong, C., and Socher, R · 2019
Cited alongside, same era.
Imitation learning from imperfect demonstration
Wu, Y.-H., Charoenphakdee, N., Bao, H., Tangkaratt, V., and Sugiyama, M · 2019
Cited alongside, same era.
Xu, J., Hou, Z., Liu, Z., and Qiao, H · 2019
Cited alongside, same era.
On covariate shift of latent confounders in imitation and reinforcement learning
Tennenholtz, G., Hallak, A., Dalal, G., Mannor, S., Chechik, G., and Shalit, U · 2021
Later among the works it cites.
Single rgb-d camera teleoperation for general robotic manipulation
Vuong, Q., Qin, Y., Guo, R., Wang, X., Su, H., and Christensen, H · 2021
Later among the works it cites.
Transporter networks: Rearranging the visual world for robotic manipulation
Zeng, A., Florence, P., Tompson, J., Welker, S., Chien, J., Attarian, M., Armstrong, T., Krasin, I., Duong, D., Sindhwani, V., et al · 2021
Later among the works it cites.
Hierarchical task learning from language instructions with unified transformers and self-monitoring
Zhang, Y. and Chai, J · 2021
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Transformers as soft reasoners over language
Clark, P., Tafjord, O., and Richardson, K · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., et al · 2022
Later among the works it cites.
Is conditional generative modeling all you need for decision-making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J., Jaakkola, T., and Agrawal, P · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Later among the works it cites.
A system for general in-hand object re-orientation
Chen, T., Xu, J., and Agrawal, P · 2022
Later among the works it cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Later among the works it cites.
Fishman, A., Murali, A., Eppner, C., Peele, B., Boots, B., and Fox, D · 2022
Later among the works it cites.
Implicit behavioral cloning
Florence, P., Lynch, C., Zeng, A., Ramirez, O. A., Wahid, A., Downs, L., Wong, A., Lee, J., Mordatch, I., and Tompson, J · 2022
Later among the works it cites.
Multi-skill mobile manipulation for object rearrangement
Gu, J., Chaplot, D. S., Su, H., and Malik, J · 2022
Later among the works it cites.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
Hansen, N., Yuan, Z., Ze, Y., Mu, T., Rajeswaran, A., Su, H., Xu, H., and Wang, X · 2022
Later among the works it cites.
Hierarchical reinforcement learning: A survey and open research challenges
Hutsebaut-Buysse, M., Mets, K., and Latré, S · 2022
Later among the works it cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
James, S. and Davison, A. J · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Later among the works it cites.
Learning options via compression
Jiang, Y., Liu, E., Eysenbach, B., Kolter, J. Z., and Finn, C · 2022
Later among the works it cites.
Masked autoencoding for scalable and generalizable decision making
Liu, F., Liu, H., Grover, A., and Abbeel, P · 2022
Later among the works it cites.
Pan, Y., Li, Y., Zhang, Y., Cai, Q., Long, F., Qiu, Z., Yao, T., and Mei, T · 2022
Later among the works it cites.
You can’t count on luck: Why decision transformers fail in stochastic environments
Paster, K., McIlraith, S., and Ba, J · 2022
Later among the works it cites.
Qin, Y., Su, H., and Wang, X · 2022
Later among the works it cites.
Mo2: Model-based offline options
Salter, S., Wulfmeier, M., Tirumala, D., Heess, N., Riedmiller, M., Hadsell, R., and Rao, D · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
Shafiullah, N. M. M., Cui, Z. J., Altanzaya, A., and Pinto, L · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D · 2022
Later among the works it cites.
Chain of thought imitation with procedure cloning
Yang, M., Schuurmans, D., Abbeel, P., and Nachum, O · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Zhou, K., Yang, J., Loy, C. C., and Liu, Z · 2022
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Closest in time.
Motion policy networks
Fishman, A., Murali, A., Eppner, C., Peele, B., Boots, B., and Fox, D · 2023
Closest in time.
Maniskill2: A unified benchmark for generalizable manipulation skills
Gu, J., Xiang, F., Li, X., Ling, Z., Liu, X., Mu, T., Tang, Y., Tao, S., Wei, X., Yao, Y., Yuan, X., Xie, P., Huang, Z., Chen, R., and Su, H · 2023
Closest in time.
Hierarchical imitation learning with vector quantized models
Kujanpää, K., Pajarinen, J., and Ilin, A · 2023
Closest in time.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Closest in time.
Foundation models for decision making: Problems, methods, and opportunities
Yang, S., Nachum, O., Du, Y., Wei, J., Abbeel, P., and Schuurmans, D · 2023
Closest in time.