Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) plays an important role in the robotic manipulation domain since it allows self-learning from trial-and-error interactions with the environment.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. Russell, et al · 2000
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W. B. Knox and P. Stone · 2009
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
A. Wilson, A. Fern, and P. Tadepalli · 2012
Earlier work this paper cites.
A gesture learning interface for simulated robot path shaping with a human teacher
P. M. Yanik, J. Manganelli, J. Merino, A. L. Threatt, J. O. Brooks, K. E. Green, and I. D. Walker · 2013
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
S. Griffith, K. Subramanian, J. Scholz, C. L. Isbell, and A. L. Thomaz · 2013
Earlier work this paper cites.
Teaching on a budget: Agents advising agents in reinforcement learning
L. Torrey and M. Taylor · 2013
Earlier work this paper cites.
Multi-modal integration of dynamic audiovisual patterns for an interactive reinforcement learning scenario
F. Cruz, G. I. Parisi, J. Twiefel, and S. Wermter · 2016
Earlier work this paper cites.
Deep reinforcement learning: A brief survey
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
J. MacGlashan, M. K. Ho, R. Loftin, B. Peng, G. Wang, D. L. Roberts, M. E. Taylor, and M. L. Littman · 2017
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces
G. Warnell, N. Waytowich, V. Lawhern, and P. Stone · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Earlier work this paper cites.
Learning to parse natural language to grounded reward functions with weak supervision
E. C. Williams, N. Gopalan, M. Rhee, and S. Tellex · 2018
Cited alongside, same era.
DQN-TAMER: human-in-the-loop reinforcement learning with intractable feedback
R. Arakawa, S. Kobayashi, Y. Unno, Y. Tsuboi, and S.-i. Maeda · 2018
Cited alongside, same era.
Improving interactive reinforcement learning: What makes a good teacher?
F. Cruz, S. Magg, Y. Nagai, and S. Wermter · 2018
Cited alongside, same era.
Using natural language for reward shaping in reinforcement learning
P. Goyal, S. Niekum, and R. J. Mooney · 2019
Cited alongside, same era.
Deep reinforcement learning from policy-dependent human feedback
D. Arumugam, J. K. Lee, S. Saskin, and M. L. Littman · 2019
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus · 2022
Later among the works it cites.
Vision-language pre-training: Basics, recent advances, and future trends
Z. Gan, L. Li, C. Li, L. Wang, Z. Liu, J. Gao, et al · 2022
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Closest in time.
ChatGPT for robotics: Design principles and model abilities
S. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
RLBench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Cited alongside, same era.
How to train your robot with deep reinforcement learning: lessons we have learned
J. Ibarz, J. Tan, C. Finn, M. Kalakrishnan, P. Pastor, and S. Levine · 2021
Cited alongside, same era.
A survey of inverse reinforcement learning: Challenges, methods and progress
S. Arora and P. Doshi · 2021
Cited alongside, same era.
Recent advances in leveraging human guidance for sequential decision-making tasks
R. Zhang, F. Torabi, G. Warnell, and P. Stone · 2021
Cited alongside, same era.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
K. Lee, L. M. Smith, and P. Abbeel · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Leveraging language for accelerated learning of tool manipulation
A. Z. Ren, B. Govil, T.-Y. Yang, K. R. Narasimhan, and A. Majumdar · 2023
Closest in time.
Chat with the environment: Interactive multimodal perception using large language models
X. Zhao, M. Li, C. Weber, M. B. Hafez, and S. Wermter · 2023
Closest in time.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Closest in time.
Language to rewards for robotic skill synthesis
W. Yu, N. Gileadi, C. Fu, S. Kirmani, K.-H. Lee, M. G. Arenas, H.-T. L. Chiang, T. Erez, L. Hasenclever, J. Humplik, et al · 2023
Closest in time.
RLAIF: Scaling reinforcement learning from human feedback with AI feedback
H. Lee, S. Phatale, H. Mansoor, K. Lu, T. Mesnard, C. Bishop, V. Carbune, and A. Rastogi · 2023
Closest in time.
Reward design with language models
M. Kwon, S. M. Xie, K. Bullard, and D. Sadigh · 2023
Closest in time.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Efficient memory management for large language model serving with PagedAttention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Closest in time.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Closest in time.