Fetching the paper…
Reading the bibliography…
Although Deep Reinforcement Learning (DRL) has achieved notable success in numerous robotic applications, designing a high-performing reward function remains a challenging task that often requires substantial manual input.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
State of the art—a survey of partially observable Markov decision processes: theory, models, and algorithms
George E Monahan. 1982 · 1982
Earlier work this paper cites.
Introduction to reinforcement learning . Vol. 135
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping. In International Conference on Machine Learning , Vol. 99. Citeseer, 278–287
Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999 · 1999
Earlier work this paper cites.
Cognition and behavior in normal-form games: An experimental study
Miguel Costa-Gomes, Vincent P Crawford, and Bruno Broseta. 2001 · 2001
Earlier work this paper cites.
Mujoco: A physics engine for model-based control. In IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 5026–5033
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
On signal temporal logic. In International Conference on Runtime Verification . Springer, 382–383
Alexandre Donzé. 2013 · 2013
Earlier work this paper cites.
Learning navigation behaviors end-to-end with autorl
Hao-Tien Lewis Chiang, Aleksandra Faust, Marek Fiser, and Anthony Francis. 2019 · 2014
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models. In Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing . 1173–1178
Joe Davison, Joshua Feldman, and Alexander M Rush. 2019 · 2019
Earlier work this paper cites.
Evolving rewards to automate reinforcement learning
Aleksandra Faust, Anthony Francis, and Dar Mehta. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning . PMLR, 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Review of deep reinforcement learning for robot manipulation. In IEEE International Conference on Robotic Computing . IEEE, 590–595
Hai Nguyen and Hung La. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019 · 2019
Earlier work this paper cites.
Regularized evolution for image classifier architecture search. In AAAI Conference on Artificial Intelligence , Vol. 33. 4780–4789
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. 2019 · 2019
Cited alongside, same era.
What matters for on-policy deep actor-critic methods? a large-scale study. In International Conference on Learning Representations
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphaël Marinier, Leonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, et al · 2020
Cited alongside, same era.
Rlbench: The robot learning benchmark & learning environment
Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J Davison. 2020 · 2020
Cited alongside, same era.
Learning locomotion for legged robots based on reinforcement learning: A survey. In International Conference on Electrical Engineering and Control Technologies . IEEE, 1–7
Jinghong Yue. 2020 · 2020
Cited alongside, same era.
robosuite: A modular simulation framework and benchmark for robot learning
Eager: Asking and answering questions for automatic reward shaping in language-guided rl
Thomas Carta, Pierre-Yves Oudeyer, Olivier Sigaud, and Sylvain Lamprier. 2022 · 2022
Later among the works it cites.
Large language models can self-improve
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Automated reinforcement learning (autorl): A survey and open problems
Jack Parker-Holder, Raghu Rajan, Xingyou Song, André Biedenkapp, Yingjie Miao, Theresa Eimer, Baohe Zhang, Vu Nguyen, Roberto Calandra, Aleksandra Faust, et al · 2022
Later among the works it cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Martín-Martín, Abhishek Joshi, Soroush Nasiriany, and Yifeng Zhu. 2020 · 2020
Cited alongside, same era.
Drone deep reinforcement learning: A review
Ahmad Taher Azar, Anis Koubaa, Nada Ali Mohamed, Habiba A Ibrahim, Zahra Fathy Ibrahim, Muhammad Kazim, Adel Ammar, Bilel Benjdira, Alaa M Khamis, Ibrahim A Hameed, et al · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Reward function design in reinforcement learning
Jonas Eschmann. 2021 · 2021
Cited alongside, same era.
AutoML: A survey of the state-of-the-art
Xin He, Kaiyong Zhao, and Xiaowen Chu. 2021 · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Cited alongside, same era.
Ella: Exploration through learned language abstraction
Suvir Mirchandani, Siddharth Karamcheti, and Dorsa Sadigh. 2021 · 2021
Cited alongside, same era.
NVIDIA Isaac Sim
NVIDIA. 2021 · 2021
Cited alongside, same era.
Crazyflie
BitCraze. 2023 · 2023
Closest in time.
Franka Emika
Franka Emika. 2023 · 2023
Closest in time.
Language instructed reinforcement learning for human-ai coordination
Hengyuan Hu and Dorsa Sadigh. 2023 · 2023
Closest in time.
Reward design with language models
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh. 2023 · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Closest in time.
Omniverse Isaac Gym Reinforcement Learning Environment
NVIDIA. 2023 · 2023
Closest in time.
Language to Rewards for Robotic Skill Synthesis
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, et al · 2023
Closest in time.
Zhehua Zhou, Jiayang Song, Xuan Xie, Zhan Shu, Lei Ma, Dikai Liu, Jianxiong Yin, and Simon See. 2023 · 2023
Closest in time.