Fetching the paper…
Reading the bibliography…
Automatically synthesizing dense rewards from natural language descriptions is a promising paradigm in reinforcement learning (RL), with applications to sparse reward problems, open-ended exploration, and hierarchical skill design.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
Jette Randlov and Preben Alstrøm · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
R-MAX - A general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Satinder Singh, Richard L Lewis, Andrew G Barto, and Jonathan Sorg · 2010
Earlier work this paper cites.
Internal rewards mitigate agent boundedness
Jonathan Sorg, Satinder P Singh, and Richard L Lewis · 2010
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2016
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Earlier work this paper cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
Go-explore: a new approach for hard-exploration problems
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2019
Cited alongside, same era.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2019
Cited alongside, same era.
The laplacian in RL: Learning representations with efficient approximations
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2020
Cited alongside, same era.
The NetHack Learning Environment
Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Motif: Intrinsic motivation from artificial intelligence feedback
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, and Mikael Henaff · 2023
Later among the works it cites.
Katakomba: Tools and benchmarks for data-driven nethack, 2023
Vladislav Kurenkov, Alexander Nikulin, Denis Tarasov, and Sergey Kolesnikov · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
Nethack is hard to hack
Ulyana Piterbarg, Lerrel Pinto, and Rob Fergus · 2023
Later among the works it cites.
Vision-language models are zero-shot reward models for reinforcement learning
Juan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez, and David Lindner · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sample factory: Egocentric 3d control from pixels at 100000 FPS with asynchronous reinforcement learning
Aleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme, and Vladlen Koltun · 2020
Cited alongside, same era.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
Roberta Raileanu and Tim Rocktäschel · 2020
Cited alongside, same era.
Nle challenge baseline using sample-factory
Anssi Myffili · 2021
Cited alongside, same era.
Towards better laplacian representation in reinforcement learning with generalized graph drawing
Kaixin Wang, Kuangqi Zhou, Qixin Zhang, Jie Shao, Bryan Hooi, and Jiashi Feng · 2021
Cited alongside, same era.
Noveld: A simple yet effective exploration criterion
Tianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu, Kurt Keutzer, Joseph E. Gonzalez, and Yuandong Tian · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances, 2022
Michael Ahn et al · 2022
Cited alongside, same era.
Text2reward: Automated dense reward function generation for reinforcement learning
Tianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu, Qian Luo, Victor Zhong, Yanchao Yang, and Tao Yu · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, et al · 2023
Later among the works it cites.
Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods
Yuji Cao, Huan Zhao, Yuheng Cheng, Ting Shu, Guolong Liu, Gaoqi Liang, Junhua Zhao, and Yun Li · 2024
Closest in time.
The llama 3 herd of models, 2024
Abhimanyu Dubey et al · 2024
Closest in time.
Proper laplacian representation learning, 2024
Diego Gomez, Michael Bowling, and Marlos C. Machado · 2024
Closest in time.
Sinan Ibrahim, Mostafa Mostafa, Ali Jnadi, and Pavel Osinenko · 2024
Closest in time.
Playing nethack with llms: Potential & limitations as zero-shot agents
Dominik Jeurissen, Diego Perez-Liebana, Jeremy Gow, Duygu Cakmak, and James Kwan · 2024
Closest in time.
Guide your agent with adaptive multimodal rewards
Changyeon Kim, Younggyo Seo, Hao Liu, Lisa Lee, Jinwoo Shin, Honglak Lee, and Kimin Lee · 2024
Closest in time.
Auto mc-reward: Automated dense reward design with large language models for minecraft
Hao Li, Xue Yang, Zhaokai Wang, Xizhou Zhu, Jie Zhou, Yu Qiao, Xiaogang Wang, Hongsheng Li, Lewei Lu, and Jifeng Dai · 2024
Closest in time.
Intelligent go-explore: Standing on the shoulders of giant foundation models
Cong Lu, Shengran Hu, and Jeff Clune · 2024
Closest in time.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2024
Closest in time.
Read and reap the rewards: Learning to play atari with the help of instruction manuals
Yue Wu, Yewen Fan, Paul Pu Liang, Amos Azaria, Yuanzhi Li, and Tom M Mitchell · 2024
Closest in time.
Maestromotif: Skill design from artificial intelligence feedback
Martin Klissarov, Mikael Henaff, Roberta Raileanu, Shagun Sodhani, Pascal Vincent, Amy Zhang, Pierre-Luc Bacon, Doina Precup, Marlos C. Machado, and Pierluca D’Oro · 2025
Closest in time.