Fetching the paper…
Reading the bibliography…
In this work, we review research studies that combine Reinforcement Learning (RL) and Large Language Models (LLMs), two areas that owe their momentum to the development of deep neural networks.
A markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
Dynamic programming and markov processes. ronald a. howard. technology press and wiley, new york, 1960. viii + 136 pp. illus. $5.75
G. Weiss · 1960
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
N. Chentanez, A. Barto, and S. Singh · 2004
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Playing atari with deep reinforcement learning, 2013
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Jumping nlp curves: A review of natural language processing research [review article]
E. Cambria and B. White · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
A survey of graphs in natural language processing
V. Nastase, R. Mihalcea, and D. R. Radev · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2015
Earlier work this paper cites.
Survey of natural language processing techniques in bioinformatics
Z. Zeng, H. Shi, Y. Wu, Z. Hong, et al · 2015
Earlier work this paper cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Solving general arithmetic word problems, 2016
S. Roy and D. Roth · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas · 2016
Earlier work this paper cites.
A brief survey of deep reinforcement learning
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath · 2017
Earlier work this paper cites.
Deal or no deal? end-to-end learning for negotiation dialogues, 2017
M. Lewis, D. Yarats, Y. N. Dauphin, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
An introduction to deep reinforcement learning
V. François-Lavet, P. Henderson, R. Islam, M. G. Bellemare, J. Pineau, et al · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
A survey on natural language processing for fake news detection
R. Oshikawa, J. Qian, and W. Y. Wang · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Teaching multiple tasks to an rl agent using ltl
R. Toro Icarte, T. Q. Klassen, R. Valenzano, and S. A. McIlraith · 2018
Earlier work this paper cites.
A survey on intrinsic motivation in reinforcement learning, 2019
A. Aubret, L. Matignon, and S. Hassas · 2019
Earlier work this paper cites.
The hanabi challenge: A new frontier for ai research
N. Bard, J. N. Foerster, S. Chandar, N. Burch, M. Lanctot, H. F. Song, E. Parisotto, V. Dumoulin, S. Moitra, E. Hughes, I. Dunning, S. Mourad, H. Larochelle, M. G. Bellemare, and M. Bowling · 2019
Earlier work this paper cites.
Hardware conditioned policies for multi-robot transfer learning, 2019
T. Chen, A. Murali, and A. Gupta · 2019
Earlier work this paper cites.
Babyai: A platform to study the sample efficiency of grounded language learning, 2019
M. Chevalier-Boisvert, D. Bahdanau, S. Lahlou, L. Willems, C. Saharia, T. H. Nguyen, and Y. Bengio · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Using natural language for reward shaping in reinforcement learning, 2019
P. Goyal, S. Niekum, and R. J. Mooney · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning, 2019
Y. Wu, G. Tucker, and O. Nachum · 2019
Earlier work this paper cites.
Generative pretraining from pixels
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever · 2020
Earlier work this paper cites.
Natural language processing
K. Chowdhary and K. Chowdhary · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning, 2020
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
A survey of the usages of deep learning for natural language processing
D. W. Otter, J. R. Medina, and J. K. Kalita · 2020
Cited alongside, same era.
Pre-trained models for natural language processing: A survey
X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, and X. Huang · 2020
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Anymal, 2023
AnyRobotics · 2023
Later among the works it cites.
Evolutionary reinforcement learning: A survey
H. Bai, R. Cheng, and Y. Jin · 2023
Later among the works it cites.
A Survey of Meta-Reinforcement Learning
J. Beck, R. Vuorio, E. Zheran Liu, Z. Xiong, L. Zintgraf, C. Finn, and S. Whiteson · 2023
Later among the works it cites.
Crazyflie, 2023
BitCraze · 2023
Later among the works it cites.
Grounding large language models in interactive environments with online reinforcement learning, 2023
T. Carta, C. Romac, T. Wolf, S. Lamprier, O. Sigaud, and P.-Y. Oudeyer · 2023
Later among the works it cites.
A survey on evaluation of large language models
Y. Chang, X. Wang, J. Wang, Y. Wu, K. Zhu, H. Chen, L. Yang, X. Yi, C. Wang, Y. Wang, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Torfi, R. A. Shirvani, Y. Keneshloo, N. Tavaf, and E. A. Fox · 2020
Cited alongside, same era.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications
W. Du and S. Ding · 2021
Cited alongside, same era.
Reward Function Design in Reinforcement Learning , pages 25–33
J. Eschmann · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning, 2021
S. Fujimoto and S. S. Gu · 2021
Cited alongside, same era.
Reinforcement learning for mobile robotics exploration: A survey
L. C. Garaffa, M. Basso, A. A. Konzen, and E. P. de Freitas · 2021
Cited alongside, same era.
The franka emika robot: A reference platform for robotics research and education
S. Haddadin, S. Parusel, L. Johannsmeier, S. Golz, S. Gabl, F. Walch, M. Sabaghian, C. Jähne, L. Hausperger, and S. Haddadin · 2021
Cited alongside, same era.
Later among the works it cites.
Collaborating with language models for embodied reasoning, 2023
I. Dasgupta, C. Kaeser-Chen, K. Marino, A. Ahuja, S. Babayan, F. Hill, and R. Fergus · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Later among the works it cites.
Large language models for software engineering: Survey and open problems, 2023
A. Fan, B. Gokkaya, M. Harman, M. Lyubarskiy, S. Sengupta, S. Yoo, and J. M. Zhang · 2023
Later among the works it cites.
Bias and fairness in large language models: A survey, 2023
I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Dernoncourt, T. Yu, R. Zhang, and N. K. Ahmed · 2023
Later among the works it cites.
Maniskill2: A unified benchmark for generalizable manipulation skills, 2023
J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y. Tang, S. Tao, X. Wei, Y. Yao, X. Yuan, P. Xie, Z. Huang, R. Chen, and H. Su · 2023
Later among the works it cites.
Evaluating large language models: A comprehensive survey
Z. Guo, R. Jin, C. Liu, Y. Huang, D. Shi, L. Yu, Y. Liu, J. Li, B. Xiong, D. Xiong, et al · 2023
Later among the works it cites.
Language instructed reinforcement learning for human-ai coordination, 2023
H. Hu and D. Sadigh · 2023
Later among the works it cites.
Aligning language models with offline reinforcement learning from human feedback
J. Hu, L. Tao, J. Yang, and C. Zhou · 2023
Later among the works it cites.
Language is not all you need: Aligning perception with language models
S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, Q. Liu, et al · 2023
Later among the works it cites.
Language-informed transfer learning for embodied household activities, 2023
Y. Jiang, Q. Gao, G. Thattai, and G. Sukhatme · 2023
Later among the works it cites.
Ai-augmented surveys: Leveraging large language models for opinion prediction in nationally representative surveys, 2023
J. Kim and B. Lee · 2023
Later among the works it cites.
Reward design with language models, 2023
M. Kwon, S. M. Xie, K. Bullard, and D. Sadigh · 2023
Later among the works it cites.
How can recommender systems benefit from large language models: A survey, 2023
J. Lin, X. Dai, Y. Xi, W. Liu, B. Chen, X. Li, C. Zhu, H. Guo, Y. Yu, R. Tang, and W. Zhang · 2023
Later among the works it cites.
Summary of ChatGPT-related research and perspective towards the future of large language models
Y. Liu, T. Han, S. Ma, J. Zhang, Y. Yang, J. Tian, H. He, A. Li, M. He, Z. Liu, Z. Wu, L. Zhao, D. Zhu, X. Li, N. Qiang, D. Shen, T. Liu, and B. Ge · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar · 2023
Later among the works it cites.
Augmented language models: a survey, 2023
G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celikyilmaz, E. Grave, Y. LeCun, and T. Scialom · 2023
Later among the works it cites.
Reinforcement learning on graphs: A survey
M. Nie, D. Chen, and D. Wang · 2023
Later among the works it cites.
Unifying large language models and knowledge graphs: A roadmap, 2023
S. Pan, L. Luo, Y. Wang, C. Chen, J. Wang, and X. Wu · 2023
Later among the works it cites.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
R. F. Prudencio, M. R. Maximo, and E. L. Colombini · 2023
Later among the works it cites.
Automatic prompt optimization with "gradient descent" and beam search, 2023
R. Pryzant, D. Iter, J. Li, Y. T. Lee, C. Zhu, and M. Zeng · 2023
Later among the works it cites.
Exploiting contextual structure to generate useful auxiliary tasks, 2023
B. Quartey, A. Shah, and G. Konidaris · 2023
Later among the works it cites.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization, 2023
R. Ramamurthy, P. Ammanabrolu, K. Brantley, J. Hessel, R. Sifa, C. Bauckhage, H. Hajishirzi, and Y. Choi · 2023
Later among the works it cites.
Syndicom: Improving conversational commonsense with error-injection and natural language feedback
C. Richardson, A. Sundar, and L. Heck · 2023
Later among the works it cites.
Large language model alignment: A survey
T. Shen, R. Jin, Y. Huang, C. Liu, W. Dong, Z. Guo, X. Wu, Y. Liu, and D. Xiong · 2023
Later among the works it cites.
J. Song, Z. Zhou, J. Liu, C. Fang, Z. Shu, and L. Ma · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, et al · 2023
Later among the works it cites.
Text2reward: Automated dense reward function generation for reinforcement learning
T. Xie, S. Zhao, C. H. Wu, Y. Liu, Q. Luo, V. Zhong, Y. Yang, and T. Yu · 2023
Later among the works it cites.
Harnessing the power of llms in practice: A survey on chatgpt and beyond, 2023
J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, B. Yin, and X. Hu · 2023
Later among the works it cites.
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks, 2023
H. Yuan, C. Zhang, H. Wang, F. Xie, P. Cai, H. Dong, and Z. Lu · 2023
Later among the works it cites.
Instruction tuning for large language models: A survey, 2023
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu, and G. Wang · 2023
Later among the works it cites.
Rladapter: Bridging large language models to reinforcement learning in open worlds, 2023
W. Zhang and Z. Lu · 2023
Later among the works it cites.
Large language models for information retrieval: A survey, 2023
Y. Zhu, H. Yuan, S. Wang, J. Liu, W. Liu, C. Deng, Z. Dou, and J.-R. Wen · 2023
Later among the works it cites.