Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have formulated a blueprint for the advancement of artificial general intelligence.
Trust region policy optimization
Schulman, J., S. Levine, P. Abbeel, et al · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., K. Kavukcuoglu, D. Silver, et al · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., A. P. Badia, M. Mirza, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., J. Leike, T. Brown, et al · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
MacGlashan, J., M. K. Ho, R. Loftin, et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., F. Wolski, P. Dhariwal, et al · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., N. Stiennon, J. Wu, et al · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., J. Buys, L. Du, et al · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., A. Ghandeharioun, J. H. Shen, et al · 2019
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Keskar, N., B. McCann, L. Varshney, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., B. Mann, N. Ryder, et al · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., L. Ouyang, J. Wu, et al · 2020
Earlier work this paper cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo, 2020
Engstrom, L., A. Ilyas, S. Santurkar, et al · 2020
Earlier work this paper cites.
Gender and representation bias in gpt-3 generated stories
Lucy, L., D. Bamman · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., T. Gebru, A. McMillan-Major, et al · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., D. A. Hudson, E. Adeli, et al · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Askell, A., Y. Bai, A. Chen, et al · 2021
Cited alongside, same era.
What matters for on-policy deep actor-critic methods? a large-scale study
Andrychowicz, M., A. Raichuk, P. Stańczyk, et al · 2021
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Wei, J., X. Wang, D. Schuurmans, et al · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Z. Li, Z. Lin, et al · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
A survey of large language models
Zhao, W. X., K. Zhou, J. Li, et al · 2023
Closest in time.
Peng, B., C. Li, P. He, et al · 2023
Closest in time.
Stanford alpaca: An instruction-following LLaMA model
Taori, R., I. Gulrajani, T. Zhang, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lamda: Language models for dialog applications
Thoppilan, R., D. De Freitas, J. Hall, et al · 2022
Cited alongside, same era.
Planning for agi and beyond
Altman, S · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., J. Wu, X. Jiang, et al · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., A. Jones, K. Ndousse, et al · 2022
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback, 2022
Bai, Y., S. Kadavath, S. Kundu, et al · 2022
Cited alongside, same era.
The 37 implementation details of proximal policy optimization
Huang, S., R. F. J. Dossa, A. Raffin, et al · 2022
Cited alongside, same era.
Easy RL: Reinforcement Learning Tutorial
Qi Wang, J. J., Yiyuan Yang · 2022
Cited alongside, same era.
Driess, D., F. Xia, M. S. Sajjadi, et al · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Park, J. S., J. C. O’Brien, C. J. Cai, et al · 2023
Closest in time.
Open-Chinese-LLaMA: Chinese large language model base generated through incremental pre-training on chinese datasets
OpenLMLab · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023
Chiang, W.-L., Z. Li, Z. Lin, et al · 2023
Closest in time.
Belle: Be everyone’s large language model engine
Ji, Y., Y. Deng, Y. Gong, et al · 2023
Closest in time.
StackLLaMA: An RL fine-tuned LLaMA model for stack exchange question and answering, 2023
Beeching, E., Y. Belkada, K. Rasul, et al · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback, 2023
Dubois, Y., X. Li, R. Taori, et al · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., W.-L. Chiang, Y. Sheng, et al · 2023
Closest in time.