Fetching the paper…
Reading the bibliography…
Training large transformers using next-token prediction has given rise to groundbreaking advancements in AI.
Nonparametric entropy estimation: An overview
J. Beirlant · 1997
Earlier work this paper cites.
Nearest neighbor estimates of entropy
H. Singh, N. Misra, V. Hnizdo, A. Fedorowicz, and E. Demchuk · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
S. Singh, R. L. Lewis, A. G. Barto, and J. Sorg · 2010
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Earlier work this paper cites.
A policy gradient method for task-agnostic exploration
M. Mutti, L. Pratissoli, and M. Restelli · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Beyond fine-tuning: Transferring behavior in reinforcement learning
V. Campos, P. Sprechmann, S. Hansen, A. Barreto, S. Kapturowski, A. Vitvitskyi, A. P. Badia, and C. Blundell · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Urlb: Unsupervised reinforcement learning benchmark
M. Laskin, D. Yarats, H. Liu, K. Lee, A. Zhan, K. Lu, C. Cang, L. Pinto, and P. Abbeel · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Earlier work this paper cites.
Submodularity in machine learning and artificial intelligence
J. Bilmes · 2022
Earlier work this paper cites.
Large language models can self-improve
J. Huang, S. S. Gu, L. Hou, Y. Wu, X. Wang, H. Yu, and J. Han · 2022
Earlier work this paper cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
H. Le, Y. Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Cited alongside, same era.
Learning math reasoning from self-sampled correct and partially-correct solutions
A. Ni, J. P. Inala, C. Wang, A. Polozov, C. Meek, D. Radev, and J. Gao · 2022
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue
J. Schulman, B. Zoph, C. Kim, J. Hilton, J. Menick, J. Weng, J. F. C. Uribe, L. Fedus, L. Metz, M. Pokorny, R. G. Lopes, S. Zhao, A. Vijayvergiya, E. Sigler, A. Perelman, C. Voss, M. Heaton, J. Parish, D. Cummings, R. Nayak, V. Balcom, D. Schnurr, T. Kaftan, C. Hallacy, N. Turley, N. Deutsch, and V. Goel · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi · 2022
Chain of hindsight aligns language models with feedback
H. Liu, C. Sferrazza, and P. Abbeel · 2023
Closest in time.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, and D. Zhang · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2023
Closest in time.
Orca: Progressive learning from complex explanation traces of gpt-4
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah · 2023
Closest in time.
Probabilistic Machine Learning: Advanced Topics
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Cited alongside, same era.
Introducing claude, 2023
Anthropic · 2023
Cited alongside, same era.
Teaching large language models to self-debug
X. Chen, M. Lin, N. Schärli, and D. Zhou · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Cited alongside, same era.
Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P. Wendell, M. Zaharia, and R. Xin · 2023
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation model alignment
H. Dong, W. Xiong, D. Goyal, R. Pan, S. Diao, J. Zhang, K. Shum, and T. Zhang · 2023
Cited alongside, same era.
K. P. Murphy · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Controllability-aware unsupervised skill discovery
S. Park, K. Lee, Y. Lee, and P. Abbeel · 2023
Closest in time.
Mastering the unsupervised reinforcement learning benchmark from pixels
S. Rajeswar, P. Mazzaglia, T. Verbelen, A. Piché, B. Dhoedt, A. Courville, and A. Lacoste · 2023
Closest in time.
Code llama: Open foundation models for code
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, et al · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Z. Sun, Y. Shen, Q. Zhou, H. Zhang, Z. Chen, D. Cox, Y. Yang, and C. Gan · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model, 2023
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Voyager: An open-ended embodied agent with large language models
G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, and D. Jiang · 2023
Closest in time.
Mammoth: Building math generalist models through hybrid instruction tuning
X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Closest in time.