Fetching the paper…
Reading the bibliography…
Adapting Large Language Models (LLMs) to downstream tasks using Reinforcement Learning (RL) has proven to be an effective approach.
Recent advances in imitation learning from observation
F. Torabi, G. Warnell, and P. Stone · 1905
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Reinforcement learning - an introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford · 2018
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Imitating latent policies from observation
A. D. Edwards, H. Sahni, Y. Schroecker, and C. L. I. Jr · 2019
Earlier work this paper cites.
On reinforcement learning for full-length game of starcraft
Z.-J. Pang, R.-Z. Liu, Z.-Y. Meng, Y. Zhang, Y. Yu, and T. Lu · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Provably efficient imitation learning from observation alone
W. Sun, A. Vemula, B. Boots, and D. Bagnell · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
M. Shridhar, X. Yuan, M.-A. Côté, Y. Bisk, A. Trischler, and M. Hausknecht · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Off-policy imitation learning from observations
Z. Zhu, K. Lin, B. Dai, and J. Zhou · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Using word clouds for fast identification of papers’ subject domain and reviewers’ competences
Y. Kalmukov · 2021
Earlier work this paper cites.
Mobile: Model-based imitation learning from observation alone
R. Kidambi, J. Chang, and W. Sun · 2021
Earlier work this paper cites.
Wudaocorpora: A super large-scale chinese corpora for pre-training language models
S. Yuan, H. Zhao, Z. Du, M. Ding, X. Liu, Y. Cen, X. Zou, Z. Yang, and J. Tang · 2021
Earlier work this paper cites.
Video pretraining (VPT): learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Earlier work this paper cites.
GLM: general language model pretraining with autoregressive blank infilling
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang · 2022
Earlier work this paper cites.
Plan your target and learn your skills: Transferable state-only imitation learning via decoupled policy optimization
M. Liu, Z. Zhu, Y. Zhuang, W. Zhang, J. Hao, Y. Yu, and J. Wang · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Reinforcement learning with action-free pre-training from videos
Y. Seo, K. Lee, S. James, and P. Abbeel · 2022
Cited alongside, same era.
Scienceworld: Is your agent smarter than a 5th grader?
R. Wang, P. Jansen, M.-A. Côté, and P. Ammanabrolu · 2022
Cited alongside, same era.
Slimpajama-627b
Cerebras · 2023
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Reinforce++: A simple and efficient approach for aligning large language models
J. Hu · 2024
Later among the works it cites.
C. Jia, P. Wang, Z. Li, Y. Li, Z. Zhang, N. Tang, and Y. Yu · 2024
Later among the works it cites.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
X. Lai, Z. Tian, Y. Chen, S. Yang, X. Peng, and J. Jia · 2024
Later among the works it cites.
Let’s verify step by step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2024
Later among the works it cites.
Mitigating the alignment tax of RLHF
Y. Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang, H. Dong, R. Pi, H. Zhao, N. Jiang, H. Ji, Y. Yao, and T. Zhang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2023
Cited alongside, same era.
L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu · 2023
Cited alongside, same era.
Starcoder: may the source be with you!
R. Li, L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, Q. Liu, E. Zheltonozhskii, T. Y. Zhuo, T. Wang, O. Dehaene, M. Davaadorj, J. Lamy-Poirier, J. Monteiro, O. Shliazhko, N. Gontier, N. Meade, A. Zebaze, M. Yee, L. K. Umapathi, J. Zhu, B. Lipkin, M. Oblokulov, Z. Wang, R. M. V, J. T. Stillerman, S. S. Patel, D. Abulkhanov, M. Zocca, M. Dey, Z. Zhang, N. Fahmy, U. Bhattacharyya, W. Yu, S. Singh, S. Luccioni, P. Villegas, M. Kunakov, F. Zhdanov, M. Romero, T. Lee, N. Timor, J. Ding, C. Schlesinger, H. Schoelkopf, J. Ebert, T. Dao, M. Mishra, A. Gu, J. Robinson, C. J. Anderson, B. Dolan-Gavitt, D. Contractor, S. Reddy, D. Fried, D. Bahdanau, Y. Jernite, C. M. Ferrandis, S. Hughes, T. Wolf, A. Guha, L. von Werra, and H. de Vries · 2023
Cited alongside, same era.
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Cited alongside, same era.
Monte carlo tree search: a review of recent modifications and applications
M. Swiechowski, K. Godlewski, B. Sawicki, and J. Mandziuk · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Cited alongside, same era.
Later among the works it cites.
Litcab: Lightweight language model calibration over short- and long-form responses
X. Liu, M. Khalifa, and L. Wang · 2024
Later among the works it cites.
Language model self-improvement by reinforcement learning contemplation
J. Pang, P. Wang, K. Li, X. Chen, J. Xu, Z. Zhang, and Y. Yu · 2024
Later among the works it cites.
Learning to act without actions
D. Schmidt and M. Jiang · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, and D. Guo · 2024
Later among the works it cites.
Trial and error: Exploration-based trajectory optimization for llm agents
Y. Song, D. Yin, X. Yue, J. Huang, S. Li, and B. Y. Lin · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al · 2024
Later among the works it cites.
Predictive inverse dynamics models are scalable learners for robotic manipulation
Y. Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang · 2024
Later among the works it cites.
Q*: Improving multi-step reasoning for llms with deliberative planning
C. Wang, Y. Deng, Z. Lv, Z. Liang, J. He, S. Yan, and B. An · 2024
Later among the works it cites.
Qurating: Selecting high-quality data for training language models
A. Wettig, A. Gupta, S. Malik, and D. Chen · 2024
Later among the works it cites.
Qwen2.5-math technical report: Toward mathematical expert model via self-improvement
A. Yang, B. Zhang, B. Hui, B. Gao, B. Yu, C. Li, D. Liu, J. Tu, J. Zhou, J. Lin, K. Lu, M. Xue, R. Lin, T. Liu, X. Ren, and Z. Zhang · 2024
Later among the works it cites.
Physics of language models: Part 2.2, how to learn from mistakes on grade-school math problems
T. Ye, Z. Xu, Y. Li, and Z. Allen-Zhu · 2024
Later among the works it cites.
Rest-mcts*: LLM self-training via process reward guided tree search
D. Zhang, S. Zhoubian, Y. Yue, Y. Dong, and J. Tang · 2024
Later among the works it cites.
C. Zheng, K. Sun, H. Wu, C. Xi, and X. Zhou · 2024
Later among the works it cites.
Dpo meets ppo: Reinforced token optimization for rlhf
H. Zhong, G. Feng, W. Xiong, X. Cheng, L. Zhao, D. He, J. Bian, and L. Wang · 2024
Later among the works it cites.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars
K. Gandhi, A. Chakravarthy, A. Singh, N. Lile, and N. D. Goodman · 2025
Closest in time.
Can better cold-start strategies improve rl training for llms?
Z. Li · 2025
Closest in time.
Preserving diversity in supervised fine-tuning of large language models
Z. Li, C. Chen, T. Xu, Z. Qin, J. Xiao, Z.-Q. Luo, and R. Sun · 2025
Closest in time.