Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have achieved remarkable success at tasks like summarization that involve a single turn of interaction.
Quinoa: a q-function you infer normalized over actions, 2019
Jonas Degrave, Abbas Abdolmaleki, Jost Tobias Springenberg, Nicolas Heess, and Martin Riedmiller · 1911
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
A study of cross-validation and bootstrap for accuracy estimation and model selection
Ron Kohavi · 1995
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Covariant policy search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
Policy search by dynamic programming
James Bagnell, Sham M Kakade, Jeff Schneider, and Andrew Ng · 2003
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J Andrew Bagnell · 2014
Earlier work this paper cites.
Contextual dueling bandits
Miroslav Dudík, Katja Hofmann, Robert E Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Michel Galley, Jianfeng Gao, and Dan Jurafsky · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods, 2018
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Earlier work this paper cites.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz · 2019
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
Wouter Kool, Herke van Hoof, and Max Welling · 2019
Earlier work this paper cites.
Continuous control with deep reinforcement learning, 2019
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning, 2021
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2024
Closest in time.
Dataset reset policy optimization for rlhf
Jonathan D Chang, Wenhao Shan, Owen Oertell, Kianté Brantley, Dipendra Misra, Jason D Lee, and Wen Sun · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
Rlhf workflow: From reward modeling to online rlhf, 2024
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang · 2024
Closest in time.
Length-controlled alpacaeval: A simple way to debias automatic evaluators, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
Efficient and optimal algorithms for contextual dueling bandits under realizability
Aadirupa Saha and Akshay Krishnamurthy · 2022
Cited alongside, same era.
Offline rl for natural language generation with implicit language q learning
Charlie Snell, Ilya Kostrikov, Yi Su, Mengjiao Yang, and Sergey Levine · 2022
Cited alongside, same era.
Lmrl gym: Benchmarks for multi-turn reinforcement learning with language models, 2023
Marwa Abdulhai, Isadora White, Charlie Snell, Charles Sun, Joey Hong, Yuexiang Zhai, Kelvin Xu, and Sergey Levine · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences, 2023
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos · 2023
Cited alongside, same era.
Learning to generate better than your llm
Jonathan D Chang, Kiante Brantley, Rajkumar Ramamurthy, Dipendra Misra, and Wen Sun · 2023
Cited alongside, same era.
Rewarding chatbots for real-world engagement with millions of users, 2023
Robert Irvine, Douglas Boubert, Vyas Raina, Adian Liusie, Ziyi Zhu, Vineet Mudupalli, Aliaksei Korshuk, Zongyi Liu, Fritz Cremer, Valentin Assassi, Christie-Carol Beauchamp, Xiaoding Lu, Thomas Rialan, and William Beauchamp · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu · 2023
Cited alongside, same era.
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B. Hashimoto · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
Rebel: Reinforcement learning via regressing relative rewards, 2024
Zhaolin Gao, Jonathan D. Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, J. Andrew Bagnell, Jason D. Lee, and Wen Sun · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Google · 2024
Closest in time.
Direct language model alignment from online ai feedback
Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao, Bilal Piot, et al · 2024
Closest in time.
Orpo: Monolithic preference optimization without reference model
Jiwoo Hong, Noah Lee, and James Thorne · 2024
Closest in time.
When is agnostic reinforcement learning statistically tractable?
Zeyu Jia, Gene Li, Alexander Rakhlin, Ayush Sekhari, and Nati Srebro · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward, 2024
Yu Meng, Mengzhou Xia, and Danqi Chen · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta · 2024
Closest in time.
Offline regularised reinforcement learning for large language models alignment, 2024
Pierre Harvey Richemond, Yunhao Tang, Daniel Guo, Daniele Calandriello, Mohammad Gheshlaghi Azar, Rafael Rafailov, Bernardo Avila Pires, Eugene Tarassov, Lucas Spangher, Will Ellsworth, Aliaksei Severyn, Jonathan Mallinson, Lior Shani, Gil Shamir, Rishabh Joshi, Tianqi Liu, Remi Munos, and Bilal Piot · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences, 2024
Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie · 2024
Closest in time.
Understanding preference fine-tuning through the lens of coverage
Yuda Song, Gokul Swamy, Aarti Singh, J Andrew Bagnell, and Wen Sun · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu, and Alekh Agarwal · 2024
Closest in time.
Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang · 2024
Closest in time.
Self-play preference optimization for language model alignment
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang, and Quanquan Gu · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees, 2024
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao Peng, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.