Fetching the paper…
Reading the bibliography…
Leveraging human preferences for steering the behavior of Large Language Models (LLMs) has demonstrated notable success in recent years.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
On a Measure of the Information Provided by an Experiment
D. V. Lindley · 1956
Earlier work this paper cites.
Sample estimate of the entropy of a random vector
L. F. Kozachenko and N. Leonenko · 1987
Earlier work this paper cites.
Information-based objective functions for active data selection
David J. C. MacKay · 1992
Earlier work this paper cites.
Pairwise preference learning and ranking
Johannes Fürnkranz and Eyke Hüllermeier · 2003
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stoegbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Preference learning with gaussian processes
Wei Chu and Zoubin Ghahramani · 2005
Earlier work this paper cites.
Batch mode active learning and its application to medical image classification
Steven C. H. Hoi, Rong Jin, Jianke Zhu, and Michael R. Lyu · 2006
Earlier work this paper cites.
Bayesian active learning for classification and preference learning, 2011
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel · 2011
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Dueling bandits: Beyond condorcet winners to general tournament solutions
Siddartha Y. Ramamohan, Arun Rajkumar, Shivani Agarwal, and Shivani Agarwal · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Deep bayesian active learning with image data
Yarin Gal, Riashat Islam, and Zoubin Ghahramani · 2017
Earlier work this paper cites.
TL;DR: Mining Reddit to learn automatic summarization
Michael Völske, Martin Potthast, Shahbaz Syed, and Benno Stein · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Deep Bayesian active learning for natural language processing: Results of a large-scale empirical study
Aditya Siddhant and Zachary C. Lipton · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks
Sylvestre-Alvise Rebuffi, Andrea Vedaldi, and Hakan Bilen · 2018
Earlier work this paper cites.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Andreas Kirsch, Joost van Amersfoort, and Yarin Gal · 2019
Earlier work this paper cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Earlier work this paper cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
On-the-fly closed-loop autonomous materials discovery via bayesian active learning
A. Gilad Kusne, Heshan Yu, Changming Wu, Huairuo Zhang, Jason Hattrick-Simpers, Brian DeCost, Suchismita Sarker, Corey Oses, Cormac Toher, Stefano Curtarolo, Albert V. Davydov, Ritesh Agarwal, Leonid A. Bendersky, Mo Li, Apurva Mehta, and Ichiro Takeuchi · 2020
Cited alongside, same era.
Quantifying aleatoric and epistemic uncertainty using density estimation in latent space
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Tong Wang, Samuel Marks, Charbel-Raphael Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Biyik, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
Deep batch active learning for drug discovery
Michael Bailey, Saeed Moayedpour, Ruijiang Li, Alejandro Corrochano-Navarro, Alexander Kötter, Lorenzo Kogler-Anele, Saleh Riahi, Christoph Grebner, Gerhard Hessler, Hans Matter, Marc Bianciotto, Pablo Mas, Ziv Bar-Joseph, and Sven Jager · 2023
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Janis Postels, Hermann Blum, César Cadena, Roland Y. Siegwart, Luc Van Gool, and Federico Tombari · 2020
Cited alongside, same era.
Energy-based out-of-distribution detection
Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li · 2020
Cited alongside, same era.
Uncertainty estimation using a single deep deterministic neural network
Joost Van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal · 2020
Cited alongside, same era.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew G Wilson and Pavel Izmailov · 2020
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Cited alongside, same era.
Bayesian deep active learning for medical image analysis
Biraja Ghoshal, Stephen Swift, and Allan Tucker · 2021
Cited alongside, same era.
Deep deterministic uncertainty: A new simple baseline
Jishnu Mukhoti, Andreas Kirsch, Joost R. van Amersfoort, Philip H. S. Torr, and Yarin Gal · 2021
Cited alongside, same era.
Behavior from the void: Unsupervised active pre-training
Hao Liu and Pieter Abbeel · 2021
Cited alongside, same era.
Advanced deep active learning and data subset selection: unifying principles with information-theory intuitions
Andreas Kirsch · 2023
Later among the works it cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Later among the works it cites.
Stochastic batch acquisition: A simple baseline for deep active learning
Andreas Kirsch, Sebastian Farquhar, Parmida Atighehchian, Andrew Jesson, Frédéric Branchaud-Charron, and Yarin Gal · 2023
Later among the works it cites.
A bayesian active learning platform for scalable combination drug screens
Christopher Tosh, Mauricio Tec, Jessica White, Jeffrey F. Quinn, Glorymar Ibanez Sanchez, Paul Calder, Andrew L. Kung, Filemon S. Dela Cruz, and Wesley Tansey · 2023
Later among the works it cites.
Epistemic neural networks
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, MORTEZA IBRAHIMI, Xiuyuan Lu, and Benjamin Van Roy · 2023
Later among the works it cites.
Accelerating reinforcement learning with value-conditional state entropy exploration
Dongyoung Kim, Jinwoo Shin, Pieter Abbeel, and Younggyo Seo · 2023
Later among the works it cites.
Probabilistic Machine Learning: Advanced Topics
Kevin P. Murphy · 2023
Later among the works it cites.
Prediction-oriented bayesian active learning
Freddie Bickford Smith, Andreas Kirsch, Sebastian Farquhar, Yarin Gal, Adam Foster, and Tom Rainforth · 2023
Later among the works it cites.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Later among the works it cites.
Demystifying prompts in language models via perplexity estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah Smith, and Luke Zettlemoyer · 2023
Later among the works it cites.
Introducing Meta Llama 3: The most capable openly available LLM to date — ai.meta.com
Meta Llama Team · 2024
Closest in time.
Reward model ensembles help mitigate overoptimization
Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger · 2024
Closest in time.
Sample efficient reinforcement learning from human feedback via active exploration, 2024
Viraj Mehta, Vikramjeet Das, Ojash Neopane, Yijia Dai, Ilija Bogunovic, Jeff Schneider, and Willie Neiswanger · 2024
Closest in time.
Reinforcement learning from human feedback with active queries, 2024
Kaixuan Ji, Jiafan He, and Quanquan Gu · 2024
Closest in time.
DUO: Diverse, uncertainty-aware, on-policy query generation and selection for reinforcement learning from human feedback, 2024
Anonymous · 2024
Closest in time.
Provably sample efficient rlhf via active preference optimization, 2024
Nirjhar Das, Souradip Chakraborty, Aldo Pacchiano, and Sayak Ray Chowdhury · 2024
Closest in time.
Efficient exploration for llms
Vikranth Reddy Dwaracherla, Seyed Mohammad Asghari, Botao Hao, and Benjamin Van Roy · 2024
Closest in time.
A sequential algorithm for training text classifiers
David D. Lewis and William A. Gale · 2099
Closest in time.