Fetching the paper…
Reading the bibliography…
Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL).
Andrea Zanette and Emma Brunskill · 1901
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time. iv
Monroe D Donsker and SR Srinivasa Varadhan · 1983
Earlier work this paper cites.
ALVINN: an autonomous land vehicle in a neural network
Dean Pomerleau · 1988
Earlier work this paper cites.
Reinforcement Learning: an Introduction
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
An asymptotic property of model selection criteria
Yuhong Yang and A.R. Barron · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara A van de Geer · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Covering number bounds of certain regularized linear function classes
Tong Zhang · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y. Ng · 2004
Earlier work this paper cites.
Empirical minimization
Peter L Bartlett and Shahar Mendelson · 2006
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2006
Earlier work this paper cites.
From ε \varepsilon -entropy to KL-entropy: Analysis of minimum information complexity density estimation
Tong Zhang · 2006
Earlier work this paper cites.
Introduction to Nonparametric Estimation
A.B. Tsybakov · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stephane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer
Y. Zhao, D. Zeng, M. A. Socinski, and M. R. Kosorok · 2011
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Ashesh Jain, Brian Wojcik, Thorsten Joachims, and Ashutosh Saxena · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: Evolutionary direct policy search using a preference-based racing algorithm
Róbert Busa-Fekete, Balázs Szörényi, Paul Weng, Weiwei Cheng, and Eyke Hüllermeier · 2014
Earlier work this paper cites.
Uniform central limit theorems , volume 142
Richard M Dudley · 2014
Earlier work this paper cites.
Logistic regression: Tight bounds for stochastic and online optimization
Elad Hazan, Tomer Koren, and Kfir Y Levy · 2014
Earlier work this paper cites.
Concentration inequalities
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Playing atari games with deep reinforcement learning and human checkpoint replay
I.-A. Hosu and T. Rebedea · 2016
Cited alongside, same era.
Reinforcement learning with few expert demonstrations
A. S. Lakshminarayanan, S. Ozair, and Y. Bengio · 2016
Cited alongside, same era.
f f -divergence inequalities
Igal Sason and Sergio Verdú · 2016
Cited alongside, same era.
Probability in high dimensions
Ramon van Handel · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Planning in markov decision processes with gap-dependent sample complexity
Anders Jonsson, Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Edouard Leurent, and Michal Valko · 2020
Later among the works it cites.
Dueling posterior sampling for preference-based reinforcement learning
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue, and Joel Burdick · 2020
Later among the works it cites.
Toward the fundamental limits of imitation learning
Nived Rajaraman, Lin Yang, Jiantao Jiao, and Kannan Ramchandran · 2020
Later among the works it cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Later among the works it cites.
Leverage the average: An analysis of kl regularization in reinforcement learning
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, and Matthieu Geist · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Vecerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin A. Riedmiller · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Cited alongside, same era.
Playing hard exploration games by watching youtube
Yusuf Aytar, Tobias Pfaff, David Budden, Thomas Paine, Ziyu Wang, and Nando de Freitas · 2018
Cited alongside, same era.
Yichong Xu, Ruosong Wang, Lin F. Yang, Aarti Singh, and Artur Dubrawski · 2020
Later among the works it cites.
Navigating to the best policy in markov decision processes
Aymen Al Marjani, Aurélien Garivier, and Alexandre Proutiere · 2021
Later among the works it cites.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Demonstration-guided reinforcement learning with learned skills
Karl Pertsch, Youngwoon Lee, Yue Wu, and Joseph J. Lim · 2021
Later among the works it cites.
On the value of interaction and function approximation in imitation learning
Nived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu, Jiantao Jiao, and Kannan Ramchandran · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Tengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong, and Yu Bai · 2021
Later among the works it cites.
Near-optimal offline reinforcement learning via double variance reduction
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Later among the works it cites.
Efficient online learning to rank for sequential music recommendation
Pedro Dalla Vecchia Chaves, Bruno L. Pereira, and Rodrygo L. T. Santos · 2022
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de las Casas, Craig Donner, Leslie Fritz, Cristian Galperti, Andrea Huber, James Keeling, Maria Tsimpoukelli, Jackie Kay, Antoine Merle, Jean-Marc Moret, Seb Noury, Federico Pesamosca, David Pfau, Olivier Sauter, Cristian Sommariva, Stefano Coda, Basil Duval, Ambrogio Fasoli, Pushmeet Kohli, Koray Kavukcuoglu, Demis Hassabis, and Martin Riedmiller · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Closest in time.
Faster sorting algorithms discovered using deep reinforcement learning
Daniel J. Mankowitz, Andrea Michi, Anton Zhernov, Marco Gelmi, Marco Selvi, Cosmin Paduraru, Edouard Leurent, Shariq Iqbal, Jean-Baptiste Lespiau, Alex Ahern, Thomas Köppe, Kevin Millikin, Stephen Gaffney, Sophie Elster, Jackson Broshear, Chris Gamble, Kieran Milan, Robert Tung, Minjae Hwang, Taylan Cemgil, Mohammadamin Barekatain, Yujia Li, Amol Mandhane, Thomas Hubert, Julian Schrittwieser, Demis Hassabis, Pushmeet Kohli, Martin Riedmiller, Oriol Vinyals, and David Silver · 2023
Closest in time.
Dueling RL: reinforcement learning with trajectory preferences
Aadirupa Saha, Aldo Pacchiano, and Jonathan Lee · 2023
Closest in time.
Best policy identification in discounted linear MDPs
Jérôme Taupin, Yassir Jedra, and Alexandre Proutiere · 2023
Closest in time.
Fast rates for maximum entropy exploration
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines, Remi Munos, Alexey Naumov, Pierre Perrault, Yunhao Tang, Michal Valko, and Pierre Menard · 2023
Closest in time.
High-probability risk bounds via sequential predictors, 2023
Dirk van der Hoeven, Nikita Zhivotovskiy, and Nicolò Cesa-Bianchi · 2023
Closest in time.
Is RLHF more difficult than standard rl?
Yuanhao Wang, Qinghua Liu, and Chi Jin · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Closest in time.