Fetching the paper…
Reading the bibliography…
Training a deep neural network to maximize a target objective has become the standard recipe for successful machine learning over the last decade.
The monte carlo method
Nicholas Metropolis and Stanislaw Ulam · 1949
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Arthur L. Samuel · 1959
Earlier work this paper cites.
Applied dynamic programming
Richard E. Bellman and Stuart E. Dreyfus · 1962
Earlier work this paper cites.
Axiomatic derivation of the principle of maximum entropy and the principle of minimum cross-entropy
John E. Shore and Rodney W. Johnson · 1980
Earlier work this paper cites.
Training and tracking in robotics
Oliver G. Selfridge, Richard S. Sutton, and Andrew G. Barto · 1985
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long Ji Lin · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Advantage updating
Leemon C Baird · 1993
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 1993
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
John N. Tsitsiklis · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Reinforcement learning: A tutorial
Mance E Harmon and Stephanie S Harmon · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L. Littman, and Andrew W. Moore · 1996
Earlier work this paper cites.
Incremental multi-step q-learning
Jing Peng and Ronald J. Williams · 1996
Earlier work this paper cites.
Actor-critic algorithms
Vijay R. Konda and John N. Tsitsiklis · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bram Bakker · 2001
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L. Bartlett, and Jonathan Baxter · 2001
Earlier work this paper cites.
Simulation-based optimization of markov reward processes
Peter Marbach and John N. Tsitsiklis · 2001
Earlier work this paper cites.
Exploration in gradient-based reinforcement learning
Nicolas Meuleau, Leonid Peshkin, and Kee-Eung Kim · 2001
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Reinforcement learning and its relationship to supervised learning
Andrew G Barto and Thomas G Dietterich · 2004
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L. Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
A short tutorial on reinforcement learning
Chengcheng Li and Larry D. Pyeatt · 2004
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Christopher M. Bishop · 2006
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li jia Li, Kai Li, and Li Fei-fei · 2009
Earlier work this paper cites.
Reinforcement learning: A tutorial survey and recent advances
Abhijit Gosavi · 2009
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd Edition
Trevor Hastie, Robert Tibshirani, and Jerome H. Friedman · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Flexible shaping: How learning in small steps helps
Kai A Krueger and Peter Dayan · 2009
Earlier work this paper cites.
Double q-learning
Hado Hasselt · 2010
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J. Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Off-policy actor-critic
Thomas Degris, Martha White, and Richard S. Sutton · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin A. Riedmiller · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Matthew J. Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
Karthik Narasimhan, Tejas D. Kulkarni, and Regina Barzilay · 2015
Earlier work this paper cites.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Lectures on reinforcement learning
David Silver · 2015
Earlier work this paper cites.
Faulty reward functions in the wild
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
The CMA evolution strategy: A tutorial
Nikolaus Hansen · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep reinforcement learning: An overview
Seyed Sajad Mousavi, Michael Schukat, and Enda Howley · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Training deep neural networks via direct loss minimization
Yang Song, Alexander G. Schwing, Richard S. Zemel, and Raquel Urtasun · 2016
Cited alongside, same era.
Deep reinforcement learning: A brief survey
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath · 2017
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Learning quadrupedal locomotion over challenging terrain
Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Sample factory: Egocentric 3d control from pixels at 100000 FPS with asynchronous reinforcement learning
Aleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme, and Vladlen Koltun · 2020
Later among the works it cites.
Automatic curriculum learning for deep RL: A short survey
Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, and Pierre-Yves Oudeyer · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy P. Lillicrap, and David Silver · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Vecerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin A. Riedmiller · 2017
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Later among the works it cites.
End-to-end model-free reinforcement learning for urban driving using implicit affordances
Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, and Marc G. Bellemare · 2021
Later among the works it cites.
Towards deeper deep reinforcement learning with spectral normalization
Nils Bjorck, Carla P Gomes, and Kilian Q Weinberger · 2021
Later among the works it cites.
GRI: general reinforced imitation and its application to vision-based autonomous driving
Raphaël Chekroun, Marin Toromanoff, Sascha Hornauer, and Fabien Moutarde · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Later among the works it cites.
Phasic policy gradient
Karl Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman · 2021
Later among the works it cites.
First return, then explore
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 2021
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Later among the works it cites.
Metricopt: Learning to optimize black-box evaluation metrics
Chen Huang, Shuangfei Zhai, Pengsheng Guo, and Josh M. Susskind · 2021
Later among the works it cites.
Prioritized level replay
Minqi Jiang, Edward Grefenstette, and Tim Rocktäschel · 2021
Later among the works it cites.
Reinforcement learning: a survey
Deepali J Joshi, Ishaan Kale, Sadanand Gandewar, Omkar Korate, Divya Patwari, and Shivkumar Patil · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2021
Later among the works it cites.
End-to-end urban driving by imitating a reinforcement learning coach
Zhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu, and Luc Van Gool · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan · 2022
Later among the works it cites.
Stabilizing off-policy deep reinforcement learning from pixels
Edoardo Cetin, Philip J Ball, Stephen Roberts, and Oya Celiktutan · 2022
Later among the works it cites.
Redeeming intrinsic rewards via constrained optimization
Eric Chen, Zhang-Wei Hong, Joni Pajarinen, and Pulkit Agrawal · 2022
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan D. Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, Craig Donner, Leslie Fritz, Cristian Galperti, Andrea Huber, James Keeling, Maria Tsimpoukelli, Jackie Kay, Antoine Merle, Jean-Marc Moret, Seb Noury, Federico Pesamosca, David Pfau, Olivier Sauter, Cristian Sommariva, Stefano Coda, Basil Duval, Ambrogio Fasoli, Pushmeet Kohli, Koray Kavukcuoglu, Demis Hassabis, and Martin A. Riedmiller · 2022
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J. R. Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, David Silver, Demis Hassabis, and Pushmeet Kohli · 2022
Later among the works it cites.
The 37 implementation details of proximal policy optimization
Shengyi Huang, Rousslan Fernand Julien Dossa, Antonin Raffin, Anssi Kanervisto, and Weixun Wang · 2022
Later among the works it cites.
Learning robust perceptive locomotion for quadrupedal robots in the wild
Takahiro Miki, Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron C. Courville · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob N. Foerster, Edward Grefenstette, and Tim Rocktäschel · 2022
Later among the works it cites.
Mastering the game of stratego with model-free multiagent reinforcement learning
Julien Perolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vincent de Boer, Paul Muller, Jerome T. Connor, Neil Burch, Thomas Anthony, Stephen McAleer, Romuald Elie, Sarah H. Cen, Zhe Wang, Audrunas Gruslys, Aleksandra Malysheva, Mina Khan, Sherjil Ozair, Finbarr Timbers, Toby Pohlen, Tom Eccles, Mark Rowland, Marc Lanctot, Jean-Baptiste Lespiau, Bilal Piot, Shayegan Omidshafiei, Edward Lockhart, Laurent Sifre, Nathalie Beauguerlange, Remi Munos, David Silver, Satinder Singh, Demis Hassabis, and Karl Tuyls · 2022
Later among the works it cites.
The phenomenon of policy churn
Tom Schaul, André Barreto, John Quan, and Georg Ostrovski · 2022
Later among the works it cites.
Causality for machine learning
Bernhard Schölkopf · 2022
Later among the works it cites.
Deep reinforcement learning: A survey
Xu Wang, Sen Wang, Xingxing Liang, Dawei Zhao, Jincai Huang, Xin Xu, Bin Dai, and Qiguang Miao · 2022
Later among the works it cites.
Outracing champion gran turismo drivers with deep reinforcement learning
Peter R. Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Thomas J. Walsh, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, Leilani Gilpin, Piyush Khandelwal, Varun Kompella, HaoChih Lin, Patrick MacAlpine, Declan Oller, Takuma Seno, Craig Sherstan, Michael D. Thomure, Houmehr Aghabozorgi, Leon Barrett, Rory Douglas, Dion Whitehead, Peter Dürr, Peter Stone, Michael Spranger, and Hiroaki Kitano · 2022
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2022
Later among the works it cites.
Training diffusion models with reinforcement learning
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine · 2023
Closest in time.
Pink noise is all you need: Colored noise exploration in deep reinforcement learning
Onno Eberhard, Jakob Hollenstein, Cristina Pinneri, and Georg Martius · 2023
Closest in time.
DPOK: reinforcement learning for fine-tuning text-to-image diffusion models
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee, and Kimin Lee · 2023
Closest in time.
Benchmarking offline reinforcement learning on real-robot hardware
Nico Gürtler, Sebastian Blaes, Pavel Kolev, Felix Widmaier, Manuel Wuthrich, Stefan Bauer, Bernhard Schölkopf, and Georg Martius · 2023
Closest in time.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Closest in time.
General intelligence requires rethinking exploration
Minqi Jiang, Tim Rocktäschel, and Edward Grefenstette · 2023
Closest in time.
Human-level atari 200x faster
Steven Kapturowski, Victor Campos, Ray Jiang, Nemanja Rakicevic, Hado van Hasselt, Charles Blundell, and Adrià Puigdomènech Badia · 2023
Closest in time.
Champion-level drone racing using deep reinforcement learning
Elia Kaufmann, Leonard Bauersfeld, Antonio Loquercio, Matthias Müller, Vladlen Koltun, and Davide Scaramuzza · 2023
Closest in time.
Cs 285: Deep reinforcement learning
Sergey Levine · 2023
Closest in time.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 2023
Closest in time.
Faster sorting algorithms discovered using deep reinforcement learning
Daniel J Mankowitz, Andrea Michi, Anton Zhernov, Marco Gelmi, Marco Selvi, Cosmin Paduraru, Edouard Leurent, Shariq Iqbal, Jean-Baptiste Lespiau, Alex Ahern, Thomas Köppe, Kevin Millikin, Stephen Gaffney, Sophie Elster, Jackson Broshear, Chris Gamble, Kieran Milan, Robert Tung, Minjae Hwang, Taylan Cemgil, Mohammadamin Barekatain, Yujia Li, Amol Mandhane, Thomas Hubert, Julian Schrittwieser, Demis Hassabis, Pushmeet Kohli, Martin Riedmiller, Oriol Vinyals, and David Silver · 2023
Closest in time.
Tuning computer vision models with task rewards
André Susano Pinto, Alexander Kolesnikov, Yuge Shi, Lucas Beyer, and Xiaohua Zhai · 2023
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Rafael Figueiredo Prudencio, Marcos R. O. A. Maximo, and Esther Luna Colombini · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Closest in time.
Bigger, better, faster: Human-level atari with human-level efficiency
Max Schwarzer, Johan Samir Obando-Ceron, Aaron C. Courville, Marc G. Bellemare, Rishabh Agarwal, and Pablo Samuel Castro · 2023
Closest in time.
A tutorial introduction to reinforcement learning
Mathukumalli Vidyasagar · 2023
Closest in time.
Diffusion model alignment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik · 2023
Closest in time.
Pairwise proximal policy optimization: Harnessing relative feedback for LLM alignment
Tianhao Wu, Banghua Zhu, Ruoyu Zhang, Zhaojin Wen, Kannan Ramchandran, and Jiantao Jiao · 2023
Closest in time.
Causally aligned curriculum learning
Mingxuan Li, Junzhe Zhang, and Elias Bareinboim · 2024
Closest in time.