Fetching the paper…
Reading the bibliography…
This manuscript gives a big-picture, up-to-date overview of the field of (deep) reinforcement learning and sequential decision making, covering value-based methods, policy-based methods, model-based methods, multi-agent RL, LLMs and RL, and various other topics (e.g., offline RL, hierarchical RL, intrinsic reward).
“Go-Explore: a New Approach for Hard-Exploration Problems”, 2019
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth Stanley and Jeff Clune · 1901
Earlier work this paper cites.
“What does the free energy principle tell us about the brain?”
Samuel Gershman · 1901
Earlier work this paper cites.
“The StarCraft Multi-Agent Challenge”
Mikayel Samvelyan, Tabish Rashid, Christian de Witt, Gregory Farquhar, Nantas Nardelli, Tim G Rudner, Chia-Man Hung, Philip H Torr, Jakob Foerster and Shimon Whiteson · 1902
Earlier work this paper cites.
“An online learning approach to model predictive control”
Nolan Wagener, Ching-An Cheng, Jacob Sacks and Byron Boots · 1902
Earlier work this paper cites.
“Model-based reinforcement learning for Atari”
Lukasz Kaiser et al · 1903
Earlier work this paper cites.
“Elements of Sequential Monte Carlo”
Christian Naesseth, Fredrik Lindsten and Thomas Sch\"on · 1903
Earlier work this paper cites.
“Introduction to Multi-Armed Bandits”
Aleksandrs Slivkins · 1904
Earlier work this paper cites.
Yun Cheung and Georgios Piliouras · 1905
Earlier work this paper cites.
“When to use parametric models in reinforcement learning?”
Hado van Hasselt, Matteo Hessel and John Aslanides · 1906
Earlier work this paper cites.
“Modern Deep Reinforcement Learning algorithms”
Sergey Ivanov and Alexander D’yakonov · 1906
Earlier work this paper cites.
“When to Trust Your Model: Model-Based Policy Optimization”
Michael Janner, Justin Fu, Marvin Zhang and Sergey Levine · 1906
Earlier work this paper cites.
“BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active Learning”
Andreas Kirsch, Joost van Amersfoort and Yarin Gal · 1906
Earlier work this paper cites.
“Rethinking formal models of partially observable multiagent decision making”
Vojtěch Kovařík, Martin Schmid, Neil Burch, Michael Bowling and Viliam Lisý · 1906
Earlier work this paper cites.
“Adapting behaviour via intrinsic reward: A survey and empirical study”
Cam Linke, Nadia Ady, Martha White, Thomas Degris and Adam White · 1906
Earlier work this paper cites.
“Robust Reinforcement Learning for Continuous Control with Model Misspecification”, 2019
Daniel Mankowitz, Nir Levine, Rae Jeong, Yuanyuan Shi, Jackie Kay, Abbas Abdolmaleki, Jost Springenberg, Timothy Mann, Todd Hester and Martin Riedmiller · 1906
Earlier work this paper cites.
“Importance Resampling for Off-policy Prediction”
Matthew Schlegel, Wesley Chung, Daniel Graves, Jian Qian and Martha White · 1906
Earlier work this paper cites.
“Deep Active Inference as Variational Policy Gradients”
Beren Millidge · 1907
Earlier work this paper cites.
“Benchmarking Model-Based Reinforcement Learning”
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel and Jimmy Ba · 1907
Earlier work this paper cites.
“A survey on intrinsic motivation in reinforcement learning”
Arthur Aubret, Laetitia Matignon and Salima Hassas · 1908
Earlier work this paper cites.
“V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control”
H Francis et al · 1909
Earlier work this paper cites.
“Introduction to online convex optimization”
Elad Hazan · 1909
Earlier work this paper cites.
“Automated curricula through setter-solver interactions”
Sebastien Racaniere, Andrew Lampinen, Adam Santoro, David Reichert, Vlad Firoiu and Timothy Lillicrap · 1909
Earlier work this paper cites.
“Soft Actor-Critic for discrete action settings”
Petros Christodoulou · 1910
Earlier work this paper cites.
“Benchmarking batch deep reinforcement learning algorithms”
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh and Joelle Pineau · 1910
Earlier work this paper cites.
“Advantage-weighted regression: Simple and scalable off-policy reinforcement learning”
Xue Peng, Aviral Kumar, Grace Zhang and Sergey Levine · 1910
Earlier work this paper cites.
“Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model”
Julian Schrittwieser et al · 1911
Earlier work this paper cites.
“SMiRL: Surprise minimizing reinforcement learning in unstable environments”
Glen Berseth, Daniel Geng, Coline Devin, Nicholas Rhinehart, Chelsea Finn, Dinesh Jayaraman and Sergey Levine · 1912
Earlier work this paper cites.
Aviral Kumar, Xue Peng and Sergey Levine · 1912
Earlier work this paper cites.
“From Reinforcement Learning to Optimal Control: A unified framework for sequential decisions”
Warren Powell · 1912
Earlier work this paper cites.
“Reinforcement learning Upside Down: Don’t predict rewards – just map them to actions”
Juergen Schmidhuber · 1912
Earlier work this paper cites.
“On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples”
William Thompson · 1933
Earlier work this paper cites.
“A SIMPLIFIED TWO-PERSON POKER”
H Kuhn · 1951
Earlier work this paper cites.
“A value for n-person games”
L Shapley · 1953
Earlier work this paper cites.
“Stochastic games”
L Shapley · 1953
Earlier work this paper cites.
“Markov Processes and the H-Theorem”
Tetsuzo Morimoto · 1963
Earlier work this paper cites.
“A formal theory of inductive inference. Part I”
R Solomonoff · 1964
Earlier work this paper cites.
“A General Class of Coefficients of Divergence of One Distribution from Another”
S Ali and S Silvey · 1966
Earlier work this paper cites.
“A Taxonomy of 2 X 2 Games”
Anatol Rapoport and Melvin Guyer · 1966
Earlier work this paper cites.
“Information-Type Measures of Difference of Probability Distributions and Indirect Observations”
I. Csiszar · 1967
Earlier work this paper cites.
“Differential Dynamic Programming”
D.. Jacobson and D.. Mayne · 1970
Earlier work this paper cites.
“The evolution of cooperation”
Robert Axelrod and William Hamilton · 1981
Earlier work this paper cites.
“Neuronlike adaptive elements that can solve difficult learning control problems”
A Barto, R Sutton and C Anderson · 1983
Earlier work this paper cites.
“The evolution of cooperation”
Robert Axelrod · 1984
Earlier work this paper cites.
“Correlated equilibrium as an expression of Bayesian rationality”
Robert Aumann · 1987
Earlier work this paper cites.
“The complexity of Markov decision processes”
C. Papadimitriou and J. Tsitsiklis · 1987
Earlier work this paper cites.
“A taxonomy of all ordinal 2 x 2 games”
D Kilgour and Niall Fraser · 1988
Earlier work this paper cites.
“Learning to predict by the methods of temporal differences”
R. Sutton · 1988
Earlier work this paper cites.
“Optimal Control: Linear Quadratic Methods”
Brian.O. Anderson and John. Moore · 1989
Earlier work this paper cites.
“Multi-armed Bandit Allocation Indices”
J. Gittins · 1989
Earlier work this paper cites.
“ALVINN: An Autonomous Land Vehicle in a Neural Network”
Dean Pomerleau · 1989
Earlier work this paper cites.
“Receding horizon control of nonlinear systems”
D Mayne and H Michalska · 1990
Earlier work this paper cites.
“Integrated Architectures for Learning, Planning, and Reacting Based on Approximating Dynamic Programming”
Richard Sutton · 1990
Earlier work this paper cites.
“Principles of metareasoning”
Stuart Russell and Eric Wefald · 1991
Earlier work this paper cites.
“Feudal Reinforcement Learning”
Peter Dayan and Geoffrey Hinton · 1992
Earlier work this paper cites.
“Self-Improving Reactive Agents Based on Reinforcement Learning, Planning and Teaching”
Long-Ji Lin · 1992
Earlier work this paper cites.
“Q-learning”
C. Watkins and P. Dayan · 1992
Earlier work this paper cites.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning”
Ronald Williams · 1992
Earlier work this paper cites.
“Improving generalization for temporal difference learning: The successor representation”
Peter Dayan · 1993
Earlier work this paper cites.
“Rational learning leads to Nash equilibrium”
Ehud Kalai and Ehud Lehrer · 1993
Earlier work this paper cites.
“Prioritized Sweeping: Reinforcement Learning with Less Data and Less Time”
A.. Moore and C.. Atkeson · 1993
Earlier work this paper cites.
“Reinforcement Learning Algorithm for Partially Observable Markov Decision Problems”
T. Jaakkola, S. Singh and M. Jordan · 1994
Earlier work this paper cites.
“Markov games as a framework for multi-agent reinforcement learning”
Michael Littman · 1994
Earlier work this paper cites.
“Fast Exact Multiplication by the Hessian”
Barak Pearlmutter · 1994
Earlier work this paper cites.
“Markov Decision Processes: Discrete Stochastic Dynamic Programming”
Martin. Puterman · 1994
Earlier work this paper cites.
“Incremental Multi-Step Q-Learning”
Jing Peng and Ronald Williams · 1994
Earlier work this paper cites.
“On-Line Q-Learning Using Connectionist Systems”, 1994
G Rummery and M Niranjan · 1994
Earlier work this paper cites.
“Residual Algorithms: Reinforcement Learning with Function Approximation”
Leemon. Baird · 1995
Earlier work this paper cites.
“Learning to act using real-time dynamic programming”
Andrew Barto, Steven Bradtke and Satinder Singh · 1995
Earlier work this paper cites.
“Stable Function Approximation in Dynamic Programming”
Geoffrey. Gordon · 1995
Earlier work this paper cites.
“Quantal response equilibria for normal form games”
Richard McKelvey and Thomas Palfrey · 1995
Earlier work this paper cites.
“TD models: Modeling the world at a mixture of time scales”
Richard Sutton · 1995
Earlier work this paper cites.
“Generalization in Reinforcement Learning: Successful Examples Using Sparse Coarse Coding”
Richard Sutton · 1996
Earlier work this paper cites.
“On-line Policy Improvement using Monte-Carlo Search”
Gerald Tesauro and Gregory Galperin · 1996
Earlier work this paper cites.
“Optimization of computer simulation models with rare events”
Reuven Rubinstein · 1997
Earlier work this paper cites.
“An analysis of temporal-difference learning with function approximation”
J. Tsitsiklis and B. Roy · 1997
Earlier work this paper cites.
“An analysis of temporal-difference learning with function approximation”
J Tsitsiklis and B Van · 1997
Earlier work this paper cites.
“Natural Gradient Works Efficiently in Learning”
S Amari · 1998
Earlier work this paper cites.
“Planning and acting in Partially Observable Stochastic Domains”
L.. Kaelbling, M. Littman and A. Cassandra · 1998
Earlier work this paper cites.
“Quantal response equilibria for extensive form games”
Richard McKelvey and Thomas Palfrey · 1998
Earlier work this paper cites.
“Mathematical Control Theory: Deterministic Finite Dimensional Systems” 6
Eduardo. Sontag · 1998
Earlier work this paper cites.
“A Sparse Sampling Algorithm for Near-Optimal Planning in Large Markov Decision Processes”
M. Kearns, Y. Mansour and A. Ng · 1999
Earlier work this paper cites.
“Policy invariance under reward transformations: Theory and application to reward shaping”
A. Ng, D. Harada and S. Russell · 1999
Earlier work this paper cites.
“Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects”
R Rao and D Ballard · 1999
Earlier work this paper cites.
“Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning”
Richard Sutton, Doina Precup and Satinder Singh · 1999
Earlier work this paper cites.
“Policy Gradient Methods for Reinforcement Learning with Function Approximation”
R. Sutton, D. McAllester, S. Singh and Y. Mansour · 1999
Earlier work this paper cites.
“Stochastic dynamic programming with factored representations”
Craig Boutilier, Richard Dearden and Moisés Goldszmidt · 2000
Earlier work this paper cites.
“Hierarchical reinforcement learning with the MAXQ value function decomposition”
T Dietterich · 2000
Earlier work this paper cites.
“A simple adaptive procedure leading to correlated equilibrium”
S Hart and A Mas-Colell · 2000
Earlier work this paper cites.
“Observable operator models for discrete stochastic time series”
H Jaeger · 2000
Earlier work this paper cites.
“A Survey of POMDP Solution Techniques”, 2000
Kevin Murphy · 2000
Earlier work this paper cites.
“Algorithms for inverse reinforcement learning”
A. Ng and S. Russell · 2000
Earlier work this paper cites.
“Eligibility Traces for Off-Policy Policy Evaluation”
Doina Precup, Richard Sutton and Satinder Singh · 2000
Earlier work this paper cites.
“Convergence Results for Single-Step On-Policy Reinforcement-Learning Algorithms”
Satinder Singh, Tommi Jaakkola, Michael Littman and Csaba Szepesv\’ari · 2000
Earlier work this paper cites.
“Nash Convergence of Gradient Dynamics in General-Sum Games”
Satinder Singh, Michael Kearns and Yishay Mansour · 2000
Earlier work this paper cites.
“A Bayesian Framework for Reinforcement Learning”
M. Strens · 2000
Earlier work this paper cites.
“Bayesian Ensemble Learning for Nonlinear Factor Analysis”, 2000
H. Valpola · 2000
Earlier work this paper cites.
“A Natural Policy Gradient”
Sham Kakade · 2001
Earlier work this paper cites.
“Predictive Representations of State”
Michael Littman and Richard Sutton · 2001
Earlier work this paper cites.
“Automatic discovery of subgoals in reinforcement learning using di- verse density”
Amy McGovern and Andrew. Barto · 2001
Earlier work this paper cites.
“Finite-time Analysis of the Multiarmed Bandit Problem”
Peter Auer, Nicol\‘o Cesa-Bianchi and Paul Fischer · 2002
Earlier work this paper cites.
“Multiagent learning using a variable learning rate”
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
“Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes”, 2002
M. Duff · 2002
Earlier work this paper cites.
“Approximately Optimal Approximate Reinforcement Learning”
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
“Near-Optimal Reinforcement Learning in Polynomial Time”
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
“R-max – A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning”
R Tennenholtz · 2002
Earlier work this paper cites.
“Planning by Probabilistic Inference”
Hagai Attias · 2003
Earlier work this paper cites.
“Learning and inference in the brain”
Karl Friston · 2003
Earlier work this paper cites.
“Equivalence notions and model minimization in Markov decision processes”
Robert Givan, Thomas Dean and Matthew Greig · 2003
Earlier work this paper cites.
“Automatic Curriculum Learning for deep RL: A short survey”
Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann and Pierre-Yves Oudeyer · 2003
Earlier work this paper cites.
“Potential-Based Shaping and Q-Value Initialization are Equivalent”
E Wiewiora · 2003
Earlier work this paper cites.
“Information theory and statistics: A tutorial”
Imre Csisz\’ar and Paul Shields · 2004
Earlier work this paper cites.
“Metrics for finite Markov decision processes”
Norman Ferns, Prakash Panangaden and Doina Precup · 2004
Earlier work this paper cites.
“Dynamic programming for partially observable stochastic games”
Eric Hansen, Daniel Bernstein and Shlomo Zilberstein · 2004
Earlier work this paper cites.
“A Tutorial on MM Algorithms”
D.. Hunter and K. Lange · 2004
Earlier work this paper cites.
“The Cross-Entropy Method: A Unified Approach to Combinatorial Optimization, Monte-Carlo Simulation, and Machine Learning”
R. Rubinstein and D. Kroese · 2004
Earlier work this paper cites.
“The reward hypothesis”, 2004
Richard Sutton · 2004
Earlier work this paper cites.
“Approximate exploitability: Learning a best response in large games”
Finbarr Timbers, Nolan Bard, Edward Lockhart, Marc Lanctot, Martin Schmid, Neil Burch, Julian Schrittwieser, Thomas Hubert and Michael Bowling · 2004
Earlier work this paper cites.
“A Tutorial on the Cross-Entropy Method”
Pieter-Tjerk de Boer, Dirk Kroese, Shie Mannor and Reuven Rubinstein · 2005
Earlier work this paper cites.
“Tree-based batch mode reinforcement learning”
D Ernst, P Geurts and L Wehenkel · 2005
Earlier work this paper cites.
“Hierarchical Predictive Coding Models in a Deep-Learning Framework”, 2020
Matin Hosseini and Anthony Maida · 2005
Earlier work this paper cites.
“Universal Artificial Intelligence: Sequential Decisions Based On Algorithmic Probability”
Marcus Hutter · 2005
Earlier work this paper cites.
“Empowerment: A universal agent-centric measure of control”
A Klyubin, D Polani and C Nehaniv · 2005
Earlier work this paper cites.
“Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems”, 2020
Sergey Levine, Aviral Kumar, George Tucker and Justin Fu · 2005
Earlier work this paper cites.
“Neural fitted Q iteration – first experiences with a data efficient neural reinforcement learning method”
Martin Riedmiller · 2005
Earlier work this paper cites.
“Planning to explore via self-supervised world models”
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner and Deepak Pathak · 2005
Earlier work this paper cites.
“A Generalized Iterative LQG Method for Locally-optimal Feedback Control of Constrained Nonlinear Stochastic Systems”
Emanuel Todorov and Weiwei Li · 2005
Earlier work this paper cites.
“Mirror descent policy optimization”
Manan Tomar, Lior Shani, Yonathan Efroni and Mohammad Ghavamzadeh · 2005
Earlier work this paper cites.
“Bootstrap your own latent: A new approach to self-supervised Learning”
Jean-Bastien Grill et al · 2006
Earlier work this paper cites.
“On divergences and informations in statistics and information theory”
Friedrich Liese and Igor Vajda · 2006
Earlier work this paper cites.
“Towards a Unified Theory of State Abstraction for MDPs”, 2006
Lihong Li, Thomas Walsh and Michael Littman · 2006
Earlier work this paper cites.
“On the Relationship Between Active Inference and Control as Inference”
Beren Millidge, Alexander Tschantz, Anil Seth and Christopher Buckley · 2006
Earlier work this paper cites.
“Model-based Reinforcement Learning: A Survey”
Thomas Moerland, Joost Broekens, Aske Plaat and Catholijn Jonker · 2006
Earlier work this paper cites.
“AWAC: Accelerating Online Reinforcement Learning with Offline Datasets”
Ashvin Nair, Abhishek Gupta, Murtaza Dalal and Sergey Levine · 2006
Earlier work this paper cites.
“Agent modelling under partial observability for deep reinforcement learning”
Georgios Papoudakis, Filippos Christianos and Stefano Albrecht · 2006
Earlier work this paper cites.
“The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis”
James Smith and Robert Winkler · 2006
Earlier work this paper cites.
“Probabilistic inference for solving discrete and continuous state Markov Decision Processes”
M. Toussaint and A. Storkey · 2006
Earlier work this paper cites.
“Learning and planning in average-reward Markov decision processes”
Yi Wan, Abhishek Naik and Richard Sutton · 2006
Earlier work this paper cites.
“An overview of bilevel optimization”
Benoît Colson, Patrice Marcotte and Gilles Savard · 2007
Earlier work this paper cites.
“Fast Direct Multiple Shooting Algorithms for Optimal Robot Control”
Moritz Diehl, Hans Bock, Holger Diedam and Pierre-Brice Wieber · 2007
Earlier work this paper cites.
“Proto-value functions: A Laplacian framework for learning representation and control in Markov decision processes”
S Mahadevan and M Maggioni · 2007
Earlier work this paper cites.
“Reinforcement Learning by Reward-Weighted Regression for Operational Space Control”
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
“Bayes-Adaptive POMDPs”
Stephane Ross, Brahim Chaib-draa and Joelle Pineau · 2007
Earlier work this paper cites.
“Regret minimization in games with incomplete information”
Martin Zinkevich, Michael Johanson, Michael Bowling and Carmelo Piccione · 2007
Earlier work this paper cites.
“Near-optimal Regret Bounds for Reinforcement Learning”
Peter Auer, Thomas Jaksch and Ronald Ortner · 2008
Earlier work this paper cites.
“Essentials of game theory: A concise, multidisciplinary introduction”, Synthesis lectures on artificial intelligence and machine learning
Kevin Leyton-Brown and Yoav Shoham · 2008
Earlier work this paper cites.
“Formulas for Discrete Time LQR, LQG, LEQG and Minimax LQG Optimal Control Problems”
A I.. Petersen · 2008
Earlier work this paper cites.
“Skill characterization based on betweenness”
Ozgur Simsek and Andrew. Barto · 2008
Earlier work this paper cites.
“Multiagent systems: Algorithmic, game-theoretic, and logical foundations”
Yoav Shoham and Kevin Leyton-Brown · 2008
Earlier work this paper cites.
“A convergent O(n) algorithm for off-policy temporal-difference learning with linear function approximation”
Richard Sutton, Csaba Szepesvári and Hamid Maei · 2008
Earlier work this paper cites.
“Dyna-style planning with linear function approximation and prioritized sweeping”
Richard Sutton, Csaba Szepesvari, Alborz Geramifard and Michael Bowling · 2008
Earlier work this paper cites.
“Maximum Entropy Inverse Reinforcement Learning”
Brian. Ziebart, Andrew. Maas, J. Bagnell and Anind. Dey · 2008
Earlier work this paper cites.
Karl Cobbe, Jacob Hilton, Oleg Klimov and John Schulman · 2009
Earlier work this paper cites.
“The free-energy principle: a rough guide to the brain?”
Karl Friston · 2009
Earlier work this paper cites.
“Deep active inference for partially observable MDPs”
Otto van Himst and Pablo Lanillos · 2009
Earlier work this paper cites.
“Revisiting design choices in proximal Policy Optimization”
Chloe Ching-Yun Hsu, Celestine Mendler-Dünner and Moritz Hardt · 2009
Earlier work this paper cites.
“Skill Discovery in Continuous Reinforcement Learning Domains using Skill Chaining”
George Konidaris and Andrew Barto · 2009
Earlier work this paper cites.
“Monte Carlo Sampling for Regret Minimization in Extensive Games”
Marc Lanctot, Kevin Waugh, Martin Zinkevich and Michael Bowling · 2009
Earlier work this paper cites.
“Convergent Temporal-Difference Learning with Arbitrary Smooth Function Approximation”
Hamid Maei, Csaba Szepesvári, Shalabh Bhatnagar, Doina Precup, David Silver and Richard Sutton · 2009
Earlier work this paper cites.
“Robot Rrajectory Optimization using Approximate Inference”
Marc Toussaint · 2009
Earlier work this paper cites.
“Best Arm Identification in Multi-Armed Bandits”
Jean-Yves Audibert, S\’ebastien Bubeck and R\’emi Munos · 2010
Earlier work this paper cites.
“Web-Scale Bayesian Click-Through Rate Prediction for Sponsored Search Advertising in Microsoft’s Bing Search Engine”
T. Graepel, J. Quinonero-Candela, T. Borchert and R. Herbrich · 2010
Earlier work this paper cites.
“Double Q-learning”
Hado van Hasselt · 2010
Earlier work this paper cites.
“Approximate Riemannian Conjugate Gradient Learning for Fixed-Form Variational Bayes”
Antti Honkela, Tapani Raiko, Mikael Kuusela, Matti Tornio and Juha Karhunen · 2010
Earlier work this paper cites.
“Near-optimal regret bounds for reinforcement learning”
Thomas Jaksch, Ronald Ortner and Peter Auer · 2010
Earlier work this paper cites.
“A contextual-bandit approach to personalized news article recommendation”
L. Li, W. Chu, J. Langford and R.. Schapire · 2010
Earlier work this paper cites.
“Deep auto-encoder neural networks in reinforcement learning”
Sascha Lange and Martin Riedmiller · 2010
Earlier work this paper cites.
“Deep learning via Hessian-free optimization”
J Martens · 2010
Earlier work this paper cites.
“Estimating Divergence Functionals and the Likelihood Ratio by Convex Risk Minimization”
X Nguyen, M Wainwright and M Jordan · 2010
Earlier work this paper cites.
“A Survey of Numerical Methods for Optimal Control”
Anil Rao · 2010
Earlier work this paper cites.
“Formal Theory of Creativity, Fun, and Intrinsic Motivation”
Jurgen Schmidhuber · 2010
Earlier work this paper cites.
“A modern Bayesian look at the multi-armed bandit”
S. Scott · 2010
Earlier work this paper cites.
“Monte-Carlo Planning in Large POMDPs”
David Silver and Joel Veness · 2010
Earlier work this paper cites.
“Algorithms for Reinforcement Learning”
Csaba Szepesvari · 2010
Earlier work this paper cites.
“Active Inference or Control as Inference? A Unifying View”
Joe Watson, Abraham Imohiosen and Jan Peters · 2010
Earlier work this paper cites.
“Modeling Interaction via the Principle of Maximum Causal Entropy”
Brian Ziebart, J Andrew and Anind Dey · 2010
Earlier work this paper cites.
“Pure Exploration in Finitely-armed and Continuous-armed Bandits”
S\’ebastien Bubeck, R\’emi Munos and Gilles Stoltz · 2011
Earlier work this paper cites.
“On planning, prediction and knowledge transfer in Fully and Partially Observable Markov Decision Processes”, 2011
Pablo Castro · 2011
Earlier work this paper cites.
“Exploring simple Siamese representation learning”
Xinlei Chen and Kaiming He · 2011
Earlier work this paper cites.
“An empirical evaluation of Thompson sampling”
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
“PILCO: A Model-Based and Data-Efficient Approach to Policy Search”
Marc Deisenroth and Carl Rasmussen · 2011
Earlier work this paper cites.
“On the role of planning in model-based deep reinforcement learning”
Jessica Hamrick, Abram Friesen, Feryal Behbahani, Arthur Guez, Fabio Viola, Sims Witherspoon, Thomas Anthony, Lars Buesing, Petar Veličković and Théophane Weber · 2011
Earlier work this paper cites.
“Bayesian active learning for classification and preference learning”
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani and Máté Lengyel · 2011
Earlier work this paper cites.
“Reinforcement learning in feedback control: Challenges and benchmarks from technical process control”
Roland Hafner and Martin Riedmiller · 2011
Earlier work this paper cites.
“Hierarchical task and motion planning in the now”
L Kaelbling and T Lozano-P\’erez · 2011
Earlier work this paper cites.
“The world of independent learners is not markovian”
Guillaume Laurent, Laëtitia Matignon and N Le-Piat · 2011
Earlier work this paper cites.
“A reduction of imitation learning and structured prediction to no-regret online learning”
Stephane Ross, Geoffrey Gordon and J Bagnell · 2011
Earlier work this paper cites.
“Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction”
R Sutton, Joseph Modayil, M Delp, T Degris, P Pilarski, Adam White and Doina Precup · 2011
Earlier work this paper cites.
“Learning to make predictions in partially observable environments without a generative model”
Erik Talvitie and Satinder Singh · 2011
Earlier work this paper cites.
“Is independent learning all you need in the StarCraft multi-agent challenge?”
Christian de Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviychuk, Philip H Torr, Mingfei Sun and Shimon Whiteson · 2011
Earlier work this paper cites.
“An overview of multi-agent reinforcement learning from game theoretical perspective”
Yaodong Yang and Jun Wang · 2011
Earlier work this paper cites.
“A Survey of Monte Carlo Tree Search Methods”
C.. Browne, E. Powley, D. Whitehouse, S.. Lucas, P.. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis and S. Colton · 2012
Earlier work this paper cites.
“Planning as inference”
Matthew Botvinick and Marc Toussaint · 2012
Earlier work this paper cites.
“Emergent complexity and zero-shot transfer via Unsupervised Environment Design”
Michael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre Bayen, Stuart Russell, Andrew Critch and Sergey Levine · 2012
Earlier work this paper cites.
Thomas Degris, Martha White and Richard Sutton · 2012
Earlier work this paper cites.
“Optimal control as a graphical model inference problem”
Hilbert Kappen, Vicen G\’omez and Manfred Opper · 2012
Earlier work this paper cites.
“Batch reinforcement learning”
Sascha Lange, Thomas Gabel and Martin Riedmiller · 2012
Earlier work this paper cites.
“Agnostic system identification for model-based reinforcement learning”
Stephane Ross and J Bagnell · 2012
Earlier work this paper cites.
“The origins of inquiry: inductive inference and exploration in early childhood”
Laura Schulz · 2012
Earlier work this paper cites.
“The Arcade Learning Environment: An Evaluation Platform for General Agents”
M.. Bellemare, Y. Naddaf, J. Veness and M. Bowling · 2013
Earlier work this paper cites.
“Model predictive control”
E.. Camacho and C.. Alba · 2013
Earlier work this paper cites.
“Gaussian Processes for Data-Efficient Learning in Robotics and Control”
Marc Deisenroth, Dieter Fox and Carl Rasmussen · 2013
Earlier work this paper cites.
“The Sample-Complexity of General Reinforcement Learning”
Tor Lattimore, Marcus Hutter and Peter Sunehag · 2013
Earlier work this paper cites.
“Ad click prediction: a view from the trenches”
H McMahan, Gary Holt, D Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov and Daniel Golovin · 2013
Earlier work this paper cites.
“(More) Efficient Reinforcement Learning via Posterior Sampling”
Ian Osband, Daniel Russo and Benjamin Van · 2013
Earlier work this paper cites.
“LASER: a scalable response prediction platform for online advertising”
Deepak Agarwal, Bo Long, Jonathan Traupman, Doris Xin and Liang Zhang · 2014
Earlier work this paper cites.
“Is it better to select or to receive? Learning via active and passive hypothesis testing”
Douglas Markant and Todd Gureckis · 2014
Earlier work this paper cites.
“Prediction driven behavior: Learning predictions that drive fixed responses”
Joseph Modayil and R Sutton · 2014
Earlier work this paper cites.
“From Bandits to Monte-Carlo Tree Search: The Optimistic Principle Applied to Optimization and Planning”
R\’emi Munos · 2014
Earlier work this paper cites.
“Multi-timescale nexting in a reinforcement learning robot”
Joseph Modayil, Adam White and Richard Sutton · 2014
Earlier work this paper cites.
“Best Response Bayesian Reinforcement Learning for Multiagent Systems with State Uncertainty”
Frans Oliehoek and Christopher Amato · 2014
Earlier work this paper cites.
“Proximal algorithms”
Neal Parikh and Stephen Boyd · 2014
Earlier work this paper cites.
“Learning to Optimize via Posterior Sampling”
Daniel Russo and Benjamin Roy · 2014
Earlier work this paper cites.
“Deterministic Policy Gradient Algorithms”
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra and Martin Riedmiller · 2014
Earlier work this paper cites.
“Bandits, Global Optimization, Active Learning, and Bayesian RL – understanding the common ground”, Autonomous Learning Summer School, 2014
Marc Toussaint · 2014
Earlier work this paper cites.
“Evolutionary dynamics of multi-agent learning: A survey”
Daan Bloembergen, Karl Tuyls, Daniel Hennes and Michael Kaisers · 2015
Earlier work this paper cites.
Springer-Verlag New York, 2015
“Handbook of Simulation Optimization” · 2015
Earlier work this paper cites.
“Bayesian Reinforcement Learning: : A Survey”
Mohammed Ghavamzadeh, Shie Mannor, Joelle Pineau and Aviv Tamar · 2015
Earlier work this paper cites.
“Contextual Markov decision processes”
Assaf Hallak, Dotan Di and Shie Mannor · 2015
Earlier work this paper cites.
“Learning Continuous Control Policies by Stochastic Value Gradients”
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez and Yuval Tassa · 2015
Earlier work this paper cites.
“Fictitious Self-Play in Extensive-Form Games”
Johannes Heinrich, Marc Lanctot and David Silver · 2015
Earlier work this paper cites.
“Risk and regret of hierarchical Bayesian learners”
Jonathan Huggins and Joshua Tenenbaum · 2015
Earlier work this paper cites.
“The Dependence of Effective Planning Horizon on Model Accuracy”
Nan Jiang, Alex Kulesza, Satinder Singh and Richard Lewis · 2015
Earlier work this paper cites.
“Optimizing Neural Networks with Kronecker-factored Approximate Curvature”
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih et al · 2015
Earlier work this paper cites.
“A new optimal stepsize for approximate dynamic programming”
Ilya Ryzhov, Peter Frazier and Warren Powell · 2015
Earlier work this paper cites.
“Universal Value Function Approximators”
Tom Schaul, Daniel Horgan, Karol Gregor and David Silver · 2015
Earlier work this paper cites.
“Trust Region Policy Optimization”
John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan and Pieter Abbeel · 2015
Earlier work this paper cites.
“A review of predictive coding algorithms”
M Spratling · 2015
Earlier work this paper cites.
“Introduction to RL with function approximation”, NIPS Tutorial, 2015
Richard Sutton · 2015
Earlier work this paper cites.
“Multi-armed Bandit Models for the Optimal Design of Clinical Trials: Benefits and Challenges”
Sof\’a Villar, Jack Bowden and James Wason · 2015
Earlier work this paper cites.
“Solving games with functional regret estimation”
Kevin Waugh, Dustin Morrill, James Bagnell and Michael Bowling · 2015
Earlier work this paper cites.
“Developing a predictive approach to knowledge”, 2015
M White Adam · 2015
Earlier work this paper cites.
“Unifying Count-Based Exploration and Intrinsic Motivation”
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton and Remi Munos · 2016
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey Hinton · 2016
Earlier work this paper cites.
“Concentration Inequalities: A Nonasymptotic Theory of Independence”
Stephane Boucheron, Gabor Lugosi and Pascal Massart · 2016
Earlier work this paper cites.
“Superintelligence: Paths, Dangers, Strategies”
Nick Bostrom · 2016
Earlier work this paper cites.
“Probabilistic inference for determining options in reinforcement learning”
Christian Daniel, Herke van Hoof, Jan Peters and Gerhard Neumann · 2016
Earlier work this paper cites.
“Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization”
Chelsea Finn, Sergey Levine and Pieter Abbeel · 2016
Earlier work this paper cites.
“Q(lambda) with Off-Policy Corrections”
A. Harutyunyan, M.. Bellemare, T. Stepleton and R. Munos · 2016
Earlier work this paper cites.
“Learning values across many orders of magnitude”
Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih and David Silver · 2016
Earlier work this paper cites.
“Generative Adversarial Imitation Learning”
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
“Deep reinforcement learning from self-play in imperfect-information games”
Johannes Heinrich and David Silver · 2016
Earlier work this paper cites.
“Categorical Reparameterization with Gumbel-Softmax”, 2016
Eric Jang, Shixiang Gu and Ben Poole · 2016
Earlier work this paper cites.
“Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation”
Tejas Kulkarni, Karthik Narasimhan, Ardavan Saeedi and Josh Tenenbaum · 2016
Earlier work this paper cites.
“Continuous control with deep reinforcement learning”
Timothy Lillicrap, Jonathan Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver and Daan Wierstra · 2016
Earlier work this paper cites.
“Second-order optimization for neural networks”, 2016
James Martens · 2016
Earlier work this paper cites.
“Asynchronous Methods for Deep Reinforcement Learning”
Volodymyr Mnih, Adri\‘a\‘enech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
“Safe and Efficient Off-Policy Reinforcement Learning”
R\’emi Munos, Tom Stepleton, Anna Harutyunyan and Marc. Bellemare · 2016
Earlier work this paper cites.
“A Concise Introduction to Decentralized POMDPs”, SpringerBriefs in Intelligent Systems
Frans Oliehoek and Christopher Amato · 2016
Earlier work this paper cites.
“Combining policy gradient and Q-learning”
Brendan O’Donoghue, Remi Munos, Koray Kavukcuoglu and Volodymyr Mnih · 2016
Earlier work this paper cites.
“Deep Exploration via Bootstrapped DQN”
Ian Osband, Charles Blundell, Alexander Pritzel and Benjamin Van · 2016
Earlier work this paper cites.
“Prioritized Experience Replay”
Tom Schaul, John Quan, Ioannis Antonoglou and David Silver · 2016
Earlier work this paper cites.
“High-Dimensional Continuous Control Using Generalized Advantage Estimation”
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan and Pieter Abbeel · 2016
Earlier work this paper cites.
“True Online Temporal-Difference Learning”
Harm van Seijen, A Rupam, Patrick Pilarski, Marlos Machado and Richard Sutton · 2016
Earlier work this paper cites.
“Mastering the game of Go with deep neural networks and tree search”
David Silver et al · 2016
Earlier work this paper cites.
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine and Pieter Abbeel · 2016
Earlier work this paper cites.
“Dueling Network Architectures for Deep Reinforcement Learning”
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot and Nando de Freitas · 2016
Earlier work this paper cites.
“Constrained Policy Optimization”
Joshua Achiam, David Held, Aviv Tamar and Pieter Abbeel · 2017
Earlier work this paper cites.
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel and Wojciech Zaremba · 2017
Earlier work this paper cites.
“A Brief Survey of Deep Reinforcement Learning”
Kai Arulkumaran, Marc Deisenroth, Miles Brundage and Anil Bharath · 2017
Earlier work this paper cites.
“Meta-reasoning: Monitoring and control of thinking and reasoning”
Rakefet Ackerman and Valerie Thompson · 2017
Earlier work this paper cites.
“Successor Features for Transfer in Reinforcement Learning”
Andre Barreto, Will Dabney, Remi Munos, Jonathan Hunt, Tom Schaul, Hado van Hasselt and David Silver · 2017
Earlier work this paper cites.
“A Distributional Perspective on Reinforcement Learning”
Marc Bellemare, Will Dabney and R\’emi Munos · 2017
Earlier work this paper cites.
“The Option-Critic Architecture”
Pierre-Luc Bacon, Jean Harb and Doina Precup · 2017
Earlier work this paper cites.
“Superhuman AI for heads-up no-limit poker: Libratus beats top professionals”
Noam Brown and Tuomas Sandholm · 2017
Earlier work this paper cites.
“The free energy principle for action and perception: A mathematical review”
Christopher Buckley, Chang Kim, Simon McGregor and Anil Seth · 2017
Earlier work this paper cites.
“Distributional reinforcement learning with quantile regression”
Will Dabney, Mark Rowland, Marc Bellemare and Rémi Munos · 2017
Earlier work this paper cites.
“Variational intrinsic control”
Karol Gregor, Danilo Rezende and Daan Wierstra · 2017
Earlier work this paper cites.
“Linear Optimal Control on Factor Graphs — A Message Passing Perspective”
Christian Hoffmann and Philipp Rostalski · 2017
Earlier work this paper cites.
“Reinforcement Learning with Unsupervised Auxiliary Tasks”
Max Jaderberg, Volodymyr Mnih, Wojciech Czarnecki, Tom Schaul, Joel Leibo, David Silver and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
“A unified game-theoretic approach to multiagent reinforcement learning”
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver and Thore Graepel · 2017
Earlier work this paper cites.
“A unified approach to interpreting model predictions”
Scott Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
“Discrete Sequential Prediction of Continuous Actions for Deep RL”, 2017
Luke Metz, Julian Ibarz, Navdeep Jaitly and James Davidson · 2017
Earlier work this paper cites.
“The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”
Chris Maddison, Andriy Mnih and Yee Teh · 2017
Earlier work this paper cites.
“DeepStack: Expert-level artificial intelligence in heads-up no-limit poker”
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson and Michael Bowling · 2017
Earlier work this paper cites.
“Value Prediction Network”
Junhyuk Oh, Satinder Singh and Honglak Lee · 2017
Earlier work this paper cites.
“Why is posterior sampling better than optimism for reinforcement learning?”
Ian Osband and Benjamin Van · 2017
Earlier work this paper cites.
“Curiosity-driven Exploration by Self-supervised Prediction”
Deepak Pathak, Pulkit Agrawal, Alexei Efros and Trevor Darrell · 2017
Earlier work this paper cites.
“Towards generalization and simplicity in continuous control”
Aravind Rajeswaran, Kendall Lowrey, Emanuel Todorov and Sham Kakade · 2017
Earlier work this paper cites.
“Evolution Strategies as a Scalable Alternative to Reinforcement Learning”, 2017
Tim Salimans, Jonathan Ho, Xi Chen and Ilya Sutskever · 2017
Earlier work this paper cites.
“Proximal Policy Optimization Algorithms”, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov · 2017
Earlier work this paper cites.
“Mastering the game of Go without human knowledge”
David Silver et al · 2017
Earlier work this paper cites.
“The predictron: end-to-end learning and planning”
David Silver et al · 2017
Earlier work this paper cites.
“Value-decomposition networks for cooperative multi-agent learning”
Peter Sunehag et al · 2017
Earlier work this paper cites.
“Human Learning in Atari”
Pedro Tsividis, Thomas Pouncy, Jaqueline Xu, Joshua Tenenbaum and Samuel Gershman · 2017
Earlier work this paper cites.
“FeUdal Networks for Hierarchical Reinforcement Learning”
Alexander Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
“Model Predictive Path Integral Control: From Theory to Parallel Computation”
Grady Williams, Andrew Aldrich and Evangelos Theodorou · 2017
Earlier work this paper cites.
“Imagination-Augmented Agents for Deep Reinforcement Learning”
Th\’eophane Weber et al · 2017
Earlier work this paper cites.
“Information theoretic MPC for model-based reinforcement learning”
Grady Williams, Nolan Wagener, Brian Goldfain, Paul Drews, James Rehg, Byron Boots and Evangelos Theodorou · 2017
Earlier work this paper cites.
Yuhuai Wu, Elman Mansimov, Shun Liao, Roger Grosse and Jimmy Ba · 2017
Earlier work this paper cites.
“Reinforcement learning for learning rate control”
Chang Xu, Tao Qin, Gang Wang and Tie-Yan Liu · 2017
Earlier work this paper cites.
“On convergence of some gradient-based temporal-differences algorithms for off-policy learning”
Huizhen Yu · 2017
Earlier work this paper cites.
“Maximum a Posteriori Policy Optimisation”
Abbas Abdolmaleki, Jost Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess and Martin Riedmiller · 2018
Earlier work this paper cites.
“Variational Option Discovery Algorithms”
Joshua Achiam, Harrison Edwards, Dario Amodei and Pieter Abbeel · 2018
Earlier work this paper cites.
“Mitigating planner overfitting in model-based reinforcement learning”
Dilip Arumugam, David Abel, Kavosh Asadi, Nakul Gopalan, Christopher Grimm, Jun Lee, Lucas Lehnert and Michael Littman · 2018
Earlier work this paper cites.
“Autonomous agents modelling other agents: A comprehensive survey and open problems”
Stefano Albrecht and Peter Stone · 2018
Earlier work this paper cites.
“Distributed Distributional Deterministic Policy Gradients”
Gabriel Barth-Maron, Matthew Hoffman, David Budden, Will Dabney, Dan Horgan, T Dhruva, Alistair Muldal, Nicolas Heess and Timothy Lillicrap · 2018
Earlier work this paper cites.
“Exploration by random network distillation”
Yuri Burda, Harrison Edwards, A Storkey and Oleg Klimov · 2018
Earlier work this paper cites.
“Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models”
Kurtland Chua, Roberto Calandra, Rowan McAllister and Sergey Levine · 2018
Cited alongside, same era.
“Knowledge Representation for Reinforcement Learning using General Value Functions”, 2018
Gheorghe Comanici, Doina Precup, Andre Barreto, Daniel Toyama, Eser Aygün, Philippe Hamel, Sasha Vezhnevets, Shaobo Hou and Shibl Mourad · 2018
Cited alongside, same era.
John Co-Reyes, Yuxuan Liu, Abhishek Gupta, Benjamin Eysenbach, Pieter Abbeel and Sergey Levine · 2018
Cited alongside, same era.
“Implicit quantile networks for distributional reinforcement learning”
Will Dabney, Georg Ostrovski, David Silver and Rémi Munos · 2018
Cited alongside, same era.
“Optimality guarantees for particle belief approximation of POMDPs”
Michael Lim, Tyler Becker, Mykel Kochenderfer, Claire Tomlin and Zachary Sunberg · 2023
Later among the works it cites.
“TiZero: Mastering multi-agent football with curriculum learning and self-play”
Fanqi Lin, Shiyu Huang, Tim Pearce, Wenze Chen and Wei-Wei Tu · 2023
Later among the works it cites.
“Reinforcement Learning, Bit by Bit”
Xiuyuan Lu, Benjamin Van, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband and Zheng Wen · 2023
Later among the works it cites.
“Temporal Abstraction in Reinforcement Learning with the Successor Representation”
Marlos Machado, Andre Barreto, Doina Precup and Michael Bowling · 2023
Later among the works it cites.
“Off-Policy Proximal Policy Optimization”
Wenjia Meng, Qian Zheng, Gang Pan and Yilong Yin · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lasse Espeholt et al · 2018
Cited alongside, same era.
“TreeQN and ATreeC: Differentiable Tree-Structured Models for Deep Reinforcement Learning”
Gregory Farquhar, Tim Rocktäschel, Maximilian Igl and Shimon Whiteson · 2018
Cited alongside, same era.
“Addressing Function Approximation Error in Actor-Critic Methods”
Scott Fujimoto, Herke van Hoof and David Meger · 2018
Cited alongside, same era.
“An Introduction to Deep Reinforcement Learning”
Vincent Francois-Lavet, Peter Henderson, Riashat Islam, Marc Bellemare and Joelle Pineau · 2018
Cited alongside, same era.
“Learning Robust Rewards with Adverserial Inverse Reinforcement Learning”
Justin Fu, Katie Luo and Sergey Levine · 2018
Cited alongside, same era.
“Noisy Networks for Exploration”
Meire Fortunato et al · 2018
Cited alongside, same era.
“Deconstructing the human algorithms for exploration”
Samuel Gershman · 2018
Cited alongside, same era.
“Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor”
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel and Sergey Levine · 2018
Cited alongside, same era.
K.. Murphy · 2023
Later among the works it cites.
“Cal-QL: Calibrated offline RL pre-training for efficient online fine-tuning”
Mitsuhiko Nakamoto, Yuexiang Zhai, Anikait Singh, Max Mark, Yi Ma, Chelsea Finn, Aviral Kumar and Sergey Levine · 2023
Later among the works it cites.
“Approximate Thompson Sampling via Epistemic Neural Networks”
Ian Osband, Zheng Wen, Seyed Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu and Benjamin Van · 2023
Later among the works it cites.
“Epistemic Neural Networks”
Ian Osband, Zheng Wen, Seyed Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu and Benjamin Van · 2023
Later among the works it cites.
“Generative agents: Interactive simulacra of human behavior”
Joon Park, Joseph O’Brien, Carrie Cai, Meredith Morris, Percy Liang and Michael Bernstein · 2023
Later among the works it cites.
“Evolving Curricula with Regret-Based Environment Design”
Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette and Tim Rocktäschel · 2023
Later among the works it cites.
“Why think step-by-step? Reasoning emerges from the locality of experience”
Ben Prystawski, Michael Li and Noah Goodman · 2023
Later among the works it cites.
“Direct Preference Optimization: Your language model is secretly a reward model”
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher Manning and Chelsea Finn · 2023
Later among the works it cites.
“A simple framework for self-supervised learning of sample-efficient world models”
Jan Robine, Marc Höftmann and Stefan Harmeling · 2023
Later among the works it cites.
“Brain-inspired computational intelligence via predictive coding”
Tommaso Salvatori, Ankur Mali, Christopher Buckley, Thomas Lukasiewicz, Rajesh P Rao, Karl Friston and Alexander Ororbia · 2023
Later among the works it cites.
“A Generalist Dynamics Model for Control”
Ingmar Schubert, Jingwei Zhang, Jake Bruce, Sarah Bechtle, Emilio Parisotto, Martin Riedmiller, Jost Springenberg, Arunkumar Byravan, Leonard Hasenclever and Nicolas Heess · 2023
Later among the works it cites.
“Bigger, Better, Faster: Human-level Atari with human-level efficiency”
Max Schwarzer, Johan Obando-Ceron, Aaron Courville, Marc Bellemare, Rishabh Agarwal and Pablo Castro · 2023
Later among the works it cites.
“Reflexion: an autonomous agent with dynamic memory and self-reflection”
Noah Shinn, Beck Labash and Ashwin Gopinath · 2023
Later among the works it cites.
“Prediction-Oriented Bayesian Active Learning”
Freddie Smith, Andreas Kirsch, Sebastian Farquhar, Yarin Gal, Adam Foster and Tom Rainforth · 2023
Later among the works it cites.
“The update-equivalence framework for decision-time planning”
Samuel Sokota, Gabriele Farina, David Wu, Hengyuan Hu, Kevin Wang, J Kolter and Noam Brown · 2023
Later among the works it cites.
“RoboCLIP: One demonstration is enough to learn robot policies”
S Sontakke, Jesse Zhang, S’ebastien M Arnold, Karl Pertsch, Erdem Biyik, Dorsa Sadigh, Chelsea Finn and Laurent Itti · 2023
Later among the works it cites.
“Long Horizon Temperature Scaling”
Andy Shih, Dorsa Sadigh and Stefano Ermon · 2023
Later among the works it cites.
“Exploration via Epistemic Value Estimation”
Simon Schmitt, John Shawe-Taylor and Hado Van · 2023
Later among the works it cites.
“Reward-Respecting Subtasks for Model-Based Reinforcement Learning”
Richard Sutton, Marlos Machado, G Zacharias, David Szepesvari, Finbarr Timbers, Brian Tanner and Adam White · 2023
Later among the works it cites.
“Understanding Self-Predictive Learning for Reinforcement Learning”
Yunhao Tang et al · 2023
Later among the works it cites.
“Probabilistic Inference in Reinforcement Learning Done Right”
Jean Tarbouriech, Tor Lattimore and Brendan O’Donoghue · 2023
Later among the works it cites.
“Learning Representations for Pixel-based Control: What Matters and Why?”
Manan Tomar, Utkarsh Mishra, Amy Zhang and Matthew Taylor · 2023
Later among the works it cites.
“Efficient exploration in continuous-time model-based reinforcement learning”
Lenart Treven, Jonas Hübotter, Bhavya Sukhija, Florian Dörfler and Andreas Krause · 2023
Later among the works it cites.
“Does Zero-Shot Reinforcement Learning Exist?”
Ahmed Touati, Jérémy Rapin and Yann Ollivier · 2023
Later among the works it cites.
“Hybrid predictive coding: Inferring, fast and slow”
Alexander Tscshantz, Beren Millidge, Anil Seth and Christopher Buckley · 2023
Later among the works it cites.
“Adversarial Policies Beat Superhuman Go AIs”
Tony Wang et al · 2023
Later among the works it cites.
“Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning”
Zhendong Wang, Jonathan Hunt and Mingyuan Zhou · 2023
Later among the works it cites.
“Learning adaptive planning representations with natural language guidance”
Lionel Wong, Jiayuan Mao, Pratyusha Sharma, Zachary Siegel, Jiahai Feng, Noa Korneev, Joshua Tenenbaum and Jacob Andreas · 2023
Later among the works it cites.
“Dyna-PPO reinforcement learning with Gaussian process for the continuous action decision-making in autonomous driving”
Guanlin Wu, Wenqi Fang, Ji Wang, Pin Ge, Jiang Cao, Yang Ping and Peng Gou · 2023
Later among the works it cites.
“Dichotomy of control: Separating what you can control from what you cannot”
Mengjiao Yang, D Schuurmans, P Abbeel and Ofir Nachum · 2023
Later among the works it cites.
Taku Yamagata, Ahmed Khalil and Raul Santos-Rodriguez · 2023
Later among the works it cites.
“Successor-Predecessor Intrinsic Exploration”
Changmin Yu, N Burgess, M Sahani and S Gershman · 2023
Later among the works it cites.
“Leveraging Jumpy Models for Planning and Fast Learning in Robotic Domains”
Jingwei Zhang, Jost Springenberg, Arunkumar Byravan, Leonard Hasenclever, Abbas Abdolmaleki, Dushyant Rao, Nicolas Heess and Martin Riedmiller · 2023
Later among the works it cites.
“STORM: Efficient Stochastic Transformer based world models for reinforcement learning”
Weipu Zhang, Gang Wang, Jian Sun, Yetian Yuan and Gao Huang · 2023
Later among the works it cites.
“Click: Controllable text generation with sequence likelihood contrastive learning”
Chujie Zheng, Pei Ke, Zheng Zhang and Minlie Huang · 2023
Later among the works it cites.
“RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control”
Brianna Zitkovich et al · 2023
Later among the works it cites.
“Multi-Agent Reinforcement Learning: Foundations and Modern Approaches”
Stefano. Albrecht, Filippos Christianos and Lukas Sch\"afer · 2024
Closest in time.
“Back to basics: Revisiting REINFORCE style optimization for learning from Human Feedback in LLMs”
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Ahmet Üstün and Sara Hooker · 2024
Closest in time.
“Genie 2: A large-scale foundation world model”, 2024
Jack Parker-Holder et al · 2024
Closest in time.
“Diffusion for world modeling: Visual details matter in Atari”
Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce and François Fleuret · 2024
Closest in time.
“A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings”
Safa Alver and Doina Precup · 2024
Closest in time.
“Partial models for building adaptive model-based reinforcement learning agents”
Safa Alver, Ali Rahimi-Kalahroudi and Doina Precup · 2024
Closest in time.
“Bayesian Reinforcement Learning With Limited Cognitive Load”
Dilip Arumugam, Mark Ho, Noah Goodman and Benjamin Van · 2024
Closest in time.
“Satisficing exploration for deep reinforcement learning”
Dilip Arumugam, Saurabh Kumar, Ramki Gummadi and Benjamin Van · 2024
Closest in time.
“Revisiting Feature Prediction for Learning Visual Representations from Video”
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann Le, Mahmoud Assran and Nicolas Ballas · 2024
Closest in time.
Maria Bauza et al · 2024
Closest in time.
“xLSTM: Extended Long Short-Term Memory”
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter and Sepp Hochreiter · 2024
Closest in time.
Dimitri Bertsekas · 2024
Closest in time.
“How to Specify Reinforcement Learning Objectives”
W Bradley and James MacGlashan · 2024
Closest in time.
“Generative AI Handbook: A Roadmap for Learning Resources”, 2024
William Brown · 2024
Closest in time.
“Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods”
Yuji Cao, Huan Zhao, Yuheng Cheng, Ting Shu, Guolong Liu, Gaoqi Liang, Junhua Zhao and Yun Li · 2024
Closest in time.
“Predictive representations: building blocks of intelligence”
Wilka Carvalho, Momchil Tomov, William de Cothi, Caswell Barry and Samuel Gershman · 2024
Closest in time.
“Simple ingredients for offline reinforcement learning”
Edoardo Cetin, Andrea Tirinzoni, Matteo Pirotta, Alessandro Lazaric, Yann Ollivier and Ahmed Touati · 2024
Closest in time.
Fengdi Che, Chenjun Xiao, Jincheng Mei, Bo Dai, Ramki Gummadi, Oscar Ramirez, Christopher Harris, A Mahmood and Dale Schuurmans · 2024
Closest in time.
“Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions”
Jiayu Chen, Bhargav Ganguly, Yang Xu, Yongsheng Mei, Tian Lan and Vaneet Aggarwal · 2024
Closest in time.
“Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions”
Jiayu Chen, Bhargav Ganguly, Yang Xu, Yongsheng Mei, Tian Lan and Vaneet Aggarwal · 2024
Closest in time.
“Vision-language models provide promptable representations for reinforcement learning”
William Chen, Oier Mees, Aviral Kumar and Sergey Levine · 2024
Closest in time.
“Learning successor Features the simple way”
Raymond Chua, Arna Ghosh, Christos Kaplanis, Blake Richards and Doina Precup · 2024
Closest in time.
“Position: Social Choice should guide AI alignment in dealing with diverse human feedback”
Vincent Conitzer et al · 2024
Closest in time.
“Generating Code World Models with large language models guided by Monte Carlo Tree Search”
Nicola Dainese, Matteo Merler, Minttu Alakuijala and Pekka Marttinen · 2024
Closest in time.
“Griffin: Mixing gated linear recurrences with local attention for efficient language models”
Soham De et al · 2024
Closest in time.
“DeepSeek-V3 Technical Report”
DeepSeek-AI · 2024
Closest in time.
“A survey on in-context learning”
Qingxiu Dong et al · 2024
Closest in time.
“Denoised Predictive Imagination: An Information-theoretic approach for learning World Models”
Vedant Dave and Elmar Rueckert · 2024
Closest in time.
“From code to play: Benchmarking program search for games using large language models”
Manuel Eberhardinger, James Goodman, Alexander Dockhorn, Diego Perez-Liebana, Raluca Gaina, Duygu Cakmak, Setareh Maghsudi and Simon Lucas · 2024
Closest in time.
“TextGenSHAP: Scalable post-hoc explanations in text generation with long documents”
James Enouen, Hootan Nakhost, Sayna Ebrahimi, Sercan Arik, Yan Liu and Tomas Pfister · 2024
Closest in time.
“Stop regressing: Training value functions via classification for scalable deep RL”
Jesse Farebrother et al · 2024
Closest in time.
“CALE: Continuous Arcade Learning Environment”
Jesse Farebrother and Pablo Castro · 2024
Closest in time.
“One step diffusion via shortcut models”
Kevin Frans, Danijar Hafner, Sergey Levine and Pieter Abbeel · 2024
Closest in time.
“Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout Adaption”
Bernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow and Sebastian Trimpe · 2024
Closest in time.
“Unifying Model-Based and Model-Free Reinforcement Learning with Equivalent Policy Sets”
Benjamin Freed, Thomas Wei, Roberto Calandra, Jeff Schneider and Howie Choset · 2024
Closest in time.
“Simplifying deep temporal difference learning”
Matteo Gallici, Mattie Fellows, Benjamin Ellis, Bartomeu Pou, Ivan Masmitja, Jakob Foerster and Mario Martin · 2024
Closest in time.
“Learning and leveraging world models in visual representation learning”
Quentin Garrido, Mahmoud Assran, Nicolas Ballas, Adrien Bardes, Laurent Najman and Yann LeCun · 2024
Closest in time.
“Axioms for AI alignment from human feedback”
Luise Ge, Daniel Halpern, Evi Micha, Ariel Procaccia, Itai Shapira, Yevgeniy Vorobeychik and Junlin Wu · 2024
Closest in time.
“AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents”
Jake Grigsby, Linxi Fan and Yuke Zhu · 2024
Closest in time.
“The development of human causal learning and reasoning”
Mariel Goddu and Alison Gopnik · 2024
Closest in time.
“BoNBoN Alignment for large language models and the sweetness of best-of-n sampling”
Lin Gui, Cristina Gârbacea and Victor Veitch · 2024
Closest in time.
“Closing the gap between TD learning and supervised learning – A generalisation Point of View”
Raj Ghugare, Matthieu Geist, Glen Berseth and Benjamin Eysenbach · 2024
Closest in time.
“Learning Universal Predictors”
Jordi Grau-Moya et al · 2024
Closest in time.
“Goal-conditioned on-policy reinforcement learning”
Xudong Gong, Dawei Feng, Kele Xu, Bo Ding and Huaimin Wang · 2024
Closest in time.
“AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers”
Jake Grigsby, Justin Sasek, Samyak Parajuli, Daniel Adebi, Amy Zhang and Yuke Zhu · 2024
Closest in time.
“Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy”
Zhenyu Guan, Xiangyu Kong, Fangwei Zhong and Yizhou Wang · 2024
Closest in time.
“LLM Multi-Agent Systems: Challenges and Open Problems”
Shanshan Han, Qifan Zhang, Yuhang Yao, Weizhao Jin and Zhaozhuo Xu · 2024
Closest in time.
“An introduction to universal artificial intelligence”
M. Hutter, D. Quarel and E. Catt · 2024
Closest in time.
“TD-MPC2: Scalable, Robust World Models for Continuous Control”
Nicklas Hansen, Hao Su and Xiaolong Wang · 2024
Closest in time.
“A Survey on Large Language Model-Based Game Agents”, 2024
Sihao Hu, Tiansheng Huang, Fatih Ilhan, Selim Tekin, Gaowen Liu, Ramana Kompella and Ling Liu · 2024
Closest in time.
“The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization”
Shengyi Huang, Michael Noukhovitch, Arian Hosseini, Kashif Rasul, Weixun Wang and Lewis Tunstall · 2024
Closest in time.
“Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling”
Sili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang, Tao Luo, Hechang Chen, Lichao Sun and Bo Yang · 2024
Closest in time.
“Bayesian online natural gradient (BONG)”
Matt Jones, Peter Chang and Kevin Murphy · 2024
Closest in time.
“An Invitation to Deep Reinforcement Learning”
Bernhard Jaeger and Andreas Geiger · 2024
Closest in time.
“Position: Benchmarking is Limited in Reinforcement Learning Research”
Scott Jordan, Adam White, Bruno da Silva, Martha White and Philip Thomas · 2024
Closest in time.
“Motif: Intrinsic motivation from artificial intelligence feedback”
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang and Mikael Henaff · 2024
Closest in time.
“Latent Plan Transformer for trajectory abstraction: Planning as latent space inference”
Deqian Kong, Dehong Xu, Minglu Zhao, Bo Pang, Jianwen Xie, Andrew Lizarraga, Yuhao Huang, Sirui Xie and Ying Wu · 2024
Closest in time.
“The Need for a Big World Simulator: A Scientific Challenge for Continual Learning”
Saurabh Kumar, Hong Jeon, Alex Lewandowski and Benjamin Van · 2024
Closest in time.
“TÜLU 3: Pushing frontiers in open language model post-training”
Nathan Lambert et al · 2024
Closest in time.
Matthias Lehmann · 2024
Closest in time.
“Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft”
Hao Li, Xue Yang, Zhaokai Wang, Xizhou Zhu, Jie Zhou, Yu Qiao, Xiaogang Wang, Hongsheng Li, Lewei Lu and Jifeng Dai · 2024
Closest in time.
Jiangmeng Li, Zehua Zang, Qirui Ji, Chuxiong Sun, Wenwen Qiang, Junge Zhang, Changwen Zheng, Fuchun Sun and Hui Xiong · 2024
Closest in time.
“Chain of thought empowers transformers to solve inherently serial problems”
Zhiyuan Li, Hong Liu, Denny Zhou and Tengyu Ma · 2024
Closest in time.
“Locality Sensitive Sparse Encoding for Learning World Models Online”
Zichen Liu, Chao Du, Wee Lee and Min Lin · 2024
Closest in time.
“Rewarded Region Replay (R3) for policy learning with discrete action space”
Bangzheng Li, Ningshan Ma and Zifan Wang · 2024
Closest in time.
“Scalable nested optimization for deep learning”
Jonathan Lorraine · 2024
Closest in time.
“Eureka: Human-Level Reward Design via Coding Large Language Models”
Yecheng Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan and Anima Anandkumar · 2024
Closest in time.
“SPO: Sequential Monte Carlo Policy Optimisation”
Matthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth and Alexandre Laterre · 2024
Closest in time.
“Efficient world models with context-aware tokenization”
Vincent Micheli, Eloi Alonso and François Fleuret · 2024
Closest in time.
“Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain”
Marcus Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar, Gail Kaiser, Suman Jana and Baishakhi Ray · 2024
Closest in time.
“Sequential Monte Carlo for Inclusive KL Minimization in Amortized Variational Inference”
Declan McNamara, Jackson Loper and Jeffrey Regier · 2024
Closest in time.
“Reinforcement Learning: Foundations”, 2024
Shie Mannor, Yishay Mansour and Aviv Tamar · 2024
Closest in time.
“BetaZero: Belief-state planning for long-horizon POMDPs using learned approximations”
Robert Moss, Anthony Corso, Jef Caers and Mykel Kochenderfer · 2024
Closest in time.
“The Expressive Power of Transformers with Chain of Thought”
William Merrill and Ashish Sabharwal · 2024
Closest in time.
“Connecting Joint-Embedding Predictive Architecture with Contrastive self-supervised learning”
Shentong Mo and Shengbang Tong · 2024
Closest in time.
Abhishek Naik, Yi Wan, Manan Tomar and Richard Sutton · 2024
Closest in time.
“Bridging State and History Representations: Understanding Self-Predictive RL”
Tianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma, Clement Gehring, Aditya Mahajan and Pierre-Luc Bacon · 2024
Closest in time.
“DINOv2: Learning Robust Visual Features without Supervision”
Maxime Oquab et al · 2024
Closest in time.
“OGBench: Benchmarking Offline Goal-Conditioned RL”
Seohong Park, Kevin Frans, Benjamin Eysenbach and Sergey Levine · 2024
Closest in time.
“Is value learning really the main bottleneck in offline RL?”
Seohong Park, Kevin Frans, Sergey Levine and Aviral Kumar · 2024
Closest in time.
“Empirical design in reinforcement learning”
Andrew Patterson, Samuel Neumann, Martha White and Adam White · 2024
Closest in time.
“Doing experiments and revising rules with natural language and probabilistic reasoning”
Wasu Piriyakulkij, Cassidy Langenfeld, Tuan Le and Kevin Ellis · 2024
Closest in time.
“The RL/LLM taxonomy tree: Reviewing synergies between Reinforcement Learning and Large Language Models”
Moschoula Pternea, Prerna Singh, Abir Chakraborty, Yagna Oruganti, Mirco Milletari, Sayli Bapat and Kebei Jiang · 2024
Closest in time.
“Generalized Policy Improvement algorithms with theoretically supported sample reuse”
James Queeney, Ioannis Paschalidis and Christos Cassandras · 2024
Closest in time.
“D5RL: Diverse datasets for data-driven deep reinforcement learning”
Rafael Rafailov et al · 2024
Closest in time.
“Diffusion Policy Policy Optimization”
Allen Ren, Justin Lidard, Lars Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai and Max Simchowitz · 2024
Closest in time.
“Vision-language models are zero-shot reward models for reinforcement learning”
Juan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez and David Lindner · 2024
Closest in time.
“Mathematical discoveries from program search with large language models”
B. Romera-Paredes et al · 2024
Closest in time.
“A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding Networks”
Tommaso Salvatori, Yuhang Song, Yordan Yordanov, Beren Millidge, Lei Sha, Cornelius Emde, Zhenghua Xu, Rafal Bogacz and Thomas Lukasiewicz · 2024
Closest in time.
“The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning”
Moritz Schneider, Robert Krug, Narunas Vaskevicius, Luigi Palmieri and Joschka Boedecker · 2024
Closest in time.
“DeepSeekMath: Pushing the limits of mathematical reasoning in open language models”
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y Li, Y Wu and Daya Guo · 2024
Closest in time.
“Scaling LLM test-time compute optimally can be more effective than scaling model parameters”
Charlie Snell, Jaehoon Lee, Kelvin Xu and Aviral Kumar · 2024
Closest in time.
“Informing Reinforcement Learning Agents by Grounding Language to Markov Decision Processes”
Benjamin Spiegel, Ziyi Yang, William Jurayj, Ben Bachmann, Stefanie Tellex and George Konidaris · 2024
Closest in time.
“Cognitive Architectures for Language Agents”
Theodore Sumers, Shunyu Yao, Karthik Narasimhan and Thomas Griffiths · 2024
Closest in time.
“FactorSim: Generative Simulation via Factorized Representation”
Fan-Yun Sun, S Harini, Angela Yi, Yihan Zhou, Alex Zook, Jonathan Tremblay, Logan Cross, Jiajun Wu and Nick Haber · 2024
Closest in time.
“To Compress or Not to Compress- Self-Supervised Learning and Information Theory: A Review”
Ravid Shwartz-Ziv and Yann LeCun · 2024
Closest in time.
“Code Repair with LLMs gives an Exploration-Exploitation Tradeoff”
Hao Tang, Keya Hu, Jin Zhou, Si Zhong, Wei-Long Zheng, Xujie Si and Kevin Ellis · 2024
Closest in time.
“Intrinsic motivation in dynamical control systems”
Stas Tiomkin, Ilya Nemenman, Daniel Polani and Naftali Tishby · 2024
Closest in time.
“WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment”
Hao Tang, Darren Key and Kevin Ellis · 2024
Closest in time.
Manan Tomar, Philippe Hansen-Estruch, Philip Bachman, Alex Lamb, John Langford, Matthew Taylor and Sergey Levine · 2024
Closest in time.
“Understanding reinforcement learning-based fine-tuning of diffusion models: A tutorial and review”
Masatoshi Uehara, Yulai Zhao, Tommaso Biancalani and Sergey Levine · 2024
Closest in time.
“Beyond The Rainbow: High Performance Deep Reinforcement Learning On A Desktop PC”, 2024
Unknown · 2024
Closest in time.
“Code as reward: Empowering reinforcement learning with VLMs”
David Venuto, Sami Islam, Martin Klissarov, Doina Precup, Sherry Yang and Ankit Anand · 2024
Closest in time.
“Voyager: An Open-Ended Embodied Agent with Large Language Models”
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan and Anima Anandkumar · 2024
Closest in time.
“EfficientZero V2: Mastering discrete and continuous control with limited data”
Shengjie Wang, Shaohuai Liu, Weirui Ye, Jiacheng You and Yang Gao · 2024
Closest in time.
“Executable Code Actions Elicit Better LLM Agents”
Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng and Heng Ji · 2024
Closest in time.
“Model-based Policy Optimization under Approximate Bayesian Inference”
Chaoqi Wang, Yuxin Chen and Kevin Murphy · 2024
Closest in time.
“A unified view on solving objective mismatch in model-based Reinforcement Learning”
Ran Wei, Nathan Lambert, Anthony McDonald, Alfredo Garcia and Roberto Calandra · 2024
Closest in time.
“Implicit bias of AdamW: ℓ ∞ \ell_{\infty} norm constrained optimization”
Shuo Xie and Zhiyuan Li · 2024
Closest in time.
“Learning Interactive Real-World Simulators”
Sherry Yang, Yilun Du, Seyed Kamyar Ghasemipour, Jonathan Tompson, Leslie Kaelbling, Dale Schuurmans and Pieter Abbeel · 2024
Closest in time.
“Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking”
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber and Noah Goodman · 2024
Closest in time.
“Fine-tuning large vision-language models as decision-making agents via reinforcement learning”
Yuexiang Zhai et al · 2024
Closest in time.
“OMNI: Open-endedness via Models of human Notions of Interestingness”
Jenny Zhang, Joel Lehman, Kenneth Stanley and Jeff Clune · 2024
Closest in time.
“A survey on the memory mechanism of large language model based agents”
Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong and Ji-Rong Wen · 2024
Closest in time.
“Probabilistic inference in language models via twisted Sequential Monte Carlo”
Stephen Zhao, Rob Brekelmans, Alireza Makhzani and Roger Grosse · 2024
Closest in time.
“Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo”
Stephen Zhao, Rob Brekelmans, Alireza Makhzani and Roger Grosse · 2024
Closest in time.
“Online intrinsic rewards for decision making agents from large language model feedback”
Qinqing Zheng, Mikael Henaff, Amy Zhang, Aditya Grover and Brandon Amos · 2024
Closest in time.
“Toward optimal LLM alignments using two-player games”
Rui Zheng et al · 2024
Closest in time.
“DINO-WM: World models on pre-trained visual features enable zero-shot planning”
Gaoyue Zhou, Hengkai Pan, Yann LeCun and Lerrel Pinto · 2024
Closest in time.
“Diffusion Model Predictive Control”
Guangyao Zhou, Sivaramakrishnan Swaminathan, Rajkumar Raju, J Guntupalli, Wolfgang Lehrach, Joseph Ortiz, Antoine Dedieu, Miguel Lázaro-Gredilla and Kevin Murphy · 2024
Closest in time.
“On representation complexity of model-based and model-free reinforcement learning”
Hanlin Zhu, Baihe Huang and Stuart Russell · 2024
Closest in time.
“Is Sora a world simulator? A comprehensive survey on General world models and beyond”
Zheng Zhu et al · 2024
Closest in time.
“Contrastive Difference Predictive Coding”
Chongyi Zheng, Ruslan Salakhutdinov and Benjamin Eysenbach · 2024
Closest in time.
“LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models”
Marwa Abdulhai, Isadora White, Charlie Snell, Charles Sun, Joey Hong, Yuexiang Zhai, Kelvin Xu and Sergey Levine · 2025
Closest in time.
“Toward efficient exploration by large language model agents”
Dilip Arumugam and Thomas Griffiths · 2025
Closest in time.
“GEPA: Reflective prompt evolution can outperform reinforcement learning”
Lakshya Agrawal et al · 2025
Closest in time.
“V-JEPA 2: Self-supervised video models enable understanding, prediction and planning”
Mido Assran et al · 2025
Closest in time.
“InfAlign: Inference-aware language model alignment”
Ananth Balashankar et al · 2025
Closest in time.
“Back to the features: DINO as a foundation for video world models”
Federico Baldassarre, Marc Szafraniec, Basile Terver, Vasil Khalidov, Francisco Massa, Yann LeCun, Patrick Labatut, Maximilian Seitzer and Piotr Bojanowski · 2025
Closest in time.
“How to build a consistency model: Learning flow maps via self-distillation”
Nicholas Boffi, Michael Albergo and Eric Vanden-Eijnden · 2025
Closest in time.
“XLSTM scaling laws: Competitive performance with linear time-complexity”
Maximilian Beck, Kajetan Schweighofer, Sebastian Böck, Sebastian Lehner and Sepp Hochreiter · 2025
Closest in time.
“ATLAS: Learning to optimally memorize the context at test time”
Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn and Vahab Mirrokni · 2025
Closest in time.
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong and Vahab Mirrokni · 2025
Closest in time.
“Theoretical guarantees on the best-of-n alignment policy”
Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D’Amour, Jacob Eisenstein, Chirag Nagpal and Ananda Suresh · 2025
Closest in time.
Liu Bo et al · 2025
Closest in time.
“The Hundred-Page Language Models Book”, 2025
Andriy Burkov · 2025
Closest in time.
“Towards empowerment gain through causal structure learning in model-based RL”
Hongye Cao, Fan Feng, Meng Fang, Shaokang Dong, Tianpei Yang, Jing Huo and Yang Gao · 2025
Closest in time.
“Causal action empowerment for efficient reinforcement learning in embodied agents”
Hongye Cao, Fan Feng, Jing Huo and Yang Gao · 2025
Closest in time.
“Causal Information Prioritization for efficient Reinforcement Learning”
Hongye Cao, Fan Feng, Tianpei Yang, Jing Huo and Yang Gao · 2025
Closest in time.
“Imitation learning in the deep learning era: A novel taxonomy and recent advances”
Iason Chrysomallis and Georgios Chalkiadakis · 2025
Closest in time.
“Verlog: A Multi-turn RL framework for LLM agents”, 2025
W. Chen, J. Chen, H. Zhu and J. Schneider · 2025
Closest in time.
“SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training”
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Sergey Levine and Yi Ma · 2025
Closest in time.
“AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms”, 2025
Google DeepMind · 2025
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”
DeepSeek-AI · 2025
Closest in time.
“Understanding world or predicting future? A comprehensive survey of world models”
Jingtao Ding et al · 2025
Closest in time.
“Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning”
Adriàópez Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu and Hao Su · 2025
Closest in time.
Jesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Rémi Munos, Alessandro Lazaric and Ahmed Touati · 2025
Closest in time.
“Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo”
Shengyu Feng, Xiang Kong, Shuang Ma, Aonan Zhang, Dong Yin, Chong Wang, Ruoming Pang and Yiming Yang · 2025
Closest in time.
Dylan Foster, Zakaria Mhammedi and Dhruv Rohatgi · 2025
Closest in time.
“Sample, don’t search: Rethinking test-time alignment for language models”
Goncalo Faria and Noah Smith · 2025
Closest in time.
Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile and Noah Goodman · 2025
Closest in time.
“BoxingGym: Benchmarking progress in automated experimental design and model discovery”
Kanishk Gandhi, Michael Li, Lyle Goodyear, Louise Li, Aditi Bhaskar, Mohammed Zaman and Noah Goodman · 2025
Closest in time.
“Seedance 1.0: Exploring the boundaries of video generation models”
Yu Gao et al · 2025
Closest in time.
“Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning”
Samuel Garcin, Trevor McInroe, Pablo Castro, Christopher Lucas, David Abel, Prakash Panangaden and Stefano Albrecht · 2025
Closest in time.
“Fourier head: Helping large language models learn complex probability distributions”
Nate Gillman, Daksh Aggarwal, Michael Freeman, Saurabh Singh and Chen Sun · 2025
Closest in time.
“Agentic design patterns: A hands-on guide to building intelligent systems”
Antonio Gulli · 2025
Closest in time.
“Mastering diverse control tasks through world models”
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba and Timothy Lillicrap · 2025
Closest in time.
“Multi-Agent Risks from Advanced AI”
Lewis Hammond et al · 2025
Closest in time.
“Hierarchical World Models as Visual Whole-Body Humanoid Controllers”
Nicklas Hansen, S Jyothir, Vlad Sobal, Yann LeCun, Xiaolong Wang and Hao Su · 2025
Closest in time.
“Reinforcement Learning in the Era of Large Language Models: Challenges and Opportunities”, 2025
Qianyue Hao et al · 2025
Closest in time.
“REINFORCE++: An efficient RLHF algorithm with robustness to both prompt and reward models”
Jian Hu, Jason Liu, Haotian Xu and Wei Shen · 2025
Closest in time.
“Is best-of-N the best of them? Coverage, scaling, and optimality in inference-time alignment”
Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Akshay Krishnamurthy and Dylan Foster · 2025
Closest in time.
“Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization”
Audrey Huang, Wenhao Zhan, Tengyang Xie, Jason Lee, Wen Sun, Akshay Krishnamurthy and Dylan Foster · 2025
Closest in time.
“Memento: Fine-tuning LLM agents without fine-tuning LLMs”
Zhou Huichi et al · 2025
Closest in time.
“Training agents inside of scalable world models”
Danijar Hafner, Wilson Yan and Timothy Lillicrap · 2025
Closest in time.
“A Clean Slate for Offline Reinforcement Learning”, 2025
Matthew Jackson, Uljad Berdica, Jarek Liesen, Shimon Whiteson and Jakob Foerster · 2025
Closest in time.
“Test-time compute: From System-1 thinking to System-2 thinking”
Yixin Ji, Juntao Li, Hai Ye, Kaixin Wu, Kai Yao, Jia Xu, Linjian Mo and Min Zhang · 2025
Closest in time.
“Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning”
Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani and Jiawei Han · 2025
Closest in time.
“Stop anthropomorphizing intermediate tokens as reasoning/thinking traces!”
Subbarao Kambhampati, Kaya Stechly, Karthik Valmeekam, Lucas Saldyt, Siddhant Bhambri, Vardhan Palod, Atharva Gundawar, Soumya Samineni, Durgesh Kalwar and Upasana Biswas · 2025
Closest in time.
“Vision-language-action models for robotics: A review towards real-world applications”
Kento Kawaharazuka, Jihoon Oh, Jun Yamada, Ingmar Posner and Yuke Zhu · 2025
Closest in time.
“VinePPO: Refining credit assignment in RL training of LLMs”
Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni, Siva Reddy, Aaron Courville and Nicolas Roux · 2025
Closest in time.
“Reasoning with sampling: Your base model is smarter than you think”
Aayush Karan and Yilun Du · 2025
Closest in time.
Zaid Khan, Archiki Prasad, Elias Stengel-Eskin, Jaemin Cho and Mohit Bansal · 2025
Closest in time.
“The art of scaling reinforcement learning compute for LLMs”
Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Duvvuri, Manzil Zaheer, Inderjit Dhillon, David Brandfonbrener and Rishabh Agarwal · 2025
Closest in time.
“A unifying framework for action-conditional self-predictive Reinforcement Learning”
Khimya Khetarpal, Zhaohan Guo, Bernardo Pires, Yunhao Tang, Clare Lyle, Mark Rowland, Nicolas Heess, Diana Borsa, Arthur Guez and Will Dabney · 2025
Closest in time.
“Discovering temporal structure: An overview of hierarchical reinforcement learning”
Martin Klissarov, Akhil Bagaria, Ziyan Luo, George Konidaris, Doina Precup and Marlos Machado · 2025
Closest in time.
“On the Modeling Capabilities of Large Language Models for Sequential Decision Making”
Martin Klissarov, R Devon, Alexander Toshev and Bogdan Mazoure · 2025
Closest in time.
“Language Self-play for data-free training”
Jakub Kuba, Mengting Gu, Qi Ma, Yuandong Tian and Vijai Mohan · 2025
Closest in time.
“LLMs Get Lost In Multi-Turn Conversation”
Philippe Laban, Hiroaki Hayashi, Yingbo Zhou and Jennifer Neville · 2025
Closest in time.
“AssistanceZero: Scalably Solving Assistance Games”
Cassidy Laidlaw, Eli Bronstein, Timothy Guo, Dylan Feng, Lukas Berglund, Justin Svegliato, Stuart Russell and Anca Dragan · 2025
Closest in time.
“Reinforcement Learning from Human Feedback”, 2025
Nathan Lambert · 2025
Closest in time.
“A view on learning robust goal-conditioned value functions: Interplay between RL and MPC”
Nathan Lawrence, Philip Loewen, Michael Forbes, R Gopaluni and Ali Mesbah · 2025
Closest in time.
“Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation”
Jingyu Liu, Beidi Chen and Ce Zhang · 2025
Closest in time.
“Intelligent Go-Explore: Standing on the shoulders of giant foundation models”
Cong Lu, Shengran Hu and Jeff Clune · 2025
Closest in time.
“Beyond single-turn: A survey on multi-turn interactions with large language models”
Yubo Li, Xiaobin Shen, Xinyu Yao, Xueying Ding, Yidi Miao, Ramayya Krishnan and Rema Padman · 2025
Closest in time.
“MARFT: Multi-Agent Reinforcement Fine-Tuning”
Junwei Liao, Muning Wen, Jun Wang and Weinan Zhang · 2025
Closest in time.
“Chasing moving targets with online self-play reinforcement learning for safer language models”
Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Du, Yejin Choi, Tim Althoff and Natasha Jaques · 2025
Closest in time.
“Understanding R1-Zero-Like Training: A Critical Perspective”
Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Lee and Min Lin · 2025
Closest in time.
“Understanding R1-zero-like training: A critical perspective”
Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Lee and Min Lin · 2025
Closest in time.
Zichen Liu et al · 2025
Closest in time.
“Part I: Tricks or traps? A deep dive into RL for LLM reasoning”
Zihe Liu et al · 2025
Closest in time.
“Syntactic and semantic control of large language models via sequential Monte Carlo”
João Loula et al · 2025
Closest in time.
“Large Language Model agent: A survey on methodology, applications and challenges”
Junyu Luo et al · 2025
Closest in time.
“Agent RL scaling law: Agent RL with spontaneous code execution for mathematical problem solving”
Xinji Mai, Haotian Xu, Zhong-Zhi Li, W, Xing, Weinong Wang, Jian Hu, Yingying Zhang and Wenqiang Zhang · 2025
Closest in time.
“Differentiable Tree Search Network”
Dixant Mittal and Wee Lee · 2025
Closest in time.
“Multi-Agent Tool-Integrated Policy Optimization”
Zhanfeng Mo, Xingxuan Li, Yuntao Chen and Lidong Bing · 2025
Closest in time.
“A survey of in-context reinforcement learning”
Amir Moeini, Jiuqi Wang, Jacob Beck, Ethan Blaser, Shimon Whiteson, Rohan Chandra and Shangtong Zhang · 2025
Closest in time.
Youssef Mroueh · 2025
Closest in time.
Xuan-Phi Nguyen, Shrey Pandit, Revanth Reddy, Austin Xu, Silvio Savarese, Caiming Xiong and Shafiq Joty · 2025
Closest in time.
“MesaNet: Sequence modeling by locally optimal test-time training”
Johannes von Oswald et al · 2025
Closest in time.
“Training deep learning models with norm-constrained LMOs”
Thomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu, Antonio Silveti-Falls and Volkan Cevher · 2025
Closest in time.
“Wasserstein Policy Optimization”
David Pfau, Ian Davies, Diana Borsa, Joao Araujo, Brendan Tracey and Hado van Hasselt · 2025
Closest in time.
“PoE-world: Compositional world modeling with products of programmatic experts”
Wasu Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller, Marta Kryven and Kevin Ellis · 2025
Closest in time.
Isha Puri, Shivchander Sudalairaj, Guangxuan Xu, Kai Xu and Akash Srivastava · 2025
Closest in time.
Xinxing Ren, Caelum Forder, Qianbo Zang, Ahsen Tahir, Roman Georgio, Suman Deb, Peter Carroll, Onder Gurcan and Zekun Guo · 2025
Closest in time.
“Simple, Good, Fast: Self-Supervised World Models Free of Baggage”
Jan Robine, Marc Hoftmann and Stefan Harmeling · 2025
Closest in time.
“General agents need world models”
Jonathan Richens, David Abel, Alexis Bellot and Tom Everitt · 2025
Closest in time.
“Reevaluating policy gradient methods for imperfect-information games”
Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour, Alexandre Bayen, J Kolter, Amy Zhang, Gabriele Farina, Eugene Vinitsky and Samuel Sokota · 2025
Closest in time.
“Evolution Strategies at the Hyperscale”
Bidipta Sarkar et al · 2025
Closest in time.
“A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks”
Thomas Schmied, Thomas Adler, Vihang Patil, Maximilian Beck, Korbinian Pöppel, Johannes Brandstetter, Günter Klambauer, Razvan Pascanu and Sepp Hochreiter · 2025
Closest in time.
“A survey of embodied world models”, 2025
Yu Shang et al · 2025
Closest in time.
“LoRA Without Regret” https://thinkingmachines.ai/blog/lora/
John Schulman and Thinking Lab · 2025
Closest in time.
“From S4 to Mamba: A comprehensive survey on Structured State Space Models”
Shriyank Somvanshi, Md Islam, Mahmuda Mimi, Sazzad Bin Polock, Gaurab Chhetri and Subasish Das · 2025
Closest in time.
“R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via reinforcement learning”
Huatong Song, Jinhao Jiang, Wenqing Tian, Zhipeng Chen, Yuhuan Wu, Jiahao Zhao, Yingqian Min, Wayne Zhao, Lei Fang and Ji-Rong Wen · 2025
Closest in time.
“Welcome to the era of experience”, 2025
David Silver and Richard Sutton · 2025
Closest in time.
“Game theory meets large language models: A systematic survey”
Haoran Sun, Yusen Wu, Yukun Cheng and Xu Chu · 2025
Closest in time.
“Python is all you need? introducing dria-agent-a”, 2025
A. Tekparmak and andthattoo · 2025
Closest in time.
“On a few pitfalls in KL divergence gradient estimation for RL”
Yunhao Tang and Rémi Munos · 2025
Closest in time.
“Learning to chain-of-thought with Jensen’s evidence lower bound”
Yunhao Tang, Sid Wang and Rémi Munos · 2025
Closest in time.
“A survey on self-supervised methods for visual representation learning”
Tobias Uelwer, Jan Robine, Stefan Wagner, Marc Höftmann, Eric Upschulte, Sebastian Konietzny, Maike Behrendt and Stefan Harmeling · 2025
Closest in time.
Hugues Van, Mark Ibrahim, Tommaso Biancalani, Aviv Regev and Randall Balestriero · 2025
Closest in time.
“Expected Free Energy-based planning as variational inference”
Bert de Vries et al · 2025
Closest in time.
“A practitioner’s guide to multi-turn agentic reinforcement learning”
Ruiyi Wang and Prithviraj Ammanabrolu · 2025
Closest in time.
“ReMA: Learning to meta-think for LLMs with multi-Agent Reinforcement Learning”
Ziyu Wan et al · 2025
Closest in time.
“Drama: Mamba-enabled model-based reinforcement learning is sample and parameter efficient”
Wenlong Wang, Ivana Dusparic, Yucheng Shi, Ke Zhang and Vinny Cahill · 2025
Closest in time.
“RAGEN: Understanding self-evolution in LLM agents via multi-turn reinforcement learning”
Zihan Wang et al · 2025
Closest in time.
“Real-time reasoning agents in evolving environments”
Yule Wen, Yixin Ye, Yanzhe Zhang, Diyi Yang and Hao Zhu · 2025
Closest in time.
“Test-time regression: a unifying framework for designing sequence models with associative memory”
Ke Wang, Jiaxin Shi and Emily Fox · 2025
Closest in time.
Zhiheng Xi et al · 2025
Closest in time.
“Simple Policy Optimization”
Zhengpeng Xie, Qiang Zhang, Fan Yang, Marco Hutter and Renjing Xu · 2025
Closest in time.
“Towards large reasoning models: A survey of reinforced reasoning with Large Language Models”
Fengli Xu et al · 2025
Closest in time.
“Foundations of large language models”
Tong Xiao and Jingbo Zhu · 2025
Closest in time.
“Tool-R1: Sample-efficient reinforcement learning for agentic tool use”
Zhang Yabo, Zeng Yihan, Li Qingyun, Hu Zhen, Han Kavin and Zuo Wangmeng · 2025
Closest in time.
“Don’t build multi-agents”, 2025
Walden Yan · 2025
Closest in time.
“Fine-tuning Diffusion Policies with backpropagation through diffusion timesteps”
Ningyuan Yang, Jiaxuan Gao, Feng Gao, Yi Wu and Chao Yu · 2025
Closest in time.
“Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions”
Eunice Yiu, Kelsey Allen, Shiry Ginosar and Alison Gopnik · 2025
Closest in time.
“DAPO: An open-source LLM reinforcement learning system at scale”
Qiying Yu et al · 2025
Closest in time.
“RLeXplore: Accelerating research in intrinsically-motivated reinforcement learning”
Mingqi Yuan, Roger Castanyer, Bo Li, Xin Jin, Glen Berseth and Wenjun Zeng · 2025
Closest in time.
“Does reinforcement Learning really incentivize reasoning capacity in LLMs beyond the base model?”
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song and Gao Huang · 2025
Closest in time.
“On the statistical complexity for offline and low-adaptive reinforcement learning with structures”
Ming Yin, Mengdi Wang and Yu-Xiang Wang · 2025
Closest in time.
“Aligning large language models with human feedback: Mathematical foundations and algorithm design”
Siliang Zeng, Luca Viano, Chenliang Li, Jiaxiang Li, Volkan Cevher, Markus Wulfmeier, Stefano Ermon, Alfredo Garcia and Mingyi Hong · 2025
Closest in time.
“The landscape of agentic reinforcement learning for LLMs: A survey”
Guibin Zhang et al · 2025
Closest in time.
“A survey of Reinforcement Learning for large reasoning models”
Kaiyan Zhang et al · 2025
Closest in time.
“A survey of Reinforcement Learning for large reasoning models”
Kaiyan Zhang et al · 2025
Closest in time.
“Agentic Context Engineering: Evolving contexts for self-improving language models”
Qizheng Zhang et al · 2025
Closest in time.
“On the design of KL-Regularized Policy Gradient algorithms for LLM reasoning”
Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Yang Yuan, Quanquan Gu and Andrew Chi-Chih Yao · 2025
Closest in time.
“Absolute Zero: Reinforced self-play reasoning with zero data”
Andrew Zhao et al · 2025
Closest in time.
“MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use”
Weikang Zhao, Xili Wang, Chengdi Ma, Lingbin Kong, Zhaohua Yang, Mingxiang Tuo, Xiaowei Shi, Yitao Zhai and Xunliang Cai · 2025
Closest in time.
“Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill Learning”
Chongyi Zheng, Jens Tuyls, Joanne Peng and Benjamin Eysenbach · 2025
Closest in time.
“Group Sequence Policy Optimization”
Chujie Zheng et al · 2025
Closest in time.
“Lifelong learning of large language model based agents: A roadmap”
Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu and Qianli Ma · 2025
Closest in time.
“Variational reasoning for language models”
Xiangxin Zhou, Zichen Liu, Haonan Wang, Chao Du, Min Lin, Chongxuan Li, Liang Wang and Tianyu Pang · 2025
Closest in time.
“Self-challenging language model agents”
Yifei Zhou, Sergey Levine, Jason Weston, Xian Li and Sainbayar Sukhbaatar · 2025
Closest in time.
Yufa Zhou, Shaobo Wang, Xingyu Dong, Xiangqi Jin, Yifang Chen, Yue Min, Kexin Yang, Xingzhang Ren, Dayiheng Liu and Linfeng Zhang · 2025
Closest in time.
“The neural coding framework for learning generative models”
Alexander Ororbia and Daniel Kifer · 2064
Closest in time.
“Deep Reinforcement Learning with Double Q-Learning”
Hado van Hasselt, Arthur Guez and David Silver · 2094
Closest in time.