Fetching the paper…
Reading the bibliography…
Most algorithms in reinforcement learning (RL) require that the objective is formalised with a Markovian reward function.
Silviu Pitis · 1902
Earlier work this paper cites.
Theory of games and economic behavior
J. Von Neumann and O. Morgenstern · 1944
Earlier work this paper cites.
Rewarding Behaviors
Fahiem Bacchus, Craig Boutilier, and Adam Grove · 1970
Earlier work this paper cites.
The Temporal Logic of Reactive and Concurrent Systems
Zohar Manna and Amir Pnueli · 1992
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Sridhar Mahadevan · 1996
Earlier work this paper cites.
Evolutionary Algorithms for Solving Multi-Objective Problems
Carlos A. Coello Coello, David A. Van Veldhuizen, and Gary B. Lamont · 2002
Earlier work this paper cites.
The reward hypothesis, 2004
Richard S. Sutton · 2004
Earlier work this paper cites.
Variational Policy Gradient Method for Reinforcement Learning with General Utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2007
Earlier work this paper cites.
Principles of model checking
Christel Baier and Joost-Pieter Katoen · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
LTL Control in Uncertain Environments with Probabilistic Satisfaction Guarantees, April 2011
Xu Chu Ding, Stephen L. Smith, Calin Belta, and Daniela Rus · 2011
Earlier work this paper cites.
Control of Markov decision processes from PCTL specifications
M. Lahijanian, S. B. Andersson, and C. Belta · 2011
Earlier work this paper cites.
Counterexamples in Topology
L.A. Steen and J.A.J. Seebach · 2012
Earlier work this paper cites.
A multiobjective reinforcement learning approach to water resources systems operation: Pareto frontier approximation in a single run
Andrea Castelletti, F. Pianosi, and Marcello Restelli · 2013
Earlier work this paper cites.
A Survey of Multi-Objective Sequential Decision-Making
D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley · 2013
Cited alongside, same era.
Cooperative inverse reinforcement learning, 2016
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences, June 2017
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Multi-objective optimization of radiotherapy: distributed Q-learning and agent-based simulation
Ammar Jalalimanesh, Hamidreza Shahabi Haghighi, Abbas Ahmadi, Hossein Hejazian, and Madjid Soltani · 2017
Cited alongside, same era.
Environment-Independent Task Specifications via GLTL, April 2017
Michael L. Littman, Ufuk Topcu, Jie Fu, Charles Isbell, Min Wen, and James MacGlashan · 2017
Cited alongside, same era.
On the Expressivity of Markov Reward, January 2022
David Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho, Michael L. Littman, Doina Precup, and Satinder Singh · 2022
Later among the works it cites.
Settling the Reward Hypothesis, December 2022
Michael Bowling, John D. Martin, David Abel, and Will Dabney · 2022
Later among the works it cites.
On the Theory of Reinforcement Learning with Once-per-Episode Feedback
Niladri S. Chatterji, Aldo Pacchiano, Peter L. Bartlett, and Michael I. Jordan · 2022
Later among the works it cites.
A Practical Guide to Multi-Objective Reinforcement Learning and Planning
Conor F. Hayes, Roxana Rădulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M. Zintgraf, Richard Dazeley, Fredrik Heintz, Enda Howley, Athirai A. Irissappane, Patrick Mannion, Ann Nowé, Gabriel Ramos, Marcello Restelli, Peter Vamplew, and Diederik M. Roijers · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Using Reward Machines for High-Level Task Specification and Decomposition in Reinforcement Learning
Rodrigo Toro Icarte, Toryn Klassen, Richard Valenzano, and Sheila McIlraith · 2018
Cited alongside, same era.
Learning Linear Temporal Properties, September 2018
Daniel Neider and Ivan Gavran · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Structured reward shaping using signal temporal logic specifications
A. Balakrishnan and J.V. Deshmukh · 2019
Cited alongside, same era.
LTL and Beyond: Formal Languages for Reward Function Specification in Reinforcement Learning
Alberto Camacho, Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith · 2019
Cited alongside, same era.
Regret Minimization for Reinforcement Learning with Vectorial Feedback and Complex Objectives
Wang Chi Cheung · 2019
Cited alongside, same era.
Provably Efficient Maximum Entropy Exploration, January 2019
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Utility Theory for Sequential Decision Making
Mehran Shakerinava and Siamak Ravanbakhsh · 2022
Later among the works it cites.
Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith · 2022
Later among the works it cites.
Scalar reward is not enough: a response to Silver, Singh, Precup and Sutton (2021)
Peter Vamplew, Benjamin J. Smith, Johan Källström, Gabriel Ramos, Roxana Rădulescu, Diederik M. Roijers, Conor F. Hayes, Fredrik Heintz, Patrick Mannion, Pieter J. K. Libin, Richard Dazeley, and Cameron Foale · 2022
Later among the works it cites.
On the Expressivity of Multidimensional Markov Reward, July 2023
Shuwa Miura · 2023
Closest in time.
Convex Reinforcement Learning in Finite Trials
Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, and Marcello Restelli · 2023
Closest in time.
Challenging Common Assumptions in Convex Reinforcement Learning, January 2023
Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, and Marcello Restelli · 2023
Closest in time.
Invariance in Policy Optimisation and Partial Identifiability in Reward Learning, June 2023
Joar Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2023
Closest in time.
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
Joar Max Viktor Skalse and Alessandro Abate · 2023
Closest in time.
Multi-Agent Reinforcement Learning Guided by Signal Temporal Logic Specifications
J. Wang, S. Yang, Z. An, S. Han, Z. Zhang, R. Mangharam, M. Ma, and F. Miao · 2023
Closest in time.
Inverse Reinforcement Learning with the Average Reward Criterion, May 2023
Feiyang Wu, Jingyang Ke, and Anqi Wu · 2023
Closest in time.
Reward is enough for convex MDPs
Tom Zahavy, Brendan O’Donoghue, Guillaume Desjardins, and Satinder Singh · 2023
Closest in time.