Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), rewards of states are typically considered additive, and following the Markov assumption, they are $\textit{independent}$ of states visited previously.
An analysis of approximations for maximizing submodular set functions—i
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
Submodular set functions, matroids and the greedy algorithm: Tight worst-case bounds and some generalizations of the rado-edmonds theorem
Michele Conforti and Gérard Cornuéjols · 1984
Earlier work this paper cites.
Tyre modelling for use in vehicle dynamics studies
Egbert Bakker, Lars Nyborg, and Hans B. Pacejka · 1987
Earlier work this paper cites.
Advantage updating
Leemon C Baird, III · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
A threshold of ln n for approximating set cover
Uriel Feige · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Optimal learning: Computational procedures for Bayes -adaptive Markov decision processes
Michael O’Gordon Duff · 2002
Earlier work this paper cites.
Polylogarithmic inapproximability
Eran Halperin and Robert Krauthgamer · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
A recursive greedy algorithm for walks in directed graphs
Chandra Chekuri and M. Pal · 2005
Earlier work this paper cites.
Chapter 19 gradient estimation
Michael C. Fu · 2006
Earlier work this paper cites.
Near-optimal sensor placements in gaussian processes: Theory, efficient algorithms and empirical studies
Andreas Krause, Ajit Singh, and Carlos Guestrin · 2008
Earlier work this paper cites.
An online algorithm for maximizing submodular functions
Matthew Streeter and Daniel Golovin · 2008
Earlier work this paper cites.
Nonmyopic adaptive informative path planning for multiple robots
Amarjeet Singh, Andreas Krause, and William J. Kaiser · 2009
Earlier work this paper cites.
Submodularity and curvature: the optimal algorithm
Jan Vondrak · 2010
Cited alongside, same era.
Learning submodular functions
Maria-Florina Balcan and Nicholas JA Harvey · 2011
Cited alongside, same era.
Understanding the nesting spatial behaviour of gorillas in the kagwene sanctuary, cameroon
Neba Funwi-gabga and Jorge Mateu · 2011
Cited alongside, same era.
Adaptive submodularity: Theory and applications in active learning and stochastic optimization
Daniel Golovin and Andreas Krause · 2011
Cited alongside, same era.
Linear submodular bandits and their application to diversified retrieval
Yisong Yue and Carlos Guestrin · 2011
Cited alongside, same era.
Submodular function maximization
Andreas Krause and Daniel Golovin · 2014
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Active model estimation in markov decision processes
Jean Tarbouriech, Shubhanshu Shekhar, Matteo Pirotta, Mohammad Ghavamzadeh, and Alessandro Lazaric · 2020
Later among the works it cites.
Submodularity in action: From machine learning to signal processing applications
Ehsan Tohidi, Rouhollah Amiri, Mario Coutino, David Gesbert, Geert Leus, and Amin Karbasi · 2020
Later among the works it cites.
Planning with submodular objective functions
Ruosong Wang, Hanrui Zhang, Devendra Singh Chaplot, Denis Garagić, and Ruslan Salakhutdinov · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimization-based autonomous racing of 1: 43 scale rc cars
Alexander Liniger, Alexander Domahidi, and Manfred Morari · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deep submodular functions: Definitions and learning
Brian W Dolhansky and Jeff A Bilmes · 2016
Cited alongside, same era.
Orienteering problem: A survey of recent variants, solution approaches and applications
Aldy Gunawan, Hoong Chuin Lau, and Pieter Vansteenwegen · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
Interactive submodular bandit
Lin Chen, Andreas Krause, and Amin Karbasi · 2017
Cited alongside, same era.
On the expressivity of markov reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K Ho, Michael Littman, Doina Precup, and Satinder Singh · 2021
Later among the works it cites.
Inverse reinforcement learning in contextual mdps
Stav Belogolovsky, Philip Korsunsky, Shie Mannor, Chen Tessler, and Tom Zahavy · 2021
Later among the works it cites.
On the theory of reinforcement learning with once-per-episode feedback
Niladri Chatterji, Aldo Pacchiano, Peter Bartlett, and Michael Jordan · 2021
Later among the works it cites.
Swarm slam: Challenges and perspectives
Miquel Kegeleirs, Giorgio Grisetti, and Mauro Birattari · 2021
Later among the works it cites.
Information directed reward learning for reinforcement learning
David Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek, and Andreas Krause · 2021
Later among the works it cites.
Sensing cox processes via posterior sampling and positive bases
Mojmír Mutný and Andreas Krause · 2021
Later among the works it cites.
Competitive policy optimization
Manish Prajapat, Kamyar Azizzadenesheli, Alexander Liniger, Yisong Yue, and Anima Anandkumar · 2021
Later among the works it cites.
Reward is enough for convex mdps
Tom Zahavy, Brendan O’Donoghue, Guillaume Desjardins, and Satinder Singh · 2021
Later among the works it cites.
Submodularity in machine learning and artificial intelligence
Jeff Bilmes · 2022
Later among the works it cites.
Challenging common assumptions in convex reinforcement learning
Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, and Marcello Restelli · 2022
Later among the works it cites.
Near-optimal multi-agent learning for safe coverage control
Manish Prajapat, Matteo Turchetta, Melanie Zeilinger, and Andreas Krause · 2022
Later among the works it cites.
Active exploration via experiment design in markov chains
Mojmir Mutny, Tadeusz Janik, and Andreas Krause · 2023
Closest in time.