Fetching the paper…
Reading the bibliography…
This paper proposes a formal approach to online learning and planning for agents operating in a priori unknown, time-varying environments.
Applied Probability Models with Optimization Applications
Sheldon M. Ross · 1970
Earlier work this paper cites.
Stabilization and sensitivity for eventually time-invariant systems
Avraham Feintuch · 1989
Earlier work this paper cites.
Sonar-based real-world mapping and navigation
Alberto Elfes · 1990
Earlier work this paper cites.
Regime switching with time-varying transition probabilities
Francis X. Diebold, Joon-Haeng Lee, and Gretchen C. Weinbach · 1994
Earlier work this paper cites.
Business-cycle phases and their transitional dynamics
Andrew J. Filardo · 1994
Earlier work this paper cites.
A short proof of the Gittins index theorem
John N. Tsitsiklis · 1994
Earlier work this paper cites.
The BATmobile: Towards a Bayesian automated taxi
Jeff Forbes, Tim Huang, Keiji Kanazawa, and Stuart Russell · 1995
Earlier work this paper cites.
Local bandit approximation for optimal learning problems
Michael O. Duff and Andrew G. Barto · 1997
Earlier work this paper cites.
Experiments with reinforcement learning in problems with continuous state and action spaces
Juan C. Santamaría, Richard S. Sutton, and Ashwin Ram · 1997
Earlier work this paper cites.
The effects of weather fronts on GPS measurements
Thierry Gregorius and Geoffrey Blewitt · 1998
Earlier work this paper cites.
Module-based reinforcement learning: Experiments with a real robot
Zsolt Kalmár, Csaba Szepesvári, and András Lőrincz · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Exact solutions to time-dependent MDPs
Justin A. Boyan and Michael L. Littman · 2001
Earlier work this paper cites.
R-MAX — a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
ε \varepsilon -MDPs: Learning in varying environments
István Szita, Bálint Takács, and András Lörincz · 2002
Earlier work this paper cites.
Handbook of Means and Their Inequalities
Peter S. Bullen · 2003
Earlier work this paper cites.
Statistical Modelling Strategies for Reliability Data on Physical Components with Possibly Multiple Causes of Failure
Ellen Andries · 2004
Earlier work this paper cites.
A Primer on Statistical Distributions
Narayanaswamy Balakrishnan and Valery B. Nevzorov · 2004
Cited alongside, same era.
Introduction to Nonlinear Optimization: Theory, Algorithms, and Applications with MATLAB
Amir Beck · 2004
Cited alongside, same era.
Multi-agent patrolling with reinforcement learning
Hugo Santana, Geber Ramalho, Vincent Corruble, and Bohdana Ratitch · 2004
Cited alongside, same era.
Solving generalized semi-Markov decision processes using continuous phase-type distributions
Håkan L. S. Younes and Reid G. Simmons · 2004
Cited alongside, same era.
Aeolian processes in Proctor Crater on Mars: Mesoscale modeling of dune-forming winds
Lori K. Fenton, Anthony D. Toigo, and Mark I. Richardson · 2005
Cited alongside, same era.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 2005
Wind-energy based path planning for unmanned aerial vehicles using Markov decision processes
Wesam H. Al-Sabban, Luis F. Gonzalez, and Ryan N. Smith · 2013
Later among the works it cites.
Probably approximately correct MDP learning and control with temporal logic constraints
Jie Fu and Ufuk Topcu · 2014
Later among the works it cites.
Adaptive motion control of wheeled mobile robot with unknown slippage
Haibo Gao, Xingguo Song, Liang Ding, Kerui Xia, Nan Li, and Zongquan Deng · 2014
Later among the works it cites.
Introduction to the Planetary Dunes special issue, and the aeolian career of Ronald Greeley
James R. Zimbelman · 2014
Later among the works it cites.
Deep recurrent Q-learning for partially observable MDPs
Matthew J. Hausknecht and Peter Stone · 2015
Later among the works it cites.
Online path planning for autonomous underwater vehicles in unknown environments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Models for Discrete Data
Daniel Zelterman · 2006
Cited alongside, same era.
Model-free Q-learning designs for linear discrete-time zero-sum games with application to H-infinity control
Asma Al-Tamimi, Frank L. Lewis, and Murad Abu-Khalaf · 2007
Cited alongside, same era.
’Infotaxis’ as a strategy for searching without gradients
Massimo Vergassola, Emmanuel Villermaux, and Boris I. Shraiman · 2007
Cited alongside, same era.
Value function based reinforcement learning in changing Markovian environments
Balázs Csanád Csáji and László Monostori · 2008
Cited alongside, same era.
An analysis of model-based interval estimation for Markov decision processes
Alexander L. Strehl and Michael L. Littman · 2008
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
J. Zico Kolter and Andrew Y. Ng · 2009
Cited alongside, same era.
Juan David Hernández, Eduard Vidal, Guillem Vallicrosa, Enric Galceran, and Marc Carreras · 2015
Later among the works it cites.
History of pilot ballooning
Alice Hickman · 2015
Later among the works it cites.
Thermophysical properties along Curiosity’s traverse in Gale crater, Mars, derived from the REMS ground temperature sensor
Ashwin R. Vasavada, Sylvain Piqueux, Kevin W. Lewis, Mark T. Lemmon, and Michael D. Smith · 2017
Later among the works it cites.
Pratik Gajane, Ronald Ortner, and Peter Auer · 2018
Later among the works it cites.
A solution to time-varying Markov decision processes
Lantao Liu and Gaurav S. Sukhatme · 2018
Later among the works it cites.
Expedited learning in MDPs with side information
Melkior Ornik, Jie Fu, Niklas T. Lauffer, W. K. Perera, Mohammed Alshiekh, Masahiro Ono, and Ufuk Topcu · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Using reward machines for high-level task specification and decomposition in reinforcement learning
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith · 2018
Later among the works it cites.
Online markov decision processes with time-varying transition probabilities and rewards
Yingying Li, Aoxiao Zhong, Guannan Qu, and Na Li · 2019
Closest in time.
In-flight air density estimation and prediction for hypersonic flight vehicles
Hamza El-Kebir and Melkior Ornik · 2020
Closest in time.
Variational regret bounds for reinforcement learning
Ronald Ortner, Pratik Gajane, and Peter Auer · 2020
Closest in time.
Trading-off static and dynamic regret in online least-squares and beyond
Jianjun Yuan and Andrew Lamperski · 2020
Closest in time.