Fetching the paper…
Reading the bibliography…
Many reinforcement learning (RL) tasks provide the agent with high-dimensional observations that can be simplified into low-dimensional continuous states.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Singh, S. P., Jaakkola, T., and Jordan, M. I · 1995
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Mueller, A · 1997
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Müller, A · 1997
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Givan, R., Dean, T., and Greig, M · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
Ferns, N., Panangaden, P., and Precup, D · 2004
Earlier work this paper cites.
Testing for equal distributions in high dimension
Székely, G. J. and Rizzo, M. L · 2004
Earlier work this paper cites.
Convex Analysis and Nonlinear Optimization
Borwein, J. and Lewis, A. S · 2005
Earlier work this paper cites.
Lipschitz continuity of value functions in markovian decision processes
Hinderer, K · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Li, L., Walsh, T. J., and Littman, M. L · 2006
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Mahadevan, S. and Maggioni, M · 2007
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L · 2008
Earlier work this paper cites.
Optimal Transport: Old and New
Villani, C · 2008
Earlier work this paper cites.
Using bisimulation for policy transfer in mdps
Castro, P. and Precup, D · 2010
Earlier work this paper cites.
Basis function discovery using spectral clustering and bisimulation metrics
Comanici, G. and Precup, D · 2011
Earlier work this paper cites.
Bisimulation metrics for continuous markov decision processes
Ferns, N., Panangaden, P., and Precup, D · 2011
Earlier work this paper cites.
A kernel two-sample test
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A. J · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Equivalence of distance-based and rkhs-based statistics in hypothesis testing
Sejdinovic, D., Sriperumbudur, B. K., Gretton, A., and Fukumizu, K · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Cited alongside, same era.
Abstraction selection in model-based reinforcement learning
Jiang, N., Kulesza, A., and Singh, S · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Policy gradient in lipschitz markov decision processes
Pirotta, M., Restelli, M., and Bascetta, L · 2015
Cited alongside, same era.
Representation discovery for mdps using bisimulation metrics
Ruan, S. S., Comanici, G., Panangaden, P., and Precup, D · 2015
Cited alongside, same era.
Dopamine: A research framework for deep reinforcement learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Later among the works it cites.
Probabilistic recurrent state-space models
Doerr, A., Daniel, C., Schiegg, M., Nguyen-Tuong, D., Schaal, S., Toussaint, M., and Trimpe, S · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Later among the works it cites.
Combined reinforcement learning via abstract representations
Francois-Lavet, V., Bengio, Y., Precup, D., and Pineau, J · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Interaction networks for learning about objects, relations and physics
Battaglia, P. W., Pascanu, R., Lai, M., Rezende, D. J., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Near optimal behavior via approximate state abstraction
Abel, D., Hershkowitz, D. E., and Littman, M. L · 2017
Cited alongside, same era.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
Learning to navigate in complex environments
Mirowski, P. W., Pascanu, R., Viola, F., Soyer, H., Ballard, A. J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., Kumaran, D., and Hadsell, R · 2017
Cited alongside, same era.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Cited alongside, same era.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H. P., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D. P., Rabinowitz, N. C., Barreto, A., and Degris, T · 2017
Cited alongside, same era.
Later among the works it cites.
Regularisation of neural networks by enforcing lipschitz continuity
Gouk, H., Frank, E., Pfahringer, B., and Cree, M. J · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
van den Oord, A., Li, Y., and Vinyals, O · 2018
Later among the works it cites.
Algorithmic framework for model-based reinforcement learning with theoretical guarantees
Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T · 2018
Later among the works it cites.
Solar: Deep structured latent representations for model-based reinforcement learning
Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M. J., and Levine, S · 2018
Later among the works it cites.
A geometric perspective on optimal representations for reinforcement learning
Bellemare, M. G., Dabney, W., Dadashi, R., Taiga, A. A., Castro, P. S., Roux, N. L., Schuurmans, D., Lattimore, T., and Lyle, C · 2019
Closest in time.
Two-timescale networks for nonlinear value function approximation
Chung, W., Nath, S., Joseph, A. G., and White, M · 2019
Closest in time.
The value function polytope in reinforcement learning
Dadashi, R., Taiga, A. A., Roux, N. L., Schuurmans, D., and Bellemare, M. G · 2019
Closest in time.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H · 2019
Closest in time.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G · 2019
Closest in time.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Sepassi, R., Tucker, G., and Michalewski, H · 2019
Closest in time.
A comparative analysis of expected and distributional reinforcement learning
Lyle, C., Castro, P. S., and Bellemare, M. G · 2019
Closest in time.