Fetching the paper…
Reading the bibliography…
Deep latent variable models have achieved significant empirical successes in model-based reinforcement learning (RL) due to their expressiveness in modeling complex transition dynamics.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Entropy, Compactness and the Approximation of Operators
Bernd Carl and Irmtraud Stephani · 1990
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Justin Boyan and Andrew Moore · 1994
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Sridhar Mahadevan and Mauro Maggioni · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Information consistency of nonparametric gaussian process methods
Matthias W Seeger, Sham M Kakade, and Dean P Foster · 2008
Earlier work this paper cites.
Support vector machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Optimal rates for regularized least squares regression
Ingo Steinwart, Don R Hush, Clint Scovel, et al · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: no regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Functional analysis
Frigyes Riesz and Béla Sz Nagy · 2012
Earlier work this paper cites.
Mercer’s theorem on general domains: On the interaction between measures, kernels, and rkhss
Ingo Steinwart and Clint Scovel · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Yuval Tassa, Tom Erez, and Emanuel Todorov · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Finite-time analysis of kernelised contextual bandits
Michal Valko, Nathan Korda, Rémi Munos, Ilias Flaounas, and Nello Cristianini · 2013
Earlier work this paper cites.
Scalable kernel methods via doubly stochastic gradients
Bo Dai, Bo Xie, Niao He, Yingyu Liang, Anant Raj, Maria-Florina F Balcan, and Le Song · 2014
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deep successor reinforcement learning
Tejas D Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Cited alongside, same era.
An introduction to the theory of reproducing kernel Hilbert spaces , volume 152
Deepmdp: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Later among the works it cites.
An introduction to variational autoencoders
Diederik P Kingma, Max Welling, et al · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vern I Paulsen and Mrinal Raghupathi · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
On some extensions of bernstein’s inequality for self-adjoint operators
Stanislav Minsker · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
The geometry of random features
Krzysztof Choromanski, Mark Rowland, Tamás Sarlós, Vikas Sindhwani, Richard Turner, and Adrian Weller · 2018
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex X Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with kernel and neural function approximations
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael Jordan · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
Denis Yarats and Ilya Kostrikov · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2020
Later among the works it cites.
Mitigating covariate shift in imitation learning via offline data with partial coverage
Jonathan Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Towards automatic evaluation of dialog systems: A model-free off-policy evaluation approach
Haoming Jiang, Bo Dai, Mengjiao Yang, Tuo Zhao, and Wei Wei · 2021
Later among the works it cites.
Online sub-sampling for reinforcement learning with general function approximation
Dingwen Kong, Ruslan Salakhutdinov, Ruosong Wang, and Lin F Yang · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Later among the works it cites.
Provable representation learning for imitation with contrastive fourier features
Ofir Nachum and Mengjiao Yang · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Cautiously optimistic policy optimization and exploration with linear function approximation
Andrea Zanette, Ching-An Cheng, and Alekh Agarwal · 2021
Later among the works it cites.
Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning
Shuang Qiu, Lingxiao Wang, Chenjia Bai, Zhuoran Yang, and Zhaoran Wang · 2022
Closest in time.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2022
Closest in time.
Daydreamer: World models for physical robot learning
Philipp Wu, Alejandro Escontrela, Danijar Hafner, Ken Goldberg, and Pieter Abbeel · 2022
Closest in time.