Fetching the paper…
Reading the bibliography…
We propose a general framework for policy representation for reinforcement learning tasks.
Zur theorie der orthogonalen funktionensysteme
Alfred Haar · 1909
Earlier work this paper cites.
Functions of positive and negative type, and their connection with the theory of integral equations
J Mercer · 1909
Earlier work this paper cites.
The theory of approximation
Dunham Jackson · 1930
Earlier work this paper cites.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator
Aryeh Dvoretzky, Jack Kiefer, and Jacob Wolfowitz · 1956
Earlier work this paper cites.
Perturbation theory and finite markov chains
Paul J Schweitzer · 1968
Earlier work this paper cites.
Bounds on the truncation error of periodic signals
C Giardina and P Chirlian · 1972
Earlier work this paper cites.
An automatic method for generating random variates with a given characteristic function
Luc Devroye · 1986
Earlier work this paper cites.
Orthonormal bases of compactly supported wavelets
Ingrid Daubechies · 1988
Earlier work this paper cites.
Optimal stopping of markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
John N Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
Comparison of perturbation bounds for the stationary distribution of a markov chain
Grace E Cho and Carl D Meyer · 2001
Earlier work this paper cites.
Predictive representations of state
Michael L Littman and Richard S Sutton · 2002
Earlier work this paper cites.
Fourier Analysis: An Introduction
E.M. Stein and R. Shakarchi · 2003
Earlier work this paper cites.
Finite mixture models
Geoffrey McLachlan and David Peel · 2004
Earlier work this paper cites.
Probability models in engineering and science , volume 193
Haym Benaroya, Seon Mi Han, and Mark Nagurka · 2005
Earlier work this paper cites.
Risk bounds for mixture density estimation
Alexander Rakhlin, Dmitry Panchenko, and Sayan Mukherjee · 2005
Earlier work this paper cites.
Hilbert space of probability density functions based on aitchison geometry
Juan José Egozcue, José Luis Díaz-Barrero, and Vera Pawlowsky-Glahn · 2006
Earlier work this paper cites.
Introduction to variance estimation
Kirk Wolter · 2007
Earlier work this paper cites.
Operator theory in function spaces
Kehe Zhu · 2007
Earlier work this paper cites.
Tailoring density estimation via reproducing kernel moment matching
Le Song, Xinhua Zhang, Alex Smola, Arthur Gretton, and Bernhard Schölkopf · 2008
Earlier work this paper cites.
Law of the unconscious statistician
Bengt Ringnér · 2009
Cited alongside, same era.
Application of integral operator for regularized least-square regression
Hongwei Sun and Qiang Wu · 2009
Cited alongside, same era.
An inequality for the trace of matrix products, using absolute values
Bernhard Baumgartner · 2011
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
Modelling transition dynamics in mdps with rkhs embeddings
Steffen Grunewalder, Guy Lever, Luca Baldassarre, Massi Pontil, and Arthur Gretton · 2012
Cited alongside, same era.
Information gathering actions over human internal state
Dorsa Sadigh, S Shankar Sastry, Sanjit A Seshia, and Anca Dragan · 2016
Later among the works it cites.
Learning the variance of the reward-to-go
Aviv Tamar, Dotan Di Castro, and Shie Mannor · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Later among the works it cites.
On the power of truncated svd for general high-rank matrix estimation problems
Simon S Du, Yining Wang, and Aarti Singh · 2017
Later among the works it cites.
Reluplex: An efficient smt solver for verifying deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu Nishiyama, Abdeslam Boularias, Arthur Gretton, and Kenji Fukumizu · 2012
Cited alongside, same era.
Wavelet analysis: the scalable structure of information
Howard L Resnikoff, O Raymond Jr, et al · 2012
Cited alongside, same era.
Provably safe and robust learning-based model predictive control
Anil Aswani, Humberto Gonzalez, S Shankar Sastry, and Claire Tomlin · 2013
Cited alongside, same era.
Hilbert space embeddings of predictive state representations
Byron Boots, Geoffrey Gordon, and Arthur Gretton · 2013
Cited alongside, same era.
Reachability-based safe learning with gaussian processes
Anayo K Akametalu, Jaime F Fisac, Jeremy H Gillula, Shahab Kaynama, Melanie N Zeilinger, and Claire J Tomlin · 2014
Cited alongside, same era.
Recovering distributions from gaussian rkhs embeddings
Motonobu Kanagawa and Kenji Fukumizu · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Guy Katz, Clark Barrett, David L Dill, Kyle Julian, and Mykel J Kochenderfer · 2017
Later among the works it cites.
The computational complexity of the fast fourier transform
Mathias Lohne · 2017
Later among the works it cites.
Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification
Yongxi Lu, Abhishek Kumar, Shuangfei Zhai, Yu Cheng, Tara Javidi, and Rogerio Feris · 2017
Later among the works it cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Later among the works it cites.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2018
Later among the works it cites.
Verifiable reinforcement learning via policy extraction
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama · 2018
Later among the works it cites.
Fourier Policy Gradients
Matthew Fellows, Kamil Ciosek, and Shimon Whiteson · 2018
Later among the works it cites.
Practical contextual bandits with regression oracles
Dylan J Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert E Schapire · 2018
Later among the works it cites.
Efficiently combining svd, pruning, clustering and retraining for enhanced neural network compression
Koen Goetschalckx, Bert Moons, Patrick Wambacq, and Marian Verhelst · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Safe policy improvement with an estimated baseline policy, 2019
Thiago D. Simão, Romain Laroche, and Rémi Tachet des Combes · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 2019
Later among the works it cites.
Efficient planning under partial observability with unnormalized q functions and spectral learning
Tianyu Li, Bogdan Mazoure, Doina Precup, and Guillaume Rabusseau · 2020
Closest in time.