Fetching the paper…
Reading the bibliography…
The transition kernel of a continuous-state-action Markov decision process (MDP) admits a natural tensor structure.
Provably efficient rl with rich observations via latent state decoding
Simon S Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudík, and John Langford · 1901
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 1910
Earlier work this paper cites.
Perturbation bounds in connection with singular value decomposition
Per-Åke Wedin · 1972
Earlier work this paper cites.
Principal components in regression analysis
Ian T Jolliffe · 1986
Earlier work this paper cites.
Variable resolution dynamic programming: Efficiently learning action maps in multivariate real-valued state-spaces
Andrew W Moore · 1991
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
John N Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Kernel-based reinforcement learning in average-cost problems
Dirk Ormoneit and Peter Glynn · 2002
Earlier work this paper cites.
State aggregation in markov decision processes
Zhiyuan Ren and Bruce H Krogh · 2002
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Proto-value functions: Developmental reinforcement learning
Sridhar Mahadevan · 2005
Earlier work this paper cites.
Diffusion maps and coarse-graining: A unied framework for dimensionality reduction, graph partitioning, and data set parameterization
Stéphane Lafon and Ann Lee · 2006
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 2007
Earlier work this paper cites.
Constructing basis functions from directed graphs for value function approximation
Jeff Johns and Sridhar Mahadevan · 2007
Earlier work this paper cites.
Analyzing feature generation for value-function approximation
Ronald Parr, Christopher Painter-Wakefield, Lihong Li, and Michael Littman · 2007
Earlier work this paper cites.
An analysis of laplacian methods for value function approximation in mdps
Marek Petrik · 2007
Earlier work this paper cites.
Diffusion maps, reduction coordinates, and low dimensional representation of stochastic systems
Ronald R. Coifman, Ioannis G. Kevrekidis, Stéphane Lafon, Mauro Maggioni, and Boaz Nadler · 2008
Earlier work this paper cites.
Tensor rank and the ill-posedness of the best low-rank approximation problem
Vin De Silva and Lek-Heng Lim · 2008
Earlier work this paper cites.
Optimal partition and effective dynamics of complex networks
Weinan E, Tiejun Li, and Eric Vanden-Eijnden · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Tensor decompositions and applications
Tamara G Kolda and Brett W Bader · 2009
Cited alongside, same era.
Markov chains and mixing times
David Asher Levin, Yuval Peres, and Elizabeth Lee Wilmer · 2009
Cited alongside, same era.
Learning representation and control in markov decision processes: New frontiers
Sridhar Mahadevan et al · 2009
Cited alongside, same era.
Markov state models based on milestoning
Christof Schütte, Frank Noe, Jianfeng Lu, Macro Sarich, and Eric Vanden-Eijnden · 2011
Cited alongside, same era.
Freedman’s inequality for matrix martingales
Joel A Tropp · 2011
Cited alongside, same era.
A new truncation strategy for the higher-order singular value decomposition
Nick Vannieuwenhoven, Raf Vandebril, and Karl Meerbergen · 2012
Cited alongside, same era.
Learning with good feature representations in bandits and in rl with a generative model
Tor Lattimore and Csaba Szepesvari · 2019
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2019
Later among the works it cites.
Learning low-dimensional state embeddings and metastable clusters from time series data
Yifan Sun, Yaqi Duan, Hao Gong, and Mengdi Wang · 2019
Later among the works it cites.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Later among the works it cites.
Limiting extrapolation in linear approximate value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Animashree Anandkumar, Rong Ge, and Majid Janzamin · 2014
Cited alongside, same era.
A statistical model for tensor pca
Emile Richard and Andrea Montanari · 2014
Cited alongside, same era.
Tensor decompositions for signal processing applications: From two-way to multiway component analysis
Andrzej Cichocki, Danilo Mandic, Lieven De Lathauwer, Guoxu Zhou, Qibin Zhao, Cesar Caiafa, and Huy Anh Phan · 2015
Cited alongside, same era.
Reinforcement learning in rich-observation mdps using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2016
Cited alongside, same era.
On the numerical approximation of the perron–frobenius and koopman operator
Stefan Klus, Péter Koltai, and Christof Schütte · 2016
Cited alongside, same era.
Sublinear time orthogonal tensor decomposition
Zhao Song, David Woodruff, and Huan Zhang · 2016
Cited alongside, same era.
Cross: Efficient low-rank tensor completion
Anru Zhang · 2019
Later among the works it cites.
Optimal sparse singular value decomposition for high-dimensional high-order data
Anru Zhang and Rungang Han · 2019
Later among the works it cites.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Later among the works it cites.
Generalized canonical polyadic tensor decomposition
David Hong, Tamara G Kolda, and Jed A Duersch · 2020
Later among the works it cites.
Eigendecompositions of transfer operators in reproducing kernel hilbert spaces
Stefan Klus, Ingmar Schuster, and Krikamol Muandet · 2020
Later among the works it cites.
Spectral state compression of markov processes
Anru Zhang and Mengdi Wang · 2020
Later among the works it cites.
Spectral thresholding for the estimation of markov chain transition operators
Matthias Löffler and Antoine Picard · 2021
Closest in time.
Tesseract: Tensorised actors for multi-agent reinforcement learning
Anuj Mahajan, Mikayel Samvelyan, Lei Mao, Viktor Makoviychuk, Animesh Garg, Jean Kossaifi, Shimon Whiteson, Yuke Zhu, and Animashree Anandkumar · 2021
Closest in time.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Closest in time.
Tensor methods in computer vision and deep learning
Yannis Panagakis, Jean Kossaifi, Grigorios G Chrysos, James Oldfield, Mihalis A Nicolaou, Anima Anandkumar, and Stefanos Zafeiriou · 2021
Closest in time.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Closest in time.
Model based multi-agent reinforcement learning with tensor decompositions
Pascal Van Der Vaart, Anuj Mahajan, and Shimon Whiteson · 2021
Closest in time.
An optimal statistical and computational framework for generalized tensor estimation
Rungang Han, Rebecca Willett, and Anru R Zhang · 2022
Closest in time.
Efficient reinforcement learning in block mdps: A model-free representation learning approach
Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang, Alekh Agarwal, and Wen Sun · 2022
Closest in time.
Learning markov models via low-rank optimization
Ziwei Zhu, Xudong Li, Mengdi Wang, and Anru Zhang · 2022
Closest in time.
Representation learning for general-sum low-rank markov games
Chengzhuo Ni, Yuda Song, Xuezhou Zhang, Chi Jin, and Mengdi Wang · 2023
Closest in time.