Fetching the paper…
Reading the bibliography…
This paper studies model-based bandit and reinforcement learning (RL) with nonlinear function approximations.
Provably efficient reinforcement learning with aggregated states
Shi Dong, Benjamin Van Roy, and Zhengyuan Zhou · 1912
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Predictive representations of state
Michael L Littman, Richard S Sutton, and Satinder P Singh · 2001
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2006
Earlier work this paper cites.
A unifying framework for computational reinforcement learning theory
Lihong Li · 2009
Earlier work this paper cites.
Nearly dimension-independent sparse linear bandit over small action spaces via best subset selection
Yining Wang, Yi Chen, Ethan X Fang, Zhaoran Wang, and Runze Li · 2009
Earlier work this paper cites.
Parametric bandits: the generalized linear case
Sarah Filippi, Olivier Cappé, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Earlier work this paper cites.
High-dimensional sparse linear bandits
Botao Hao, Tor Lattimore, and Mengdi Wang · 2011
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
Sham M Kakade, Varun Kanade, Ohad Shamir, and Adam Kalai · 2011
Earlier work this paper cites.
Online learning: stochastic, constrained, and smoothed adversaries
Alexander Rakhlin, Karthik Sridharan, and Ambuj Tewari · 2011
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2011
Earlier work this paper cites.
Bandit theory meets compressed sensing for high dimensional stochastic linear bandit
Alexandra Carpentier and Rémi Munos · 2012
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, Tong Zhang, et al · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Finite-time analysis of kernelised contextual bandits
Michal Valko, Nathan Korda, Rémi Munos, Ilias Flaounas, and Nello Cristianini · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Elad Hazan, Kfir Levy, and Shai Shalev-Shwartz · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network, 2015
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2016
Cited alongside, same era.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Later among the works it cites.
Improved optimistic algorithms for logistic bandits
Louis Faury, Marc Abeille, Clément Calauzènes, and Olivier Fercoq · 2020
Later among the works it cites.
Beyond UCB: Optimal and efficient contextual bandits with regression oracles
Dylan Foster and Alexander Rakhlin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Input convex neural networks
Brandon Amos, Lei Xu, and J Zico Kolter · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
Efficient regret minimization in non-convex games
Elad Hazan, Karan Singh, and Cyril Zhang · 2017
Cited alongside, same era.
Provably optimal algorithms for generalized linear contextual bandits
Lihong Li, Yu Lu, and Dengyong Zhou · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Dylan J Foster, Alexander Rakhlin, David Simchi-Levi, and Yunzong Xu · 2020
Later among the works it cites.
On the optimization landscape of tensor decompositions
Rong Ge and Tengyu Ma · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Information theoretic regret bounds for online nonlinear control
Sham Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi, and Wen Sun · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
A primer on zeroth-order optimization in signal processing and machine learning: Principals, recent advances, and applications
Sijia Liu, Pin-Yu Chen, Bhavya Kailkhura, Gaoyuan Zhang, Alfred O Hero III, and Pramod K Varshney · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Aditya Modi, Nan Jiang, Ambuj Tewari, and Satinder Singh · 2020
Later among the works it cites.
Efficient planning in large mdps with weak linear function approximation
Roshan Shariff and Csaba Szepesvári · 2020
Later among the works it cites.
David Simchi-Levi and Yunzong Xu · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael I Jordan · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Learning near optimal policies with low inherent Bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Neural contextual bandits with ucb-based exploration
Dongruo Zhou, Lihong Li, and Quanquan Gu · 2020
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
Online sparse reinforcement learning
Botao Hao, Tor Lattimore, Csaba Szepesvári, and Mengdi Wang · 2021
Closest in time.
Eluder dimension and generalized rank, 2021
Gene Li, Pritish Kamath, Dylan J. Foster, and Nathan Srebro · 2021
Closest in time.