Fetching the paper…
Reading the bibliography…
Our goal is to learn control policies for robots that provably generalize well to novel environments given a dataset of example environments.
Chance-constrained programming
A. Charnes and W. W. Cooper · 1959
Earlier work this paper cites.
Asymptotic evaluation of certain Markov process expectations for large time
M. D. Donsker and S. R. S. Varadhan · 1975
Earlier work this paper cites.
A Course in H-Infinity Control Theory
B. A. Francis · 1987
Earlier work this paper cites.
Interior-point polynomial algorithms in convex programming , volume 13
Y. Nesterov and A. Nemirovskii · 1994
Earlier work this paper cites.
Sample compression, learnability, and the Vapnik-Chervonenkis dimension
S. Floyd and M. Warmuth · 1995
Earlier work this paper cites.
Behavior-Based Robotics
R. C. Arkin · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Algorithmic stability and sanity-check bounds for leave-one-out cross-validation
M. Kearns and D. Ron · 1999
Earlier work this paper cites.
Some PAC-Bayesian theorems
D. A. McAllester · 1999
Earlier work this paper cites.
Approximate planning in large POMDPs via reusable trajectories
M. J. Kearns, Y. Mansour, and A. Y. Ng · 2000
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
J. A. Bagnell and J. G. Schneider · 2001
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
R-max – A general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
(Not) bounding the true error
J. Langford and R. Caruana · 2002
Earlier work this paper cites.
PAC-Bayesian generalisation error bounds for gaussian process classification
M. Seeger · 2002
Earlier work this paper cites.
PAC-Bayes & margins
J. Langford and J. Shawe-Taylor · 2003
Earlier work this paper cites.
Learning Decisions: Robustness, Uncertainty, and Approximation
J. A. Bagnell · 2004
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
A note on the PAC Bayesian theorem
A. Maurer · 2004
Earlier work this paper cites.
Receding horizon path planning with implicit safety guarantees
T. Schouwenaars, J. How, and E. Feron · 2004
Earlier work this paper cites.
Tutorial on practical prediction theory for classification
J. Langford · 2005
Earlier work this paper cites.
A probabilistic approach to optimal robust path planning with obstacles
L. Blackmore, H. Li, and B. Williams · 2006
Earlier work this paper cites.
A short paper about motion safety
T. Fraichard · 2007
Earlier work this paper cites.
PAC-Bayes bounds for the risk of the majority vote and the variance of the Gibbs classifier
A. Lacasse, F. Laviolette, M. Marchand, P. Germain, and N. Usunier · 2007
Earlier work this paper cites.
Vision-based control of near-obstacle flight
A. Beyeler, J.-C. Zufferey, and D. Floreano · 2009
Cited alongside, same era.
Implementation of wide-field integration of optic flow for autonomous quadrotor navigation
J. Conroy, G. Gremillion, B. Ranganathan, and J. S. Humbert · 2009
Cited alongside, same era.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
PAC-Bayesian learning of linear classifiers
P. Germain, A. Lacasse, F. Laviolette, and M. Marchand · 2009
Cited alongside, same era.
PAC-Bayesian model selection for reinforcement learning
M. M. Fard and J. Pineau · 2010
Cited alongside, same era.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
X. Nguyen, M. J. Wainwright, and M. I. Jordan · 2010
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Later among the works it cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Later among the works it cites.
CVXPY: A Python-embedded modeling language for convex optimization
S. Diamond and S. Boyd · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Later among the works it cites.
Conic optimization via operator splitting and homogeneous self-dual embedding
B. O’Donoghue, E. Chu, N. Parikh, and S. Boyd · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Entropy and Information Theory
R. M. Gray · 2011
Cited alongside, same era.
Closed-loop belief space planning for linear, Gaussian systems
M. P. Vitus and C. J. Tomlin · 2011
Cited alongside, same era.
PAC-Bayesian policy evaluation for reinforcement learning
M. M. Fard, J. Pineau, and C. Szepesvári · 2012
Cited alongside, same era.
ECOS: An SOCP solver for embedded systems
A. Domahidi, E. Chu, and S. Boyd · 2013
Cited alongside, same era.
A PAC-Bayesian tutorial with a dropout bound
D. McAllester · 2013
Cited alongside, same era.
Learning monocular reactive uav control in cluttered natural environments
S. Ross, N. Melik-Barkhudarov, K. S. Shankar, A. Wendel, D. Dey, J. A. Bagnell, and M. Hebert · 2013
Cited alongside, same era.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
Relative entropy optimization and its applications
V. Chandrasekaran and P. Shah · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
J. Mahler, J. Liang, S. Niyaz, M. Laskey, R. Doan, X. Liu, J. A. Ojea, and K. Goldberg · 2017
Later among the works it cites.
Funnel libraries for real-time robust feedback motion planning
A. Majumdar and R. Tedrake · 2017
Later among the works it cites.
A unified view of entropy-regularized markov decision processes
G. Neu, A. Jonsson, and V. Gómez · 2017
Later among the works it cites.
SCS: Splitting conic solver, version 2.0.2
B. O’Donoghue, E. Chu, N. Parikh, and S. Boyd · 2017
Later among the works it cites.
Safe visual navigation via deep learning and novelty detection
C. Richter and N. Roy · 2017
Later among the works it cites.
Domain randomization and generative models for robotic grasping
J. Tobin, W. Zaremba, and P. Abbeel · 2017
Later among the works it cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi · 2017
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Closest in time.
Understanding batch normalization
N. Bjorck, C. P. Gomes, B. Selman, and K. Q. Weinberger · 2018
Closest in time.
Pybullet, a python module for physics simulation for games, robotics and machine learning, 2018
E. Coumans and Y. Bai · 2018
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Closest in time.
PAC-Bayes Control: synthesizing controllers that provably generalize to novel environments
A. Majumdar and M. Goldstein · 2018
Closest in time.
The limits and potentials of deep learning for robotics
N. Sünderhauf, O. Brock, W. Scheirer, R. Hadsell, D. Fox, J. Leitner, B. Upcroft, P. Abbeel, W. Burgard, M. Milford, and P. Corke · 2018
Closest in time.
Pyparrot 1.5.21, 2019
A. McGovern · 2019
Closest in time.
Mosek fusion api for python 9.0.84(beta), 2019
MOSEK ApS · 2019
Closest in time.
Mathematica, Version 12.0, 2019
Wolfram Research, Inc · 2019
Closest in time.