Fetching the paper…
Reading the bibliography…
We propose a theoretical framework for studying behavior cloning of complex expert demonstrations using generative modeling.
The existence of probability measures with given marginals
V. Strassen · 1965
Earlier work this paper cites.
Differential dynamic programming. number 24, 1970
D. H. Jacobson and D. Q. Mayne · 1970
Earlier work this paper cites.
A lyapunov approach to incremental stability properties
D. Angeli · 2002
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Generative modeling with denoising auto-encoders and langevin sampling
A. Block, Y. Mroueh, and A. Rakhlin · 2002
Earlier work this paper cites.
Nonlinear Systems
H. Khalil · 2002
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
A. Hyvärinen and P. Dayan · 2005
Earlier work this paper cites.
Fast mixing of multi-scale langevin dynamics under the manifold hypothesis
A. Block, Y. Mroueh, A. Rakhlin, and J. Ross · 2006
Earlier work this paper cites.
Recovering a function from a dini derivative
J. W. Hagood and B. S. Thomson · 2006
Earlier work this paper cites.
Optimal control: linear quadratic methods
B. D. Anderson and J. B. Moore · 2007
Earlier work this paper cites.
Information, physics, and computation
M. Mezard and A. Montanari · 2009
Earlier work this paper cites.
Lqr-trees: Feedback motion planning on sparse randomized trees
R. Tedrake · 2009
Earlier work this paper cites.
Optimal transport: old and new , volume 338
C. Villani et al · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2010
Earlier work this paper cites.
Hierarchical task and motion planning in the now
L. P. Kaelbling and T. Lozano-Pérez · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
P. Vincent · 2011
Earlier work this paper cites.
Probability in high dimension
R. Van Handel · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Earlier work this paper cites.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Earlier work this paper cites.
A vector-contraction inequality for rademacher complexities
A. Maurer · 2016
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Imitation learning: A survey of learning methods
A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne · 2017
Earlier work this paper cites.
Dart: Noise injection for robust imitation learning
M. Laskey, J. Lee, R. Fox, A. Dragan, and K. Goldberg · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Cited alongside, same era.
Empirical entropy, minimax regret and minimax risk
A. Rakhlin, K. Sridharan, and A. B. Tsybakov · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst
M. Bansal, A. Krizhevsky, and A. Ogale · 2018
Cited alongside, same era.
Deep imitation learning for 3d navigation tasks
A. Hussein, E. Elyan, M. M. Gaber, and C. Jayne · 2018
On imitation learning of linear control policies: Enforcing stability and robustness constraints via lmi conditions
A. Havens and B. Hu · 2021
Later among the works it cites.
Grasping with chopsticks: Combating covariate shift in model-free imitation learning for fine manipulation
L. Ke, J. Wang, T. Bhattacharjee, B. Boots, and S. Srinivasa · 2021
Later among the works it cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Later among the works it cites.
Shortest paths in graphs of convex sets
T. Marcucci, J. Umenberger, P. A. Parrilo, and R. Tedrake · 2021
Later among the works it cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Cited alongside, same era.
Group normalization
Y. Wu and K. He · 2018
Cited alongside, same era.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Cited alongside, same era.
Pairwise optimal coupling of multiple random variables
O. Angel and Y. Spinka · 2019
Cited alongside, same era.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Cited alongside, same era.
Later among the works it cites.
Topics in optimal transportation , volume 58
C. Villani · 2021
Later among the works it cites.
On the stability of nonlinear receding horizon control: a geometric perspective
T. Westenbroek, M. Simchowitz, M. I. Jordan, and S. S. Sastry · 2021
Later among the works it cites.
Privacy of noisy stochastic gradient descent: More iterations without more privacy loss
J. Altschuler and K. Talwar · 2022
Later among the works it cites.
Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions
S. Chen, S. Chewi, J. Li, Y. Li, A. Salim, and A. Zhang · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine · 2022
Later among the works it cites.
Tasil: Taylor series imitation learning
D. Pfrommer, T. Zhang, S. Tu, and N. Matni · 2022
Later among the works it cites.
Information Theory: From Coding to Learning
Y. Polyanskiy and Y. Wu · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Later among the works it cites.
On the sample complexity of stability constrained imitation learning
S. Tu, A. Robey, T. Zhang, and N. Matni · 2022
Later among the works it cites.
Compositional foundation models for hierarchical planning
A. Ajay, S. Han, Y. Du, S. Li, A. Gupta, T. Jaakkola, J. Tenenbaum, L. Kaelbling, A. Srivastava, and P. Agrawal · 2023
Closest in time.
Faster high-accuracy log-concave sampling via algorithmic warm starts
J. M. Altschuler and S. Chewi · 2023
Closest in time.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Closest in time.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
P. Hansen-Estruch, I. Kostrikov, M. Janner, J. G. Kuba, and S. Levine · 2023
Closest in time.
Convergence of score-based generative modeling for general data distributions
H. Lee, J. Lu, and Y. Tan · 2023
Closest in time.
Imitating human behaviour with diffusion models
T. Pearce, T. Rashid, A. Kanervisto, D. Bignell, M. Sun, R. Georgescu, S. V. Macua, S. Z. Tan, I. Momennejad, K. Hofmann, et al · 2023
Closest in time.
The power of learned locally linear models for nonlinear policy optimization
D. Pfrommer, M. Simchowitz, T. Westenbroek, N. Matni, and S. Tu · 2023
Closest in time.
Mega-dagger: Imitation learning with multiple imperfect experts
X. Sun, S. Yang, and R. Mangharam · 2023
Closest in time.
Decision stacks: Flexible reinforcement learning via modular generative models
S. Zhao and A. Grover · 2023
Closest in time.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Closest in time.