Fetching the paper…
Reading the bibliography…
This work studies training instabilities of behavior cloning with deep neural networks.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
Modern wiener-hopf design of optimal controllers–part ii: The multivariable case
Dante Youla, Hamid Jabr, and Jr Bongiorno · 1976
Earlier work this paper cites.
On the lower tail of gaussian seminorms
Jorgen Hoffmann-Jorgensen, Lawrence A Shepp, and Richard M Dudley · 1979
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Distributional and Lq norm inequalities for polynomials over convex bodies in rn
Anthony Carbery and James Wright · 2001
Earlier work this paper cites.
A Lyapunov approach to incremental stability properties
David Angeli · 2002
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
An introduction to hybrid dynamical systems , volume 251
Arjan J Van Der Schaft and Hans Schumacher · 2007
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
A simpler approach to obtaining an o (1/t) convergence rate for the projected stochastic subgradient method
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Brownian motion, martingales, and stochastic calculus
Jean-François Le Gall · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Dart: Noise injection for robust imitation learning
Michael Laskey, Jonathan Lee, Roy Fox, Anca Dragan, and Ken Goldberg · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
End-to-end driving via conditional imitation learning
Felipe Codevilla, Matthias Müller, Antonio López, Vladlen Koltun, and Alexey Dosovitskiy · 2018
Earlier work this paper cites.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Earlier work this paper cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Yurii Nesterov et al · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun, and Dawn Song · 2018
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Later among the works it cites.
Rvs: What is essential for offline rl via supervised learning?
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, and Sergey Levine · 2021
Later among the works it cites.
On imitation learning of linear control policies: Enforcing stability and robustness constraints via lmi conditions
Aaron Havens and Bin Hu · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of RL problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pipps: Flexible model-based policy search robust to the curse of chaos
Paavo Parmas, Carl Edward Rasmussen, Jan Peters, and Kenji Doya · 2018
Cited alongside, same era.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Cited alongside, same era.
The unusual effectiveness of averaging in gan training
Yasin Yaz, Chuan-Sheng Foo, Stefan Winkler, Kim-Hui Yap, Georgios Piliouras, Vijay Chandrasekhar, et al · 2018
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
Chiyuan Zhang, Oriol Vinyals, Remi Munos, and Samy Bengio · 2018
Cited alongside, same era.
Online control with adversarial disturbances
Naman Agarwal, Brian Bullins, Elad Hazan, Sham Kakade, and Karan Singh · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
On the ineffectiveness of variance reduced optimization for deep learning
Aaron Defazio and Léon Bottou · 2019
Cited alongside, same era.
Grasping with chopsticks: Combating covariate shift in model-free imitation learning for fine manipulation
Liyiming Ke, Jingqiang Wang, Tapomayukh Bhattacharjee, Byron Boots, and Siddhartha Srinivasa · 2021
Later among the works it cites.
Stabilizing dynamical systems via policy gradient methods
Juan Perdomo, Jack Umenberger, and Max Simchowitz · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Descending through a crowded valley-benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2021
Later among the works it cites.
The multiberts: Bert reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, et al · 2021
Later among the works it cites.
Instabilities of offline rl with pre-trained neural representation
Ruosong Wang, Yifan Wu, Ruslan Salakhutdinov, and Sham Kakade · 2021
Later among the works it cites.
On the stability of nonlinear receding horizon control: a geometric perspective
Tyler Westenbroek, Max Simchowitz, Michael I Jordan, and S Shankar Sastry · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch rl can be exponentially harder than online rl
Andrea Zanette · 2021
Later among the works it cites.
Kushal Arora, Layla El Asri, Hareesh Bahuleyan, and Jackie Chi Kit Cheung · 2022
Later among the works it cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Later among the works it cites.
Efficient and near-optimal smoothed online learning for generalized linear functions
Adam Block and Max Simchowitz · 2022
Later among the works it cites.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
Preference dynamics under personalized recommendations
Sarah Dean and Jamie Morgenstern · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Later among the works it cites.
Convex formulation of overparameterized deep neural networks
Cong Fang, Yihong Gu, Weizhong Zhang, and Tong Zhang · 2022
Later among the works it cites.
Implicit behavioral cloning
Pete Florence, Corey Lynch, Andy Zeng, Oscar A Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson · 2022
Later among the works it cites.
Introduction to online nonstochastic control
Elad Hazan and Karan Singh · 2022
Later among the works it cites.
Stop wasting my time! Saving days of ImageNet and BERT training with latest weight averaging
Jean Kaddour · 2022
Later among the works it cites.
On the sdes and scaling rules for adaptive gradient algorithms
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora · 2022
Later among the works it cites.
Tasil: Taylor series imitation learning
Daniel Pfrommer, Thomas Zhang, Stephen Tu, and Nikolai Matni · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
Nur Muhammad Shafiullah, Zichen Cui, Ariuntuya Arty Altanzaya, and Lerrel Pinto · 2022
Later among the works it cites.
Do differentiable simulators give better policy gradients?
Hyung Ju Suh, Max Simchowitz, Kaiqing Zhang, and Russ Tedrake · 2022
Later among the works it cites.
Feature selection with gradient descent on two-layer networks in low-rotation regimes
Matus Telgarsky · 2022
Later among the works it cites.
On the sample complexity of stability constrained imitation learning
Stephen Tu, Alexander Robey, Tingnan Zhang, and Nikolai Matni · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Later among the works it cites.
SGD with large step sizes learns sparse features
Maksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
Dan Busbridge, Jason Ramapuram, Pierre Ablin, Tatiana Likhomanenko, Eeshan Gunesh Dhekane, Xavier Suau, and Russ Webb · 2023
Closest in time.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song · 2023
Closest in time.
TinyStories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Closest in time.
No train no gain: Revisiting efficient training algorithms for Transformer-based language models
Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, and Matt J Kusner · 2023
Closest in time.
The power of learned locally linear models for nonlinear policy optimization
Daniel Pfrommer, Max Simchowitz, Tyler Westenbroek, Nikolai Matni, and Stephen Tu · 2023
Closest in time.
Training trajectories, mini-batch losses and the curious role of the learning rate
Mark Sandler, Andrey Zhmoginov, Max Vladymyrov, and Nolan Miller · 2023
Closest in time.
Understanding the effectiveness of early weight averaging for training large language models
Sunny Sanyal, Jean Kaddour, Abhishek Kumar, and Sujay Sanghavi · 2023
Closest in time.
Gymnasium, March 2023
Mark Towers, Jordan K. Terry, Ariel Kwiatkowski, John U. Balis, Gianluca de Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Arjun KG, Markus Krimmel, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Andrew Tan Jin Shen, and Omar G. Younis · 2023
Closest in time.