Fetching the paper…
Reading the bibliography…
A commonly cited inefficiency of neural network training by back-propagation is the update locking problem: each layer must wait for the signal to propagate through the full network before updating.
Cybernetic Predicting Devices. CCM Information Corporation. , 1965
Ivakhnenko, A. G. and Lapa, V. G · 1965
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Distributed optimization of deeply nested systems
Carreira-Perpinan, M. and Wang, W · 2014
Earlier work this paper cites.
Lee, D., Zhang, S., Biard, A., and Bengio, Y · 2014
Earlier work this paper cites.
Random feedback weights support learning in deep neural networks
Lillicrap, T. P., Cownden, D., Tweed, D. B., and Akerman, C. J · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Deeply-supervised nets
Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Zhang, S., Choromanska, A. E., and LeCun, Y · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Direct feedback alignment provides learning in deep neural networks
Nø kland, A · 2016
Cited alongside, same era.
Training neural networks without gradients: A scalable admm approach
Taylor, G., Burmeister, R., Xu, Z., Singh, B., Patel, A., and Goldstein, T · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Understanding synthetic gradients and decoupled neural interfaces
Czarnecki, W. M., Swirszcz, G., Jaderberg, M., Osindero, S., Vinyals, O., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Beyond backprop: Alternating minimization with co-activation memory
Choromanska, A., Kumaravel, S., Luss, R., Rish, I., Kingsbury, B., Tejwani, R., and Bouneffouf, D · 2018
Later among the works it cites.
Latent space policies for hierarchical reinforcement learning
Haarnoja, T., Hartikainen, K., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Leblond, R., Pedregosa, F., and Lacoste-Julien, S · 2018
Later among the works it cites.
Deep cascade learning
Marquez, E. S., Hare, J. S., and Niranjan, M · 2018
Later among the works it cites.
Deep supervised learning using local errors
Mostafa, H., Ramesh, V., and Cauwenberghs, G · 2018
Later among the works it cites.
Compressing the input for cnns with the first-order scattering transform
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Decoupled neural interfaces using synthetic gradients
Jaderberg, M., Czarnecki, W. M., Osindero, S., Vinyals, O., Graves, A., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Asaga: Asynchronous parallel saga
Leblond, R., Pedregosa, F., and Lacoste-Julien, S · 2017
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Lian, X., Zhang, W., Zhang, C., and Liu, J · 2017
Cited alongside, same era.
Building a regular decision boundary with deep networks
Oyallon, E · 2017
Cited alongside, same era.
Assessing the scalability of biologically-motivated deep learning algorithms and architectures
Bartunov, S., Santoro, A., Richards, B., Marris, L., Hinton, G. E., and Lillicrap, T · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Cited alongside, same era.
Learning deep resnet blocks sequentially using boosting theory
Huang, F., Ash, J., Langford, J., and Schapire, R
Cited in the paper.
Oyallon, E., Belilovsky, E., Zagoruyko, S., and Valko, M · 2018
Later among the works it cites.
Manifold mixup: Learning better representations by interpolating hidden states
Verma, V., Lamb, A., Beckham, C., Najafi, A., Courville, A., Mitliagkas, I., and Bengio, Y · 2018
Later among the works it cites.
Greedy layerwise learning can scale to imagenet
Belilovsky, E., Eickenberg, M., and Oyallon, E · 2019
Closest in time.
Training neural networks with local error signals
Nøkland, A. and Eidnes, L. H · 2019
Closest in time.
Biologically-plausible learning algorithms can scale to large datasets
Xiao, W., Chen, H., Liao, Q., and Poggio, T. A · 2019
Closest in time.
Decoupled parallel backpropagation with convergence guarantee
Huo, Z., Gu, B., qian Yang, and Huang, H · 2098
Closest in time.