Fetching the paper…
Reading the bibliography…
Gating is a key feature in modern neural networks including LSTMs, GRUs and sparsely-gated deep neural networks.
Can sgd learn recurrent neural networks with provable generalization?
Allen-Zhu, Z. and Li, Y. (2019) · 1902
Earlier work this paper cites.
Convergence rates for gaussian mixtures of experts
Ho, N., Yang, C.-Y., and Jordan, M. I. (2019) · 1907
Earlier work this paper cites.
Note on n-dimensional hermite polynomials
Grad, H. (1949) · 1949
Earlier work this paper cites.
Learning one-hidden-layer neural networks under general input distributions
Gao, W., Makkuva, A. V., Oh, S., and Viswanath, P. (2019) · 1959
Earlier work this paper cites.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Stein, C. (1972) · 1972
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E. (1991) · 1991
Earlier work this paper cites.
Probability in Banach Spaces: isoperimetry and processes
Ledoux, M. and Talagrand, M. (1991) · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Jordan, M. I. and Jacobs, R. A. (1994) · 1994
Earlier work this paper cites.
Convergence results for the EM approach to mixtures of experts architectures
Jordan, M. I. and Xu, L. (1995) · 1995
Earlier work this paper cites.
The d-variate vector hermite polynomial of order k
Holmquist, B. (1996) · 1996
Earlier work this paper cites.
Weak convergence and empirical processes: with applications to statistics
Vaart, A. W. and Wellner, J. A. (1996) · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Mixtures of gaussian processes
Tresp, V. (2001) · 2001
Earlier work this paper cites.
A parallel mixture of SVMs for very large scale problems
Collobert, R., Bengio, S., and Bengio, Y. (2002) · 2002
Earlier work this paper cites.
Infinite mixtures of gaussian process experts
Rasmussen, C. E. and Ghahramani, Z. (2002) · 2002
Earlier work this paper cites.
Twenty years of mixture of experts
Yuksel, S. E., Wilson, J. N., and Gader, P. D. (2012) · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A. (2013) · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., rahman Mohamed, A., and Hinton, G. (2013) · 2013
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
Cho, K., van Merrienboer, B., Bahdanau, D., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Learning factored representations in a deep mixture of experts
Eigen, D., Ranzato, M., and Sutskever, I. (2014) · 2014
Cited alongside, same era.
Score function features for discriminative learning: Matrix and tensor framework
Janzamin, M., Sedghi, H., and Anandkumar, A. (2014) · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O. (2014) · 2014
Hard mixtures of experts for large scale weakly supervised vision
Gross, S., Ranzato, M., and Szlam, A. (2017) · 2017
Later among the works it cites.
Identity matters in deep learning
Hardt, M. and Ma, T. (2017) · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Li, Y. and Yuan, Y. (2017) · 2017
Later among the works it cites.
Convergence results for neural networks via electrodynamics
Panigrahy, R., Rahimi, A., Sachdeva, S., and Zhang, Q. (2017) · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mixture of experts: a literature survey
Masoudnia, S. and Ebrahimpour, R. (2014) · 2014
Cited alongside, same era.
Hierarchical mixture-of-experts model for large-scale gaussian process regression
Ng, J. W. and Deisenroth, M. P. (2014) · 2014
Cited alongside, same era.
Provable tensor methods for learning mixtures of classifiers
Sedghi, H., Janzamin, M., and Anandkumar, A. (2014) · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2014) · 2014
Cited alongside, same era.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y. (2015) · 2015
Cited alongside, same era.
Draw: A recurrent neural network for image generation
Gregor, K., Danihelka, I., Graves, A., Rezende, D. J., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S. (2017) · 2017
Later among the works it cites.
Linear model regression on time-series data: Non-asymptotic error bounds and applications
Alaeddini, A., Alemzadeh, S., Mesbahit, A., and Mesbahi, M. (2018) · 2018
Later among the works it cites.
On the convergence rate of training recurrent neural networks
Allen-Zhu, Z., Li, Y., and Song, Z. (2018) · 2018
Later among the works it cites.
Towards provable control for unknown linear dynamical systems
Arora, S., Hazan, E., Lee, H., Singh, K., Zhang, C., and Zhang, Y. (2018) · 2018
Later among the works it cites.
Safely learning to control the constrained linear quadratic regulator
Dean, S., Tu, S., Matni, N., and Recht, B. (2018) · 2018
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Ge, R., Lee, J. D., and Ma, T. (2018) · 2018
Later among the works it cites.
Gradient descent learns linear dynamical systems
Hardt, M., Ma, T., and Recht, B. (2018) · 2018
Later among the works it cites.
Towards binary-valued gates for robust lstm training
Li, Z., He, D., Tian, F., Chen, W., Qin, T., Wang, L., and Liu, T.-Y. (2018) · 2018
Later among the works it cites.
Robust spectral filtering and anomaly detection
Marecek, J. and Tchrakian, T. (2018) · 2018
Later among the works it cites.
Non-asymptotic identification of lti systems from a single trajectory
Oymak, S. and Ozay, N. (2018) · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Simchowitz, M., Mania, H., Tu, S., Jordan, M. I., and Recht, B. (2018) · 2018
Later among the works it cites.
Breaking the gridlock in mixture-of-experts: Consistent and efficient algorithms
Makkuva, A. V., Oh, S., Kannan, S., and Viswanath, P. (2019) · 2019
Closest in time.