Transformer-XL: Attentive language models beyond a fixed-length context
Original
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. V. Le, and R. Salakhutdinov · 1901
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Minimization of functions having Lipschitz continuous first partial derivatives
L. Armijo · 1966
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Introduction to optimization. optimization software
B. T. Polyak · 1987
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
A. Beck and M. Teboulle · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
S. Ghadimi and G. Lan · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Original
R. Pascanu, T. Mikolov, and Y. Bengio · 2012
Earlier work this paper cites.
LSTM neural networks for language modeling
M. Sundermeyer, R. Schlüter, and H. Ney · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o ( 1 / n ) o(1/n)
F. Bach and E. Moulines · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Semi-stochastic gradient descent methods
Original
J. Konečnỳ and P. Richtárik · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.