Fetching the paper…
Reading the bibliography…
Many popular first-order optimization methods (e.g., Momentum, AdaGrad, Adam) accelerate the convergence rate of deep learning models.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak. 1964 · 1964
Earlier work this paper cites.
Finding frequent items in data streams. In Intl. Colloquium on Automata, Languages, and Programming
M. Charikar, K. Chen, and M. Farach-Colton. 2002 · 2002
Earlier work this paper cites.
An improved data stream summary: the count-min sketch and its applications
Graham Cormode and Shan Muthukrishnan. 2005 · 2005
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Hokusai-sketching streams in real time
Sergiy Matusevych, Alex Smola, and Amr Ahmed. 2012 · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning. In International conference on machine learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. 2013 · 2013
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Asymmetric LSH (ALSH) for sublinear time maximum inner product search (MIPS). In Advances in Neural Information Processing Systems
Anshumali Shrivastava and Ping Li. 2014 · 2014
Earlier work this paper cites.
Deep networks with large output spaces
Sudheendra Vijayanarasimhan, Jonathon Shlens, Rajat Monga, and Jay Yagnik. 2014 · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015 · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition
Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015 · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Cited alongside, same era.
Multiplicative LSTM for sequence modelling
Ben Krause, Liang Lu, Iain Murray, and Steve Renals. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Time adaptive sketches (ada-sketches) for summarizing data streams. In Proceedings of the 2016 International Conference on Management of Data
MISSION: Ultra Large-Scale Feature Selection using Count-Sketches. In Proceedings of the 35th International Conference on Machine Learning
Amirali Aghazadeh, Ryan Spring, Daniel Lejeune, Gautam Dasarathy, Anshumali Shrivastava, and richard baraniuk. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Extreme Classification in Log Memory
Qixuan Huang, Yiqiu Wang, Tharun Medini, and Anshumali Shrivastava. 2018 · 2018
Later among the works it cites.
Mixed Precision Training. In International Conference on Learning Representations
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anshumali Shrivastava, Arnd Christian Konig, and Mikhail Bilenko. 2016 · 2016
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry. 2017 · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Cited alongside, same era.
Scalable and sustainable deep learning via randomized hashing. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Ryan Spring and Anshumali Shrivastava. 2017 · 2017
Cited alongside, same era.
Hash Embeddings for Efficient Word Representations
Dan Tito Svenstrup, Jonas Hansen, and Ole Winther. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning intrinsic sparse structures within long short-term memory
Wei Wen, Yuxiong He, Samyam Rajbhandari, Minjia Zhang, Wenhan Wang, Fang Liu, Bin Hu, Yiran Chen, and Hai Li. 2017 · 2017
Cited alongside, same era.
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Later among the works it cites.
Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
Raul Puri, Robert Kirby, Nikolai Yakovenko, and Bryan Catanzaro. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost. In Proceedings of the 35th International Conference on Machine Learning
Noam Shazeer and Mitchell Stern. 2018 · 2018
Later among the works it cites.
Divide-and-conquer checkpointing for arbitrary programs with no user annotation
Jeffrey Mark Siskind and Barak A Pearlmutter. 2018 · 2018
Later among the works it cites.
Sketching Linear Classifiers over Data Streams. In Proceedings of the 2018 International Conference on Management of Data
Kai Sheng Tai, Vatsal Sharan, Peter Bailis, and Gregory Valiant. 2018 · 2018
Later among the works it cites.
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model. In International Conference on Learning Representations
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018 · 2018
Later among the works it cites.
Adaptive Methods for Nonconvex Optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar. 2018 · 2018
Later among the works it cites.