Fetching the paper…
Reading the bibliography…
Recent work in language modeling demonstrates that training large transformer models advances the state of the art in Natural Language Processing applications.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J. G., Le, Q. V., and Salakhutdinov, R · 1901
Earlier work this paper cites.
Multi-task deep neural networks for natural language understanding
Liu, X., He, P., Chen, W., and Gao, J · 1901
Earlier work this paper cites.
Defending against neural fake news
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y · 1905
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J. G., Salakhutdinov, R., and Le, Q. V · 1906
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
A bridging model for parallel computation
Valiant, L. G · 1990
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Turian, J., Ratinov, L., and Bengio, Y · 2010
Earlier work this paper cites.
Empirical evaluation and combination of advanced language modeling techniques
Mikolov, T., Deoras, A., Kombrink, S., Burget, L., and Černockỳ, J · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server, 2014
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation, 2014
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
context2vec: Learning generic context embedding with bidirectional lstm
Melamud, O., Goldberger, J., and Dagan, I · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
Unsupervised pretraining for sequence to sequence learning
Ramachandran, P., Liu, P. J., and Le, Q. V · 2016
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q. V., and Chen, Z · 2018
Later among the works it cites.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2018
Later among the works it cites.
Mixed precision training: Choosing a scaling factor, 2018
NVIDIA · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Later among the works it cites.
Improving language understanding by generative pre-training, 2018
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
On large- batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Race: Large-scale reading comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E · 2017
Cited alongside, same era.
Learned in translation: Contextualized word vectors
McCann, B., Bradbury, J., Xiong, C., and Socher, R · 2017
Cited alongside, same era.
Micikevicius, P., Narang, S., Alben, J., Diamos, G. F., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2017
Cited alongside, same era.
Learning to generate reviews and discovering sentiment
Radford, A., Józefowicz, R., and Sutskever, I · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P · 2018
Later among the works it cites.
Mesh-TensorFlow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B · 2018
Later among the works it cites.
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V · 2018
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., and Levy, O · 2019
Closest in time.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2019
Closest in time.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., and Soricut, P. S. R · 2019
Closest in time.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Closest in time.
Better language models and their implications, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Closest in time.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Closest in time.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., and Hsieh, C.-J · 2019
Closest in time.
Turing-nlg: A 17-billion-parameter language model by microsoft, 2020
Microsoft · 2020
Closest in time.