Fetching the paper…
Reading the bibliography…
Transformer-based neural models are used in many AI applications.
P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,” Journal of Parallel and Distributed Computing , vol. 69, no. 2, pp. 117–124, 2009
2009
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Tech. Rep., 2009
2009
Earlier work this paper cites.
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, R. Soricut, L. Specia, and A. Tamchyna, “Findings of the 2014 workshop on statistical machine translation,” in Proceedings of the Ninth Workshop on Statistical Machine Translation . Baltimore, Maryland, USA: Association for Computational Linguistics, June 2014, pp. 12–58. [Online]. Available: http://www.aclweb.org/anthology/W/W14/W14-3302
2014
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. of NeurIPS , I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, Eds., 2017, pp. 5998–6008
2017
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings . OpenReview.net, 2017. [Online]. Available: https://openreview.net/forum?id=Byj72udxe
2017
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2018
2018
Earlier work this paper cites.
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Q. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy, “TVM: an automated end-to-end optimizing compiler for deep learning,” in PRoc. of OSDI , A. C. Arpaci-Dusseau and G. Voelker, Eds., 2018, pp. 578–594
2018
Earlier work this paper cites.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. G. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proc. of CVPR , 2018, pp. 2704–2713
2018
Earlier work this paper cites.
N. Wang, J. Choi, D. Brand, C. Chen, and K. Gopalakrishnan, “Training deep neural networks with 8-bit floating point numbers,” in Proc. of NeurIPS , S. Bengio, H. M. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., 2018, pp. 7686–7695
2018
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. F. Diamos, E. Elsen, D. García, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in Proc. of ICLR , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. of NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. G. Carbonell, R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” in Proc. of NeurIPS , H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alché-Buc, E. B. Fox, and R. Garnett, Eds., 2019, pp. 5754–5764
2019
Earlier work this paper cites.
L. Gong, D. He, Z. Li, T. Qin, L. Wang, and T. Liu, “Efficient training of BERT by progressively stacking,” in Proc. of ICML , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97, 2019, pp. 2337–2346
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
X. Sun, J. Choi, C. Chen, N. Wang, S. Venkataramani, V. Srinivasan, X. Cui, W. Zhang, and K. Gopalakrishnan, “Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks,” in Proc. of NeurIPS , H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alché-Buc, E. B. Fox, and R. Garnett, Eds., 2019, pp. 4901–4910
2019
Cited alongside, same era.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. [Online]. Available: https://openreview.net/forum?id=rJ4km2R5t7
2019
Cited alongside, same era.
2020
Later among the works it cites.
X. Zhang, S. Liu, R. Zhang, C. Liu, D. Huang, S. Zhou, J. Guo, Q. Guo, Z. Du, T. Zhi, and Y. Chen, “Fixed-point back-propagation training,” in Proc. of CVPR , 2020, pp. 2327–2335
2020
Later among the works it cites.
Z. Li, E. Wallace, S. Shen, K. Lin, K. Keutzer, D. Klein, and J. Gonzalez, “Train big, then compress: Rethinking model size for efficient training and inference of transformers,” in Proc. of ICML , ser. Proceedings of Machine Learning Research, vol. 119, 2020, pp. 5958–5968
2020
Later among the works it cites.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in Proc. of KDD , R. Gupta, Y. Liu, J. Tang, and B. A. Prakash, Eds., 2020, pp. 3505–3506
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in Proc. of NAACL-Demonstrations , 2019, pp. 48–53
2019
Cited alongside, same era.
2020
Cited alongside, same era.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Proc. of NeurIPS , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020
2020
Cited alongside, same era.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: memory optimizations toward training trillion parameter models,” in Proc. of SC , C. Cuicchi, I. Qualters, and W. T. Kramer, Eds., 2020, p. 20
2020
Cited alongside, same era.
S. Li, Y. Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, and S. Chintala, “Pytorch distributed: Experiences on accelerating data parallel training,” Proc. VLDB Endow. , vol. 13, no. 12, pp. 3005–3018, 2020
2020
Cited alongside, same era.
A. Vyas, A. Katharopoulos, and F. Fleuret, “Fast transformers with clustered attention,” in Proc. of NeurIPS , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Fan, E. Grave, and A. Joulin, “Reducing transformer depth on demand with structured dropout,” in Proc. of ICLR , 2020
2020
Cited alongside, same era.
M. Zhang and Y. He, “Accelerating training of transformer-based language models with progressive layer dropping,” in Proc. of NeurIPS , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020
2020
Cited alongside, same era.
Later among the works it cites.
2021
Closest in time.
X. Pan, L. Wu, M. Wang, and L. Li, “Contrastive learning for many-to-many multilingual neural machine translation,” in Proc. of ACL , 2021
2021
Closest in time.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. of ICLR , 2021
2021
Closest in time.
X. Wang, Y. Xiong, Y. Wei, M. Wang, and L. Li, “LightSeq: A high performance inference library for transformers,” in Proc. of NAACL-HLT: Industry Papers , 2021, pp. 113–120
2021
Closest in time.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient GPU serving system for transformer models,” in Proc. of PPoPP , J. Lee and E. Petrank, Eds., 2021, pp. 389–402
2021
Closest in time.
2021
Closest in time.
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong, “Random feature attention,” in Proc. of ICLR , 2021
2021
Closest in time.
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlós, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, D. Belanger, L. J. Colwell, and A. Weller, “Rethinking attention with performers,” in Proc. of ICLR , 2021
2021
Closest in time.
B. Li, Z. Wang, H. Liu, Q. Du, T. Xiao, C. Zhang, and J. Zhu, “Learning light-weight translation models from deep transformer,” in Proc. of AAAI , 2021, pp. 13 217–13 225
2021
Closest in time.