Fetching the paper…
Reading the bibliography…
Transformers are one of the most important machine learning workloads today.
I/O complexity: The red-blue pebble game
Jia-Wei, H. and Kung, H.-T · 1981
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
High performance convolutional neural networks for document processing
Chellapilla, K., Puri, S., and Simard, P · 2006
Earlier work this paper cites.
Automatic transformations for communication-minimized parallelization and locality optimization in the polyhedral model
Bondhugula, U., Baskaran, M., Krishnamoorthy, S., Ramanujam, J., Rountev, A., and Sadayappan, P · 2008
Earlier work this paper cites.
Legion: Expressing locality and independence with logical regions
Bauer, M., Treichler, S., Slaughter, E., and Aiken, A · 2012
Earlier work this paper cites.
Pipelined back-propagation for context-dependent deep neural networks
Chen, X., Eversole, A., Li, G., Yu, D., and Seide, F · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., et al · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, A · 2012
Earlier work this paper cites.
Polly - performing polyhedral optimizations on a low-level intermediate representation
Grosser, T., Größlinger, A., and Lengauer, C · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Deep learning with COTS HPC systems
Coates, A., Huval, B., Wang, T., Wu, D., Catanzaro, B., and Andrew, N · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Kalchbrenner, N. and Blunsom, P · 2013
Earlier work this paper cites.
Fast training of convolutional networks through FFTs
Mathieu, M., Henaff, M., and LeCun, Y · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Mikolov, T., Yih, W.-t., and Zweig, G · 2013
Earlier work this paper cites.
Halide: A language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Ragan-Kelley, J., Barnes, C., Adams, A., Paris, S., Durand, F., and Amarasinghe, S · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
cuDNN: Efficient primitives for deep learning
Chetlur, S., Woolley, C., Vandermersch, P., Cohen, J., Tran, J., Catanzaro, B., and Shelhamer, E · 2014
Earlier work this paper cites.
Project Adam: Building an efficient and scalable deep learning training system
Chilimbi, T., Suzue, Y., Apacible, J., and Kalyanaraman, K · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
LBANN: Livermore big artificial neural network HPC toolkit
Van Essen, B., Kim, H., Pearce, R., Boakye, K., and Chen, B · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Fast algorithms for convolutional neural networks
Lavin, A. and Gray, S · 2016
Earlier work this paper cites.
Optimizing memory efficiency for deep convolutional neural networks on GPUs
Li, C., Yang, Y., Feng, M., Chakradhar, S., and Zhou, H · 2016
Earlier work this paper cites.
Latte: A language, compiler, and runtime for elegant and efficient deep neural networks
Truong, L., Barik, R., Totoni, E., Liu, H., Markley, C., Fox, A., and Shpeisman, T · 2016
Earlier work this paper cites.
Extremely large minibatch SGD: Training ResNet-50 on ImageNet in 15 minutes
Akiba, T., Suzuki, S., and Fukuda, K · 2017
Earlier work this paper cites.
Geometric deep learning: going beyond Euclidean data
Bronstein, M. M., Bruna, J., LeCun, Y., Szlam, A., and Vandergheynst, P · 2017
Earlier work this paper cites.
The reversible residual network: Backpropagation without storing activations
Gomez, A. N., Ren, M., Urtasun, R., and Grosse, R. B · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Earlier work this paper cites.
Polyhedral optimization of tensorflow computation graphs
Lethin, R · 2017
Earlier work this paper cites.
Device placement optimization with reinforcement learning
Mirhoseini, A., Pham, H., Le, Q. V., Steiner, B., Larsen, R., Zhou, Y., Kumar, N., Norouzi, M., Bengio, S., and Dean, J · 2017
Earlier work this paper cites.
NNVM compiler: Open compiler for AI frameworks, 2017
Paul G. Allen School of Computer Science & Engineering, University of Washington, Amazon Web Service AI team, and DMLC open-source community · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sabour, S., Frosst, N., and Hinton, G. E · 2017
Earlier work this paper cites.
Lift: A functional data-parallel ir for high-performance gpu code generation
Steuwer, M., Remmelg, T., and Dubach, C · 2017
Earlier work this paper cites.
Trends in Data Locality Abstractions for HPC Systems
Unat, D., Dubey, A., Hoefler, T., Shalf, J., Abraham, M., Bianco, M., Chamberlain, B. L., Cledat, R., Edwards, H. C., Finkel, H., Fuerlinger, K., Hannig, F., Jeannot, E., Kamil, A., Keasler, J., Kelly, P. H. J., Leung, V., Ltaief, H., Maruyama, N., Newburn, C. J., , and Pericas, M · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Cited alongside, same era.
TVM: An end-to-end optimization stack for deep learning
Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Shen, H., Cowan, M., Wang, L., Hu, Y., Ceze, L., Guestrin, C., and Krishnamurthy, A · 2018
Cited alongside, same era.
Intel nGraph: An intermediate representation, compiler, and executor for deep learning
Cyphers, S., Bansal, A. K., Bhiwandiwalla, A., Bobba, J., Brookhart, M., Chakraborty, A., Constable, W., Convey, C., Cook, L., Kanawi, O., et al · 2018
Cited alongside, same era.
Diesel: DSL for linear algebra and neural net computations on GPUs
Elango, V., Rubin, N., Ravishankar, M., Sandanagobalane, H., and Grover, V · 2018
Cited alongside, same era.
Microsoft invests in and partners with OpenAI to support us building beneficial AGI, 2019
OpenAI · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, H. F., Rae, J. W., Pascanu, R., Gulcehre, C., Jayakumar, S. M., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al · 2019
Later among the works it cites.
Stand-alone self-attention in vision models
Parmar, N., Ramachandran, P., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Compiling machine learning programs via high-level tracing
Frostig, R., Johnson, M. J., and Leary, C · 2018
Cited alongside, same era.
Near-global climate simulation at 1 km resolution: establishing a performance baseline on 4888 GPUs with COSMO 5.0
Fuhrer, O., Chadha, T., Hoefler, T., Kwasniewski, G., Lapillonne, X., Leutwyler, D., Lüthi, D., Osuna, C., Schär, C., Schulthess, T. C., et al · 2018
Cited alongside, same era.
Integrated model, batch, and domain parallelism in training neural networks
Gholami, A., Azad, A., Jin, P., Keutzer, K., and Buluc, A · 2018
Cited alongside, same era.
Pipe-SGD: A decentralized pipelined SGD framework for distributed deep net training
Li, Y., Yu, M., Li, S., Avestimehr, S., Kim, N. S., and Schwing, A · 2018
Cited alongside, same era.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al · 2018
Cited alongside, same era.
Massively distributed SGD: ImageNet/ResNet-50 training in a flash
Mikami, H., Suganuma, H., U-chupala, P., Tanaka, Y., and Kageyama, Y · 2018
Cited alongside, same era.
AI and compute, 2018
OpenAI · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
ZeRO: Memory optimization towards training a trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Later among the works it cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Later among the works it cites.
Megatron-LM: Training multi-billion parameter language models using GPU model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Later among the works it cites.
Astra: Exploiting predictability to optimize deep learning
Sivathanu, M., Chugh, T., Singapuram, S. S., and Zhou, L · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Strubell, E., Ganesh, A., and McCallum, A · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sukhbaatar, S., Grave, E., Bojanowski, P., and Joulin, A · 2019
Later among the works it cites.
Simple and effective curriculum pointer-generator networks for reading comprehension over long narratives
Tay, Y., Wang, S., Tuan, L. A., Fu, J., Phan, M. C., Yuan, X., Rao, J., Hui, S. C., and Zhang, A · 2019
Later among the works it cites.
SWIRL: High-performance many-core CPU code generation for deep neural networks
Venkat, A., Rusira, T., Barik, R., Hall, M., and Truong, L · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Later among the works it cites.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Later among the works it cites.
HuggingFace’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Later among the works it cites.
Transformer-transducer: End-to-end speech recognition with self-attention
Yeh, C.-F., Mahadeokar, J., Kalgaonkar, K., Wang, Y., Le, D., Jain, M., Schubert, K., Fuegen, C., and Seltzer, M. L · 2019
Later among the works it cites.
A data-centric approach to extreme-scale ab initio dissipative quantum transport simulations
Ziogas, A. N., Ben-Nun, T., Fernández, G. I., Schneider, T., Luisier, M., and Hoefler, T · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Closest in time.
Cerebras CS-1 Product Overview, 2020
Cerebras · 2020
Closest in time.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2020
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Closest in time.
Caffe2, 2020
Facebook · 2020
Closest in time.
XLA: Optimizing compiler for machine learning, 2020
Google · 2020
Closest in time.
Checkmate: Breaking the memory wall with optimal tensor rematerialization
Jain, P., Jain, A., Nrusimha, A., Gholami, A., Abbeel, P., Keutzer, K., Stoica, I., and Gonzalez, J. E · 2020
Closest in time.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Closest in time.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Closest in time.
MLIR: A compiler infrastructure for the end of moore’s law
Lattner, C., Pienaar, J., Amini, M., Bondhugula, U., Riddle, R., Cohen, A., Shpeisman, T., Davis, A., Vasilache, N., and Zinenko, O · 2020
Closest in time.
Li, Z., Wallace, E., Shen, S., Lin, K., Keutzer, K., Klein, D., and Gonzalez, J. E · 2020
Closest in time.
Lassen, 2020
Livermore Computing Center · 2020
Closest in time.
Molecule attention transformer
Maziarka, Ł., Danel, T., Mucha, S., Rataj, K., Tabor, J., and Jastrzębski, S · 2020
Closest in time.
ONNX Runtime, 2020
Microsoft · 2020
Closest in time.
TensorFloat-32 in the A100 GPU Accelerates AI Training, HPC up to 20x, 2020
Nvidia · 2020
Closest in time.
TorchScript, 2020
PyTorch Team · 2020
Closest in time.
A constructive prediction of the generalization error across scales
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 2020
Closest in time.
Tay, Y., Bahri, D., Yang, L., Metzler, D., and Juan, D.-C · 2020
Closest in time.
OpenAI’s massive GPT-3 model is impressive, but size isn’t everything, 2020
Wiggers, K · 2020
Closest in time.
Large batch optimization for deep learning: Training BERT in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Closest in time.