Fetching the paper…
Reading the bibliography…
Attention mechanisms, primarily designed to capture pairwise correlations between words, have become the backbone of machine learning, expanding beyond natural language processing into other domains.
Optimizing Supercompilers for Supercomputers
Michael Joseph Wolfe · 1982
Earlier work this paper cites.
Vector Register Allocation
Randy Allen and Ken Kennedy · 1992
Earlier work this paper cites.
Collective Loop Fusion for Array Contraction
Guang Gao, Russ Olsen, Vivek Sarkar, and Radhika Thekkath · 1992
Earlier work this paper cites.
Maximizing Loop Parallelism and Improving Data Locality via Loop Fusion and Distribution
Ken Kennedy and Kathryn S McKinley · 1993
Earlier work this paper cites.
Improving Effective Bandwidth Through Compiler Enhancement of Global Cache Reuse
Chen Ding and Ken Kennedy · 2004
Earlier work this paper cites.
Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
ShiDianNao: Shifting Vision Processing Closer to the Sensor
Zidong Du, Robert Fasthuber, Tianshi Chen, Paolo Ienne, Ling Li, Tao Luo, Xiaobing Feng, Yunji Chen, and Olivier Temam · 2015
Earlier work this paper cites.
Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks
Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong · 2015
Earlier work this paper cites.
Fused-layer CNN Accelerators
Manoj Alwani, Han Chen, Michael Ferdman, and Peter Milder · 2016
Earlier work this paper cites.
Eyeriss: An Energy-efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks
Yu-Hsin Chen et al · 2016
Earlier work this paper cites.
Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks
Chen, Yu-Hsin and Krishna, Tushar and Emer, Joel and Sze, Vivienne · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory
Mingyu Gao, Jing Pu, Xuan Yang, Mark Horowitz, and Christos Kozyrakis · 2017
Earlier work this paper cites.
In-Datacenter Performance Analysis of a Tensor Processing Unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Earlier work this paper cites.
The Tensor Algebra Compiler
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amarasinghe · 2017
Earlier work this paper cites.
FlexFlow: A Flexible Dataflow Accelerator Architecture for Convolutional Neural Networks
Wenyan Lu et al · 2017
Earlier work this paper cites.
NVDLA Deep Learning Accelerator
Nvidia · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
SCALEDEEP: A Scalable Compute Architecture for Learning and Evaluating Deep Networks
Swagath Venkataramani, Ashish Ranjan, Subarno Banerjee, Dipankar Das, Sasikanth Avancha, Ashok Jagannathan, Ajaya Durg, Dheemanth Nagaraj, Bharat Kaul, Pradeep Dubey, and Anand Raghunathan · 2017
Earlier work this paper cites.
Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAs
Xuechao Wei, Cody Hao Yu, Peng Zhang, Youxiang Chen, Yuxin Wang, Han Hu, Yun Liang, and Jason Cong · 2017
Earlier work this paper cites.
Learning to Optimize Tensor Programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Music Transformer: Generating Music with Long-Term Structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck · 2018
Earlier work this paper cites.
Beyond Data and Model Parallelism for Deep Neural Networks
Zhihao Jia, Matei Zaharia, and Alex Aiken · 2018
Earlier work this paper cites.
MAERI: Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects
Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna · 2018
Earlier work this paper cites.
Generating Wikipedia by Summarizing Long Sequences
Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer · 2018
Earlier work this paper cites.
Image Transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
SCALE-Sim: Systolic CNN Accelerator Simulator
Ananda Samajdar, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna · 2018
Earlier work this paper cites.
Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang · 2018
Earlier work this paper cites.
Nvidia T4 Tensor Core GPU
Tesla, Nvidia · 2018
Earlier work this paper cites.
V100 GPU Architecture
Tesla, Nvidia · 2018
Earlier work this paper cites.
TGPA: Tile-Grained Pipeline Architecture for Low Latency CNN Inference
Xuechao Wei, Yun Liang, Xiuhong Li, Cody Hao Yu, Peng Zhang, and Jason Cong · 2018
Earlier work this paper cites.
GANAX: A Unified MIMD-SIMD Acceleration for Generative Adversarial Networks
Amir Yazdanbakhsh, Kambiz Samadi, Nam Sung Kim, and Hadi Esmaeilzadeh · 2018
Earlier work this paper cites.
Tiramisu: A Polyhedral Compiler for Expressing Fast and Portable Code
Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Abdurrahman Akkas, Yunming Zhang, Patricia Suriana, Shoaib Kamil, and Saman Amarasinghe · 2019
Earlier work this paper cites.
Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices
Yu-Hsin Chen, Tien-Ju Yang, Joel Emer, and Vivienne Sze · 2019
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
UniProt: A Worldwide Hub of Protein Knowledge
UniProt Consortium · 2019
Earlier work this paper cites.
Adaptively Sparse Transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 2019
Earlier work this paper cites.
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
dMazeRunner: Executing Perfectly Nested Loops on Dataflow Accelerators
Shail Dave, Youngbin Kim, Sasikanth Avancha, Kyoungwoo Lee, and Aviral Shrivastava · 2019
Cited alongside, same era.
TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators
Mingyu Gao, Xuan Yang, Jing Pu, Mark Horowitz, and Christos Kozyrakis · 2019
Cited alongside, same era.
Reweighted Proximal Pruning for Large-scale Language Representation
Fu-Ming Guo, Sijia Liu, Finlay S Mungall, Xue Lin, and Yanzhi Wang · 2019
Cited alongside, same era.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
TransTrack: Multiple Object Tracking with Transformer
Peize Sun, Yi Jiang, Rufeng Zhang, Enze Xie, Jinkun Cao, Xinting Hu, Tao Kong, Zehuan Yuan, Changhu Wang, and Ping Luo · 2020
Later among the works it cites.
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2020
Later among the works it cites.
Efficient Processing of Deep Neural Networks
Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S. Emer · 2020
Later among the works it cites.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding Reuse, Performance, and Hardware Cost of DNN Dataflow: A Data-Centric Approach
Hyoukjun Kwon, Prasanth Chatarasi, Michael Pellauer, Angshuman Parashar, Vivek Sarkar, and Tushar Krishna · 2019
Cited alongside, same era.
Cross-Lingual Language Model Pretraining
Guillaume Lample and Alexis Conneau · 2019
Cited alongside, same era.
Flaubert: Unsupervised Language Model Pre-training for French
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoît Crabbé, Laurent Besacier, and Didier Schwab · 2019
Cited alongside, same era.
Timeloop: A Systematic Approach to DNN Accelerator Evaluation
Angshuman Parashar, Priyanka Raina, Yakun Sophia Shao, Yu-Hsin Chen, Victor A Ying, Anurag Mukkara, Rangharajan Venkatesan, Brucek Khailany, Stephen W Keckler, and Joel Emer · 2019
Cited alongside, same era.
Blockwise Self-Attention for Long Document Understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang · 2019
Cited alongside, same era.
Compressive Transformers for Long-Range Sequence Modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2019
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
Later among the works it cites.
Linformer: Self-Attention with Linear Complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2020
Later among the works it cites.
Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators
Xuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter, Jing Pu, Ankita Nayak, Steven Bell, Kaidi Cao, Heonjae Ha, Priyanka Raina, Christos Kozyrakis, and Mark Horowitz · 2020
Later among the works it cites.
TernaryBERT: Distillation-aware Ultra-low Bit BERT
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu · 2020
Later among the works it cites.
Self-attention based context-aware 3d object detection
Prarthana Bhattacharyya, Chengjie Huang, and Krzysztof Czarnecki · 2021
Closest in time.
TensorFlow XLA
Google · 2021
Closest in time.
LeViT: a Vision Transformer in ConvNet’s Clothing for Faster Inference
Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze · 2021
Closest in time.
ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee · 2021
Closest in time.
Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search Extended Abstract
Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, and Christopher W Fletcher · 2021
Closest in time.
Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs
Wen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, and Yi-Hsuan Yang · 2021
Closest in time.
CoSA: Scheduling by Constrained Optimization for Spatial Accelerators
Qijing Huang, Minwoo Kang, Grace Dinh, Thomas Norell, Aravind Kalaiah, James Demmel, John Wawrzynek, and Yakun Sophia Shao · 2021
Closest in time.
Data Movement Is All You Need: A Case Study on Optimizing Transformers
Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler · 2021
Closest in time.
I-BERT: Integer-only BERT Quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2021
Closest in time.
GCNAX: A Flexible and Energy-Efficient Accelerator for Graph Convolutional Neural Networks
Jiajun Li, Ahmed Louri, Avinash Karanth, and Razvan Bunescu · 2021
Closest in time.
Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang · 2021
Closest in time.
Boosting NVIDIA MLPerf Training v1.1 Performance with Full Stack Optimization
Vinh Nguyen, Sukru Burc Eryilmax, Karthik Mandakolathur, and Shar Narasimhan · 2021
Closest in time.
DNNFusion: Accelerating Deep Neural Networks Execution with Advanced Operator Fusion
Wei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal, and Bin Ren · 2021
Closest in time.
FasterTransforemr
Nvidia · 2021
Closest in time.
Self-attention Does Not Need O ( n 2 ) O(n^{2}) Memory
Markus N. Rabe and Charles Staats · 2021
Closest in time.
Efficient Content-Based Sparse Attention with Routing Transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Closest in time.
SM6: A 16nm System-on-Chip for Accurate and Noise-Robust Attention-Based NLP Applications
Thierry Tambe, En-Yu Yang, Glenn G Ko, Yuji Chai, Coleman Hooper, Marco Donato, Paul N Whatmough, Alexander M Rush, David Brooks, and Gu-Yeon Wei · 2021
Closest in time.
Long Range Arena: A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Closest in time.
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
Hanrui Wang, Zhekai Zhang, and Song Han · 2021
Closest in time.
CvT: Introducing Convolutions to Vision Transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Closest in time.
Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao · 2021
Closest in time.
Long-Short Transformer: Efficient Transformers for Language and Vision
Chen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi, Tom Goldstein, Anima Anandkumar, and Bryan Catanzaro · 2021
Closest in time.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2021
Closest in time.
Data-Driven Offline Optimization for Architecting Hardware Accelerators
Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky, and Sergey Levine · 2022
Closest in time.
Accelerating Attention through Gradient-Based Learned Runtime Pruning
Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, and Mingu Kang · 2022
Closest in time.
An Evaluation of Edge TPU Accelerators for Convolutional Neural Networks
Kiran Seshadri, Berkin Akin, James Laudon, Ravi Narayanaswami, and Amir Yazdanbakhsh · 2022
Closest in time.
Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation
Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, and Mingu Kang · 2022
Closest in time.