Fetching the paper…
Reading the bibliography…
The attention mechanism is becoming increasingly popular in Natural Language Processing (NLP) applications, showing superior performance than convolutional and recurrent architectures.
M. P. Marcus et al. , “Building a large annotated corpus of English: The Penn Treebank,” Computational Linguistics , 1993
1993
Earlier work this paper cites.
D. E. Knuth, The Art of Computer Programming, Volume 3: (2nd Ed.) Sorting and Searching . Addison Wesley Longman Publishing Co., Inc., 1998
1998
Earlier work this paper cites.
E. F. Tjong Kim Sang et al. , “Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition,” in HLT-NAACL , 2003
2003
Earlier work this paper cites.
S. Sukhbaatar et al. , “End-to-end memory networks,” in NeurIPS , 2015
2012
Earlier work this paper cites.
C. Chelba et al. , “One billion word benchmark for measuring progress in statistical language modeling,” arXiv , 2013
2013
Earlier work this paper cites.
R. Socher et al. , “Recursive deep models for semantic compositionality over a sentiment treebank,” in EMNLP , 2013
2013
Earlier work this paper cites.
P. Nilsson et al. , “Hardware implementation of the exponential function using taylor series,” in NORCHIP , 2014
2014
Earlier work this paper cites.
Y. Kim et al. , “Ramulator: A fast and extensible dram simulator,” IEEE Computer architecture letters , 2015
2015
Earlier work this paper cites.
S. Salehi et al. , “Energy and area analysis of a floating-point unit in 15nm cmos process technology,” in SoutheastCon , 2015
2015
Earlier work this paper cites.
N. Muralimanohar et al. , “Cacti 6.0: A tool to model large caches,” HPL , 2015
2015
Earlier work this paper cites.
P. Rajpurkar et al. , “Squad: 100,000+ questions for machine comprehension of text,” EMNLP , 2016
2016
Earlier work this paper cites.
S. Merity et al. , “Pointer sentinel mixture models,” arXiv , 2016
2016
Earlier work this paper cites.
S. Han et al. , “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” in ICLR , 2016
2016
Earlier work this paper cites.
J. Park et al. , “Faster cnns with direct sparse convolutions and guided pruning,” arXiv , 2016
2016
Earlier work this paper cites.
S. Han et al. , “Eie: Efficient inference engine on compressed deep neural network,” in ISCA , 2016
2016
Earlier work this paper cites.
S. Zhang et al. , “Cambricon-x: An accelerator for sparse neural networks,” in MICRO , 2016
2016
Earlier work this paper cites.
P. Judd et al. , “Stripes: Bit-serial deep neural network computing,” in MICRO , 2016
2016
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
M. O’Connor et al. , “Fine-grained dram: energy-efficient dram for extreme bandwidth systems,” in MICRO , 2017
2017
Earlier work this paper cites.
Y. He, X. Zhang, and J. Sun, “Channel pruning for accelerating very deep neural networks,” ICCV , 2017
2017
Earlier work this paper cites.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” ACM SIGARCH Computer Architecture News , vol. 45, no. 2, pp. 27–40, 2017
2017
Earlier work this paper cites.
T. Rzayev et al. , “Deeprecon: Dynamically reconfigurable architecture for accelerating deep neural networks,” in IJCNN , 2017
2017
Earlier work this paper cites.
N. P. Jouppi et al. , “In-datacenter performance analysis of a tensor processing unit,” in ISCA , 2017
2017
Earlier work this paper cites.
J. Devlin et al. , “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv , 2018
2018
Earlier work this paper cites.
A. Wang et al. , “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in EMNLP Workshop BlackboxNLP , 2018
2018
Earlier work this paper cites.
S. Pal et al. , “Outerspace: An outer product based sparse matrix multiplication accelerator,” in HPCA , 2018
2018
Earlier work this paper cites.
H. Wang, J. Yang, H.-S. Lee, and S. Han, “Learning to design circuits,” NeurIPS MLSys Workshop , 2018
2018
Earlier work this paper cites.
Y. He et al. , “Amc: Automl for model compression and acceleration on mobile devices,” in ECCV , 2018
2018
Earlier work this paper cites.
X. Zhou et al. , “Cambricon-s: Addressing irregularity in sparse neural networks through a cooperative software/hardware approach,” in MICRO , 2018
2018
Earlier work this paper cites.
J. Cong et al. , “Understanding performance differences of fpgas and gpus,” in FCCM , 2018
2018
Earlier work this paper cites.
J. Lee et al. , “Unpu: A 50.6 tops/w unified deep neural network accelerator with 1b-to-16b fully-variable weight bit-precision,” in ISSCC , 2018
2018
Earlier work this paper cites.
S. Sharify et al. , “Loom: Exploiting weight and activation precisions to accelerate convolutional neural networks,” in DAC , 2018
2018
Earlier work this paper cites.
Y. Umuroglu et al. , “Bismo: A scalable bit-serial matrix multiplication overlay for reconfigurable computing,” in FPL , 2018
2018
Earlier work this paper cites.
H. Sharma et al. , “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,” in ISCA , 2018
2018
Earlier work this paper cites.
Nvidia, “Nvidia tensor cores,” in Nvidia , 2018
2018
Cited alongside, same era.
J. Gu et al. , “Non-autoregressive neural machine translation,” in ICLR , 2018
2018
Cited alongside, same era.
A. Radford et al. , “Language models are unsupervised multitask learners,” OpenAI Blog , 2019
2019
Cited alongside, same era.
E. Voita et al. , “Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,” ACL , 2019
2019
Cited alongside, same era.
H. Jang et al. , “Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,” in ISCA , 2019
2019
Cited alongside, same era.
H. Mao et al. , “Park: An open platform for learning-augmented computer systems,” NeurIPS , 2019
J. Weng, S. Liu, V. Dadu, Z. Wang, P. Shah, and T. Nowatzki, “Dsagen: Synthesizing programmable spatial accelerators,” in ISCA , 2020
2020
Closest in time.
M. Vilim et al. , “Gorgon: accelerating machine learning from relational data,” in ISCA , 2020
2020
Closest in time.
L. Ke et al. , “Recnmp: Accelerating personalized recommendation with near-memory processing,” in ISCA , 2020
2020
Closest in time.
R. D. Evans et al. , “Jpeg-act: accelerating deep learning via transform-based lossy compression,” in ISCA , 2020
2020
Closest in time.
U. Gupta et al. , “Deeprecsys: A system for optimizing end-to-end at-scale neural recommendation inference,” ISCA , 2020
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Y. Ji et al. , “Fpsa: A full system stack solution for reconfigurable reram-based nn accelerator architecture,” in ASPLOS , 2019
2019
Cited alongside, same era.
S. Gudaparthi et al. , “Wire-aware architecture and dataflow for cnn accelerators,” in MICRO , 2019
2019
Cited alongside, same era.
Y. S. Shao et al. , “Simba: Scaling deep-learning inference with multi-chip-module-based architecture,” in MICRO , 2019
2019
Cited alongside, same era.
B. Akin et al. , “Zcomp: Reducing dnn cross-layer memory footprint using vector extensions,” in MICRO , 2019
2019
Cited alongside, same era.
C.-T. Huang et al. , “ecnn: A block-based and highly-parallel cnn accelerator for edge inference,” in MICRO , 2019
2019
Cited alongside, same era.
T.-H. Yang et al. , “Sparse reram engine: joint exploration of activation and weight sparsity in compressed neural networks,” in ISCA , 2019
2019
Cited alongside, same era.
2020
Closest in time.
E. Baek et al. , “A multi-neural network acceleration architecture,” in ISCA , 2020
2020
Closest in time.
K. Ishida et al. , “Supernpu: An extremely fast neural processing unit using superconducting logic devices,” in MICRO , 2020
2020
Closest in time.
Y. Gan et al. , “Ptolemy: Architecture support for robust deep learning,” in MICRO , 2020
2020
Closest in time.
M. He et al. , “Newton: A dram-maker’s accelerator-in-memory (aim) architecture for machine learning,” in MICRO , 2020
2020
Closest in time.
S. o. Ghodrati, “Planaria: Dynamic architecture fission for spatial multi-tenant acceleration of deep neural networks,” in MICRO , 2020
2020
Closest in time.
Y. Feng et al. , “Mesorasi: Architecture support for point cloud analytics via delayed-aggregation,” in MICRO , 2020
2020
Closest in time.
S. Singh et al. , “Nebula: a neuromorphic spin-based ultra-low power architecture for snns and anns,” in ISCA , 2020
2020
Closest in time.
M. Imani et al. , “Deep learning acceleration with neuron-to-memory transformation,” in HPCA , 2020
2020
Closest in time.
K. Shiflett et al. , “Pixel: Photonic neural network accelerator,” in HPCA , 2020
2020
Closest in time.
D. Yang et al. , “Procrustes: a dataflow and accelerator for sparse deep neural network training,” in MICRO , 2020
2020
Closest in time.
Z. Gong et al. , “Save: Sparsity-aware vector engine for accelerating dnn training and inference on cpus,” in MICRO , 2020
2020
Closest in time.
R. Hwang et al. , “Centaur: A chiplet-based, hybrid sparse-dense accelerator for personalized recommendations,” ISCA , 2020
2020
Closest in time.
E. Qin et al. , “Sigma: A sparse and irregular gemm accelerator with flexible interconnects for dnn training,” in HPCA , 2020
2020
Closest in time.
M. Yan et al. , “Hygcn: A gcn accelerator with hybrid architecture,” in HPCA , 2020
2020
Closest in time.
L. Song et al. , “Accpar: Tensor partitioning for heterogeneous deep learning accelerators,” in HPCA , 2020
2020
Closest in time.
X. Zhang et al. , “Enabling highly efficient capsule networks processing through a pim-based architecture design,” in HPCA , 2020
2020
Closest in time.
Y. Chen et al. , “tpspmv: A two-phase large-scale sparse matrix-vector multiplication kernel for manycore architectures,” Information Sciences , 2020
2020
Closest in time.
M. Mahmoud et al. , “Tensordash: Exploiting sparsity to accelerate deep neural network training,” in MICRO , 2020
2020
Closest in time.
N. Srivastava et al. , “Matraptor: A sparse-sparse matrix multiplication accelerator based on row-wise product,” in MICRO , 2020
2020
Closest in time.
B. Asgari et al. , “Alrescha: A lightweight reconfigurable sparse-computation accelerator,” in HPCA , 2020
2020
Closest in time.
N. Srivastava et al. , “Tensaurus: A versatile accelerator for mixed sparse-dense tensor computations,” in HPCA , 2020
2020
Closest in time.
J. Weng et al. , “A hybrid systolic-dataflow architecture for inductive matrix algorithms,” in HPCA , 2020
2020
Closest in time.
Z. Song et al. , “Drq: dynamic region-based quantization for deep neural network acceleration,” in ISCA , 2020
2020
Closest in time.
A. H. o. Zadeh, “Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference,” in MICRO , 2020
2020
Closest in time.
Z. Yan et al. , “Micronet for efficient language modeling,” JMLR , 2020
2020
Closest in time.
H. Wang, “Efficient algorithms and hardware for natural language processing,” Massachusetts Institute of Technology , 2020
2020
Closest in time.
S. o. Goyal, “Power-bert: Accelerating bert inference for classification tasks,” ICML , 2020
2020
Closest in time.