Fetching the paper…
Reading the bibliography…
Transformers have achieved great success in a wide variety of natural language processing (NLP) tasks due to the attention mechanism, which assigns an importance score for every word relative to other words in a sequence.
Phase change memory — opportunities and challenges
B. Rajendran, H. Lung, and C. Lam · 2007
Earlier work this paper cites.
Optimizing nuca organizations and wiring alternatives for large caches with cacti 6.0
N. Muralimanohar, R. Balasubramonian, and N. Jouppi · 2007
Earlier work this paper cites.
Resistive random access memory (reram) based on metal oxides
H. Akinaga and H. Shima · 2010
Earlier work this paper cites.
Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks
Y.H. Chen, J. Emer, and V. Sze · 2016
Earlier work this paper cites.
Isaac: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J.P. Strachan, M. Hu, R.S. Williams, and V. Srikumar · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
N.P. Jouppi et al · 2017
Earlier work this paper cites.
Scaledeep: A scalable compute architecture for learning and evaluating deep networks
S. Venkataramani, A. Ranjan, S. Banerjee, D. Das, S. Avancha, A. Jagannathan, A. Durg, D. Nagaraj, B. Kaul, P. Dubey, and A. Raghunathan · 2017
Earlier work this paper cites.
Pipelayer: A pipelined reram-based accelerator for deep learning
L. Song, X. Qian, H. Li, and Y. Chen · 2017
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
In-memory computing: Advances and prospects
N. Verma, H. Jia, H. Valavi, Y. Tang, M. Ozatay, L.Y. Chen, B. Zhang, and P. Deaville · 2019
Cited alongside, same era.
Puma: A programmable ultra-efficient memristor-based accelerator for machine learning inference
A. Ankit et al · 2019
Cited alongside, same era.
Cxdnn: Hardware-software compensation methods for deep neural networks on resistive crossbar systems
Shubham Jain and Anand Raghunathan · 2019
Cited alongside, same era.
Language models are few-shot learners
T.B. Brown et al · 2020
Cited alongside, same era.
Panther: A programmable architecture for neural network training harnessing energy-efficient reram
A. Ankit, I. Hajj, S. Chalamalasetti, S. Agarwal, M. Marinella, M. Foltin, J.P. Strachan, D. Milojicic, W.M. Hwu, and K. Roy · 2020
Cited alongside, same era.
Retransformer: Reram-based processing-in-memory architecture for transformer acceleration
X. Yang, B. Yan, H. Li, and Y. Chen · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
T. Wolf et al · 2020
Later among the works it cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
H. Wang, Z. Zhang, and S. Han · 2021
Later among the works it cites.
Softermax: Hardware/software co-design of an efficient softmax for transformers
J.R. Stevens, R. Venkatesan, S. Dai, B. Khailany, and A. Raghunathan · 2021
Later among the works it cites.
Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference
T. Tambe et al · 2021
Later among the works it cites.
In-memory computing based accelerator for transformer networks for long sequences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Zahoor, T. Zulkifli, and F. Khanday · 2020
Cited alongside, same era.
Resistive crossbars as approximate hardware building blocks for machine learning: Opportunities and challenges
I. Chakraborty, M. Ali, A. Ankit, S. Jain, S. Roy, S. Sridharan, A. Agrawal, A. Raghunathan, and K. Roy · 2020
Cited alongside, same era.
Efficient transformers: A survey
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler · 2020
Cited alongside, same era.
Optimizing transformers with approximate computing for faster, smaller and more accurate NLP models
A. Nagarajan, S. Sen, J.R. Stevens, and A. Raghunathan · 2020
Cited alongside, same era.
GOBO: Quantizing attention-based NLP models for low latency and energy efficient inference
A. Hadizadeh, I. Edo, O.M. Awadand, and A. Moshovos · 2020
Cited alongside, same era.
Power-bert: Accelerating BERT inference for classification tasks
S. Goyal, A.R. Choudhury, V. Chakaravarthy, S. ManishRaje, Y. Sabharwal, and A. Verma · 2020
Cited alongside, same era.
Optimus: Optimized matrix multiplication structure for transformer neural network accelerator
J. Park, H. Yoon, D. Ahn, J. Choi, and J. Kim · 2020
Cited alongside, same era.
A.F. Laguna, A. Kazemi, M. Niemier, and X. Hu · 2021
Later among the works it cites.
Towards fully 8-bit integer inference for the transformer model
Ye Lin, Yanyang Li, Tengbo Liu, Tong Xiao, Tongran Liu, and Jingbo Zhu · 2021
Later among the works it cites.
Txsim: Modeling training of deep neural networks on resistive crossbar systems
Sourjya Roy, Shrihari Sridharan, Shubham Jain, and Anand Raghunathan · 2021
Later among the works it cites.
Neurosim simulator for compute-in-memory hardware accelerator: Validation and benchmark
Anni Lu, Xiaochen Peng, Wantong Li, Hongwu Jiang, and Shimeng Yu · 2021
Later among the works it cites.
Mustafa Ali, Indranil Chakraborty, Utkarsh Saxena, Amogh Agrawal, Aayush Ankit, and Kaushik Roy · 2021
Later among the works it cites.
Adc performance survey 1997-2022
Boris Murmann · 2022
Later among the works it cites.