Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have recently transformed natural language processing, enabling machines to generate human-like text and engage in meaningful conversations.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
M. Iyyer, J. Boyd-Graber, L. Claudino, R. Socher, and H. Daumé III, “A neural network for factoid question answering over paragraphs,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 633–644
2014
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 2704–2713
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. Sutton, “The bitter lesson,” Incomplete Ideas (blog) , vol. 13, no. 1, p. 38, 2019
2019
Earlier work this paper cites.
S. Zhang, G. L. Zhang, B. Li, H. H. Li, and U. Schlichtmann, “Aging-aware lifetime enhancement for memristor-based neuromorphic computing,” in 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE) . IEEE, 2019, pp. 1751–1756
2019
Earlier work this paper cites.
A. Ankit, I. E. Hajj, S. R. Chalamalasetti, G. Ndu, M. Foltin, R. S. Williams, P. Faraboschi, W.-m. W. Hwu, J. P. Strachan, K. Roy et al. , “Puma: A programmable ultra-efficient memristor-based accelerator for machine learning inference,” in Proceedings of the twenty-fourth international conference on architectural support for programming languages and operating systems , 2019, pp. 715–731
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. E. Leiserson, N. C. Thompson, J. S. Emer, B. C. Kuszmaul, B. W. Lampson, D. Sanchez, and T. B. Schardl, “There’s plenty of room at the top: What will drive computer performance after moore’s law?” Science , vol. 368, no. 6495, p. eaam9744, 2020
2020
Earlier work this paper cites.
A. Sebastian, M. Le Gallo, R. Khaddam-Aljameh, and E. Eleftheriou, “Memory devices and applications for in-memory computing,” Nature nanotechnology , vol. 15, no. 7, pp. 529–544, 2020
2020
Earlier work this paper cites.
X. Yang, B. Yan, H. Li, and Y. Chen, “Retransformer: Reram-based processing-in-memory architecture for transformer acceleration,” in Proceedings of the 39th International Conference on Computer-Aided Design , 2020, pp. 1–9
2020
Cited alongside, same era.
H. Guo, L. Peng, J. Zhang, Q. Chen, and T. D. LeCompte, “Att: A fault-tolerant reram accelerator for attention-based neural networks,” in 2020 IEEE 38th International Conference on Computer Design (ICCD) . IEEE, 2020, pp. 213–221
2020
Cited alongside, same era.
A. F. Laguna, A. Kazemi, M. Niemier, and X. S. Hu, “In-memory computing based accelerator for transformer networks for long sequences,” in 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE) . IEEE, 2021, pp. 1839–1844
2021
Cited alongside, same era.
M. Kang, H. Shin, and L.-S. Kim, “A framework for accelerating transformer-based language model on reram-based architecture,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 41, no. 9, pp. 3026–3039, 2021
2023
Later among the works it cites.
G. Burr, H. Tsai, W. Simon, I. Boybat, S. Ambrogio, C.-E. Ho, Z.-W. Liou, M. Rasch, J. Büchel, P. Narayanan et al. , “Design of analog-ai hardware accelerators for transformer-based language models,” in 2023 International Electron Devices Meeting (IEDM) . IEEE, 2023, pp. 1–4
2023
Later among the works it cites.
S. Sridharan, J. R. Stevens, K. Roy, and A. Raghunathan, “X-former: In-memory acceleration of transformers,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , 2023
2023
Later among the works it cites.
Z. Yang, K. Liu, Y. Duan, M. Fan, Q. Zhang, and Z. Jin, “Three challenges in reram-based process-in-memory for neural network,” in 2023 IEEE 5th International Conference on Artificial Intelligence Circuits and Systems (AICAS) . IEEE, 2023, pp. 1–5
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
K. Spoon, H. Tsai, A. Chen, M. J. Rasch, S. Ambrogio, C. Mackin, A. Fasoli, A. M. Friz, P. Narayanan, M. Stanisavljevic et al. , “Toward software-equivalent accuracy on transformer-based deep neural networks with analog memory devices,” Frontiers in Computational Neuroscience , vol. 15, p. 675741, 2021
2021
Cited alongside, same era.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al. , “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2022, pp. 1–15
2022
Cited alongside, same era.
2022
Cited alongside, same era.
A. F. Laguna, M. M. Sharifi, A. Kazemi, X. Yin, M. Niemier, and X. S. Hu, “Hardware-software co-design of an in-memory transformer network accelerator,” Frontiers in Electronics , vol. 3, p. 847069, 2022
2022
Cited alongside, same era.
M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memory-based acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2022, pp. 1071–1085
2022
Cited alongside, same era.
M. Ali, S. Roy, U. Saxena, T. Sharma, A. Raghunathan, and K. Roy, “Compute-in-memory technologies and architectures for deep learning workloads,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 30, no. 11, pp. 1615–1630, 2022
2022
Cited alongside, same era.
A. Okazaki, P. Narayanan, S. Ambrogio, K. Hosokawa, H. Tsai, A. Nomura, T. Yasuda, C. Mackin, A. Friz, M. Ishii et al. , “Analog-memory-based 14nm hardware accelerator for dense deep neural networks including transformers,” in 2022 IEEE International Symposium on Circuits and Systems (ISCAS) . IEEE, 2022, pp. 3319–3323
2022
Cited alongside, same era.
C. Yang, X. Wang, and Z. Zeng, “Full-circuit implementation of transformer network based on memristor,” IEEE Transactions on Circuits and Systems I: Regular Papers , vol. 69, no. 4, pp. 1395–1407, 2022
2022
Cited alongside, same era.
Later among the works it cites.
M. Le Gallo, R. Khaddam-Aljameh, M. Stanisavljevic, A. Vasilopoulos, B. Kersting, M. Dazzi, G. Karunaratne, M. Brändli, A. Singh, S. M. Mueller et al. , “A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference,” Nature Electronics , vol. 6, no. 9, pp. 680–693, 2023
2023
Later among the works it cites.
D. Reis, A. F. Laguna, M. Niemier, and X. S. Hu, “In-memory computing accelerators for emerging learning paradigms,” in Proceedings of the 28th Asia and South Pacific Design Automation Conference , 2023, pp. 606–611
2023
Later among the works it cites.
2023
Later among the works it cites.
M. J. Rasch, C. Mackin, M. Le Gallo, A. Chen, A. Fasoli, F. Odermatt, N. Li, S. Nandakumar, P. Narayanan, H. Tsai et al. , “Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators,” Nature communications , vol. 14, no. 1, p. 5282, 2023
2023
Later among the works it cites.
M. Kang, H. Shin, J. Kim, and L.-S. Kim, “Mgen: A framework for energy-efficient in-reram acceleration of multi-task bert,” IEEE Transactions on Computers , 2023
2023
Later among the works it cites.
H. Zhou, J. Chen, J. Li, L. Yang, Y. Li, and X. Miao, “Bring memristive in-memory computing into general-purpose machine learning: A perspective,” APL Machine Learning , vol. 1, no. 4, 2023
2023
Later among the works it cites.
S. Liu, C. Mu, H. Jiang, Y. Wang, J. Zhang, F. Lin, K. Zhou, Q. Liu, and C. Chen, “Hardsea: Hybrid analog-reram clustering and digital-sram in-memory computing accelerator for dynamic sparse self-attention in transformer,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , 2023
2023
Later among the works it cites.
F. Tu, Z. Wu, Y. Wang, W. Wu, L. Liu, Y. Hu, S. Wei, and S. Yin, “Multcim: Digital computing-in-memory-based multimodal transformer accelerator with attention-token-bit hybrid sparsity,” IEEE Journal of Solid-State Circuits , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
D. S. Modha, F. Akopyan, A. Andreopoulos, R. Appuswamy, J. V. Arthur, A. S. Cassidy, P. Datta, M. V. DeBole, S. K. Esser, C. O. Otero et al. , “Neural inference at the frontier of energy, space, and time,” Science , vol. 382, no. 6668, pp. 329–335, 2023
2023
Later among the works it cites.
D. Castelvecchi, “’mind-blowing’ibm chip speeds up ai.” Nature , 2023
2023
Later among the works it cites.
S. Zeng, J. Liu, G. Dai, X. Yang, T. Fu, H. Wang, W. Ma, H. Sun, S. Li, Z. Huang et al. , “Flightllm: Efficient large language model inference with a complete mapping flow on fpgas,” in Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays , 2024, pp. 223–234
2024
Closest in time.
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
R. F. De Moura and L. Carro, “Reprogrammable non-linear circuits using reram for nn accelerators,” ACM Transactions on Reconfigurable Technology and Systems , vol. 17, no. 1, pp. 1–19, 2024
2024
Closest in time.