Fetching the paper…
Reading the bibliography…
Text generation is a compelling sub-field of natural language processing, aiming to generate human-readable text from input words.
Y.-B. Kim and T. W. Chen, “Assessing merged dram/logic technology,” Integration, the VLSI Journal , vol. 2, no. 27, pp. 179–194, 1999
1999
Earlier work this paper cites.
Y. Kim, V. Seshadri, D. Lee, J. Liu, and O. Mutlu, “A case for exploiting subarray-level parallelism (salp) in dram,” in 2012 39th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2012, pp. 368–379
2012
Earlier work this paper cites.
J. Standard, “High bandwidth memory (hbm) dram,” Jesd235 , vol. 16, 2013
2013
Earlier work this paper cites.
Y. Kim, W. Yang, and O. Mutlu, “Ramulator: A fast and extensible dram simulator,” IEEE Computer architecture letters , vol. 15, no. 1, pp. 45–49, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
V. Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Ambit: In-memory accelerator for bulk bitwise operations using commodity dram technology,” in 2017 50th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2017, pp. 273–287
2017
Earlier work this paper cites.
S. Li, D. Niu, K. T. Malladi, H. Zheng, B. Brennan, and Y. Xie, “Drisa: A dram-based reconfigurable in-situ accelerator,” in 2017 50th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2017, pp. 288–301
2017
Earlier work this paper cites.
M. O’Connor, N. Chatterjee, D. Lee, J. Wilson, A. Agrawal, S. W. Keckler, and W. J. Dally, “Fine-grained dram: Energy-efficient dram for extreme bandwidth systems,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture , 2017, pp. 41–54
2017
Earlier work this paper cites.
H. H. Shin, Y. M. Park, D. Choi, B. J. Kim, D.-H. Cho, and E.-Y. Chung, “Extreme: Exploiting page table for reducing refresh power of 3d-stacked dram memory,” IEEE Transactions on Computers , vol. 67, no. 1, pp. 32–44, 2017
2017
Earlier work this paper cites.
R. Sanchis, Ó. García-Perales, F. Fraile, and R. Poler, “Low-code as enabler of digital transformation in manufacturing industry,” Applied Sciences , vol. 10, no. 1, p. 12, 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
F. Gao, G. Tziantzioulis, and D. Wentzlaff, “Computedram: In-memory compute using off-the-shelf drams,” in Proceedings of the 52nd annual IEEE/ACM international symposium on microarchitecture , 2019, pp. 100–113
2019
Earlier work this paper cites.
F. Devaux, “The true processing in memory accelerator,” in 2019 IEEE Hot Chips 31 Symposium (HCS) . IEEE Computer Society, 2019, pp. 1–24
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Q. Deng, Y. Zhang, M. Zhang, and J. Yang, “Lacc: Exploiting lookup table-based fast and accurate vector multiplication in dram-based cnn accelerator,” in Proceedings of the 56th Annual Design Automation Conference 2019 , 2019, pp. 1–6
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Y.-C. Kwon, S. H. Lee, J. Lee, S.-H. Kwon, J. M. Ryu, J.-P. Son, O. Seongil, H.-S. Yu, H. Lee, S. Y. Kim et al. , “25.4 a 20nm 6gb function-in-memory dram, based on hbm2 with a 1.2 tflops programmable computing unit using bank-level parallelism, for machine learning applications,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 64. IEEE, 2021, pp. 350–352
2021
Later among the works it cites.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2021, pp. 97–110
2021
Later among the works it cites.
T. J. Ham, Y. Lee, S. H. Seo, S. Kim, H. Choi, S. J. Jung, and J. W. Lee, “Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2021, pp. 692–705
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I , 2020, pp. 213–229
2020
Cited alongside, same era.
“A robot wrote this entire article. are you scared yet, human? — gpt-3,” Sep 2020. [Online]. Available: https://www.theguardian.com/commentisfree/2020/sep/08/robot-wrote-this-article-gpt-3
2020
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
M. He, C. Song, I. Kim, C. Jeong, S. Kim, I. Park, M. Thottethodi, and T. Vijaykumar, “Newton: A dram-maker’s accelerator-in-memory (aim) architecture for machine learning,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 372–385
2020
Cited alongside, same era.
M. Lenjani, P. Gonzalez, E. Sadredini, S. Li, Y. Xie, A. Akel, S. Eilert, M. R. Stan, and K. Skadron, “Fulcrum: a simplified control and access mechanism toward flexible and practical in-situ accelerators,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 556–569
2020
Cited alongside, same era.
A. H. Zadeh, I. Edo, O. M. Awad, and A. Moshovos, “Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 811–824
2020
Cited alongside, same era.
S. Cho, H. Choi, E. Park, H. Shin, and S. Yoo, “Mcdram v2: In-dynamic random access memory systolic array accelerator to address the large model problem in deep neural networks on the edge,” IEEE Access , vol. 8, pp. 135 223–135 243, 2020
2020
Cited alongside, same era.
S. Lee, S.-h. Kang, J. Lee, H. Kim, E. Lee, S. Seo, H. Yoon, S. Lee, K. Lim, H. Shin et al. , “Hardware architecture and software stack for pim based on commercial dram technology: Industrial product,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2021, pp. 43–56
2021
Cited alongside, same era.
J. Schulman, B. Zoph, C. Kim, J. Hilton, J. Menick, J. Weng, J. Uribe, L. Fedus, L. Metz, M. Pokorny et al. , “Chatgpt: Optimizing language models for dialogue,” 2022
2022
Later among the works it cites.
M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memory-based acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2022, pp. 1071–1085
2022
Later among the works it cites.
D. Kim, C. Yu, S. Xie, Y. Chen, J.-Y. Kim, B. Kim, J. Kulkarni, and T. T.-H. Kim, “An overview of processing-in-memory circuits for artificial intelligence and machine learning,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems , 2022
2022
Later among the works it cites.
S. Lee, K. Kim, S. Oh, J. Park, G. Hong, D. Ka, K. Hwang, J. Park, K. Kang, J. Kim et al. , “A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications,” in 2022 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 65. IEEE, 2022, pp. 1–3
2022
Later among the works it cites.
J.-H. Kim, S. Lee, S. Moon, S. Yoo, and J.-Y. Kim, “19.2 a 409.6 gops and 204.8 gflops mixed-precision vector processor system for general-purpose machine learning acceleration,” in 2022 IEEE Asian Solid-State Circuits Conference (A-SSCC) . IEEE, 2022
2022
Later among the works it cites.
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y. Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2022, pp. 616–630
2022
Later among the works it cites.
J. D. Ferreira, G. Falcao, J. Gómez-Luna, M. Alser, L. Orosa, M. Sadrosadati, J. S. Kim, G. F. Oliveira, T. Shahroodi, A. Nori et al. , “Pluto: Enabling massively parallel computation in dram via lookup tables,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2022, pp. 900–919
2022
Later among the works it cites.
2023
Later among the works it cites.
R. Zhou, S. Tabrizchi, A. Roohi, and S. Angizi, “Lt-pim: An lut-based processing-in-dram architecture with rowhammer self-tracking,” IEEE Computer Architecture Letters , 2022
2023
Later among the works it cites.