Fetching the paper…
Reading the bibliography…
Large language Models (LLMs), though growing exceedingly powerful, comprises of orders of magnitude less neurons and synapses than the human brain.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Distilling task-specific knowledge from bert into simple neural networks
Tang, R.; Lu, Y.; Liu, L.; Mou, L.; Vechtomova, O.; and Lin, J. 2019 · 1903
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2019 · 1909
Earlier work this paper cites.
NAS-BERT: task-agnostic and adaptive-size BERT compression with neural architecture search
Xu, J.; Tan, X.; Luo, R.; Song, K.; Li, J.; Qin, T.; and Liu, T.-Y. 2021 · 1943
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Spiking neural networks
Ghosh-Dastidar, S.; and Adeli, H. 2009 · 2009
Earlier work this paper cites.
Cifar-10 and cifar-100 datasets
Krizhevsky, A.; Nair, V.; and Hinton, G. 2009 · 2009
Earlier work this paper cites.
Learning both Weights and Connections for Efficient Neural Networks
Han, S.; Pool, J.; Tran, J.; and Dally, W. J. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Converting static image datasets to spiking neuromorphic datasets using saccades
Orchard, G.; Jayawant, A.; Cohen, G. K.; and Thakor, N. 2015 · 2015
Earlier work this paper cites.
Energy efficient stochastic-based deep spiking neural networks for sparse datasets
Alawad, M.; Yoon, H.-J.; and Tourassi, G. 2017 · 2017
Earlier work this paper cites.
A low power, fully event-based gesture recognition system
Amir, A.; Taba, B.; Berg, D.; Melano, T.; McKinstry, J.; Di Nolfo, C.; Nayak, T.; Andreopoulos, A.; Garreau, G.; Mendoza, M.; et al. 2017 · 2017
Earlier work this paper cites.
Equilibrium propagation: Bridging the gap between energy-based models and backpropagation
Scellier, B.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
M-FAC: Efficient matrix-free approximations of second-order information
Frantar, E.; Kurtic, E.; and Alistarh, D. 2021 · 2021
Later among the works it cites.
I-bert: Integer-only bert quantization
Kim, S.; Gholami, A.; Yao, Z.; Mahoney, M. W.; and Keutzer, K. 2021 · 2021
Later among the works it cites.
Training low-latency spiking neural network through knowledge distillation
Takuya, S.; Zhang, R.; and Nakashima, Y. 2021 · 2021
Later among the works it cites.
Training feedback spiking neural networks by implicit differentiation on the equilibrium state
Xiao, M.; Meng, Q.; Zhang, Z.; Wang, Y.; and Lin, Z. 2021 · 2021
Later among the works it cites.
Sequence Learning using Equilibrium Propagation
Bal, M.; and Sengupta, A. 2022 · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep equilibrium models
Bai, S.; Kolter, J. Z.; and Koltun, V. 2019 · 2019
Cited alongside, same era.
Brain-like object recognition with high-performing shallow recurrent ANNs
Kubilius, J.; Schrimpf, M.; Kar, K.; Rajalingham, R.; Hong, H.; Majaj, N.; Issa, E.; Bashivan, P.; Prescott-Roy, J.; Schmidt, K.; et al. 2019 · 2019
Cited alongside, same era.
Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks
Neftci, E. O.; Mostafa, H.; and Zenke, F. 2019 · 2019
Cited alongside, same era.
Going deeper in spiking neural networks: VGG and residual architectures
Sengupta, A.; Ye, Y.; Wang, R.; Liu, C.; and Roy, K. 2019 · 2019
Cited alongside, same era.
Enabling spike-based backpropagation for training deep neural network architectures
Lee, C.; Sarwar, S. S.; Panda, P.; Srinivasan, G.; and Roy, K. 2020 · 2020
Cited alongside, same era.
Exploring the connection between binary and spiking neural networks
Lu, S.; and Sengupta, A. 2020 · 2020
Cited alongside, same era.
Advancing neuromorphic computing with loihi: A survey of results and outlook
Davies, M.; Wild, A.; Orchard, G.; Sandamirskaya, Y.; Guerra, G. A. F.; Joshi, P.; Plank, P.; and Risbud, S. R. 2021 · 2021
Cited alongside, same era.
Kurtic, E.; Campos, D.; Nguyen, T.; Frantar, E.; Kurtz, M.; Fineran, B.; Goin, M.; and Alistarh, D. 2022 · 2022
Later among the works it cites.
Emergent Abilities of Large Language Models
Wei, J.; Tay, Y.; Bommasani, R.; Raffel, C.; Zoph, B.; Borgeaud, S.; Yogatama, D.; Bosma, M.; Zhou, D.; Metzler, D.; Chi, E. H.; Hashimoto, T.; Vinyals, O.; Liang, P.; Dean, J.; and Fedus, W. 2022 · 2022
Later among the works it cites.
Spikformer: When spiking neural network meets transformer
Zhou, Z.; Zhu, Y.; He, C.; Wang, Y.; Yan, S.; Tian, Y.; and Yuan, L. 2022 · 2022
Later among the works it cites.
Hong, D.; Shen, J.; Qi, Y.; and Wang, Y. 2023 · 2023
Closest in time.
Constructing deep spiking neural networks from artificial neural networks with knowledge distillation
Xu, Q.; Li, Y.; Shen, J.; Liu, J. K.; Tang, H.; and Pan, G. 2023 · 2023
Closest in time.
Spikegpt: Generative pre-trained language model with spiking neural networks
Zhu, R.-J.; Zhao, Q.; and Eshraghian, J. K. 2023 · 2023
Closest in time.