Fetching the paper…
Reading the bibliography…
The power of large language models (LLMs) has been demonstrated through numerous data and computing resources.
Investigating prior knowledge for challenging chinese machine reading comprehension
Sun, K., Yu, D., Yu, D., and Cardie, C · 1904
Earlier work this paper cites.
Byte pair encoding: A text compression scheme that accelerates pattern matching
Shibata, Y., Kida, T., Fukamachi, S., Takeda, M., Shinohara, A., Shinohara, T., and Arikawa, S · 1999
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Dynamic network surgery for efficient dnns
Guo, Y., Yao, A., and Chen, Y · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Earlier work this paper cites.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization, 2018
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., Combes, R. T. d., Trischler, A., Bengio, Y., and Gordon, G. J · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Snip: single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H. S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Cited alongside, same era.
Opencompass: A universal evaluation platform for foundation models
Contributors, O · 2023
Later among the works it cites.
Efficient and effective text encoding for chinese llama and alpaca
Cui, Y., Yang, Z., and Yao, X · 2023
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Later among the works it cites.
Openllama: An open reproduction of llama, May 2023
Geng, X. and Liu, H · 2023
Later among the works it cites.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Huang, Y., Bai, Y., Zhu, Z., Zhang, J., Zhang, J., Su, T., Liu, J., Lv, C., Zhang, Y., Lei, J., Fu, Y., Sun, M., and He, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., Kunin, D., Yamins, D. L. K., and Ganguli, S · 2020
Cited alongside, same era.
Scop: Scientific control for reliable neural network pruning
Tang, Y., Wang, Y., Xu, Y., Tao, D., Xu, C., Xu, C., and Xu, C · 2020
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems, 2020
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Layer-adaptive sparsity for the magnitude-based pruning
Lee, J., Park, S., Mo, S., Ahn, S., and Shin, J · 2021
Cited alongside, same era.
Fewclue: A chinese few-shot learning evaluation benchmark
Xu, L., Lu, X., Yuan, C., Zhang, X., Xu, H., Yuan, H., Wei, G., Pan, X., Tian, X., Qin, L., et al · 2021
Cited alongside, same era.
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
Tinyllama, Sep 2023
Peiyuan Zhang, Guangtao Zeng, T. W. and Lu, W · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al · 2023
Later among the works it cites.
Pangu- Σ \Sigma : Towards trillion parameter language model with sparse heterogeneous computing
Ren, X., Zhou, P., Meng, X., Huang, X., Wang, Y., Wang, W., Li, P., Zhang, X., Podolskiy, A., Arshinov, G., et al · 2023
Later among the works it cites.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Pangu- π \pi : Enhancing language model architectures via nonlinearity compensation
Wang, Y., Chen, H., Tang, Y., Guo, T., Han, K., Nie, Y., Wang, X., Hu, H., Bai, Z., Wang, Y., et al · 2023
Later among the works it cites.
Skywork: A more open bilingual foundation model, 2023
Wei, T., Zhao, L., Zhang, L., Zhu, B., Wang, L., Yang, H., Li, B., Cheng, C., Lü, W., Hu, R., Li, C., Yang, L., Luo, X., Wu, X., Liu, L., Cheng, W., Cheng, P., Zhang, J., Zhang, X., Lin, L., Wang, X., Ma, Y., Dong, C., Sun, Y., Chen, Y., Peng, Y., Liang, X., Yan, S., Fang, H., and Zhou, Y · 2023
Later among the works it cites.
Overcoming catastrophic forgetting in massively multilingual continual learning
Winata, G. I., Xie, L., Radhakrishnan, K., Wu, S., Jin, X., Cheng, P., Kulkarni, M., and Preotiuc-Pietro, D · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
Xia, M., Gao, T., Zeng, Z., and Chen, D · 2023
Later among the works it cites.
Baichuan 2: Open large-scale language models
Yang, A., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., et al · 2023
Later among the works it cites.
A series of large language models trained from scratch by developers at 01-ai
Yi · 2023
Later among the works it cites.
A survey on transformer compression
Tang, Y., Wang, Y., Guo, J., Tu, Z., Han, K., Hu, H., and Tao, D · 2024
Closest in time.