Fetching the paper…
Reading the bibliography…
Current large language models (LLMs) primarily utilize next-token prediction method for inference, which significantly impedes their processing speed.
The Curious Case of Neural Text Degeneration
Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; and Choi, Y. 2019 · 1904
Earlier work this paper cites.
Modèles connexionnistes de l’apprentissage
Le Cun, Y. 1987 · 1987
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Sequence Transduction with Recurrent Neural Networks
Graves, A. 2012 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Cho, K.; Van Merriënboer, B.; Bahdanau, D.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014 · 2014
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Segnet: A deep convolutional encoder-decoder architecture for image segmentation
Badrinarayanan, V.; Kendall, A.; and Cipolla, R. 2017 · 2017
Earlier work this paper cites.
Focal Loss for Dense Object Detection
Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017 · 2017
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Cited alongside, same era.
Micikevicius, P.; Narang, S.; Alben, J.; Diamos, G.; Elsen, E.; Garcia, D.; Ginsburg, B.; Houston, M.; Kuchaiev, O.; Venkatesh, G.; et al. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Hierarchical Neural Story Generation
Fan, A.; Lewis, M.; and Dauphin, Y. 2018 · 2018
LMDeploy: A Toolkit for Compressing, Deploying, and Serving LLM
Contributors, L. 2023 · 2023
Later among the works it cites.
WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
He, C.; Jin, Z.; Xu, C.; Qiu, J.; Wang, B.; Li, W.; Yan, H.; Wang, J.; and Lin, D. 2023 · 2023
Later among the works it cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C. H.; Gonzalez, J.; Zhang, H.; and Stoica, I. 2023 · 2023
Later among the works it cites.
Towards expert-level medical question answering with large language models
Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Hou, L.; Clark, K.; Pfohl, S.; Cole-Lewis, H.; Neal, D.; et al. 2023 · 2023
Later among the works it cites.
Document-Level Machine Translation with Large Language Models
Wang, L.; Lyu, C.; Ji, T.; Zhang, Z.; Yu, D.; Shi, S.; and Tu, Z. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robutrans: A robust transformer-based text-to-speech model
Li, N.; Liu, Y.; Wu, Y.; Liu, S.; Zhao, S.; and Liu, M. 2020 · 2020
Cited alongside, same era.
Machine translation of cortical activity to text with an encoder–decoder framework
Makin, J. G.; Moses, D. A.; and Chang, E. F. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
Internet-augmented language models through few-shot prompting for open-domain question answering
Lazaridou, A.; Gribovskaya, E.; Stokowiec, W.; and Grigorev, N. 2022 · 2022
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Cited alongside, same era.
Computational ghost imaging in turbulent water based on self-supervised information extraction network
Chen, Y.; Sun, Z.; Li, C.; and Li, X. 2023 · 2023
Cited alongside, same era.
Part-based image-loop network for single-pixel imaging
Li, X.; Chen, Y.; Tian, T.; and Sun, Z. 2024a
Cited in the paper.
Later among the works it cites.
Prompting Large Language Model for Machine Translation: A Case Study
Zhang, B.; Haddow, B.; and Birch, A. 2023 · 2023
Later among the works it cites.
Cai, Z.; Cao, M.; Chen, H.; Chen, K.; Chen, K.; Chen, X.; Chen, X.; Chen, Z.; Chen, Z.; Chu, P.; Dong, X.; Duan, H.; Fan, Q.; Fei, Z.; Gao, Y.; Ge, J.; Gu, C.; Gu, Y.; Gui, T.; Guo, A.; Guo, Q.; He, C.; Hu, Y.; Huang, T.; Jiang, T.; Jiao, P.; Jin, Z.; Lei, Z.; Li, J.; Li, J.; Li, L.; Li, S.; Li, W.; Li, Y.; Liu, H.; Liu, J.; Hong, J.; Liu, K.; Liu, K.; Liu, X.; Lv, C.; Lv, H.; Lv, K.; Ma, L.; Ma, R.; Ma, Z.; Ning, W.; Ouyang, L.; Qiu, J.; Qu, Y.; Shang, F.; Shao, Y.; Song, D.; Song, Z.; Sui, Z.; Sun, P.; Sun, Y.; Tang, H.; Wang, B.; Wang, G.; Wang, J.; Wang, J.; Wang, R.; Wang, Y.; Wang, Z.; Wei, X.; Weng, Q.; Wu, F.; Xiong, Y.; Xu, C.; Xu, R.; Yan, H.; Yan, Y.; Yang, X.; Ye, H.; Ying, H.; Yu, J.; Yu, J.; Zang, Y.; Zhang, C.; Zhang, L.; Zhang, P.; Zhang, P.; Zhang, R.; Zhang, S.; Zhang, S.; Zhang, W.; Zhang, W.; Zhang, X.; Zhang, X.; Zhao, H.; Zhao, Q.; Zhao, X.; Zhou, F.; Zhou, Z.; Zhuo, J.; Zou, Y.; Qiu, X.; Qiao, Y.; and Lin, D. 2024 · 2024
Closest in time.
Better & Faster Large Language Models via Multi-token Prediction
Gloeckle, F.; Idrissi, B. Y.; Rozière, B.; Lopez-Paz, D.; and Synnaeve, G. 2024 · 2024
Closest in time.
GPT-4o: The Cutting-Edge Advancement in Multimodal LLM
Islam, R.; and Moushi, O. M. 2024 · 2024
Closest in time.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Su, J.; Ahmed, M.; Lu, Y.; Pan, S.; Bo, W.; and Liu, Y. 2024 · 2024
Closest in time.
Wang, Z.; Liu, X.; Liu, S.; Yao, Y.; Huang, Y.; He, Z.; Li, X.; Li, Y.; Che, Z.; Zhang, Z.; Wang, Y.; Wang, X.; Pu, L.; Xu, H.; Fang, R.; Zhao, Y.; Zhang, J.; Huang, X.; Lu, Z.; Peng, J.; Zheng, W.; Wang, S.; Yang, B.; he, X.; Jiang, Z.; Xie, Q.; Zhang, Y.; Li, Z.; Shi, L.; Fu, W.; Zhang, Y.; Huang, Z.; Xiong, S.; Zhang, Y.; Wang, C.; and Song, S. 2024 · 2024
Closest in time.