Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved remarkable success in contextual knowledge understanding.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training. openai
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
What does BERT look at? an analysis of BERT‘s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J. G., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Rethinking positional encoding in language pre-training
Ke, G., He, D., and Liu, T.-Y · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Thomas unterthiner mostafa dehghani matthias minderer georg heigold sylvain gelly jakob uszkoreit and neil houlsby. an image isworth 16 × \times 16 words: Transformers for image recognition atscale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., and Zhai, X · 2021
Earlier work this paper cites.
Entity-based knowledge conflicts in question answering
Longpre, S., Perisetla, K., Chen, A., Ramesh, N., DuBois, C., and Singh, S · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Evaluating the evaluation of diversity in natural language generation
Tevet, G. and Berant, J · 2021
Earlier work this paper cites.
Outlier suppression: Pushing the limit of low-bit transformer language models
Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X · 2021
Earlier work this paper cites.
GPT-NeoX-20B: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Earlier work this paper cites.
GPTQ: Accurate post-training compression for generative pretrained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Transformer language models without positional encodings still learn positional information
Haviv, A., Ram, O., Press, O., Izsak, P., and Levy, O · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Gpt-4 technical report
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Intriguing properties of quantization at scale
Ahmadian, A., Dash, S., Chen, H., Venkitesh, B., Gou, Z. S., Blunsom, P., Üstün, A., and Hooker, S · 2023
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Cited alongside, same era.
Long context prompting for claude 2.1
Anthropic · 2023
Cited alongside, same era.
Rethinking channel dimensions to isolate outliers for low-bit weight quantization of large language models
Heo, J. H., Kim, J., Kwon, B., Kim, B., Kwon, S. J., and Lee, D · 2024
Later among the works it cites.
KVQuant: Towards 10 million context length LLM inference with KV cache quantization
Hooper, C. R. C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, S., Keutzer, K., and Gholami, A · 2024
Later among the works it cites.
Pre-rmsnorm and pre-crmsnorm transformers: equivalent and efficient pre-ln transformers
Jiang, Z., Gu, J., Zhu, H., and Pan, D · 2024
Later among the works it cites.
The impact of reasoning step length on large language models
Jin, M., Yu, Q., Shu, D., Zhao, H., Hua, W., Meng, Y., Zhang, Y., and Du, M · 2024
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Natesan Ramamurthy, K., Das, P., and Reddy, S · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The internal state of an llm knows when it’s lying
Azaria, A. and Mitchell, T · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Chiang, C.-H. and Lee, H.-Y · 2023
Cited alongside, same era.
A survey of knowledge enhanced pre-trained language models
Hu, L., Liu, Z., Zhao, Z., Hou, L., Nie, L., and Li, J · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Marks, S. and Tegmark, M · 2023
Cited alongside, same era.
Random-access infinite context length for transformers
Mohtashami, A. and Jaggi, M · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Cited alongside, same era.
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al · 2024
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Later among the works it cites.
Lut-gemm: Quantized matrix multiplication based on luts for efficient inference in large-scale generative language models
Park, G., Park, B., Kim, M., Lee, S., Kim, J., Kwon, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
Massive activations in large language models
Sun, M., Chen, X., Kolter, J. Z., and Liu, Z · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., and Lin, J · 2024
Later among the works it cites.
Rnns are not transformers (yet): The key bottleneck on in-context retrieval
Wen, K., Dang, X., and Lyu, K · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H · 2024
Later among the works it cites.
Knowledge conflicts for LLMs: A survey
Xu, R., Qi, Z., Guo, Z., Wang, C., Wang, H., Zhang, Y., and Xu, W · 2024
Later among the works it cites.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., and Fan, Z · 2024
Later among the works it cites.
Atom: Low-bit quantization for efficient and accurate llm serving
Zhao, Y., Lin, C.-Y., Zhu, K., Ye, Z., Chen, L., Zheng, S., Ceze, L., Krishnamurthy, A., Chen, T., and Kasikci, B · 2024
Later among the works it cites.
Generalization vs. memorization: Tracing language models’ capabilities back to pretraining data
Antoniades, A., Wang, X., Elazar, Y., Amayuelas, A., Albalak, A., Zhang, K., and Wang, W. Y · 2025
Closest in time.
Round and round we go! what makes rotary positional encodings useful?
Barbero, F., Vitvitskyi, A., Perivolaropoulos, C., Pascanu, R., and Veličković, P · 2025
Closest in time.
Exploring concept depth: How large language models acquire knowledge and concept at different layers?
Jin, M., Yu, Q., Huang, J., Zeng, Q., Wang, Z., Hua, W., Zhao, H., Mei, K., Meng, Y., Ding, K., Yang, F., Du, M., and Zhang, Y · 2025
Closest in time.
Time series forecasting with llms: Understanding and enhancing model capabilities
Tang, H., Zhang, C., Jin, M., Yu, Q., Wang, Z., Jin, X., Zhang, Y., and Du, M · 2025
Closest in time.
Do larger language models imply better reasoning? a pretraining scaling law for reasoning
Wang, X., Tan, S., Jin, M., Wang, W. Y., Panda, R., and Shen, Y · 2025
Closest in time.