Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved remarkable performance on various natural language tasks.
Long short-term memory
Hochreiter, S · 1997
Earlier work this paper cites.
Elements of information theory
Cover, T. M · 1999
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., and Vincent, P · 2000
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Bishop, C. M. and Nasrabadi, N. M · 2006
Earlier work this paper cites.
On early stopping in gradient descent learning
Yao, Y., Rosasco, L., and Caponnetto, A · 2007
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I · 2014
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks, 2016
Lamb, A., Goyal, A., Zhang, Y., Zhang, S., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2019
Earlier work this paper cites.
When does label smoothing help?
Müller, R., Kornblith, S., and Hinton, G. E · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Do massively pretrained language models make better storytellers?
See, A., Pappu, A., Saxena, R., Yerukola, A., and Manning, C. D · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Self-distillation as instance-specific label smoothing
Zhang, Z. and Sabuncu, M · 2020
Earlier work this paper cites.
Modifying memories in transformer models
Zhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F., and Kumar, S · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2021
Cited alongside, same era.
Editing factual knowledge in language models
De Cao, N., Aziz, W., and Titov, I · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2021
Cited alongside, same era.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E · 2022
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Zhou, C., Li, Q., Li, C., Yu, J., Liu, Y., Wang, G., Zhang, K., Ji, C., Yan, Q., He, L., et al · 2023
Later among the works it cites.
Large convolutional model tuning via filter subspace
Chen, W., Miao, Z., and Qiu, Q · 2024
Later among the works it cites.
Evaluating the ripple effects of knowledge editing in language models
Cohen, R., Biran, E., Yoran, O., Globerson, A., and Geva, M · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Aging with grace: Lifelong model editing with discrete key-value adaptors
Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y., and Ghassemi, M · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Truncation sampling as language model desmoothing
Hewitt, J., Manning, C. D., and Liang, P · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
Adaptive label smoothing with self-knowledge in natural language generation
Lee, D., Cheung, K. C., and Zhang, N. L · 2022
Cited alongside, same era.
Memory-based model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Later among the works it cites.
Learning to edit: Aligning llms with knowledge editing
Jiang, Y., Wang, Y., Wu, C., Zhong, W., Zeng, X., Gao, J., Li, L., Jiang, X., Shang, L., Tang, R., et al · 2024
Later among the works it cites.
Neighboring perturbations of knowledge editing on large language models
Ma, J.-Y., Ling, Z.-H., Zhang, N., and Gu, J.-C · 2024
Later among the works it cites.
In-context editing: Learning knowledge from self-induced distributions
Qi, S., Yang, B., Jiang, K., Wang, X., Li, J., Zhong, Y., Yang, Y., and Zheng, Z · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Knowledge editing in language models via adapted direct preference optimization
Rozner, A., Battash, B., Wolf, L., and Lindenbaum, O · 2024
Later among the works it cites.
Top- n σ n\sigma : Not all logits are you need
Tang, C., Liu, J., Xu, H., and Huang, L · 2024
Later among the works it cites.
Stable knowledge editing in large language models
Wei, Z., Pang, L., Ding, H., Deng, J., Shen, H., and Cheng, X · 2024
Later among the works it cites.
Reft: Representation finetuning for language models
Wu, Z., Arora, A., Wang, Z., Geiger, A., Jurafsky, D., Manning, C. D., and Potts, C · 2024
Later among the works it cites.
Xu, R., Liu, H., Nag, S., Dai, Z., Xie, Y., Tang, X., Luo, C., Li, Y., Ho, J. C., Yang, C., et al · 2024
Later among the works it cites.
Melo: Enhancing model editing with neuron-indexed dynamic lora
Yu, L., Chen, Q., Zhou, J., and He, L · 2024
Later among the works it cites.
On the generalization of language models from in-context learning and finetuning: a controlled study
Lampinen, A. K., Chaudhry, A., Chan, S. C., Wild, C., Wan, D., Ku, A., Bornschein, J., Pascanu, R., Shanahan, M., and McClelland, J. L · 2025
Closest in time.
Coeff-tuning: A graph filter subspace view for tuning attention-based large models
Miao, Z., Chen, W., and Qiu, Q · 2025
Closest in time.
Xu, R., Shi, W., Zhuang, Y., Yu, Y., Ho, J. C., Wang, H., and Yang, C · 2025
Closest in time.
Rankrag: Unifying context ranking with retrieval-augmented generation in llms
Yu, Y., Ping, W., Liu, Z., Wang, B., You, J., Zhang, C., Shoeybi, M., and Catanzaro, B · 2025
Closest in time.