Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) have exhibited remarkable emergent capabilities through extensive pre-training, they still face critical limitations in generalizing to specialized domains and handling diverse linguistic variations, known as distribution shifts.
A neural probabilistic language model
Bengio, Y., Ducharme, R., and Vincent, P · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Earlier work this paper cites.
Evaluating prediction-time batch normalization for robustness under covariate shift
Nado, Z., Padhy, S., Sculley, D., D’Amour, A., Lakshminarayanan, B., and Snoek, J · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Improving robustness against common corruptions by covariate shift adaptation
Schneider, S., Rusak, E., Eck, L., Bringmann, O., Brendel, W., and Bethge, M · 2020
Earlier work this paper cites.
Test-time training with self-supervision for generalization under distribution shifts
Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A., and Hardt, M · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Test time adaptation through perturbation robustness
Fleuret, F. et al · 2021
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H · 2021
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Earlier work this paper cites.
Test-time classifier adjustment module for model-agnostic domain generalization
Iwasawa, Y. and Matsuo, Y · 2021
Earlier work this paper cites.
Ttt++: When does self-supervised test-time training fail or thrive?
Liu, Y., Kothari, P., Van Delft, B., Bellot-Gurlet, B., Mordan, T., and Alahi, A · 2021
Earlier work this paper cites.
Tent: Fully test-time adaptation by entropy minimization
Wang, D., Shelhamer, E., Liu, S., Olshausen, B., and Darrell, T · 2021
Earlier work this paper cites.
Mt3: Meta test-time training for self-supervised test-time adaption
Bartler, A., Bühler, A., Wiewel, F., Döbler, M., and Yang, B · 2022
Earlier work this paper cites.
Parameter-free online test-time adaptation
Boudiaf, M., Mueller, R., Ben Ayed, I., and Bertinetto, L · 2022
Cited alongside, same era.
GLM: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Cited alongside, same era.
Test-time training with masked autoencoders
Gandelsman, Y., Sun, Y., Chen, X., and Efros, A · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2022
Cited alongside, same era.
Test-time prompt tuning for zero-shot generalization in vision-language models
Shu, M., Nie, W., Huang, D.-A., Yu, Z., Goldstein, T., Anandkumar, A., and Xiao, C · 2022
Cited alongside, same era.
Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions
Trivedi, H., Balasubramanian, N., Khot, T., and Sabharwal, A · 2023
Later among the works it cites.
Self-knowledge guided retrieval augmentation for large language models
Wang, Y., Li, P., Sun, M., and Liu, Y · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.
The surprising effectiveness of test-time training for abstract reasoning
Akyürek, E., Damani, M., Qiu, L., Guo, H., Kim, Y., and Andreas, J · 2024
Later among the works it cites.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Cited alongside, same era.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Back to the source: Diffusion-driven adaptation to test-time corruption
Gao, J., Zhang, J., Liu, X., Darrell, T., Shelhamer, E., and Wang, D · 2023
Cited alongside, same era.
Mecta: Memory-economic continual test-time model adaptation
Hong, J., Lyu, L., Zhou, J., and Spranger, M · 2023
Cited alongside, same era.
Active retrieval augmented generation
Jiang, Z., Xu, F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
Tackling language modelling bias in support of linguistic diversity
Bella, G., Helm, P., Koch, G., and Giunchiglia, F · 2024
Later among the works it cites.
A survey on evaluation of large language models
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al · 2024
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
A survey on rag meeting llms: Towards retrieval-augmented large language models
Fan, W., Ding, Y., Ning, L., Wang, S., Li, H., Yin, D., Chua, T.-S., and Li, Q · 2024
Later among the works it cites.
Test-time training on nearest neighbors for large language models
Hardt, M. and Sun, Y · 2024
Later among the works it cites.
Efficiently learning at test-time: Active fine-tuning of llms
Hübotter, J., Bongni, S., Hakimi, I., and Krause, A · 2024
Later among the works it cites.
Longrag: Enhancing retrieval-augmented generation with long-context llms
Jiang, Z., Ma, X., and Chen, W · 2024
Later among the works it cites.
Entropy is not enough for test-time adaptation: From the perspective of disentangled factors
Lee, J., Jung, D., Lee, S., Park, J., Shin, J., Hwang, U., and Yoon, S · 2024
Later among the works it cites.
A comprehensive survey on test-time adaptation under distribution shifts
Liang, J., He, R., and Tan, T · 2024
Later among the works it cites.
RA-DIT: Retrieval-augmented dual instruction tuning
Lin, X. V., Chen, X., Chen, M., Shi, W., Lomeli, M., James, R., Rodriguez, P., Kahn, J., Szilvasy, G., Lewis, M., Zettlemoyer, L., and tau Yih, W · 2024
Later among the works it cites.
Test-time model adaptation with only forward passes
Niu, S., Miao, C., Chen, G., Wu, P., and Zhao, P · 2024
Later among the works it cites.
Memorag: Moving towards next-gen rag via memory-inspired knowledge discovery
Qian, H., Zhang, P., Liu, Z., Mao, K., and Dou, Z · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
From question to exploration: Can classic test-time adaptation strategies be effectively applied in semantic segmentation?
Yi, C., Chen, H., Zhang, Y., Xu, Y., Zhou, Y., and Cui, L · 2024
Later among the works it cites.
Memorybank: Enhancing large language models with long-term memory
Zhong, W., Guo, L., Gao, Q., Ye, H., and Wang, Y · 2024
Later among the works it cites.
Efficient diffusion-driven corruption editor for test-time adaptation
Oh, Y., Lee, J., Choi, J., Jung, D., Hwang, U., and Yoon, S · 2025
Closest in time.
Generating long-form story using dynamic hierarchical outlining with memory-enhancement
Wang, Q., Hu, J., Li, Z., Wang, Y., Hu, Y., Tan, M., et al · 2025
Closest in time.
Come: Test-time adaption by conservatively minimizing entropy
Zhang, Q., Bian, Y., Kong, X., Zhao, P., and Zhang, C · 2025
Closest in time.