Fetching the paper…
Reading the bibliography…
Recent progress in Language Models (LMs) has dramatically advanced the field of natural language processing (NLP), excelling at tasks like text generation, summarization, and question answering.
H. S. Seung, M. Opper, and H. Sompolinsky, “Query by committee,” in Proceedings of the fifth annual workshop on Computational learning theory , 1992, pp. 287–294
1992
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, A. Y. Ng et al. , “Multimodal deep learning.” in ICML , vol. 11, 2011, pp. 689–696
2011
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012
2012
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3156–3164
2015
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
I. Goodfellow, “Deep learning,” 2016
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 2, pp. 423–443, 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proceedings of the conference. Association for computational linguistics. Meeting , vol. 2019. NIH Public Access, 2019, p. 6558
2019
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
R. Dale, “Gpt-3: What is it good for,” Nat. Lang. Eng , vol. 27, pp. 113–118, 2021
2021
Earlier work this paper cites.
G. Soyalp, A. Alar, K. Ozkanli, and B. Yildiz, “Improving text classification with transformer,” in 2021 6th International Conference on Computer Science and Engineering (UBMK) . IEEE, 2021, pp. 707–712
2021
Earlier work this paper cites.
S. Chaudhari, V. Mithal, G. Polatkan, and R. Ramanath, “An attentive survey of attention models,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 12, no. 5, pp. 1–32, 2021
2021
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International conference on machine learning . PMLR, 2021, pp. 10 347–10 357
2021
Earlier work this paper cites.
Z. Liu, Y. Xu, T. Yu, W. Dai, Z. Ji, S. Cahyawijaya, A. Madotto, and P. Fung, “Crossner: Evaluating cross-domain named entity recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 15, 2021, pp. 13 452–13 460
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” Advances in neural information processing systems , vol. 34, pp. 9694–9705, 2021
2021
Earlier work this paper cites.
Y. Guan, Z. Li, Z. Lin, Y. Zhu, J. Leng, and M. Guo, “Block-skim: Efficient question answering for transformer,” in Proceedings of the AAAI conference on artificial intelligence , vol. 36, no. 10, 2022, pp. 10 710–10 719
2022
Earlier work this paper cites.
A. de Santana Correia and E. L. Colombini, “Attention, please! a survey of neural attention models in deep learning,” Artificial Intelligence Review , vol. 55, no. 8, pp. 6037–6124, 2022
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
OpenAI and Microsoft, “Chatgpt,” https://chatgpt.com/ , 2022, accessed: date-of-access
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Sanderson, “Gpt-4 is here: what scientists think,” Nature , vol. 615, no. 7954, p. 773, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
E. Foundation, “Xtext language engineering for everyone,” https://eclipse.dev/Xtext/ , 2023, accessed: date-of-access
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Amazon, “Amazon titan in amazon bedrock,” https://aws.amazon.com/bedrock/amazon-models/ , 2023, accessed: date-of-access
2023
Cited alongside, same era.
Anthropic, “Claude,” https://claude.ai/ , 2023, accessed: date-of-access
2023
Cited alongside, same era.
MetaAI, “Llama,” https://www.llama.com/ , 2023, accessed: date-of-access
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
K. North, M. Zampieri, and M. Shardlow, “Lexical complexity prediction: An overview,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–42, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
R. AlSaad, A. Abd-Alrazaq, S. Boughorbel, A. Ahmed, M.-A. Renault, R. Damseh, and J. Sheikh, “Multimodal large language models in health care: applications, challenges, and future outlook,” Journal of medical Internet research , vol. 26, p. e59505, 2024
2024
Later among the works it cites.
B. He, H. Li, Y. K. Jang, M. Jia, X. Cao, A. Shah, A. Shrivastava, and S.-N. Lim, “Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 504–13 514
2024
Later among the works it cites.
Z. Liu, H. Zhang, K. Dong, and Y. Fang, “Collaborative cross-modal fusion with large language model for recommendation,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management , 2024, pp. 1565–1574
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
R. Karanjai and W. Shi, “Trusted llm inference on the edge with smart contracts,” in 2024 IEEE International Conference on Blockchain and Cryptocurrency (ICBC) . IEEE, 2024, pp. 1–7
2024
Later among the works it cites.
2024
Later among the works it cites.
I. C. Wiest, M.-E. Leßmann, F. Wolf, D. Ferber, M. Van Treeck, J. Zhu, M. P. Ebert, C. B. Westphalen, M. Wermke, and J. N. Kather, “Anonymizing medical documents with local, privacy preserving large language models: The llm-anonymizer,” medRxiv , pp. 2024–06, 2024
2024
Later among the works it cites.
W. Kuang, B. Qian, Z. Li, D. Chen, D. Gao, X. Pan, Y. Xie, Y. Li, B. Ding, and J. Zhou, “Federatedscope-llm: A comprehensive package for fine-tuning large language models in federated learning,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 5260–5271
2024
Later among the works it cites.
W. Yuan, C. Yang, G. Ye, T. Chen, N. Q. V. Hung, and H. Yin, “Fellas: Enhancing federated sequential recommendation with llm as external services,” ACM Transactions on Information Systems , 2024
2024
Later among the works it cites.
G. I. Kim, S. Hwang, and B. Jang, “Efficient compressing and tuning methods for large language models: A systematic literature review,” ACM Computing Surveys , vol. 57, no. 10, pp. 1–39, 2025
2025
Closest in time.
2025
Closest in time.
O. Dorémus, D. Russon, B. Contrand, A. Guerra-Adames, M. Avalos-Fernandez, C. Gil-Jardiné, E. Lagarde et al. , “Harnessing moderate-sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study,” JMIR AI , vol. 4, no. 1, p. e57828, 2025
2025
Closest in time.
R. O. Popov, N. V. Karpenko, and V. V. Gerasimov, “Overview of small language models in practice,” in CEUR Workshop Proceedings , 2025, pp. 164–182
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Zheng, Y. Chen, B. Qian, X. Shi, Y. Shu, and J. Chen, “A review on edge large language models: Design, execution, and applications,” ACM Computing Surveys , vol. 57, no. 8, pp. 1–35, 2025
2025
Closest in time.
I. Research. (2024) Llm routers: Efficient routing of large language models. Accessed: 2025-06-06. [Online]. Available: https://research.ibm.com/blog/LLM-routers?utm_source=chatgpt.com
2025
Closest in time.
P. AI. (2024) Llm routing: Ai costs optimization without sacrificing quality. Accessed: 2025-06-06. [Online]. Available: https://blog.premai.io/llm-routing-ai-costs-optimisation-without-sacrificing-quality/?utm_source=chatgpt.com
2025
Closest in time.
S. Jang and R. Morabito, “Edge-first language model inference: Models, metrics, and tradeoffs,” in Proceedings of the 2025 IEEE 45th International Conference on Distributed Computing Systems Workshops (ICDCSW) , 2025, to appear
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Rossello, “Llm hallucinations and personal data accuracy: can they really co-exist?” European Law Blog (https://www. europeanlawblog. eu/pub/2klfhf06/release/1) , 2025
2025
Closest in time.