Fetching the paper…
Reading the bibliography…
Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models such as LLaMA and DeepSeek.
Episodic memory reader: Learning what to remember for question answering from streaming data, 2019
Moonsu Han, Minki Kang, Hyunwoo Jung, et al · 1903
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, et al · 1907
Earlier work this paper cites.
Amazonqa: A review-based question answering task, 2019
Mansi Gupta, Nitish Kulkarni, Raghuveer Chanda, et al · 1908
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering, 2019
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, et al · 1909
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models, 2020
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, et al · 1910
Earlier work this paper cites.
Root mean square layer normalization, 2019
Biao Zhang and Rico Sennrich · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale, 2020
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, et al · 1911
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, et al · 2001
Earlier work this paper cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2002
Earlier work this paper cites.
On layer normalization in the transformer architecture, 2020
Ruibin Xiong, Yunchang Yang, Di He, et al · 2002
Earlier work this paper cites.
Overcoming the lack of parallel data in sentence compression, October 2013
Katja Filippova and Yasemin Altun · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference, 2015
Samuel R. Bowman, Gabor Angeli, Christopher Potts, et al · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Richard Zemel, et al · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification, 2016
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts, 2017
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset, 2018
Payal Bajaj, Daniel Campos, Nick Craswell, et al · 2018
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset, 2018
Mahnaz Koupaee and William Yang Wang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, et al · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification, 2018
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, et al · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference, 2018
Adina Williams, Nikita Nangia, and Samuel R. Bowman · 2018
Cited alongside, same era.
Cloze-driven pretraining of self-attention networks
Alexei Baevski, Sergey Edunov, Yinhan Liu, et al · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, et al · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
Alex Wang, Amanpreet Singh, Julian Michael, et al · 2019
Yarn: Efficient context window extension of large language models, 2023
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, et al · 2023
Later among the works it cites.
In-context retrieval-augmented language models, 2023
Ori Ram, Yoav Levine, Itay Dalmedigos, et al · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, et al · 2023
Later among the works it cites.
Should You Mask 15% in Masked Language Modeling?, February 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Limits to Depth Efficiencies of Self-Attention
Yoav Levine, Noam Wies, Or Sharir, et al · 2020
Cited alongside, same era.
Gooaq: Open question answering with diverse answer types, 2021
Daniel Khashabi, Amos Ng, Tushar Khot, et al · 2021
Cited alongside, same era.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, et al · 2022
Cited alongside, same era.
Reasoning over public and private data in retrieval-based systems, 2022
Simran Arora, Patrick Lewis, Angela Fan, et al · 2022
Cited alongside, same era.
Simcse: Simple contrastive learning of sentence embeddings, 2022
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2022
Cited alongside, same era.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, et al · 2022
Cited alongside, same era.
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu · 2022
Cited alongside, same era.
Alexander Wettig, Tianyu Gao, Zexuan Zhong, et al · 2023
Later among the works it cites.
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al · 2024
Later among the works it cites.
Olmo: Accelerating the science of language models, 2024
Dirk Groeneveld, Iz Beltagy, Pete Walsh, et al · 2024
Later among the works it cites.
Enhancing training efficiency using packing with flash attention, 2024
Achintya Kundu, Rhui Dih Lee, Laura Wynter, et al · 2024
Later among the works it cites.
Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024
Rui Meng, Ye Liu, Shafiq Rayhan Joty, et al · 2024
Later among the works it cites.
Contextual document embeddings, 2024
John X. Morris and Alexander M. Rush · 2024
Later among the works it cites.
Nomic embed: Training a reproducible long context text embedder, 2024
Zach Nussbaum, John X. Morris, Brandon Duderstadt, et al · 2024
Later among the works it cites.
Gistembed: Guided in-sample selection of training negatives for text embedding fine-tuning, 2024
Aivin V. Solatorio · 2024
Later among the works it cites.
jina-embeddings-v3: Multilingual embeddings with task lora, 2024
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, et al · 2024
Later among the works it cites.
Benjamin Warner, Antoine Chaffin, Benjamin Clavié, et al · 2024
Later among the works it cites.
mGTE: Generalized long-context text representation and reranking models for multilingual text retrieval
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Daya Guo, Dejian Yang, et al · 2025
Closest in time.
Soap: Improving and stabilizing shampoo using adam, 2025
Nikhil Vyas, Depen Morwani, Rosie Zhao, et al · 2025
Closest in time.