Flashattention-2: Faster attention with better parallelism and work partitioning
Original
Tri Dao · 2023
Later among the works it cites.
Exaranker: Explanation-augmented neural ranker
Original
Fernando Ferraretto, Thiago Soares Laitz, Roberto de Alencar Lotufo, and Rodrigo Nogueira · 2023
Later among the works it cites.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Original
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister · 2023
Later among the works it cites.
Query expansion by prompting large language models
Original
Rolf Jagerman, Honglei Zhuang, Zhen Qin, Xuanhui Wang, and Michael Bendersky · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls, 2023
Original
Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma, Guoliang Li, Kevin C. C. Chang, Fei Huang, Reynold Cheng, and Yongbin Li · 2023
Later among the works it cites.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Later among the works it cites.
Query rewriting in retrieval-augmented large language models
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan · 2023
Later among the works it cites.
Generative relevance feedback with large language models
Iain Mackie, Shubham Chatterjee, and Jeffrey Dalton · 2023
Later among the works it cites.
Din-sql: Decomposed in-context learning of text-to-sql with self-correction, 2023
Original
Mohammadreza Pourreza and Davood Rafiei · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Original
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Original
Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Enhancing conversational search: Large language model-aided informative query rewriting
Original
Fanghua Ye, Meng Fang, Shenghui Li, and Emine Yilmaz · 2023
Later among the works it cites.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Later among the works it cites.
The faiss library
Original
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Later among the works it cites.
A survey on rag meeting llms: Towards retrieval-augmented large language models
Wenqi Fan, Yujuan Ding, Liang bo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li · 2024
Later among the works it cites.
The llama 3 herd of models
Original
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Grounding by trying: Llms with reinforcement learning-enhanced retrieval
Original
Sheryl Hsu, Omar Khattab, Chelsea Finn, and Archit Sharma · 2024
Later among the works it cites.
Gpt-4o system card
Original
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
Pet-sql: A prompt-enhanced two-round refinement of text-to-sql with cross-consistency, 2024
Original
Zhishuai Li, Xiang Wang, Jingjing Zhao, Sun Yang, Guoqing Du, Xiaoru Hu, Bin Zhang, Yuxiao Ye, Ziyue Li, Rui Zhao, and Hangyu Mao · 2024
Later among the works it cites.
Unleashing the power of llms as multi-modal encoders for text and graph-structured data
Original
Jiacheng Lin, Kun Qian, Haoyu Han, Nurendra Choudhary, Tianxin Wei, Zhongruo Wang, Sahika Genc, Edward W Huang, Sheng Wang, Karthik Subbian, et al · 2024
Later among the works it cites.
Simpo: Simple preference optimization with a reference-free reward
Original
Yu Meng, Mengzhou Xia, and Danqi Chen · 2024
Later among the works it cites.
Hybridflow: A flexible and efficient rlhf framework
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology, 2024
Original
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Later among the works it cites.
C-pack: Packed resources for general chinese embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie · 2024
Later among the works it cites.
Is dpo superior to ppo for llm alignment? a comprehensive study
Original
Shusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye, Weiling Liu, Zhiyu Mei, Guangju Wang, Chao Yu, and Yi Wu · 2024
Later among the works it cites.
Qwen2. 5 technical report
Original
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely
Original
Siyun Zhao, Yuqing Yang, Zilong Wang, Zhiyuan He, Luna K. Qiu, and Lili Qiu · 2024
Later among the works it cites.
Introducing claude 3.5 sonnet
Anthropic · 2025
Closest in time.
Sft memorizes, rl generalizes: A comparative study of foundation model post-training
Original
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V Le, Sergey Levine, and Yi Ma · 2025
Closest in time.
Flagembedding dataset
FlagOpen · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Original
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Next-generation database interfaces: A survey of llm-based text-to-sql, 2025
Original
Zijin Hong, Zheng Yuan, Qinggang Zhang, Hao Chen, Junnan Dong, Feiran Huang, and Xiao Huang · 2025
Closest in time.
Reinforce++: A simple and efficient approach for aligning large language models
Original
Jian Hu · 2025
Closest in time.
Llm post-training: A deep dive into reasoning large language models
Original
Komal Kumar, Tajamul Ashraf, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, Phillip HS Torr, Salman Khan, and Fahad Shahbaz Khan · 2025
Closest in time.
A foundation model for human-ai collaboration in medical literature mining, 2025
Original
Zifeng Wang, Lang Cao, Qiao Jin, Joey Chan, Nicholas Wan, Behdad Afzali, Hyun-Jin Cho, Chang-In Choi, Mehdi Emamverdi, Manjot K. Gill, Sun-Hyung Kim, Yijia Li, Yi Liu, Hanley Ong, Justin Rousseau, Irfan Sheikh, Jenny J. Wei, Ziyang Xu, Christopher M. Zallek, Kyungsang Kim, Yifan Peng, Zhiyong Lu, and Jimeng Sun · 2025
Closest in time.
When more is less: Understanding chain-of-thought length in llms, 2025
Original
Yuyang Wu, Yifei Wang, Tianqi Du, Stefanie Jegelka, and Yisen Wang · 2025
Closest in time.