Fetching the paper…
Reading the bibliography…
Large language models have achieved remarkable success in various tasks but suffer from high computational costs during inference, limiting their deployment in resource-constrained applications.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Dynamic Programming
Richard Bellman · 1957
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Deep learning using rectified linear units (relu), 2019
Abien Fred Agarap · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Specification gaming examples in ai
Jan Leike, Victoria Krakovna, et al · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
AdapterFusion: Non-destructive task composition for transfer learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
Token-level adaptation of lora adapters for downstream task generalization, 2023
Joshua Belofsky · 2023
Earlier work this paper cites.
Adaptersoup: Weight averaging to improve generalization of pretrained language models, 2023
Alexandra Chronopoulou, Matthew E. Peters, Alexander Fraser, and Jesse Dodge · 2023
Earlier work this paper cites.
Shizhe Diao, Tianyang Xu, Ruijia Xu, Jiawei Wang, and Tong Zhang · 2023
Cited alongside, same era.
Exploring the benefits of training expert language models over instruction tuning
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding, 2023
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
Routing to the expert: Efficient reward-guided ensemble of large language models
Towards robust qa evaluation via open llms
Ehsan Kamalloo, Shivani Upadhyay, and Jimmy Lin · 2024
Later among the works it cites.
Cllms: Consistency large language models, 2024
Siqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng, and Hao Zhang · 2024
Later among the works it cites.
Twin-merging: Dynamic integration of modular expertise in model merging, 2024
Zhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu, Dangyang Chen, and Yu Cheng · 2024
Later among the works it cites.
Routoo: Learning to route to large language models effectively, 2024
Alireza Mohammadshahi, Arshad Rafiq Shaikh, and Majid Yazdani · 2024
Later among the works it cites.
Learning to route among specialized experts for zero-shot generalization, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
Dynamic context pruning for efficient and interpretable autoregressive transformers, 2024
Sotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci, Aurelien Lucchi, and Thomas Hofmann · 2024
Cited alongside, same era.
Speculative streaming: Fast llm inference without auxiliary models, 2024
Nikhil Bhendawade, Irina Belousova, Qichen Fu, Henry Mason, Mohammad Rastegari, and Mahyar Najibi · 2024
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao · 2024
Cited alongside, same era.
Dam: Dynamic adapter merging for continual video qa learning, 2024
Feng Cheng, Ziyang Wang, Yi-Lin Sung, Yan-Bo Lin, Mohit Bansal, and Gedas Bertasius · 2024
Cited alongside, same era.
Llm-assisted rule based machine translation for low/no-resource languages
Jared Coleman, Bhaskar Krishnamachari, Khalil Iskarous, and Ruben Rosales · 2024
Cited alongside, same era.
Mohammed Muqeeth, Haokun Liu, Yufan Liu, and Colin Raffel · 2024
Later among the works it cites.
Faster cascades via speculative decoding, 2024
Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Seungyeon Kim, Neha Gupta, Aditya Krishna Menon, and Sanjiv Kumar · 2024
Later among the works it cites.
Routellm: Learning to route llms with preference data, 2024
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica · 2024
Later among the works it cites.
Towards modular llms by building and reusing a library of loras, 2024
Oleksiy Ostapenko, Zhan Su, Edoardo Maria Ponti, Laurent Charlin, Nicolas Le Roux, Matheus Pereira, Lucas Caccia, and Alessandro Sordoni · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
Learning to decode collaboratively with multiple language models
Shannon Zejiang Shen, Hunter Lang, Bailin Wang, Yoon Kim, and David Sontag · 2024
Later among the works it cites.
Harnessing the power of multiple minds: Lessons learned from LLM routing
Kv Aditya Srivatsa, Kaushal Maurya, and Ekaterina Kochmar · 2024
Later among the works it cites.
TensorOpera router: A multi-model router for efficient LLM inference
Dimitris Stripelis, Zhaozhuo Xu, Zijian Hu, Alay Dilipbhai Shah, Han Jin, Yuhang Yao, Jipeng Zhang, Tong Zhang, Salman Avestimehr, and Chaoyang He · 2024
Later among the works it cites.
Branch-train-mix: Mixing expert llms into a mixture-of-experts llm, 2024
Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma, Hu Xu, Xi Victoria Lin, Baptiste Rozière, Jacob Kahn, Daniel Li, Wen tau Yih, Jason Weston, and Xian Li · 2024
Later among the works it cites.
Xun Wu, Shaohan Huang, and Furu Wei · 2024
Later among the works it cites.
Meteora: Multiple-tasks embedded lora for large language models, 2024
Jingwei Xu, Junyu Lai, and Yunpeng Huang · 2024
Later among the works it cites.