Fetching the paper…
Reading the bibliography…
In Large Language Model (LLM) inference, the output length of an LLM request is typically regarded as not known a priori.
Pranking with ranking
Koby Crammer and Yoram Singer · 2001
Earlier work this paper cites.
An efficient boosting algorithm for combining preferences
Yoav Freund, Raj Iyer, Robert E Schapire, and Yoram Singer · 2003
Earlier work this paper cites.
Learning to rank using gradient descent
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender · 2005
Earlier work this paper cites.
Subset ranking using regression
David Cossock and Tong Zhang · 2006
Earlier work this paper cites.
Learning to rank with nonsmooth cost functions
Christopher Burges, Robert Ragno, and Quoc Le · 2006
Earlier work this paper cites.
Mcrank: Learning to rank using multiple classification and gradient boosting
Ping Li, Qiang Wu, and Christopher Burges · 2007
Earlier work this paper cites.
A regression framework for learning ranking functions using relative relevance judgments
Zhaohui Zheng, Keke Chen, Gordon Sun, and Hongyuan Zha · 2007
Earlier work this paper cites.
Adarank: a boosting algorithm for information retrieval
Jun Xu and Hang Li · 2007
Earlier work this paper cites.
Learning to rank: from pairwise approach to listwise approach
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li · 2007
Earlier work this paper cites.
Softrank: optimizing non-smooth rank metrics
Michael Taylor, John Guiver, Stephen Robertson, and Tom Minka · 2008
Earlier work this paper cites.
Listwise approach to learning to rank: theory and algorithm
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li · 2008
Earlier work this paper cites.
Learning to rank for information retrieval
Tie-Yan Liu et al · 2009
Earlier work this paper cites.
Adapting boosting for information retrieval measures
Qiang Wu, Christopher JC Burges, Krysta M Svore, and Jianfeng Gao · 2010
Earlier work this paper cites.
From ranknet to lambdarank to lambdamart: An overview
Christopher JC Burges · 2010
Earlier work this paper cites.
Learning to rank for recommender systems
Alexandros Karatzoglou, Linas Baltrunas, and Yue Shi · 2013
Earlier work this paper cites.
Ranking relevance in yahoo search
Dawei Yin, Yuening Hu, Jiliang Tang, Tim Daly, Mianwei Zhou, Hua Ouyang, Jianhui Chen, Changsung Kang, Hongbo Deng, Chikashi Nobata, et al · 2016
Earlier work this paper cites.
Learning to optimize tensor programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Neuralndcg: Direct optimisation of a ranking metric via differentiable relaxation of sorting
Przemysław Pobrotyn and Radosław Białobrzeski · 2021
Cited alongside, same era.
Introducing chatgpt
OpenAI · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
s 3 s^{3} : Increasing gpu utilization during generative inference for higher throughput
Yunho Jin, Chun-Feng Wu, David Brooks, and Gu-Yeon Wei · 2024
Closest in time.
Foteini Strati, Sara Mcallister, Amar Phanishayee, Jakub Tarnawski, and Ana Klimovic · 2024
Closest in time.
Inference without interference: Disaggregate llm inference for mixed downstream workloads
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, et al · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You · 2023
Cited alongside, same era.
Splitwise: Efficient generative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini · 2023
Cited alongside, same era.
Fast distributed inference serving for large language models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin · 2023
Cited alongside, same era.
Fairness in serving large language models
Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
https://sharegpt.com/, 2023
ShareGPT Team · 2023
Cited alongside, same era.
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, Alexey Tumanov, and Ramachandran Ramjee · 2024
Closest in time.
Andes: Defining and enhancing quality-of-experience in llm-based text streaming services
Jiachen Liu, Zhiyu Wu, Jae-Won Chung, Fan Lai, Myungjin Lee, and Mosharaf Chowdhury · 2024
Closest in time.
Dynamollm: Designing llm inference clusters for performance and energy efficiency
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse · 2024
Closest in time.
Enabling efficient batch serving for lmaas via generation length prediction
Ke Cheng, Wen Hu, Zhi Wang, Peng Du, Jianguo Li, and Sheng Zhang · 2024
Closest in time.
Efficient interactive llm serving with proxy model-based sequence length prediction
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T Kalbarczyk, Tamer Başar, and Ravishankar K Iyer · 2024
Closest in time.
Power-aware deep learning model serving with { \{ μ \mu -Serve } \}
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew Kalbarczyk, Tamer Başar, and Ravishankar K Iyer · 2024
Closest in time.
Lero: applying learning-to-rank in query optimizer
Xingguang Chen, Rong Zhu, Bolin Ding, Sibo Wang, and Jingren Zhou · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
AI Meta · 2024
Closest in time.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You · 2024
Closest in time.
Towards efficient and reliable llm serving: A real-world workload study
Yuxin Wang, Yuhan Chen, Zeyu Li, Zhenheng Tang, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, and Xiaowen Chu · 2024
Closest in time.
Cosmopedia, February 2024
Loubna Ben Allal, Anton Lozhkov, Guilherme Penedo, Thomas Wolf, and Leandro von Werra · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2024
Closest in time.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto · 2024
Closest in time.