Fetching the paper…
Reading the bibliography…
As large language models gain widespread adoption, running them efficiently becomes crucial.
A parallel priority queue with fast updates for GPU architectures
John Iacono, Ben Karsin, and Nodari Sitchinava · 1908
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 1912
Earlier work this paper cites.
A note on two problems in connexion with graphs
Edsger W. Dijkstra · 1959
Earlier work this paper cites.
High-probability parallel transitive closure algorithms
Jeffery Ullman and Mihalis Yannakakis · 1990
Earlier work this paper cites.
Using selective path-doubling for parallel shortest-path computations
Edith Cohen · 1993
Earlier work this paper cites.
A randomized parallel algorithm for single-source shortest paths
Philip N Klein and Sairam Subramanian · 1997
Earlier work this paper cites.
Time-work tradeoffs for parallel algorithms
Thomas H Spencer · 1997
Earlier work this paper cites.
Time–work tradeoffs of the single-source shortest paths problem
Hanmao Shi and Thomas H Spencer · 1999
Earlier work this paper cites.
Polylog-time and near-linear work approximation scheme for undirected shortest paths
Edith Cohen · 2000
Earlier work this paper cites.
Heaps are better than buckets: parallel shortest paths on unbalanced graphs
Ulrich Meyer · 2001
Earlier work this paper cites.
Training large neural networks with constant memory using a new execution algorithm
Bharadwaj Pudipeddi, Maral Mesmakhosroshahi, Jinwen Xi, and Sujeeth Bharadwaj · 2002
Earlier work this paper cites.
Lazy and speculative execution in computer systems
Butler Lampson · 2006
Earlier work this paper cites.
Accelerating large graph algorithms on the gpu using cuda
Pawan Harish and Petter J Narayanan · 2007
Earlier work this paper cites.
Pregel: a system for large-scale graph processing
Grzegorz Malewicz, Matthew H Austern, Aart JC Bik, James C Dehnert, Ilan Horn, Naty Leiser, and Grzegorz Czajkowski · 2010
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Audio chord recognition with recurrent neural networks
Nicolas Boulanger-Lewandowski, Yoshua Bengio, and Pascal Vincent · 2013
Earlier work this paper cites.
A lightweight infrastructure for graph analytics
Donald Nguyen, Andrew Lenharth, and Keshav Pingali · 2013
Earlier work this paper cites.
Work-efficient parallel gpu methods for single-source shortest paths
Andrew Davidson, Sean Baxter, Michael Garland, and John D. Owens · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Scott Beamer, Krste Asanović, and David Patterson · 2015
Earlier work this paper cites.
A top-down parallel semisort
Yan Gu, Julian Shun, Yihan Sun, and Guy E Blelloch · 2015
Earlier work this paper cites.
Parallel shortest paths using radius stepping
Guy E Blelloch, Yan Gu, Yihan Sun, and Kanat Tangwongsan · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Gunrock: A high-performance graph processing library on the gpu
Yangzihao Wang, Andrew Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, and John D Owens · 2016
Cited alongside, same era.
Gemini: A { \{ Computation-Centric } \} distributed graph processing system
Xiaowei Zhu, Wenguang Chen, Weimin Zheng, and Xiaosong Ma · 2016
Cited alongside, same era.
To push or to pull: On reducing communication and synchronization in graph computations
Maciej Besta, Michał Podstawski, Linus Groner, Edgar Solomonik, and Torsten Hoefler · 2017
Cited alongside, same era.
Julienne: A framework for parallel graph algorithms using work-efficient bucketing
Laxman Dhulipala, Guy Blelloch, and Julian Shun · 2017
Cited alongside, same era.
Rest: Retrieval-based speculative decoding
Zhenyu He, Zexuan Zhong, Tianle Cai, Jason D Lee, and Di He · 2023
Later among the works it cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Openassistant conversations – democratizing large language model alignment, 2023
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding, 2023
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit · 2018
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter Liu · 2020
Cited alongside, same era.
Optimizing ordered graph algorithms with graphit
Yunming Zhang, Ajay Brahmakshatriya, Xinyi Chen, Laxman Dhulipala, Shoaib Kamil, Saman Amarasinghe, and Julian Shun · 2020
Cited alongside, same era.
Prevent the language model from being overconfident in neural machine translation
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, and Jie Zhou · 2021
Cited alongside, same era.
Zero-offload: Democratizing billion-scale model training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2021
Cited alongside, same era.
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han · 2023
Later among the works it cites.
Xiaoxuan Liu, Lanxiang Hu, Peter Bailis, Ion Stoica, Zhijie Deng, Alvin Cheung, and Hao Zhang · 2023
Later among the works it cites.
Localllama, 2023
LocalLlama · 2023
Later among the works it cites.
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia · 2023
Later among the works it cites.
Flexgen: High-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang · 2023
Later among the works it cites.
Accelerating llm inference with staged speculative decoding
Benjamin Spector and Chris Re · 2023
Later among the works it cites.
Spectr: Fast speculative decoding via optimal transport
Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami, Himanshu Jain, and Felix Yu · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Llmcad: Fast and scalable on-device large language model inference
Daliang Xu, Wangsong Yin, Xin Jin, Ying Zhang, Shiyun Wei, Mengwei Xu, and Xuanzhe Liu · 2023
Later among the works it cites.
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date
Meta AI · 2024
Closest in time.
Sequoia: Sequoia: Serving exact llama2-70b on an rtx4090 with half-second per token latency, 2024b
Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen · 2024
Closest in time.
Which gpu(s) to get for deep learning: My experience and advice for using gpus in deep learning, 2023
Tim Dettmers · 2024
Closest in time.
Meta llama 2-70b chat generation config at hf.co/meta-llama/Llama-2-70b-chat-hf/blob/e1ce257/generation_config.json , 2024
Hugging Face · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Quip#: Quip with lattice codebooks, 2023
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa · 2024
Closest in time.