Fetching the paper…
Reading the bibliography…
As the field of Large Language Models (LLMs) continues to evolve, the context length in inference is steadily growing.
A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor Search
Mengzhao Wang, Xiaoliang Xu, Qiang Yue, and Yuxiang Wang. 2021 · 1978
Earlier work this paper cites.
Towards Efficient Index Construction and Approximate Nearest Neighbor Search in High-Dimensional Spaces
Xi Zhao, Yao Tian, Kai Huang, Bolong Zheng, and Xiaofang Zhou. 2023 · 1991
Earlier work this paper cites.
Product Quantization for Nearest Neighbor Search
Hervé Jégou, Matthijs Douze, and Cordelia Schmid. 2011 · 2011
Earlier work this paper cites.
Optimized Product Quantization for Approximate Nearest Neighbor Search. In IEEE Conference on Computer Vision and Pattern Recognition
Tiezheng Ge, Kaiming He, Qifa Ke, and Jian Sun. 2013 · 2013
Earlier work this paper cites.
Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan
Fabien André, Anne-Marie Kermarrec, and Nicolas Le Scouarnec. 2015 · 2015
Earlier work this paper cites.
TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016 · 2016
Earlier work this paper cites.
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design. In 49th Annual IEEE/ACM International Symposium on Microarchitecture, MICRO
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W. Keckler. 2016 · 2016
Earlier work this paper cites.
PQBF: I/O-Efficient Approximate Nearest Neighbor Search by Product Quantization. In Proceedings of the Conference on Information and Knowledge Management, CIKM
Yingfan Liu, Hong Cheng, and Jiangtao Cui. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning. In Advances in Neural Information Processing Systems, NeurIPS
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Revisiting the Inverted Indices for Billion-Scale Approximate Nearest Neighbors. In Computer Vision - ECCV - 15th European Conference
Dmitry Baranchuk, Artem Babenko, and Yury Malkov. 2018 · 2018
Earlier work this paper cites.
Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
A Generic Inverted Index Framework for Similarity Search on the GPU. In 34th IEEE International Conference on Data Engineering, ICDE
Jingbo Zhou, Qi Guo, H. V. Jagadish, Lubos Krcál, Siyuan Liu, Wenhao Luan, Anthony K. H. Tung, Yueji Yang, and Yuxin Zheng. 2018 · 2018
Earlier work this paper cites.
Dynamic Memory Management for GPU-Based Training of Deep Neural Networks. In IEEE International Parallel and Distributed Processing Symposium, IPDPS
Shriram S. B, Anshuj Garg, and Purushottam Kulkarni. 2019 · 2019
Earlier work this paper cites.
Fast Approximate Nearest Neighbor Search With The Navigating Spreading-out Graph
Cong Fu, Chao Xiang, Changxu Wang, and Deng Cai. 2019 · 2019
Earlier work this paper cites.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In Advances in Neural Information Processing Systems, NeurIPS
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019 · 2019
Earlier work this paper cites.
Diskann: Fast accurate billion-point nearest neighbor search on a single node. In Advances in Neural Information Processing Systems, NeurIPS
Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi. 2019 · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems, NeurIPS
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
Differentiable Product Quantization for End-to-End Embedding Compression. In International Conference on Machine Learning, ICML
Ting Chen, Lala Li, and Yizhou Sun. 2020 · 2020
Earlier work this paper cites.
Effective and Efficient Retrieval of Structured Entities
Ruihong Huang, Shaoxu Song, Yunsu Lee, Jungho Park, Soo-Hyung Kim, and Sungmin Yi. 2020 · 2020
Earlier work this paper cites.
Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs
Yury A. Malkov and Dmitry A. Yashunin. 2020 · 2020
Earlier work this paper cites.
Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems. In The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
Hao-Jun Michael Shi, Dheevatsa Mudigere, Maxim Naumov, and Jiyan Yang. 2020 · 2020
Earlier work this paper cites.
DeltaPQ: Lossless Product Quantization Code Compression for High Dimensional Similarity Search
Runhui Wang and Dong Deng. 2020 · 2020
Earlier work this paper cites.
Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
Xiaofei Wang, Yiwen Han, Victor C. M. Leung, Dusit Niyato, Xueqiang Yan, and Xu Chen. 2020 · 2020
Earlier work this paper cites.
PASE: PostgreSQL Ultra-High-Dimensional Approximate Nearest Neighbor Search Extension. In International Conference on Management of Data, SIGMOD
Wen Yang, Tao Li, Gai Fang, and Hong Wei. 2020 · 2020
Earlier work this paper cites.
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, and et al. 2021 · 2021
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In 9th International Conference on Learning Representations, ICLR
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Earlier work this paper cites.
Accelerating Product Quantization Query Execution Runtime. In International Conference on Management of Data, SIGMOD
Ikraduya Edian. 2021 · 2021
Earlier work this paper cites.
Billion-Scale Similarity Search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2021 · 2021
Earlier work this paper cites.
HVS: Hierarchical Graph Structure Based on Voronoi Diagrams for Solving Approximate Nearest Neighbor Search
Kejing Lu, Mineichi Kudo, Chuan Xiao, and Yoshiharu Ishikawa. 2021 · 2021
Earlier work this paper cites.
ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning. In International Conference for High Performance Computing, Networking, Storage and Analysis, SC
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021 · 2021
Earlier work this paper cites.
ZeRO-Offload: Democratizing Billion-Scale Model Training. In USENIX Annual Technical Conference, ATC
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021 · 2021
Earlier work this paper cites.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In Advances in Neural Information Processing Systems, NeurIPS
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan. 2022 · 2022
Earlier work this paper cites.
Exploring Query Processing on CPU-GPU Integrated Edge Device
Jiesong Liu, Feng Zhang, Hourun Li, Dalin Wang, Weitao Wan, Xiaokun Fang, Jidong Zhai, and Xiaoyong Du. 2022 · 2022
Earlier work this paper cites.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
TSPLIT: Fine-grained GPU Memory Management for Efficient DNN Training via Tensor Splitting. In 38th IEEE International Conference on Data Engineering, ICDE
Xiaonan Nie, Xupeng Miao, Zhi Yang, and Bin Cui. 2022 · 2022
Earlier work this paper cites.
Survey: Transformer based video-language pre-training
Ludan Ruan and Qin Jin. 2022 · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems, NeurIPS
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Implementation, NSDI
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022 · 2022
Cited alongside, same era.
Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense Embeddings. In The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval
Shitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang, Defu Lian, Yeyun Gong, Qi Chen, Fan Yang, Hao Sun, Yingxia Shao, and Xing Xie. 2022 · 2022
Cited alongside, same era.
The Llama 3 Herd of Models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurélien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Rozière, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Grégoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel M. Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, and et al. 2024 · 2024
Closest in time.
Data Engineering for Scaling Language Models to 128K Context. In International Conference on Machine Learning, ICML
Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Yoon Kim, and Hao Peng. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation, OSDI
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Cited alongside, same era.
The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4
Microsoft Research AI4Science and Microsoft Azure Quantum. 2023 · 2023
Cited alongside, same era.
Long context prompting for Claude 2.1
Anthropic. 2023 · 2023
Cited alongside, same era.
Elpis: Graph-Based Similarity Search for Scalable Data Science
Ilias Azizi, Karima Echihabi, and Themis Palpanas. 2023 · 2023
Cited alongside, same era.
Accurate medium-range global weather forecasting with 3D neural networks
Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. 2023 · 2023
Cited alongside, same era.
Transformer-based deep learning for predicting protein properties in the life sciences
Abel Chandra, Laura Tünnermann, Tommy Löfstedt, and Regina Gratz. 2023 · 2023
Cited alongside, same era.
Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models’ Reasoning Performance
Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot. 2023 · 2023
Cited alongside, same era.
LM-Infinite: Simple On-the-Fly Length Generalization for Large Language Models
Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. 2023 · 2023
Cited alongside, same era.
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024 · 2024
Closest in time.
RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search. In International Conference on Management of Data, SIGMOD
Jianyang Gao and Cheng Long. 2024 · 2024
Closest in time.
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs. In 12th International Conference on Learning Representations, ICLR
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. 2024 · 2024
Closest in time.
Advances of pipeline model parallelism for deep learning training: an overview
Lei Guan, Dong-Sheng Li, Ji-Ye Liang, Wen-Jian Wang, Ke-Shi Ge, and Xi-Cheng Lu. 2024 · 2024
Closest in time.
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization. In Advances in Neural Information Processing Systems, NeurIPS
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024 · 2024
Closest in time.
Large Language Models for Software Engineering: A Systematic Literature Review
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024 · 2024
Closest in time.
Characterization of Large Language Model Development in the Datacenter. In 21st USENIX Symposium on Networked Systems Design and Implementation, NSDI
Qinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang, Meng Zhang, Qiaoling Chen, Peng Sun, Dahua Lin, Xiaolin Wang, Yingwei Luo, Yonggang Wen, and Tianwei Zhang. 2024 · 2024
Closest in time.
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2024 · 2024
Closest in time.
Needle-in-a-Haystack
Greg Kamradt. 2024 · 2024
Closest in time.
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, and Tuo Zhao. 2024 · 2024
Closest in time.
SnapKV: LLM Knows What You are Looking for Before Generation
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. 2024 · 2024
Closest in time.
Transformer-VQ: Linear-Time Transformers via Vector Quantization. In 12th International Conference on Learning Representations, ICLR
Lucas D. Lingle. 2024 · 2024
Closest in time.
World Model on Million-Length Video And Language With Blockwise RingAttention
Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. 2024b · 2024
Closest in time.
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression. In Findings of the Association for Computational Linguistics, ACL
Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. 2024 · 2024
Closest in time.
Qwen2.5: A Party of Foundation Models
Qwen Team. 2024 · 2024
Closest in time.
On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
Siyu Ren and Kenny Q. Zhu. 2024 · 2024
Closest in time.
SparQ Attention: Bandwidth-Efficient LLM Inference. In International Conference on Machine Learning, ICML
Luka Ribar, Ivan Chelombiev, Luke Hudlass-Galley, Charlie Blake, Carlo Luschi, and Douglas Orr. 2024 · 2024
Closest in time.
A communication theory perspective on prompting engineering methods for large language models
Yuan-Feng Song, Yuan-Qin He, Xue-Fang Zhao, Han-Lin Gu, Di Jiang, Hai-Jun Yang, and Li-Xin Fan. 2024 · 2024
Closest in time.
DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving. In International Conference on Machine Learning, ICML
Foteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski, and Ana Klimovic. 2024 · 2024
Closest in time.
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
Hanlin Tang, Yang Lin, Jing Lin, Qingsen Han, Shikuan Hong, Yiwu Yao, and Gongyi Wang. 2024 · 2024
Closest in time.
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
Zihao Wang, Bin Cui, and Shaoduo Gan. 2024 · 2024
Closest in time.
Scalable Distributed Inverted List Indexes in Disaggregated Memory. In International Conference on Management of Data, SIGMOD
Manuel Widmoser, Daniel Kocher, and Nikolaus Augsten. 2024 · 2024
Closest in time.
A survey on large language models for recommendation
Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, Hui Xiong, and Enhong Chen. 2024 · 2024
Closest in time.
Chaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao, Yankai Lin, Zhengyan Zhang, Zhiyuan Liu, Song Han, and Maosong Sun. 2024b · 2024
Closest in time.
June Yong Yang, Byeongwook Kim, Jeongin Bae, Beomseok Kwon, Gunho Park, Eunho Yang, Se Jung Kwon, and Dongsoo Lee. 2024 · 2024
Closest in time.
Deep Learning Workload Scheduling in GPU Datacenters: A Survey
Zhisheng Ye, Wei Gao, Qinghao Hu, Peng Sun, Xiaolin Wang, Yingwei Luo, Tianwei Zhang, and Yonggang Wen. 2024 · 2024
Closest in time.
Deep learning for code generation: a survey
Huangzhao Zhang, Kechi Zhang, Zhuo Li, Jia Li, Jia Li, Yongmin Li, Yunfei Zhao, Yuqi Zhu, Fang Liu, Ge Li, et al · 2024
Closest in time.
Open-Sora: Democratizing Efficient Video Production for All
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. 2024 · 2024
Closest in time.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024 · 2024
Closest in time.
Db-gpt: Large language model meets database
Xuanhe Zhou, Zhaoyan Sun, and Guoliang Li. 2024 · 2024
Closest in time.
Relational Data Cleaning Meets Artificial Intelligence: A Survey
Jingyu Zhu, Xintong Zhao, Yu Sun, Shaoxu Song, and Xiaojie Yuan. 2024 · 2024
Closest in time.
Deep learning-based software engineering: progress, challenges, and opportunities
Xiangping Chen, Xing Hu, Yuan Huang, He Jiang, Weixing Ji, Yanjie Jiang, Yanyan Jiang, Bo Liu, Hui Liu, Xiaochen Li, et al · 2025
Closest in time.
SlimDB: A Space-Efficient Key-Value Storage Engine For Semi-Sorted Data
Kai Ren, Qing Zheng, Joy Arulraj, and Garth Gibson. 2017 · 2048
Closest in time.