Fetching the paper…
Reading the bibliography…
An evolving solution to address hallucination and enhance accuracy in large language models (LLMs) is Retrieval-Augmented Generation (RAG), which involves augmenting LLMs with information retrieved from an external knowledge source, such as the web.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 1906
Earlier work this paper cites.
A comprehensive survey and experimental comparison of graph-based approximate nearest neighbor search
Mengzhao Wang, Xiaoliang Xu, Qiang Yue, and Yuxiang Wang. 2021 · 1978
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Retrieval-Augmented Knowledge-Intensive Dialogue. In Natural Language Processing and Chinese Computing , Fei Liu, Nan Duan, Qingting Xu, and Yu Hong (Eds.). Springer Nature Switzerland, Cham, 16–28
Zelin Wang, Ping Gong, Yibo Zhang, Jihao Gu, and Xuanyuan Yang. 2023 · 2005
Earlier work this paper cites.
Product Quantization for Nearest Neighbor Search
Herve Jégou, Matthijs Douze, and Cordelia Schmid. 2011 · 2010
Earlier work this paper cites.
The inverted multi-index. In 2012 IEEE Conference on Computer Vision and Pattern Recognition . 3069–3076
Artem Babenko and Victor Lempitsky. 2012 · 2012
Earlier work this paper cites.
1.1 Computing’s energy problem (and what we can do about it). In 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) . 10–14
Mark Horowitz. 2014 · 2014
Earlier work this paper cites.
Interconnect-Memory Challenges for Multi-chip, Silicon Interposer Systems. In Proceedings of the 2015 International Symposium on Memory Systems (Washington DC, DC, USA) (MEMSYS ’15) . Association for Computing Machinery, New York, NY, USA, 3–10
Gabriel H. Loh, Natalie Enright Jerger, Ajaykumar Kannan, and Yasuko Eckert. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Carreras (Eds.). Association for Computational Linguistics, Austin, Texas, 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Multi-chip technologies to unleash computing performance gains over the next decade. In 2017 IEEE International Electron Devices Meeting (IEDM) . 1.1.1–1.1.8
Lisa T. Su, Samuel Naffziger, and Mark Papermaster. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Application-Transparent Near-Memory Processing Architecture with Memory Channel Network. In 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 802–814
Mohammad Alian, Seung Won Min, Hadi Asgharimoghaddam, Ashutosh Dhar, Dong Kai Wang, Thomas Roewer, Adam McPadden, Oliver O’Halloran, Deming Chen, Jinjun Xiong, et al · 2018
Earlier work this paper cites.
Search engine guided neural machine translation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32
Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li. 2018 · 2018
Earlier work this paper cites.
A Retrieve-and-Edit Framework for Predicting Structured Outputs. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Montréal, Canada) (NIPS’18) . Curran Associates Inc., Red Hook, NY, USA, 10073–10083
Tatsunori B. Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs
Yu A. Malkov and D. A. Yashunin. 2020 · 2018
Earlier work this paper cites.
Retrieve and Refine: Improved Sequence Generation Models For Dialogue. In Proceedings of the 2018 EMNLP Workshop SCAI: The 2nd International Workshop on Search-Oriented Conversational AI , Aleksandr Chuklin, Jeff Dalton, Julia Kiseleva, Alexey Borisov, and Mikhail Burtsev (Eds.). Association for Computational Linguistics, Brussels, Belgium, 87–92
Jason Weston, Emily Dinan, and Alexander Miller. 2018 · 2018
Earlier work this paper cites.
Guiding Neural Machine Translation with Retrieved Translation Pieces. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , Marilyn Walker, Heng Ji, and Amanda Stent (Eds.). Association for Computational Linguistics, New Orleans, Louisiana, 1325–1335
Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, and Satoshi Nakamura. 2018 · 2018
Earlier work this paper cites.
Skeleton-to-Response: Dialogue Generation Guided by Retrieval Memory. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 1219–1228
Deng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Wai Lam, and Shuming Shi. 2019 · 2019
Earlier work this paper cites.
Natural Questions: a Benchmark for Question Answering Research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Approximate Nearest Neighbor Search on High Dimensional Data — Experiments, Analyses, and Improvement
Wen Li, Ying Zhang, Yifang Sun, Wei Wang, Mingjie Li, Wenjie Zhang, and Xuemin Lin. 2020b · 2019
Earlier work this paper cites.
Text Generation with Exemplar-based Adaptive Decoding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 2555–2565
Hao Peng, Ankur Parikh, Manaal Faruqui, Bhuwan Dhingra, and Dipanjan Das. 2019 · 2019
Earlier work this paper cites.
Learning to Abstract for Memory-augmented Conversational Response Generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, Florence, Italy, 3816–3825
Zhiliang Tian, Wei Bi, Xiaopeng Li, and Nevin L. Zhang. 2019 · 2019
Earlier work this paper cites.
Response Generation by Context-Aware Prototype Editing
Yu Wu, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li, and Ming Zhou. 2019 · 2019
Earlier work this paper cites.
ANN-Benchmarks: A benchmarking tool for approximate nearest neighbor algorithms
Martin Aumüller, Erik Bernhardsson, and Alexander Faithfull. 2020 · 2020
Earlier work this paper cites.
Domain-specific hardware accelerators
William J. Dally, Yatish Turakhia, and Song Han. 2020 · 2020
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 6769–6781
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Paraphrase Generation by Learning How to Edit from Samples. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 6010–6021
Amirhossein Kazemnejad, Mohammadreza Salehi, and Mahdieh Soleymani Baghshah. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20) . Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
Improving Approximate Nearest Neighbor Search through Learned Adaptive Early Termination. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data (Portland, OR, USA) (SIGMOD ’20) . Association for Computing Machinery, New York, NY, USA, 2539–2554
Conglong Li, Minjia Zhang, David G. Andersen, and Yuxiong He. 2020a · 2020
Earlier work this paper cites.
2.2 AMD Chiplet Architecture for High-Performance Server and Desktop Products. In 2020 IEEE International Solid- State Circuits Conference - (ISSCC) . 44–45
Samuel Naffziger, Kevin Lepak, Milam Paraschou, and Mahesh Subramony. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Boosting Neural Machine Translation with Similar Translations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 1580–1590
Jitao Xu, Josep Crego, and Jean Senellart. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Transformer-XL for Close-Domain Dialog Generation
Giovanni Bonetta, Rossella Cancelliere, Ding Liu, and Paul Vozila. 2021 · 2021
Earlier work this paper cites.
Memory-Augmented Image Captioning
Zhengcong Fei. 2021 · 2021
Earlier work this paper cites.
Fast and Accurate Neural Machine Translation with Translation Memory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 3170–3180
Qiuxiang He, Guoping Huang, Qu Cui, Li Li, and Lemao Liu. 2021 · 2021
Earlier work this paper cites.
Analyzing and leveraging decoupled L1 caches in GPUs. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 467–478
Mohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H Loh, and Adwait Jog. 2021 · 2021
Cited alongside, same era.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty (Eds.). Association for Computational Linguistics, Online, 874–880
Gautier Izacard and Edouard Grave. 2021b · 2021
Cited alongside, same era.
Near-Memory Processing in Action: Accelerating Personalized Recommendation With AxDIMM
Liu Ke, Xuan Zhang, Jinin So, Jong-Geon Lee, Shin-Haeng Kang, Sukhan Lee, Songyi Han, YeonGon Cho, Jin Hyun Kim, Yongsuk Kwon, KyungSoo Kim, Jin Jung, Ilkwon Yun, Sung Joo Park, Hyunsun Park, Joonho Song, Jeonghyeon Cho, Kyomin Sohn, Nam Sung Kim, and Hsien-Hsin S. Lee. 2022 · 2021
Cited alongside, same era.
Retrieval-Augmented Generation for Code Summarization via Hybrid GNN. In International Conference on Learning Representations
Shangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow, and Yang Liu. 2021 · 2021
Retrieval-augmented Image Captioning. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , Andreas Vlachos and Isabelle Augenstein (Eds.). Association for Computational Linguistics, Dubrovnik, Croatia, 3666–3681
Rita Ramos, Desmond Elliott, and Bruno Martins. 2023 · 2023
Later among the works it cites.
Pre-Training Multi-Modal Dense Retrievers for Outside-Knowledge Visual Question Answering. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval (Taipei, Taiwan) (ICTIR ’23) . Association for Computing Machinery, New York, NY, USA, 169–176
Alireza Salemi, Mahta Rafiee, and Hamed Zamani. 2023b · 2023
Later among the works it cites.
Efficient llm inference on cpus
Haihao Shen, Hanwen Chang, Bo Dong, Yu Luo, and Hengyu Meng. 2023 · 2023
Later among the works it cites.
Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering
Shamane Siriwardhana, Rivindu Weerasekera, Elliott Wen, Tharindu Kaluarachchi, Rajib Rana, and Suranga Nanayakkara. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Retrieval Augmented Code Generation and Summarization. In Findings of the Association for Computational Linguistics: EMNLP 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, Punta Cana, Dominican Republic, 2719–2734
Md Rizwan Parvez, Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Cited alongside, same era.
KILT: a Benchmark for Knowledge Intensive Language Tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (Eds.). Association for Computational Linguistics, Online, 2523–2544
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021 · 2021
Cited alongside, same era.
RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (Eds.). Association for Computational Linguistics, Online, 5835–5847
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021 · 2021
Cited alongside, same era.
DeepScaleTool: A Tool for the Accurate Estimation of Technology Scaling in the Deep-Submicron Era. In 2021 IEEE International Symposium on Circuits and Systems (ISCAS) . 1–5
Satyabrata Sarangi and Bevan Baas. 2021 · 2021
Cited alongside, same era.
Retrieval Augmentation Reduces Hallucination in Conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, Punta Cana, Dominican Republic, 3784–3803
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021 · 2021
Cited alongside, same era.
Keep the Primary, Rewrite the Secondary: A Two-Stage Approach for Paraphrase Generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 560–569
Yixuan Su, David Vandyke, Simon Baker, Yan Wang, and Nigel Collier. 2021 · 2021
Cited alongside, same era.
Efficient retrieval augmented generation from unstructured knowledge for task-oriented dialog
David Thulke, Nico Daheim, Christian Dugast, and Hermann Ney. 2021 · 2021
Cited alongside, same era.
Retrieval-Augmented Response Generation for Knowledge-Grounded Conversation in the Wild
Yeonchan Ahn, Sang-Goo Lee, Junho Shim, and Jaehui Park. 2022 · 2022
Cited alongside, same era.
Retrieval Augumented Generation Overview
Heidi Steen and Dan Wahlin. 2023 · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning . PMLR, 38087–38099
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023 · 2023
Later among the works it cites.
Inference with reference: Lossless acceleration of large language models
Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023 · 2023
Later among the works it cites.
Rambda: RDMA-driven Acceleration Framework for Memory-intensive µs-scale Datacenter Applications. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 499–515
Yifan Yuan, Jinghan Huang, Yan Sun, Tianchen Wang, Jacob Nelson, Dan R. K. Ports, Yipeng Wang, Ren Wang, Charlie Tai, and Nam Sung Kim. 2023 · 2023
Later among the works it cites.
DIMM-Link: Enabling Efficient Inter-DIMM Communication for Near-Memory Processing. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 302–316
Zhe Zhou, Cong Li, Fan Yang, and Guangyu Sun. 2023 · 2023
Later among the works it cites.
Early Exit Strategies for Approximate k-NN Search in Dense Retrieval. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (Boise, ID, USA) (CIKM ’24) . Association for Computing Machinery, New York, NY, USA, 3647–3652
Francesco Busolin, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando, Raffaele Perego, and Salvatore Trani. 2024 · 2024
Closest in time.
Faiss: A Library for Efficient Similarity Search
Hervé Jégou, Matthijs Douze, and Jeff Johnson. 2017 · 2024
Closest in time.
REALTIME QA: what’s the answer right now?. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23) . Curran Associates Inc., Red Hook, NY, USA, Article 2130, 19 pages
Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Velocity Yu, Dragomir Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2024 · 2024
Closest in time.
Retrieval-Enhanced Machine Learning: Synthesis and Opportunities
To Eun Kim, Alireza Salemi, Andrew Drozdov, Fernando Diaz, and Hamed Zamani. 2024 · 2024
Closest in time.
LongLaMP: A Benchmark for Personalized Long-form Text Generation
Ishita Kumar, Snigdha Viswanathan, Sushrita Yerra, Alireza Salemi, Ryan A. Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, Nedim Lipka, Chien Van Nguyen, Thien Huu Nguyen, and Hamed Zamani. 2024 · 2024
Closest in time.
LLM Inference Series: 4. KV caching, a deeper look
Pierre Lienhart. 2024 · 2024
Closest in time.
Die Analysis: Samsung Exynos 2200 with RDNA2 Graphics
Locuza. 2022 · 2024
Closest in time.
SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (La Jolla, CA, USA) (ASPLOS ’24) . Association for Computing Machinery, New York, NY, USA, 932–949
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024 · 2024
Closest in time.
He Who Can Pay Top Dollar For HBM Memory Controls AI Training
Timothy Prickett Morgan. 2024 · 2024
Closest in time.
An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 970–982
Sang-Soo Park, KyungSoo Kim, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Jieun Lee, YeonGon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, and Nam Sung Kim. 2024 · 2024
Closest in time.
CXL Is Dead In The AI Era
Dylan Patel and Jeremie Eliahou Ontiveros. 2024 · 2024
Closest in time.
SmartDIMM: In-Memory Acceleration of Upper Layer Protocols. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE Computer Society, Los Alamitos, CA, USA, 312–329
N. Patel, A. Mamandipoor, M. Nouri, and M. Alian. 2024b · 2024
Closest in time.
Splitwise: Efficient Generative LLM Inference Using Phase Splitting . In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) . IEEE Computer Society, Los Alamitos, CA, USA, 118–132
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024a · 2024
Closest in time.
Introducing KILT, a new unified benchmark for knowledge-intensive NLP tasks
Fabio Petroni, Aleksandra Piktus, and Angela Fan. 2020 · 2024
Closest in time.
LaMP: When Large Language Models Meet Personalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 7370–7392
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024b · 2024
Closest in time.
CC-NIC: a Cache-Coherent Interface to the NIC. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 (La Jolla, CA, USA) (ASPLOS ’24) . Association for Computing Machinery, New York, NY, USA, 52–68
Henry N. Schuh, Arvind Krishnamurthy, David Culler, Henry M. Levy, Luigi Rizzo, Samira Khan, and Brent E. Stephens. 2024 · 2024
Closest in time.
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
Jovan Stojkovic, Esha Choukse, Chaojie Zhang, Inigo Goiri, and Josep Torrellas. 2024 · 2024
Closest in time.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team. 2024 · 2024
Closest in time.
Yitu Wang, Shiyu Li, Qilin Zheng, Linghao Song, Zongwang Li, Andrew Chang, Hai "Helen" Li, and Yiran Chen. 2024 · 2024
Closest in time.
Mask / Reticle
WikiChip. 2024 · 2024
Closest in time.
Efficient Streaming Language Models with Attention Sinks. In The Twelfth International Conference on Learning Representations
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024 · 2024
Closest in time.
Corrective Retrieval Augmented Generation
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024 · 2024
Closest in time.
The Shift from Models to Compound AI Systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024 · 2024
Closest in time.
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24) . Association for Computing Machinery, New York, NY, USA, 2641–2646
Hamed Zamani and Michael Bendersky. 2024 · 2024
Closest in time.
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
Yun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko, Yinxiao Liu, Chu-Cheng Lin, Lei Shu, Liangchen Luo, Lei Meng, Bang Liu, and Jindong Chen. 2024 · 2024
Closest in time.