Fetching the paper…
Reading the bibliography…
Self-attention in transformer models is an incremental associative memory that maps key vectors to value vectors.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap · 1911
Earlier work this paper cites.
Multidimensional binary search trees used for associative searching
Jon Louis Bentley · 1975
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Managing gigabytes: compressing and indexing documents and images
Ian H Witten, Alistair Moffat, and Timothy C Bell · 1999
Earlier work this paper cites.
Locality-sensitive hashing scheme based on p-stable distributions
Mayur Datar, Nicole Immorlica, Piotr Indyk, and Vahab S Mirrokni · 2004
Earlier work this paper cites.
Multi-probe lsh: efficient indexing for high-dimensional similarity search
Qin Lv, William Josephson, Zhe Wang, Moses Charikar, and Kai Li · 2007
Earlier work this paper cites.
Improving regularized singular value decomposition for collaborative filtering
Arkadiusz Paterek · 2007
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2010
Earlier work this paper cites.
Locality sensitive hashing: A comparison of hash function types and querying mechanisms
Loïc Paulevé, Hervé Jégou, and Laurent Amsaleg · 2010
Earlier work this paper cites.
Efficient k-nearest neighbor graph construction for generic similarity measures
Wei Dong, Charikar Moses, and Kai Li · 2011
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Fast approximate nearest neighbor search with the navigating spreading-out graph
Cong Fu, Chao Xiang, Changxu Wang, and Deng Cai · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2018
Earlier work this paper cites.
Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Yu A Malkov and Dmitry A Yashunin · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Learning space partitions for nearest neighbor search
Yihe Dong, Piotr Indyk, Ilya Razenshteyn, and Tal Wagner · 2019
Cited alongside, same era.
Diskann: Fast accurate billion-point nearest neighbor search on a single node
Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Later among the works it cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Later among the works it cites.
Framequant: Flexible low-bit quantization for transformers
Harshavardhan Adepu, Zhanpeng Zeng, Li Zhang, and Vikas Singh · 2024
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2024
Later among the works it cites.
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2022
Cited alongside, same era.
Ood-diskann: Efficient and scalable graph anns for out-of-distribution queries
Shikhar Jaiswal, Ravishankar Krishnaswamy, Ankit Garg, Harsha Vardhan Simhadri, and Sheshansh Agrawal · 2022
Cited alongside, same era.
xformers: A modular and hackable transformer modelling library
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, Daniel Haziza, Luca Wehrstedt, Jeremy Reizenstein, and Grigory Sizov · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Cited alongside, same era.
Dedrift: Robust similarity search under content drift
Dmitry Baranchuk, Matthijs Douze, Yash Upadhyay, and I Zeki Yalniz · 2023
Cited alongside, same era.
Unlimiformer: Long-range transformers with unlimited length input
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew Gormley · 2023
Cited alongside, same era.
Flash-decoding for long-context inference, 2023
Tri Dao, Daniel Haziza, Francisco Massa, and Grigory Sizov · 2023
Cited alongside, same era.
Later among the works it cites.
The llama 3 herd of models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Ruler: What’s the real context size of your long-context language models?, 2024
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg · 2024
Later among the works it cites.
Needle in a haystack - pressure testing llms
Greg Kamradt · 2024
Later among the works it cites.
Gear: An efficient kv cache compression recipefor near-lossless generative inference of llm
Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, and Tuo Zhao · 2024
Later among the works it cites.
Cagra: Highly parallel graph construction and approximate nearest neighbor search for gpus
Hiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang, Tamas Feher, and Yong Wang · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Later among the works it cites.
Keep the cost down: A review on methods to optimize llm’s kv-cache consumption
Luohe Shi, Hongyi Zhang, Yao Yao, Zuchao Li, and Hai Zhao · 2024
Later among the works it cites.
Results of the big ann: Neurips’23 competition
Harsha Vardhan Simhadri, Martin Aumüller, Amir Ingber, Matthijs Douze, George Williams, Magdalen Dobson Manohar, Dmitry Baranchuk, Edo Liberty, Frank Liu, Ben Landrum, et al · 2024
Later among the works it cites.
Quest: Query-aware sparsity for efficient long-context llm inference
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han · 2024
Later among the works it cites.
∞ \infty Bench: Extending long context evaluation beyond 100K tokens
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun · 2024
Later among the works it cites.