Fetching the paper…
Reading the bibliography…
This paper presents a context key/value compression method for Transformer language models in online scenarios, where the context continually expands.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Memory networks
Jason Weston, Sumit Chopra, and Antoine Bordes · 2015
Earlier work this paper cites.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Earlier work this paper cites.
Memory-efficient backpropagation through time
Audrunas Gruslys, Rémi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves · 2016
Earlier work this paper cites.
Dailydialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Never-ending learning
Tom Mitchell, William Cohen, Estevam Hruschka, et al · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, et al · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, et al · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, et al · 2020
Earlier work this paper cites.
Mind the gap: Assessing temporal generalization in neural language models
Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, et al · 2021
Cited alongside, same era.
Recurrent memory transformer
Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, et al · 2022
Cited alongside, same era.
Meta-learning fast weight language models
Kevin Clark, Kelvin Guu, Ming-Wei Chang, Panupong Pasupat, Geoffrey Hinton, and Mohammad Norouzi · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Block-recurrent transformers
DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam Neyshabur · 2022
Meta-learning online adaptation of language models
Nathan Hu, Eric Mitchell, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
Soda: Million-scale dialogue distillation with social commonsense contextualization
Hyunwoo Kim, Jack Hessel, Liwei Jiang, Ximing Lu, et al · 2023
Closest in time.
An overview of bard: an early experiment with generative ai
James Manyika · 2023
Closest in time.
Landmark attention: Random-access infinite context length for transformers
Amirkeivan Mohtashami and Martin Jaggi · 2023
Closest in time.
Learning to compress prompts with gist tokens
Jesse Mu, Xiang Lisa Li, and Noah Goodman · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Metaicl: Learning to learn in context
Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
Learning by distilling context
Charlie Snell, Dan Klein, and Ruiqi Zhong · 2022
Cited alongside, same era.
Prompt compression and contrastive conditioning for controllability and toxicity reduction in language models
David Wingate, Mohammad Shoeybi, and Taylor Sorensen · 2022
Cited alongside, same era.
Memorizing transformers
Yuhuai Wu, Markus N Rabe, DeLesley Hutchins, and Christian Szegedy · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, et al · 2022
Cited alongside, same era.
Adapting language models to compress contexts
Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen · 2023
Cited alongside, same era.
OpenAI · 2023
Closest in time.
Hypertuning: Toward adapting large language models without back-propagation
Jason Phang, Yi Mao, Pengcheng He, and Weizhu Chen · 2023
Closest in time.
Lamp: When large language models meet personalization
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al · 2023
Closest in time.
Focused transformer: Contrastive training for context scaling
Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłoś · 2023
Closest in time.
Augmenting language models with long-term memory
Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei · 2023
Closest in time.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Closest in time.
Lmsys-chat-1m: A large-scale real-world llm conversation dataset
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al · 2023
Closest in time.
Memorybank: Enhancing large language models with long-term memory
Wanjun Zhong, Lianghong Guo, Qiqi Gao, and Yanlin Wang · 2023
Closest in time.