Fetching the paper…
Reading the bibliography…
Prompt compression is crucial for enhancing inference speed, reducing costs, and improving user experience.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
RACE: Large-scale ReAding Comprehension Dataset From Examinations
Lai, G.; Xie, Q.; Liu, H.; Yang, Y.; and Hovy, E. 2017 · 2017
Earlier work this paper cites.
Zero-Shot Relation Extraction via Reading Comprehension
Levy, O.; Seo, M.; Choi, E.; and Zettlemoyer, L. 2017 · 2017
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Earlier work this paper cites.
Prompt Compression and Contrastive Conditioning for Controllability and Toxicity Reduction in Language Models
Wingate, D.; Shoeybi, M.; and Sorensen, T. 2022 · 2022
Cited alongside, same era.
Adapting Language Models to Compress Contexts
Chevalier, A.; Wettig, A.; Ajith, A.; and Chen, D. 2023 · 2023
Cited alongside, same era.
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Jiang, H.; Wu, Q.; Lin, C.-Y.; Yang, Y.; and Qiu, L. 2023a · 2023
Cited alongside, same era.
Compressing Context to Enhance Inference Efficiency of Large Language Models
Li, Y.; Dong, B.; Guerin, F.; and Lin, C. 2023 · 2023
Cited alongside, same era.
Context Compression for Auto-regressive Transformers with Sentinel Tokens
Ren, S.; Jia, Q.; and Zhu, K. Q. 2023 · 2023
Cited alongside, same era.
xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
In-context Autoencoder for Context Compression in a Large Language Model
Ge, T.; Jing, H.; Wang, L.; Wang, X.; Chen, S.-Q.; and Wei, F. 2024 · 2024
Closest in time.
Hierarchical and Dynamic Prompt Compression for Efficient Zero-shot API Usage
Jiang, Y.; Vecchio, M.; Bansal, M.; and Johannsen, A. 2024 · 2024
Closest in time.
Learning to compress prompts with gist tokens
Mu, J.; Li, X. L.; and Goodman, N. 2024 · 2024
Closest in time.
Context Embeddings for Efficient Answer Generation in RAG
Rau, D.; Wang, S.; Déjean, H.; and Clinchant, S. 2024 · 2024
Closest in time.
Tag-LLM: Repurposing General-Purpose LLMs for Specialized Domains
Shen, J.; Tenenholtz, N.; Hall, J. B.; Alvarez-Melis, D.; and Fusi, N. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheng, X.; Wang, X.; Zhang, X.; Ge, T.; Chen, S.-Q.; Wei, F.; Zhang, H.; and Zhao, D. 2024 · 2024
Cited alongside, same era.
Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression
Jiang, H.; Wu, Q.; Luo, X.; Li, D.; Lin, C.-Y.; Yang, Y.; and Qiu, L. 2023b
Cited in the paper.
Function Vectors in Large Language Models
Todd, E.; Li, M.; Sharma, A. S.; Mueller, A.; Wallace, B. C.; and Bau, D. 2024 · 2024
Closest in time.