Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved remarkable success due to their exceptional generative capabilities.
F. Jelinek, “Interpolated estimation of markov source parameters from sparse data,” in Proc. Workshop on Pattern Recognition in Practice , 1980
1980
Earlier work this paper cites.
B. Biggio, B. Nelson, and P. Laskov, “Poisoning attacks against support vector machines,” in ICML , 2012
2012
Earlier work this paper cites.
T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng, “Ms marco: A human generated machine reading comprehension dataset,” choice , vol. 2640, p. 660, 2016
2016
Earlier work this paper cites.
T. Gu, B. Dolan-Gavitt, and S. Garg, “Badnets: Identifying vulnerabilities in the machine learning model supply chain,” IEEE Access , 2017
2017
Earlier work this paper cites.
X. Chen, C. Liu, B. Li, K. Lu, and D. Song, “Targeted backdoor attacks on deep learning systems using data poisoning,” arXiv , 2017
2017
Earlier work this paper cites.
J. Steinhardt, P. W. W. Koh, and P. S. Liang, “Certified defenses for data poisoning attacks,” NeurIPS , 2017
2017
Earlier work this paper cites.
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning, “Hotpotqa: A dataset for diverse, explainable multi-hop question answering,” in EMNLP , 2018
2018
Earlier work this paper cites.
Y. Liu, S. Ma, Y. Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang, “Trojaning attack on neural networks,” in NDSS , 2018
2018
Earlier work this paper cites.
A. Shafahi, W. R. Huang, M. Najibi, O. Suciu, C. Studer, T. Dumitras, and T. Goldstein, “Poison frogs! targeted clean-label poisoning attacks on neural networks,” in NeurIPS , 2018
2018
Earlier work this paper cites.
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou, “Hotflip: White-box adversarial examples for text classification,” in ACL , 2018
2018
Earlier work this paper cites.
J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi, “Black-box generation of adversarial text sequences to evade deep learning classifiers,” in SPW , 2018
2018
Earlier work this paper cites.
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, “Fever: a large-scale dataset for fact extraction and verification,” arXiv , 2018
2018
Earlier work this paper cites.
I. Soboroff, S. Huang, and D. Harman, “Trec 2019 news track overview.” in TREC , 2019
2019
Earlier work this paper cites.
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee et al. , “Natural questions: a benchmark for question answering research,” TACL , vol. 7, pp. 452–466, 2019
2019
Earlier work this paper cites.
J. Li, S. Ji, T. Du, B. Li, and T. Wang, “Textbugger: Generating adversarial text against real-world applications,” in NDSS , 2019
2019
Earlier work this paper cites.
B. Wang, Y. Yao, S. Shan, H. Li, B. Viswanath, H. Zheng, and B. Y. Zhao, “Neural cleanse: Identifying and mitigating backdoor attacks in neural networks,” in IEEE S& P , 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” NeurIPS , 2020
2020
Earlier work this paper cites.
V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, “Dense passage retrieval for open-domain question answering,” in EMNLP , 2020
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” NeurIPS , 2020
2020
Earlier work this paper cites.
N. Kassner and H. Schütze, “Bert-knn: Adding a knn search component to pretrained language models for better qa,” in Findings of ACL: EMNLP , 2020
2020
Earlier work this paper cites.
V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, “Dense passage retrieval for open-domain question answering,” in EMNLP , 2020
2020
Earlier work this paper cites.
X. Pan, M. Zhang, S. Ji, and M. Yang, “Privacy risks of general-purpose language models,” in IEEE S & P , 2020
2020
Earlier work this paper cites.
E. Bagdasaryan, A. Veit, Y. Hua, D. Estrin, and V. Shmatikov, “How to backdoor federated learning,” in AISTATS , 2020
2020
Earlier work this paper cites.
M. Fang, X. Cao, J. Jia, and N. Gong, “Local model poisoning attacks to byzantine-robust federated learning,” in USENIX Security Symposium , 2020
2020
Earlier work this paper cites.
J. Morris, E. Lifland, J. Y. Yoo, J. Grigsby, D. Jin, and Y. Qi, “Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp,” in EMNLP , 2020
2020
Earlier work this paper cites.
D. Jin, Z. Jin, J. T. Zhou, and P. Szolovits, “Is bert really robust? a strong baseline for natural language attack on text classification and entailment,” in AAAI , 2020
2020
Earlier work this paper cites.
L. Li, R. Ma, Q. Guo, X. Xue, and X. Qiu, “Bert-attack: Adversarial attack against bert using bert,” in EMNLP , 2020
2020
Earlier work this paper cites.
R. Z. Mahari, “Autolaw: Augmented legal reasoning through legal precedent prediction,” arXiv , 2021
2021
Earlier work this paper cites.
N. Thakur, N. Reimers, A. Rücklé, A. Srivastava, and I. Gurevych, “Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models,” in NeurIPS , 2021
2021
Earlier work this paper cites.
E. Voorhees, T. Alam, S. Bedrick, D. Demner-Fushman, W. R. Hersh, K. Lo, K. Roberts, I. Soboroff, and L. L. Wang, “Trec-covid: constructing a pandemic information retrieval test collection,” in ACM SIGIR Forum , vol. 54, no. 1, 2021, pp. 1–12
2021
Earlier work this paper cites.
L. Xiong, C. Xiong, Y. Li, K.-F. Tang, J. Liu, P. N. Bennett, J. Ahmed, and A. Overwijk, “Approximate nearest neighbor negative contrastive learning for dense text retrieval,” in ICLR , 2021
2021
Earlier work this paper cites.
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson et al. , “Extracting training data from large language models,” in Usenix Security , 2021
2021
Earlier work this paper cites.
N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, Ú. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” in Usenix Security , 2021
2021
Earlier work this paper cites.
Z. Zhang, J. Jia, B. Wang, and N. Z. Gong, “Backdoor attacks to graph neural networks,” in SACMAT , 2021
2021
Cited alongside, same era.
N. Carlini, “Poisoning the unlabeled dataset of semi-supervised learning,” in USENIX Security , 2021
2021
Cited alongside, same era.
J. Jia, X. Cao, and N. Z. Gong, “Intrinsic certified robustness of bagging against data poisoning attacks,” in AAAI , 2021
2021
Cited alongside, same era.
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark et al. , “Improving language models by retrieving from trillions of tokens,” in ICML , 2022
2022
Cited alongside, same era.
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du et al. , “Lamda: Language models for dialog applications,” arXiv , 2022
N. Carlini, M. Jagielski, C. A. Choquette-Choo, D. Paleka, W. Pearce, H. Anderson, A. Terzis, K. Thomas, and F. Tramèr, “Poisoning web-scale training datasets is practical,” arXiv , 2023
2023
Later among the works it cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv , 2023
2023
Later among the works it cites.
Z. Zhong, Z. Huang, A. Wettig, and D. Chen, “Poisoning retrieval corpora by injecting adversarial passages,” in EMNLP , 2023
2023
Later among the works it cites.
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P.-y. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein, “Baseline defenses for adversarial attacks against aligned language models,” arXiv , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
J. Liu, “LlamaIndex,” 11 2022. [Online]. Available: https://github.com/jerryjliu/llama_index
2022
Cited alongside, same era.
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” arXiv , 2022
2022
Cited alongside, same era.
G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave, “Unsupervised dense information retrieval with contrastive learning,” Transactions on Machine Learning Research , 2022
2022
Cited alongside, same era.
H. Gonen, S. Iyer, T. Blevins, N. A. Smith, and L. Zettlemoyer, “Demystifying prompts in language models via perplexity estimation,” arXiv , 2022
2022
Cited alongside, same era.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” in NeurIPS ML Safety Workshop , 2022
2022
Cited alongside, same era.
H. J. Branch, J. R. Cefalu, J. McHugh, L. Hujer, A. Bahl, D. d. C. Iglesias, R. Heichman, and R. Darwishi, “Evaluating the susceptibility of pre-trained language models via handcrafted adversarial examples,” arXiv , 2022
2022
Cited alongside, same era.
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, “Quantifying memorization across neural language models,” in ICLR , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
Y. Liu, G. Deng, Y. Li, K. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng, and Y. Liu, “Prompt injection attack against llm-integrated applications,” arXiv , 2023
2023
Later among the works it cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” in AISec , 2023
2023
Later among the works it cites.
R. Pedro, D. Castro, P. Carreira, and N. Santos, “From prompt injections to sql injection attacks: How protected is your llm-integrated web application?” arXiv , 2023
2023
Later among the works it cites.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” in NeurIPS , 2023
2023
Later among the works it cites.
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” arXiv , 2023
2023
Later among the works it cites.
X. Qi, K. Huang, A. Panda, P. Henderson, M. Wang, and P. Mittal, “Visual adversarial examples jailbreak aligned large language models,” arXiv , 2023
2023
Later among the works it cites.
H. Li, D. Guo, W. Fan, M. Xu, J. Huang, F. Meng, and Y. Song, “Multi-step jailbreaking privacy attacks on chatgpt,” arXiv , 2023
2023
Later among the works it cites.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, “"do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models,” arXiv , 2023
2023
Later among the works it cites.
N. Kandpal, M. Jagielski, F. Tramèr, and N. Carlini, “Backdoor attacks for in-context learning with language models,” in ICML Workshop , 2023
2023
Later among the works it cites.
A. Wan, E. Wallace, S. Shen, and D. Klein, “Poisoning language models during instruction tuning,” in ICML , 2023
2023
Later among the works it cites.
J. Mattern, F. Mireshghallah, Z. Jin, B. Schoelkopf, M. Sachan, and T. Berg-Kirkpatrick, “Membership inference attacks against language models via neighbourhood comparison,” in Findings of ACL: EMNLP , 2023
2023
Later among the works it cites.
X. Li and J. Li, “Angle-optimized text embeddings,” arXiv preprint arXiv:2309.12871 , 2023
2023
Later among the works it cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” 2023
2023
Later among the works it cites.
M. R. Rizqullah, A. Purwarianti, and A. F. Aji, “Qasina: Religious domain question answering using sirah nabawiyah,” in ICAICTA , 2023
2023
Later among the works it cites.
Y. Huang, S. Gupta, M. Xia, K. Li, and D. Chen, “Catastrophic jailbreak of open-source llms via exploiting generation,” arXiv , 2023
2023
Later among the works it cites.
Y. Pan, L. Pan, W. Chen, P. Nakov, M.-Y. Kan, and W. Y. Wang, “On the risk of misinformation pollution with large language models,” in EMNLP , 2023
2023
Later among the works it cites.
O. Yoran, T. Wolfson, O. Ram, and J. Berant, “Making retrieval-augmented language models robust to irrelevant context,” in ICLR , 2023
2023
Later among the works it cites.
H. Luo, T. Zhang, Y.-S. Chuang, Y. Gong, Y. Kim, X. Wu, H. Meng, and J. Glass, “Search augmented instruction learning,” in EMNLP , 2023, pp. 3717–3729
2023
Later among the works it cites.
J. Jia, Y. Liu, Y. Hu, and N. Z. Gong, “Pore: Provably robust recommender systems against data poisoning attacks,” in USENIX Security Symposium , 2023
2023
Later among the works it cites.
“Bing search,” https://www.microsoft.com/en-us/bing?form=MG0AUO&OCID=MG0AUO#faq , 2024
2024
Closest in time.
“Generative ai in search: Let google do the searching for you,” https://blog.google/products/search/generative-ai-google-search-may-2024/
2024
Closest in time.
A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-rag: Learning to retrieve, generate, and critique through self-reflection,” in ICLR , 2024
2024
Closest in time.
Y. Liu, Y. Jia, R. Geng, J. Jia, and N. Z. Gong, “Formalizing and benchmarking prompt injection attacks and defenses,” arXiv , 2024
2024
Closest in time.
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu, “Masterkey: Automated jailbreaking of large language model chatbots,” in NDSS , 2024
2024
Closest in time.
S.-Q. Yan, J.-C. Gu, Y. Zhu, and Z.-H. Ling, “Corrective retrieval augmented generation,” arXiv , 2024
2024
Closest in time.
T. Zhang, S. G. Patil, N. Jain, S. Shen, M. Zaharia, I. Stoica, and J. E. Gonzalez, “Raft: Adapting language model to domain specific rag,” arXiv , 2024
2024
Closest in time.
Y. Wang, W. Zou, and J. Jia, “Fcert: Certifiably robust few-shot classification in the era of foundation models,” in IEEE S & P , 2024
2024
Closest in time.
E. Kortukov, A. Rubinstein, E. Nguyen, and S. J. Oh, “Studying large language model behaviors under realistic knowledge conflicts,” arXiv , 2024
2024
Closest in time.