Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have strong capabilities in solving diverse natural language processing tasks.
H. Jia, C. A. Choquette-Choo, V. Chandrasekaran, and N. Papernot, “Entangled watermarks as a defense against model extraction,” in USENIX Security , 2021, pp. 1937–1954
1954
Earlier work this paper cites.
M. F. Medress, F. S. Cooper, J. W. Forgie, C. C. Green, D. H. Klatt, M. H. O’Malley, E. P. Neuburg, A. Newell, R. Reddy, H. B. Ritea, J. E. Shoup-Hummel, D. E. Walker, and W. A. Woods, “Speech understanding systems,” Artif. Intell. , vol. 9, no. 3, pp. 307–316, 1977
1977
Earlier work this paper cites.
J. T. Brassil, S. Low, N. F. Maxemchuk, and L. O’Gorman, “Electronic marking and identification techniques to discourage document copying,” IEEE Journal on Selected Areas in Communications , vol. 13, no. 8, pp. 1495–1504, 1995
1995
Earlier work this paper cites.
A. B. Kahng, J. C. Lach, W. H. Mangione-Smith, S. Mantik, I. L. Markov, M. Potkonjak, P. Tucker, H. Wang, and G. Wolfe, “Watermarking techniques for intellectual property protection,” in DAC , 1998, pp. 776–781
1998
Earlier work this paper cites.
P. Ruch, R. H. Baud, A. Rassinoux, P. Bouillon, and G. Robert, “Medical document anonymization with a semantic lexicon,” in AMIA , 2000
2000
Earlier work this paper cites.
J. D. Lafferty, A. McCallum, and F. C. N. Pereira, “Conditional random fields: Probabilistic models for segmenting and labeling sequence data,” in ICML , 2001, pp. 282–289
2001
Earlier work this paper cites.
M. J. Atallah, V. Raskin, M. Crogan, C. Hempelmann, F. Kerschbaum, D. Mohamed, and S. Naik, “Natural language watermarking: Design, analysis, and a proof-of-concept implementation,” in Information Hiding , 2001, pp. 185–200
2001
Earlier work this paper cites.
X. Luo and R. K. C. Chang, “On a new class of pulsing denial-of-service attacks and the defense,” in NDSS , 2005
2005
Earlier work this paper cites.
U. Topkara, M. Topkara, and M. J. Atallah, “The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions,” in MM&Sec , 2006, pp. 164–174
2006
Earlier work this paper cites.
C. Dwork, “Differential privacy: A survey of results,” in TAMC , 2008, pp. 1–19
2008
Earlier work this paper cites.
W. Fish, “Perception, hallucination, and illusion.” OUP USA , 2009
2009
Earlier work this paper cites.
Z. Jalil and A. M. Mirza, “A review of digital watermarking techniques for text documents,” in ICIMT , 2009, pp. 230–234
2009
Earlier work this paper cites.
J. D. Blom, A dictionary of hallucinations . Springer, 2010
2010
Earlier work this paper cites.
C. Dwork, “A firm foundation for private data analysis,” Commun. ACM , vol. 54, no. 1, pp. 86–95, 2011
2011
Earlier work this paper cites.
T. Scholte, W. Robertson, D. Balzarotti, and E. Kirda, “Preventing input validation vulnerabilities in web applications through automated type analysis,” in CSA , 2012, pp. 233–243
2012
Earlier work this paper cites.
L. Deléger, K. Molnár, G. Savova, F. Xia, T. Lingren, Q. Li, K. Marsolo, A. G. Jegga, M. Kaiser, L. Stoutenborough, and I. Solti, “Large-scale evaluation of automated clinical note de-identification and its impact on information extraction,” J. Am. Medical Informatics Assoc. , vol. 20, no. 1, pp. 84–94, 2013
2013
Earlier work this paper cites.
C. Zhang, T. Wei, Z. Chen, L. Duan, L. Szekeres, S. McCamant, D. Song, and W. Zou, “Practical control flow integrity and randomization for binary executables,” in SP , 2013, pp. 559–573
2013
Earlier work this paper cites.
D. Sánchez, M. Batet, and A. Viejo, “Automatic general-purpose sanitization of textual documents,” IEEE Transactions on Information Forensics and Security , vol. 8, no. 6, pp. 853–862, 2013
2013
Earlier work this paper cites.
C. Dwork and A. Roth, “The algorithmic foundations of differential privacy,” Found. Trends Theor. Comput. Sci. , vol. 9, no. 3-4, pp. 211–407, 2014
2014
Earlier work this paper cites.
I. M. Alabdulmohsin, X. Gao, and X. Zhang, “Adding robustness to support vector machines against adversarial reverse engineering,” in CIKM , 2014, pp. 231–240
2014
Earlier work this paper cites.
E. Göktas, E. Athanasopoulos, H. Bos, and G. Portokalidis, “Out of control: Overcoming control-flow integrity,” in SP , 2014, pp. 575–589
2014
Earlier work this paper cites.
A. Blankstein and M. J. Freedman, “Automating isolation and least privilege in web services,” in SP , 2014, pp. 133–148
2014
Earlier work this paper cites.
Y. Zhu, R. Kiros, R. S. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, “Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,” in ICCV , 2015, pp. 19–27
2015
Earlier work this paper cites.
M. Fredrikson, S. Jha, and T. Ristenpart, “Model inversion attacks that exploit confidence information and basic countermeasures,” in CCS , 2015, pp. 1322–1333
2015
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in ICLR , 2015
2015
Earlier work this paper cites.
H. Ebadi, D. Sands, and G. Schneider, “Differential privacy: Now it’s getting personal,” in POPL , 2015, pp. 69–81
2015
Earlier work this paper cites.
S. Gu and L. Rigazio, “Towards deep neural network architectures robust to adversarial examples,” in ICLR workshop , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
N. Carlini, A. Barresi, M. Payer, D. A. Wagner, and T. R. Gross, “Control-flow bending: On the effectiveness of control-flow integrity,” in USENIX Security , 2015, pp. 161–176
2015
Earlier work this paper cites.
A. Bates, D. Tian, K. R. B. Butler, and T. Moyer, “Trustworthy whole-system provenance for the linux kernel,” in USENIX Security , 2015, pp. 319–334
2015
Earlier work this paper cites.
I. J. Goodfellow, Y. Bengio, and A. C. Courville, Deep Learning , ser. Adaptive computation and machine learning. MIT Press, 2016
2016
Earlier work this paper cites.
C. Dwork, F. McSherry, K. Nissim, and A. D. Smith, “Calibrating noise to sensitivity in private data analysis,” J. Priv. Confidentiality , vol. 7, no. 3, pp. 17–51, 2016
2016
Earlier work this paper cites.
T. Bolukbasi, K. Chang, J. Y. Zou, V. Saligrama, and A. T. Kalai, “Man is to computer programmer as woman is to homemaker? debiasing word embeddings,” in NeurIPS , 2016, pp. 4349–4357
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
F. Tramèr, F. Zhang, A. Juels, M. K. Reiter, and T. Ristenpart, “Stealing machine learning models via prediction apis,” in USENIX Security , 2016, pp. 601–618
2016
Earlier work this paper cites.
M. Abadi, A. Chu, I. J. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep learning with differential privacy,” in SIGSAC , 2016, pp. 308–318
2016
Earlier work this paper cites.
N. Papernot, P. D. McDaniel, X. Wu, S. Jha, and A. Swami, “Distillation as a defense to adversarial perturbations against deep neural networks,” in S&P , 2016, pp. 582–597
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in NeurIPS , 2017, pp. 4299–4307
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Hilton, N. Nelson, T. Tunnell, D. Marinov, and D. Dig, “Trade-offs in continuous integration: assurance, security, and flexibility,” in ESEC/FSE , 2017, pp. 197–207
2017
Earlier work this paper cites.
F. Dernoncourt, J. Y. Lee, Ö. Uzuner, and P. Szolovits, “De-identification of patient notes with recurrent neural networks,” J. Am. Medical Informatics Assoc. , vol. 24, no. 3, pp. 596–606, 2017
2017
Earlier work this paper cites.
J. Pavlopoulos, P. Malakasiotis, and I. Androutsopoulos, “Deeper attention to abusive user content moderation,” in EMNLP , 2017, pp. 1125–1135
2017
Earlier work this paper cites.
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in ICML , 2017, pp. 1321–1330
2017
Earlier work this paper cites.
G. Pereyra, G. Tucker, J. Chorowski, L. Kaiser, and G. E. Hinton, “Regularizing neural networks by penalizing confident output distributions,” in ICLR workshop , 2017
2017
Earlier work this paper cites.
J. Lu, T. Issaranon, and D. A. Forsyth, “Safetynet: Detecting and rejecting adversarial examples robustly,” in ICCV , 2017, pp. 446–454
2017
Earlier work this paper cites.
J. H. Metzen, T. Genewein, V. Fischer, and B. Bischoff, “On detecting adversarial perturbations,” in ICLR , 2017, p. 105978
2017
Earlier work this paper cites.
D. Meng and H. Chen, “Magnet: A two-pronged defense against adversarial examples,” in Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, CCS 2017, Dallas, TX, USA, October 30 - November 03, 2017 , 2017, pp. 135–147
2017
Earlier work this paper cites.
G. Katz, C. W. Barrett, D. L. Dill, K. Julian, and M. J. Kochenderfer, “Reluplex: An efficient SMT solver for verifying deep neural networks,” in CAV , 2017, pp. 97–117
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Fan, M. Lewis, and Y. N. Dauphin, “Hierarchical neural story generation,” in ACL , 2018, pp. 889–898
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. V. Georgakopoulos, S. K. Tasoulis, A. G. Vrahatis, and V. P. Plagianakos, “Convolutional neural networks for toxic comment classification,” in SETN , 2018, pp. 35:1–35:6
2018
Earlier work this paper cites.
C. N. dos Santos, I. Melnyk, and I. Padhi, “Fighting offensive language on social media with unsupervised text style transfer,” in ACL , 2018, pp. 189–194
2018
Earlier work this paper cites.
J. Zhao, Y. Zhou, Z. Li, W. Wang, and K. Chang, “Learning gender-neutral word embeddings,” in EMNLP , 2018, pp. 4847–4853
2018
Earlier work this paper cites.
N. Akhtar and A. S. Mian, “Threat of adversarial attacks on deep learning in computer vision: A survey,” IEEE Access , vol. 6, pp. 14 410–14 430, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Nasr, R. Shokri, and A. Houmansadr, “Machine learning with membership privacy using adversarial regularization,” in CCS , 2018, pp. 634–646
2018
Earlier work this paper cites.
J. Jia and N. Z. Gong, “Attriguard: A practical defense against attribute inference attacks via adversarial machine learning,” in USENIX Security , 2018, pp. 513–529
2018
Earlier work this paper cites.
W. U. Hassan, M. Lemay, N. Aguse, A. Bates, and T. Moyer, “Towards scalable cluster auditing through grammatical inference over provenance graphs,” in NDSS , 2018
2018
Earlier work this paper cites.
Y. Mirsky, T. Doitshman, Y. Elovici, and A. Shabtai, “Kitsune: An ensemble of autoencoders for online network intrusion detection,” in NDSS , 2018
2018
Earlier work this paper cites.
R. Zellers, A. Holtzman, H. Rashkin, Y. Bisk, A. Farhadi, F. Roesner, and Y. Choi, “Defending against neural fake news,” in NeurIPS , 2019, pp. 9051–9062
2019
Earlier work this paper cites.
S. Bordia and S. R. Bowman, “Identifying and reducing gender bias in word-level language models,” in NAACL-HLT , 2019, pp. 7–15
2019
Earlier work this paper cites.
J. Jia, A. Salem, M. Backes, Y. Zhang, and N. Z. Gong, “Memguard: Defending against black-box membership inference attacks via adversarial examples,” in CCS , 2019, pp. 259–274
2019
Earlier work this paper cites.
Q. Xiao, Y. Chen, C. Shen, Y. Chen, and K. Li, “Seeing is not believing: Camouflage attacks on image scaling algorithms,” in USENIX Security , 2019, pp. 443–460
2019
Earlier work this paper cites.
A. S. Rakin, Z. He, and D. Fan, “Bit-flip attack: Crushing neural network with progressive bit search,” in ICCV , 2019, pp. 1211–1220
2019
Earlier work this paper cites.
Y. Peng, Y. Zhu, Y. Chen, Y. Bao, B. Yi, C. Lan, C. Wu, and C. Guo, “A generic communication scheduler for distributed DNN training acceleration,” in SOSP , T. Brecht and C. Williamson, Eds. ACM, 2019, pp. 16–29
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Zhao, T. Wang, M. Yatskar, R. Cotterell, V. Ordonez, and K. Chang, “Gender bias in contextualized word embeddings,” in NAACL-HLT , 2019, pp. 629–634
2019
Earlier work this paper cites.
R. H. Maudslay, H. Gonen, R. Cotterell, and S. Teufel, “It’s all in the name: Mitigating gender bias with name-based counterfactual data substitution,” in EMNLP-IJCNLP , 2019, pp. 5266–5274
2019
Earlier work this paper cites.
R. L. L. IV, N. F. Liu, M. E. Peters, M. Gardner, and S. Singh, “Barack’s wife hillary: Using knowledge graphs for fact-aware language modeling,” in ACL , 2019, pp. 5962–5971
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Lee, B. Edwards, I. M. Molloy, and D. Su, “Defending against neural network model stealing attacks using deceptive perturbations,” in S&P Workshop , 2019, pp. 43–49
2019
Earlier work this paper cites.
M. Juuti, S. Szyller, S. Marchal, and N. Asokan, “PRADA: protecting against DNN model stealing attacks,” in EuroS&P , 2019, pp. 512–527
2019
Earlier work this paper cites.
B. Wang, Y. Yao, S. Shan, H. Li, B. Viswanath, H. Zheng, and B. Y. Zhao, “Neural cleanse: Identifying and mitigating backdoor attacks in neural networks,” in S&P , 2019, pp. 707–723
2019
Earlier work this paper cites.
Y. Liu, W. Lee, G. Tao, S. Ma, Y. Aafer, and X. Zhang, “ABS: scanning neural networks for back-doors by artificial brain stimulation,” in CCS , 2019, pp. 1265–1282
2019
Earlier work this paper cites.
S. M. Milajerdi, R. Gjomemo, B. Eshete, R. Sekar, and V. N. Venkatakrishnan, “HOLMES: real-time APT detection through correlation of suspicious information flows,” in SP , 2019, pp. 1137–1152
2019
Earlier work this paper cites.
M. Tran et al. , “On the feasibility of rerouting-based ddos defenses,” in SP , 2019, pp. 1169–1184
2019
Earlier work this paper cites.
F. Hassan, D. Sánchez, J. Soria-Comas, and J. Domingo-Ferrer, “Automatic anonymization of textual documents: detecting sensitive information via word embeddings,” in TrustCom/BigDataSE , 2019, pp. 358–365
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Goodrich, V. Rao, P. J. Liu, and M. Saleh, “Assessing the factual accuracy of generated text,” in SIGKDD , 2019, pp. 166–175
2019
Earlier work this paper cites.
T. Falke, L. F. Ribeiro, P. A. Utama, I. Dagan, and I. Gurevych, “Ranking generated summaries by correctness: An interesting but challenging application for natural language inference,” in ACL , 2019, pp. 2214–2220
2019
Earlier work this paper cites.
F. Nie, J.-G. Yao, J. Wang, R. Pan, and C.-Y. Lin, “A simple recipe towards reducing hallucination in neural surface realisation,” in ACL , 2019, pp. 2673–2679
2019
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” in ICLR , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Baumgartner, S. Zannettou, B. Keegan, M. Squire, and J. Blackburn, “The pushshift reddit dataset,” in ICWSM , 2020, pp. 830–839
2020
Earlier work this paper cites.
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “Realtoxicityprompts: Evaluating neural toxic degeneration in language models,” in EMNLP , 2020, pp. 3356–3369
2020
Earlier work this paper cites.
J. Li, T. Du, S. Ji, R. Zhang, Q. Lu, M. Yang, and T. Wang, “Textshield: Robust text classification based on multimodal embedding and neural machine translation,” in USENIX Security , 2020, pp. 1381–1398
2020
Earlier work this paper cites.
S. Lee, H. Han, S. K. Cha, and S. Son, “Montage: A neural network language model-guided javascript engine fuzzer,” in USENIX , 2020, pp. 2613–2630
2020
Earlier work this paper cites.
E. Quiring, D. Klein, D. Arp, M. Johns, and K. Rieck, “Adversarial preprocessing: Understanding and preventing image-scaling attacks in machine learning,” in USENIX Security , 2020, pp. 1363–1380
2020
Earlier work this paper cites.
F. Yao, A. S. Rakin, and D. Fan, “Deephammer: Depleting the intelligence of deep neural networks through targeted chain of bit flips,” in USENIX , 2020, pp. 1463–1480
2020
Earlier work this paper cites.
Y. Jiang, Y. Zhu, C. Lan, B. Yi, Y. Cui, and C. Guo, “A unified architecture for accelerating distributed DNN training in heterogeneous GPU/CPU clusters,” in OSDI , 2020, pp. 463–479
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “Realtoxicityprompts: Evaluating neural toxic degeneration in language models,” in Findings , 2020
2020
Earlier work this paper cites.
nostalgebraist, “interpreting gpt: the logit lens,” https://www.alignmentforum.org/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens , 2020
2020
Earlier work this paper cites.
A. E. W. Johnson, L. Bulgarelli, and T. J. Pollard, “Deidentification of free-text medical records using pre-trained bidirectional transformers,” in CHIL , 2020, pp. 214–221
2020
Earlier work this paper cites.
I. Kotsogiannis, S. Doudalis, S. Haney, A. Machanavajjhala, and S. Mehrotra, “One-sided differential privacy,” in ICDE , 2020, pp. 493–504
2020
Earlier work this paper cites.
X. Peng, S. Li, S. Frazier, and M. O. Riedl, “Reducing non-normative text generation from language models,” in INLG , 2020, pp. 374–383
2020
Earlier work this paper cites.
T. Orekondy, B. Schiele, and M. Fritz, “Prediction poisoning: Towards defenses against DNN model stealing attacks,” in ICLR , 2020
2020
Earlier work this paper cites.
X. Han, T. F. J. Pasquier, A. Bates, J. Mickens, and M. I. Seltzer, “Unicorn: Runtime provenance-based detector for advanced persistent threats,” in NDSS , 2020
2020
Earlier work this paper cites.
Q. Wang, W. U. Hassan, D. Li, K. Jee, X. Yu, K. Zou, J. Rhee, Z. Chen, W. Cheng, C. A. Gunter, and H. Chen, “You are what you do: Hunting stealthy malware via data provenance analysis,” in NDSS , 2020
2020
Earlier work this paper cites.
X. Han, T. F. J. Pasquier, A. Bates, J. Mickens, and M. I. Seltzer, “Unicorn: Runtime provenance-based detector for advanced persistent threats,” in NDSS , 2020
2020
Earlier work this paper cites.
Q. Wang, W. U. Hassan, D. Li, K. Jee, X. Yu, K. Zou, J. Rhee, Z. Chen, W. Cheng, C. A. Gunter, and H. Chen, “You are what you do: Hunting stealthy malware via data provenance analysis,” in NDSS , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. B. Brown, D. Song, Ú. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” in USENIX Security , 2021, pp. 2633–2650
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
N. Ousidhoum, X. Zhao, T. Fang, Y. Song, and D. Yeung, “Probing toxic content in large pre-trained language models,” in ACL , 2021, pp. 4262–4274
2021
Earlier work this paper cites.
J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P. Huang, “Challenges in detoxifying language models,” in EMNLP , 2021, pp. 2447–2469
2021
Earlier work this paper cites.
M. Nadeem, A. Bethke, and S. Reddy, “Stereoset: Measuring stereotypical bias in pretrained language models,” in ACL , 2021, pp. 5356–5371
2021
Earlier work this paper cites.
K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston, “Retrieval augmentation reduces hallucination in conversation,” in Findings of EMNLP , 2021, pp. 3784–3803
2021
Earlier work this paper cites.
N. Dziri, A. Madotto, O. Zaïane, and A. J. Bose, “Neural path hunter: Reducing hallucination in dialogue systems via path grounding,” in EMNLP , 2021, pp. 2197–2214
2021
Earlier work this paper cites.
I. Shumailov, Y. Zhao, D. Bates, N. Papernot, R. D. Mullins, and R. Anderson, “Sponge examples: Energy-latency attacks on neural networks,” in SP , 2021, pp. 212–231
2021
Earlier work this paper cites.
C. Lao, Y. Le, K. Mahajan, Y. Chen, W. Wu, A. Akella, and M. M. Swift, “ATP: in-network aggregation for multi-tenant learning,” in NSDI , 2021, pp. 741–761
2021
Earlier work this paper cites.
S. Tan, B. Knott, Y. Tian, and D. J. Wu, “Cryptgpu: Fast privacy-preserving machine learning on the GPU,” in SP , 2021, pp. 1021–1038
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Hoory, A. Feder, A. Tendler, S. Erell, A. Peled-Cohen, I. Laish, H. Nakhost, U. Stemmer, A. Benjamini, A. Hassidim, and Y. Matias, “Learning and evaluating a differentially private pre-trained language model,” in EMNLP , 2021, pp. 1178–1189
2021
Earlier work this paper cites.
Z. Zhao, Z. Zhang, and F. Hopfgartner, “A comparative study of using pre-trained language models for toxic comment classification,” in WWW , 2021, pp. 500–507
2021
Earlier work this paper cites.
C. AI, “Perspective api documentation,” https://github.com/conversationai/perspectiveapi , 2021
2021
Earlier work this paper cites.
L. Laugier, J. Pavlopoulos, J. Sorensen, and L. Dixon, “Civil rephrases of toxic texts with self-supervised transformers,” in EACL , 2021, pp. 1442–1461
2021
Earlier work this paper cites.
S. Dev, T. Li, J. M. Phillips, and V. Srikumar, “Oscar: Orthogonal subspace correction and rectification of biases in word embeddings,” in EMNLP , 2021, pp. 5034–5050
2021
Earlier work this paper cites.
S. Awan, B. Luo, and F. Li, “CONTRA: defending against poisoning attacks in federated learning,” in ESORICS , 2021, pp. 455–475
2021
Earlier work this paper cites.
F. Qi, M. Li, Y. Chen, Z. Zhang, Z. Liu, Y. Wang, and M. Sun, “Hidden killer: Invisible textual backdoor attacks with syntactic trigger,” in ACL/IJCNLP , 2021, pp. 443–453
2021
Earlier work this paper cites.
W. Yang, Y. Lin, P. Li, J. Zhou, and X. Sun, “Rethinking stealthiness of backdoor attack against NLP models,” in ACL/IJCNLP , 2021, pp. 5543–5557
2021
Earlier work this paper cites.
L. Yu, S. Ma, Z. Zhang, G. Tao, X. Zhang, D. Xu, V. E. Urias, H. W. Lin, G. F. Ciocarlie, V. Yegneswaran, and A. Gehani, “Alchemist: Fusing application and audit logs for precise attack provenance without instrumentation,” in NDSS , 2021
2021
Earlier work this paper cites.
A. Alsaheel, Y. Nan, S. Ma, L. Yu, G. Walkup, Z. B. Celik, X. Zhang, and D. Xu, “ATLAS: A sequence-based learning approach for attack investigation,” in USENIX Security , 2021, pp. 3005–3022
2021
Earlier work this paper cites.
L. Yu, S. Ma, Z. Zhang, G. Tao, X. Zhang, D. Xu, V. E. Urias, H. W. Lin, G. F. Ciocarlie, V. Yegneswaran, and A. Gehani, “Alchemist: Fusing application and audit logs for precise attack provenance without instrumentation,” in NDSS , 2021
2021
Earlier work this paper cites.
C. Fu, Q. Li, M. Shen, and K. Xu, “Realtime robust malicious traffic detection via frequency domain analysis,” in CCS , 2021, pp. 3431–3446
2021
Earlier work this paper cites.
D. Barradas, N. Santos, L. Rodrigues, S. Signorello, F. M. V. Ramos, and A. Madeira, “Flowlens: Enabling efficient flow classification for ml-based network security applications,” in NDSS , 2021
2021
Earlier work this paper cites.
S. Panda et al. , “Smartwatch: accurate traffic analysis and flow-state tracking for intrusion prevention using smartnics,” in CoNEXT , 2021, pp. 60–75
2021
Earlier work this paper cites.
J. Holland, P. Schmitt, N. Feamster, and P. Mittal, “New directions in automated traffic analysis,” in CCS , 2021, pp. 3366–3383
2021
Earlier work this paper cites.
D. Wagner et al. , “United we stand: Collaborative detection and mitigation of amplification ddos attacks at scale,” in CCS , 2021, pp. 970–987
2021
Earlier work this paper cites.
Y. Guo, J. Liu, W. Tang, and C. Huang, “Exsense: Extract sensitive information from unstructured data,” Computers & Security , vol. 102, p. 102156, 2021
2021
Earlier work this paper cites.
K. Gémes and G. Recski, “Tuw-inf at germeval2021: Rule-based and hybrid methods for detecting toxic, engaging, and fact-claiming comments,” in GermEval KONVENS , 2021, pp. 69–75
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. Abdelnabi and M. Fritz, “Adversarial watermarking transformer: Towards tracing text provenance with data hiding,” in S&P , 2021, pp. 121–140
2021
Cited alongside, same era.
A. Chakraborty, M. Alam, V. Dey, A. Chattopadhyay, and D. Mukhopadhyay, “A survey on adversarial attacks and defences,” CAAI Trans. Intell. Technol. , vol. 6, no. 1, pp. 25–45, 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
D. Li, A. S. Rawat, M. Zaheer, X. Wang, M. Lukasik, A. Veit, F. X. Yu, and S. Kumar, “Large language models with controllable working memory,” in Findings of ACL , 2023, pp. 1774–1793
2023
Later among the works it cites.
A. Mallen, A. Asai, V. Zhong, R. Das, D. Khashabi, and H. Hajishirzi, “When not to trust language models: Investigating effectiveness of parametric and non-parametric memories,” in ACL , 2023, pp. 9802–9822
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
B. Mathew, P. Saha, S. M. Yimam, C. Biemann, P. Goyal, and A. Mukherjee, “Hatexplain: A benchmark dataset for explainable hate speech detection,” in AAAI , 2021, pp. 14 867–14 875
2021
Cited alongside, same era.
J. Dhamala, T. Sun, V. Kumar, S. Krishna, Y. Pruksachatkun, K.-W. Chang, and R. Gupta, “Bold: Dataset and metrics for measuring biases in open-ended language generation,” in FAccT , 2021, pp. 862–872
2021
Cited alongside, same era.
L. Reynolds and K. McDonell, “Prompt programming for large language models: Beyond the few-shot paradigm,” in CHI Extended Abstracts , 2021, pp. 1–7
2021
Cited alongside, same era.
2021
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in NeurIPS , 2022
2022
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in NeurIPS , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Kandpal, H. Deng, A. Roberts, E. Wallace, and C. Raffel, “Large language models struggle to learn long-tail knowledge,” in ICML , 2023, p. 15696–15707
2023
Later among the works it cites.
N. McKenna, T. Li, L. Cheng, M. J. Hosseini, M. Johnson, and M. Steedman, “Sources of hallucination by large language models on inference tasks,” in Findings of EMNLP , 2023, pp. 2758–2774
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Chen, R. Guan, X. Gong, J. Dong, and M. Xue, “D-DAE: defense-penetrating model extraction attacks,” in SP , 2023, pp. 382–399
2023
Later among the works it cites.
J. Mattern, F. Mireshghallah, Z. Jin, B. Schölkopf, M. Sachan, and T. Berg-Kirkpatrick, “Membership inference attacks against language models via neighbourhood comparison,” in ACL , 2023, pp. 11 330–11 343
2023
Later among the works it cites.
H. Yang, M. Ge, and K. X. andF Jingwei Li, “Using highly compressed gradients in federated learning for data reconstruction attacks,” IEEE Trans. Inf. Forensics Secur. , vol. 18, pp. 818–830, 2023
2023
Later among the works it cites.
G. Xia, J. Chen, C. Yu, and J. Ma, “Poisoning attacks in federated learning: A survey,” IEEE Access , vol. 11, pp. 10 708–10 722, 2023
2023
Later among the works it cites.
E. O. Soremekun, S. Udeshi, and S. Chattopadhyay, “Towards backdoor attacks and defense in robust machine learning models,” Comput. Secur. , vol. 127, p. 103101, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Zhao, J. Wen, A. T. Luu, J. Zhao, and J. Fu, “Prompt as triggers for backdoor attack: Examining the vulnerability in language models,” in EMNLP , 2023, pp. 12 303–12 317
2023
Later among the works it cites.
H. Mai, J. Zhao, H. Zheng, Y. Zhao, Z. Liu, M. Gao, C. Wang, H. Cui, X. Feng, and C. Kozyrakis, “Honeycomb: Secure and efficient GPU executions via static validation,” in OSDI , 2023, pp. 155–172
2023
Later among the works it cites.
J. Wang, Z. Zhang, M. Wang, H. Qiu, T. Zhang, Q. Li, Z. Li, T. Wei, and C. Zhang, “Aegis: Mitigating targeted bit-flip attacks against deep neural networks,” in USENIX , 2023, pp. 2329–2346
2023
Later among the works it cites.
Q. Liu, J. Yin, W. Wen, C. Yang, and S. Sha, “Neuropots: Realtime proactive defense against bit-flip attacks in neural networks,” in USENIX , 2023, pp. 6347–6364
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
T. Gao, H. Yen, J. Yu, and D. Chen, “Enabling large language models to generate text with citations,” in EMNLP , 2023, pp. 6465–6488
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” in ICLR , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith, and M. Lewis, “Measuring and narrowing the compositionality gap in language models,” in Findings of EMNLP , 2023, pp. 5687–5711
2023
Later among the works it cites.
R. W. McGee, “Is chat gpt biased against conservatives? an empirical study,” An Empirical Study (February 15, 2023) , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
O. Oviedo-Trespalacios, A. E. Peden, T. Cole-Hunter, A. Costantini, M. Haghani, J. Rod, S. Kelly, H. Torkamaan, A. Tariq, J. D. A. Newton et al. , “The risks of using chatgpt to obtain common safety-related information and advice,” Safety Science , vol. 167, p. 106244, 2023
2023
Later among the works it cites.
OPC, “Opc to investigate chatgpt jointly with provincial privacy authorities,” https://www.priv.gc.ca/en/opc-news/news-and-announcements/2023/an_230525-2/ , 2023
2023
Later among the works it cites.
M. Gurman, “Samsung bans staff’s ai use after spotting chatgpt data leak,” https://www.bloomberg.com/news/articles/2023-05-02/samsung-bans-chatgpt-and-other-generative-ai-use-by-staff-after-leak?srnd=technology-vp&in_source=embedded-checkout-banner/
2023
Later among the works it cites.
S. Sabin, “Companies are struggling to keep corporate secrets out of chatgpt,” https://www.axios.com/2023/03/10/chatgpt-ai-cybersecurity-secrets/
2023
Later among the works it cites.
H. Alkaissi and S. I. McFarlane, “Artificial hallucinations in chatgpt: implications in scientific writing,” Cureus , vol. 15, no. 2, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Vincent, “Google’s ai chatbot bard makes factual error in first demo.” https://www.theverge.com/2023/2/8/23590864/google-ai-chatbot-bard-mistake-error-exoplanet-demo
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Elsen-Rooney, “Nyc education department blocks chatgpt on school devices, networks,” https://ny.chalkbeat.org/2023/1/3/23537987/nyc-schools-ban-chatgpt-writing-artificial-intelligence
2023
Later among the works it cites.
J. Lee, T. Le, J. Chen, and D. Lee, “Do language models plagiarize?” in Proceedings of the ACM Web Conference 2023 , 2023, pp. 3637–3647
2023
Later among the works it cites.
P. Sharma and B. Dash, “Impact of big data analytics and chatgpt on cybersecurity,” in 2023 4th International Conference on Computing and Communication Systems (I3CS) , 2023, pp. 1–6
2023
Later among the works it cites.
2023
Later among the works it cites.
Github, “Github copilot,” https://github.com/features/copilot , 2023
2023
Later among the works it cites.
E. Crothers, N. Japkowicz, and H. L. Viktor, “Machine-generated text: A comprehensive survey of threat models and detection methods,” IEEE Access , 2023
2023
Later among the works it cites.
L. Prompting, “Defensive measures,” https://learnprompting.org/docs/category/-defensive-measures , 2023
2023
Later among the works it cites.
A. Volkov, “Discovery of sandwich defense,” https://twitter.com/altryne?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Eauthor , 2023
2023
Later among the works it cites.
OpenAI, “GPT-4 Technical Report,” CoRR , vol. abs/2303.08774, 2023
2023
Later among the works it cites.
NVIDIA, “Nemo guardrails,” https://github.com/NVIDIA/NeMo-Guardrails , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Azure, “Azure ai content safety,” https://azure.microsoft.com/en-us/products/ai-services/ai-content-safety , 2023
2023
Later among the works it cites.
H. Thakur, A. Jain, P. Vaddamanu, P. P. Liang, and L. Morency, “Language models get a gender makeover: Mitigating gender bias with few-shot data interventions,” in ACL , 2023, pp. 340–351
2023
Later among the works it cites.
Z. Xie and T. Lukasiewicz, “An empirical analysis of parameter-efficient methods for debiasing pre-trained language models,” in ACL , 2023, pp. 15 730–15 745
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Huang, H. P. Chan, and H. Ji, “Zero-shot faithful factual error correction,” in ACL , 2023, pp. 5660–5676
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Zhao, X. Li, S. Joty, C. Qin, and L. Bing, “Verify-and-edit: A knowledge-enhanced chain-of-thought framework,” in ACL , 2023, pp. 5823–5840
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Shao, Y. Gong, Y. Shen, M. Huang, N. Duan, and W. Chen, “Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy,” in Findings of EMNLP , 2023, pp. 9248–9274
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. L. Li, A. Holtzman, D. Fried, P. Liang, J. Eisner, T. Hashimoto, L. Zettlemoyer, and M. Lewis, “Contrastive decoding: Open-ended text generation as optimization,” in ACL , 2023, pp. 12 286–12 312
2023
Later among the works it cites.
S. Willison, “Reducing sycophancy and improving honesty via activation steering,” https://www.alignmentforum.org/posts/zt6hRsDE84HeBKh7E/reducing-sycophancy-and-improving-honesty-via-activation , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Cohen, M. Hamri, M. Geva, and A. Globerson, “LM vs LM: detecting factual errors via cross examination,” in EMNLP , 2023, pp. 12 621–12 640
2023
Later among the works it cites.
OWASP, “Owasp top 10 for llm applications,” https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-2023-v1_0_1.pdf , 2023
2023
Later among the works it cites.
R. T. Gollapudi, G. Yuksek, D. Demicco, M. Cole, G. Kothari, R. Kulkarni, X. Zhang, K. Ghose, A. Prakash, and Z. Umrigar, “Control flow and pointer integrity enforcement in a secure tagged architecture,” in SP , 2023, pp. 2974–2989
2023
Later among the works it cites.
H. Ding, J. Zhai, D. Deng, and S. Ma, “The case for learned provenance graph storage systems,” in USENIX Security , 2023
2023
Later among the works it cites.
F. Yang, J. Xu, C. Xiong, Z. Li, and K. Zhang, “PROGRAPHER: an anomaly detection system based on provenance graph embedding,” in USENIX Security , 2023
2023
Later among the works it cites.
K. Mukherjee, J. Wiedemeier, T. Wang, J. Wei, F. Chen, M. Kim, M. Kantarcioglu, and K. Jee, “Evading provenance-based ML detectors with adversarial system actions,” in USENIX Security , 2023, pp. 1199–1216
2023
Later among the works it cites.
M. A. Inam, Y. Chen, A. Goyal, J. Liu, J. Mink, N. Michael, S. Gaur, A. Bates, and W. U. Hassan, “Sok: History is a vast early warning system: Auditing the provenance of system intrusions,” in SP , 2023, pp. 2620–2638
2023
Later among the works it cites.
G. Zhou, Z. Liu, C. Fu, Q. Li, and K. Xu, “An efficient design of intelligent network data plane,” in USENIX Security , 2023
2023
Later among the works it cites.
C. Fu, Q. Li, and K. Xu, “Detecting unknown encrypted malicious traffic in real time via flow interaction graph analysis,” in NDSS , 2023
2023
Later among the works it cites.
VirusTotal, “Virustotal,” https://www.virustotal.com/gui/home/upload , 2023
2023
Later among the works it cites.
W. G. D. Note, “Ethical principles for web machine learning,” https://www.w3.org/TR/webmachinelearning-ethics , 2023
2023
Later among the works it cites.
G. AI, “Guardrails ai,” https://www.guardrailsai.com/docs/ , 2023
2023
Later among the works it cites.
Laiyer.ai, “Llm guard - the security toolkit for llm interactions,” https://github.com/laiyer-ai/llm-guard/ , 2023
2023
Later among the works it cites.
Azure, “Content filtering,” https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/content-filter?tabs=warning%2Cpython , 2023
2023
Later among the works it cites.
A. V. Lilian Weng, Vik Goel, “Using gpt-4 for content moderation,” https://searchengineland.com/openai-ai-classifier-no-longer-available-429912/ , 2023
2023
Later among the works it cites.
M. AI, “Llama 2 responsible use guide,” https://ai.meta.com/llama/responsible-use-guide/ , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
B. A. Galitsky, “Truth-o-meter: Collaborating with llm in fighting its hallucinations,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Gao, Z. Dai, P. Pasupat, A. Chen, A. T. Chaganty, Y. Fan, V. Zhao, N. Lao, H. Lee, D.-C. Juan et al. , “Rarr: Researching and revising what language models say, using language models,” in ACL , 2023, pp. 16 477–16 508
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Fang, Z. Tan, and X. Shi, “Cosywa: Enhancing semantic integrity in watermarking natural language generation,” in NLPCC , 2023, pp. 708–720
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Chen, Y. Li, Z. Wu, M. Xu, R. Wang, and Z. Zheng, “Towards reliable utilization of AIGC: blockchain-empowered ownership verification mechanism,” IEEE Open J. Comput. Soc. , vol. 4, pp. 326–337, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Vaghani and M. Thummar, “Flipkart product reviews with sentiment dataset,” https://www.kaggle.com/dsv/4940809 , 2023
2023
Later among the works it cites.
P. Manakul, A. Liusie, and M. J. F. Gales, “Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models,” in EMNLP , H. Bouamor, J. Pino, and K. Bali, Eds., 2023, pp. 9004–9017
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Li, X. Cheng, W. X. Zhao, J.-Y. Nie, and J.-R. Wen, “Halueval: A large-scale hallucination evaluation benchmark for large language models,” in EMNLP , 2023, pp. 6449–6464
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
OpenAI, “Open AI Privacy Policy,” https://openai.com/policies/privacy-policy , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Huang, H. Shao, and K. C. Chang, “Are large pre-trained language models leaking your personal information?” in EMNLP , 2022, pp. 2038–2047
2047
Closest in time.