Fetching the paper…
Reading the bibliography…
Methods for watermarking large language models have been proposed that distinguish AI-generated text from human-generated text by slightly altering the model output distribution, but they also distort the quality of the text, exposing the watermark to adversarial detection.
Probability, Random Variables, and Stochastic Processes
Papoulis, A. and Pillai, S · 2002
Earlier work this paper cites.
Detecting fake content with relative entropy scoring
Lavergne, T., Urvoy, T., and Yvon, F · 2008
Earlier work this paper cites.
Computer-generated text detection using machine learning: A systematic review
Beresneva, D · 2016
Earlier work this paper cites.
Real or fake? learning to discriminate machine from human generated text
Bakhtin, A., Gross, S., Ott, M., Deng, Y., Ranzato, M., and Szlam, A · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors · 2020
Earlier work this paper cites.
Adversarial watermarking transformer: Towards tracing text provenance with data hiding
Abdelnabi, S. and Fritz, M · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
TweepFake: About Detecting Deepfake Tweets
Fagni, T., Falchi, F., Gambini, M., Martella, A., and Tesconi, M · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Earlier work this paper cites.
Understanding the capabilities, limitations, and societal impact of large language models
Tamkin, A., Brundage, M., Clark, J., and Ganguli, D · 2021
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
My AI Safety Lecture for UT Effective Altruism., Nov. 2023
Aaronson, S · 2023
Cited alongside, same era.
ChatEval: Towards better LLM-based evaluators through multi-agent debate
Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z · 2023
Cited alongside, same era.
Undetectable Watermarks for Language Models
Christ, M., Gunn, S., and Zamir, O · 2023
Cited alongside, same era.
Towards Next-Generation Intelligent Assistants Leveraging LLM Techniques
Dong, X. L., Moon, S., Xu, Y. E., Malik, K., and Yu, Z · 2023
Cited alongside, same era.
Augmented language models: A survey
Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al · 2023
Later among the works it cites.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Gpt-2: 1.5b release
OpenAI · 2023
Later among the works it cites.
HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Later among the works it cites.
Gptzero update v1
Tian, E · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Three bricks to consolidate watermarks for large language models
Fernandez, P., Chaffin, A., Tit, K., Chappelier, V., and Furon, T · 2023
Cited alongside, same era.
Openagi: When llm meets domain experts
Ge, Y., Hua, W., Ji, J., Tan, J., Xu, S., and Zhang, Y · 2023
Cited alongside, same era.
Goldstein, J. A., Sastry, G., Musser, M., DiResta, R., Gentzel, M., and Sedova, K · 2023
Cited alongside, same era.
Robust distortion-free watermarks for language models
Kuditipudi, R., Thickstun, J., Hashimoto, T., and Liang, P · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models
Lapid, R., Langberg, R., and Sipper, M · 2023
Cited alongside, same era.
A semantic invariant robust watermark for large language models
Liu, A., Pan, L., Hu, X., Meng, S., and Wen, L · 2023
Cited alongside, same era.
A Watermark for Large Language Models
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T
Cited in the paper.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Chatgpt for robotics: Design principles and model abilities
Vemprala, S., Bonatti, R., Bucker, A., and Kapoor, A · 2023
Later among the works it cites.
Towards Codable Text Watermarking for Large Language Models
Wang, L., Yang, W., Chen, D., Zhou, H., Lin, Y., Meng, F., Zhou, J., and Sun, X · 2023
Later among the works it cites.
Anatomy of an AI-powered malicious social botnet
Yang, K.-C. and Menczer, F · 2023
Later among the works it cites.
Natural language is all a graph needs
Ye, R., Zhang, C., Wang, R., Xu, S., and Zhang, Y · 2023
Later among the works it cites.
Robust multi-bit natural language watermarking through invariant features
Yoo, K., Ahn, W., Jang, J., and Kwak, N · 2023
Later among the works it cites.