Fetching the paper…
Reading the bibliography…
Potential harms of large language models can be mitigated by watermarking model output, i.e., embedding signals into generated text that are invisible to humans but algorithmically detectable from a short span of tokens.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 1910
Earlier work this paper cites.
Information hiding-a survey
Petitcolas, F., Anderson, R., and Kuhn, M · 1999
Earlier work this paper cites.
Information Hiding Techniques for Steganography and Digital Watermarking
Katzenbeisser, S. and Petitcolas, F. A · 2000
Earlier work this paper cites.
Natural Language Watermarking: Design, Analysis, and a Proof-of-Concept Implementation
Atallah, M. J., Raskin, V., Crogan, M., Hempelmann, C., Kerschbaum, F., Mohamed, D., and Naik, S · 2001
Earlier work this paper cites.
The homograph attack
Gabrilovich, E. and Gontmakher, A · 2002
Earlier work this paper cites.
Attacking Neural Text Detectors
Wolff, M. and Wolff, S · 2002
Earlier work this paper cites.
Natural Language Watermarking and Tamperproofing
Atallah, M. J., Raskin, V., Hempelmann, C. F., Karahan, M., Sion, R., Topkara, U., and Triezenberg, K. E · 2003
Earlier work this paper cites.
Natural Language Watermarking Using Semantic Substitution for Chinese Text
Chiang, Y.-L., Chang, L.-P., Hsieh, W.-T., and Chen, W.-C · 2004
Earlier work this paper cites.
Natural language watermarking via morphosyntactic alterations
Meral, H. M., Sankur, B., Sumru Özsoy, A., Güngör, T., and Sevinç, E · 2008
Earlier work this paper cites.
A Review of Digital Watermarking Techniques for Text Documents
Jalil, Z. and Mirza, A. M · 2009
Earlier work this paper cites.
Watermarking the Outputs of Structured Prediction with an application in Statistical Machine Translation
Venugopal, A., Uszkoreit, J., Talbot, D., Och, F., and Ganitkevitch, J · 2011
Earlier work this paper cites.
Dual canonicalization: An answer to the homograph attack
Helfrich, J. N. and Neff, R · 2012
Earlier work this paper cites.
Linguistic steganography on Twitter: Hierarchical language modeling with manual interaction
Wilson, A., Blunsom, P., and Ker, A. D · 2014
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Generating Steganographic Text with LSTMs
Fang, T., Jaggi, M., and Argyraki, K · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Turning your weakness into a strength: Watermarking deep neural networks by backdooring
Adi, Y., Baum, C., Cisse, M., Pinkas, B., and Keshet, J · 2018
Cited alongside, same era.
HiDDeN: Hiding Data with Deep Networks
Zhu, J., Kaplan, R., Johnson, J., and Fei-Fei, L · 2018
Cited alongside, same era.
Towards Near-imperceptible Steganographic Text
Dai, F. and Cai, Z · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Academic Plagiarism Detection: A Systematic Literature Review
Foltýnek, T., Meuschke, N., and Gipp, B · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
My AI Safety Lecture for UT Effective Altruism, November 2022
Aaronson, S · 2022
Later among the works it cites.
Constitutional AI: Harmlessness from AI Feedback, December 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosiute, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Later among the works it cites.
Guiding the Release of Safer E2E Conversational AI through Value Sensitive Design
Bergman, A. S., Abercrombie, G., Spruit, S., Hovy, D., Dinan, E., Boureau, Y.-L., and Rieser, V · 2022
Later among the works it cites.
Bad Characters: Imperceptible NLP Attacks
Boucher, N., Shumailov, I., Anderson, R., and Papernot, N · 2022
Later among the works it cites.
Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
Defending Against Neural Fake News
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y · 2019
Cited alongside, same era.
Neural Linguistic Steganography
Ziegler, Z., Deng, Y., and Rush, A · 2019
Cited alongside, same era.
Automatic Detection of Machine Generated Text: A Critical Survey
Jawahar, G., Abdul-Mageed, M., and Lakshmanan, V.S., L · 2020
Cited alongside, same era.
Detecting Cross-Modal Inconsistency to Defend Against Neural Fake News
Tan, R., Plummer, B., and Saenko, K · 2020
Cited alongside, same era.
Reverse Engineering Configurations of Neural Text Generation Models
Tay, Y., Bahri, D., Zheng, C., Brunk, C., Metzler, D., and Tomkins, A · 2020
Cited alongside, same era.
Crothers, E., Japkowicz, N., and Viktor, H · 2022
Later among the works it cites.
On pushing DeepFake Tweet Detection capabilities to the limits
Gambini, M., Fagni, T., Falchi, F., and Tesconi, M · 2022
Later among the works it cites.
The Ethical Need for Watermarks in Machine-Generated Language
Grinbaum, A. and Adomaitis, L · 2022
Later among the works it cites.
Watermarking Pre-trained Language Models with Backdooring
Gu, C., Huang, C., Zheng, X., Chang, K.-W., and Hsieh, C.-J · 2022
Later among the works it cites.
The Threat of Offensive AI to Organizations
Mirsky, Y., Demontis, A., Kotak, J., Shankar, R., Gelei, D., Yang, L., Zhang, X., Pintor, M., Lee, W., Elovici, Y., and Biggio, B · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Muennighoff, N., Wang, T., Sutawika, L., Roberts, A., Biderman, S., Scao, T. L., Bari, M. S., Shen, S., Yong, Z.-X., Schoelkopf, H., et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Robust Speech Recognition via Large-Scale Weak Supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Later among the works it cites.
ChatGPT: Optimizing Language Models for Dialogue, November 2022
Schulman, J., Zoph, B., Kim, C., Hilton, J., Menick, J., Weng, J., Uribe, J. F. C., Fedus, L., Metz, L., Pokorny, M., Gontijo-Lopes, R., Zhao, S., Vijayvergiya, A., Sigler, E., Perelman, A., Voss, C., Heaton, M., Parish, J., Cummings, D., Nayak, R., Balcom, V., Schnurr, D., Kaftan, T., Hallacy, C., Turley, N., Deutsch, N., Goel, V., Ward, J., Konstantinidis, A., Zaremba, W., Ouyang, L., Bogndonoff, L., Gross, J., Medina, D., Yoo, S., Lee, T., Lowe, R., Mossing, D., Huizinga, J., Jiang, R., Wainwright, C., Almeida, D., Lin, S., Zhang, M., Xiao, K., Slama, K., Bills, S., Gray, A., Leike, J., Pachocki, J., Tillet, P., Jain, S., Brockman, G., and Ryder, N · 2022
Later among the works it cites.
Unifying language learning paradigms
Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Bahri, D., Schuster, T., Zheng, H. S., Houlsby, N., and Metzler, D · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
Things like GPTZero are scary. I’m sure the creator didn’t have the intention, but the fact that it’s marketed as ”a solution to detecting AI written responses” even though there’s no evidence to show it works consistently and is nevertheless being EMPLOYED by schools, is crazy., January 2023
Butoi, V · 2023
Closest in time.
There are adversarial attacks for that proposal as well — in particular, generating with emojis after words and then removing them before submitting defeats it., January 2023
Goodside, R · 2023
Closest in time.
Gptzero update v1, January 2023
Tian, E · 2023
Closest in time.