Fetching the paper…
Reading the bibliography…
Since the remarkable generation performance of large language models raised ethical and legal concerns, approaches to detect machine-generated text by embedding watermarks are being developed.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019 · 1908
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 1909
Earlier work this paper cites.
Natural language watermarking: Design, analysis, and a proof-of-concept implementation
Mikhail J. Atallah, Victor Raskin, Michael Crogan, Christian Hempelmann, Florian Kerschbaum, Dina Mohamed, and Sanket Naik. 2001 · 2001
Earlier work this paper cites.
Natural language watermarking and tamperproofing
Mikhail J. Atallah, Victor Raskin, Christian F. Hempelmann, Mercan Karahan, Radu Sion, Umut Topkara, and Katrina E. Triezenberg. 2002 · 2002
Earlier work this paper cites.
A text watermarking algorithm based on word classification and inter-word space statistics
Young-Won Kim, Kyung-Ae Moon, and Il-Seok Oh. 2003 · 2003
Earlier work this paper cites.
The evaluation of two software watermarking algorithms
Ginger Myles, Christian Collberg, Zachary Heidepriem, and Armand Navabi. 2005 · 2005
Earlier work this paper cites.
The hiding virtues of ambiguity
Umut Topkara, Mercan Topkara, and Mikhail J. Atallah. 2006 · 2006
Earlier work this paper cites.
A review of digital watermarking techniques for text documents
Zunera Jalil and Anwar M. Mirza. 2009 · 2009
Earlier work this paper cites.
Natural language watermarking via morphosyntactic alterations
Hasan Mesut Meral, Bülent Sankur, A. Sumru Özsoy, Tunga Güngör, and Emre Sevinç. 2009 · 2009
Earlier work this paper cites.
Design of a software watermarking algorithm based on register allocation
Jun Li and Quan Liu. 2010 · 2010
Earlier work this paper cites.
A survey of static software watermarking
James Hamilton and Sebastian Danicic. 2011 · 2011
Earlier work this paper cites.
Software watermarking: Progress and challenges
Ayan Dey, Sukriti Bhattacharya, and Nabendu Chaki. 2018 · 2018
Earlier work this paper cites.
Exception handling-based dynamic software watermarking
Yilong Wang, Daofu Gong, Bin Lu, Fei Xiang, and Fenlin Liu. 2018 · 2018
Earlier work this paper cites.
GLTR: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander Rush. 2019 · 2019
Earlier work this paper cites.
Xmark: Dynamic software watermarking using collatz conjecture
Haoyu Ma, Chunfu Jia, Shijia Li, Wantong Zheng, and Dinghao Wu. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Automatic detection of generated text is easiest when humans are fooled
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020 · 2020
Earlier work this paper cites.
Adversarial watermarking transformer: Towards tracing text provenance with data hiding
Sahar Abdelnabi and Mario Fritz. 2021 · 2021
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Cited alongside, same era.
Efficient training of language models to fill in the middle
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. 2022 · 2022
Cited alongside, same era.
Asleep at the keyboard? assessing the security of GitHub copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2022 · 2022
Cited alongside, same era.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023 · 2023
Closest in time.
The threat of offensive AI to organizations
Yisroel Mirsky, Ambra Demontis, Jaidip Kotak, Ram Shankar, Deng Gelei, Liu Yang, Xiangyu Zhang, Maura Pintor, Wenke Lee, Yuval Elovici, and Battista Biggio. 2023 · 2023
Closest in time.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Closest in time.
Sandra Mitrović, Davide Andreoletti, and Omran Ayoub. 2023 · 2023
Closest in time.
Codegen: An open large language model for code with multi-turn program synthesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models
Priyan Vaithilingam, Tianyi Zhang, and Elena L Glassman. 2022 · 2022
Cited alongside, same era.
Tracing text provenance via context-aware lexical substitution
Xi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang, Zehua Ma, Feng Wang, and Nenghai Yu. 2022 · 2022
Cited alongside, same era.
Undetectable watermarks for language models
Miranda Christ, Sam Gunn, and Or Zamir. 2023 · 2023
Cited alongside, same era.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2023 · 2023
Cited alongside, same era.
Incoder: A generative model for code infilling and synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen-tau Yih, Luke Zettlemoyer, and Mike Lewis. 2023 · 2023
Cited alongside, same era.
starcoder-gpteacher-code-instruct
GeorgiaTechResearchInstitute. 2023 · 2023
Cited alongside, same era.
On the learnability of watermarks for language models
Chenchen Gu, Xiang Lisa Li, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Cited alongside, same era.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al. 2023 · 2023
Cited alongside, same era.
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023 · 2023
Closest in time.
A robust semantics-based watermark for large language model against paraphrasing
Jie Ren, Han Xu, Yiding Liu, Yingqian Cui, Shuaiqiang Wang, Dawei Yin, and Jiliang Tang. 2023 · 2023
Closest in time.
Can ai-generated text be reliably detected?
Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023 · 2023
Closest in time.
Lost at c: A user study on the security implications of large language model code assistants
Gustavo Sandoval, Hammond A. Pearce, Teo Nys, Ramesh Karri, Siddharth Garg, and Brendan Dolan-Gavitt. 2023 · 2023
Closest in time.
Necessary and sufficient watermark for large language models
Yuki Takezawa, Ryoma Sato, Han Bao, Kenta Niwa, and Makoto Yamada. 2023 · 2023
Closest in time.
Gptzero: Towards detection of ai-generated text using zero-shot and supervised methods
Edward Tian and Alexander Cui. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Towards codable text watermarking for large language models
Lean Wang, Wenkai Yang, Deli Chen, Hao Zhou, Yankai Lin, Fandong Meng, Jie Zhou, and Xu Sun. 2023 · 2023
Closest in time.
LLMDet: A third party large language models generated text detection tool
Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua. 2023 · 2023
Closest in time.
Dna-gpt: Divergent n-gram analysis for training-free detection of gpt-generated text
Xianjun Yang, Wei Cheng, Linda Petzold, William Yang Wang, and Haifeng Chen. 2023 · 2023
Closest in time.
Robust natural language watermarking through invariant features
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak. 2023 · 2023
Closest in time.
Gpt paternity test: Gpt generated text detection with gpt genetic inheritance
Xiao Yu, Yuang Qi, Kejiang Chen, Guoqiang Chen, Xi Yang, Pengyuan Zhu, Weiming Zhang, and Nenghai Yu. 2023 · 2023
Closest in time.
Provable robust watermarking for ai-generated text
Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yu-Xiang Wang. 2023 · 2023
Closest in time.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, et al. 2023 · 2023
Closest in time.
A survey of text watermarking in the era of large language models
Aiwei Liu, Leyi Pan, Yijian Lu, Jingjing Li, Xuming Hu, Lijie Wen, Irwin King, and Philip S. Yu. 2024 · 2024
Closest in time.
Octopack: Instruction tuning code large language models
Niklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro Von Werra, and Shayne Longpre. 2024 · 2024
Closest in time.