Fetching the paper…
Reading the bibliography…
We study how to watermark LLM outputs, i.e.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Deep reinforcement learning: A brief survey
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Earlier work this paper cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization, 2019
Zhang, J., Zhao, Y., Saleh, M., and Liu, P. J · 2019
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Performance trade-offs of watermarking large language models
Ajith, A., Singh, S., and Pruthi, D · 2023
Cited alongside, same era.
Undetectable watermarks for language models
Christ, M., Gunn, S., and Zamir, O · 2023
Cited alongside, same era.
Publicly detectable watermarking for language models
Fairoze, J., Garg, S., Jha, S., Mahloujifar, S., Mahmoody, M., and Wang, M · 2023
Cited alongside, same era.
Three bricks to consolidate watermarks for large language models
Fernandez, P., Chaffin, A., Tit, K., Chappelier, V., and Furon, T · 2023
Who wrote this code? watermarking for code generation
Lee, T., Hong, S., Ahn, J., Hong, I., Lee, H., Yun, S., Shin, J., and Kim, G · 2023
Later among the works it cites.
A semantic invariant robust watermark for large language models
Liu, A., Pan, L., Hu, X., Meng, S., and Wen, L · 2023
Later among the works it cites.
On the fragility of learned reward functions
McKinney, L., Duan, Y., Krueger, D., and Gleave, A · 2023
Later among the works it cites.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the learnability of watermarks for language models
Gu, C., Li, X. L., Liang, P., and Hashimoto, T · 2023
Cited alongside, same era.
Semstamp: A semantic watermark with paraphrastic robustness for text generation
Hou, A. B., Zhang, J., He, T., Wang, Y., Chuang, Y.-S., Wang, H., Shen, L., Van Durme, B., Khashabi, D., and Tsvetkov, Y · 2023
Cited alongside, same era.
Unbiased watermark for large language models
Hu, Z., Chen, L., Wu, X., Wu, Y., Zhang, H., and Huang, H · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Sun, R., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense
Krishna, K., Song, Y., Karpinska, M., Wieting, J., and Iyyer, M · 2023
Cited alongside, same era.
Robust distortion-free watermarks for language models
Kuditipudi, R., Thickstun, J., Hashimoto, T., and Liang, P · 2023
Cited alongside, same era.
A watermark for large language models
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T
Cited in the paper.
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Can ai-generated text be reliably detected?
Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
A survey on llm-gernerated text detection: Necessity, methods, and future directions
Wu, J., Yang, S., Zhan, R., Yuan, Y., Wong, D. F., and Chao, L. S · 2023
Later among the works it cites.
Provable robust watermarking for ai-generated text
Zhao, X., Ananth, P., Li, L., and Wang, Y.-X · 2023
Later among the works it cites.
Secrets of rlhf in large language models part i: Ppo
Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Liu, Q., Zhou, Y., et al · 2023
Later among the works it cites.