Fetching the paper…
Reading the bibliography…
Instruction-tuned LMs such as ChatGPT, FLAN, and InstructGPT are finetuned on datasets that contain user-submitted examples, e.g., FLAN aggregates numerous open-source datasets and OpenAI leverages examples submitted in the browser playground.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Ensemble adversarial training: Attacks and defenses
Tramèr, F., Kurakin, A., Papernot, N., Goodfellow, I., Boneh, D., and McDaniel, P · 2018
Earlier work this paper cites.
Data poisoning attacks on multi-task relationship learning
Zhao, M., An, B., Yu, Y., Liu, S., and Pan, S · 2018
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Poison attacks against text datasets with conditional adversarially regularized autoencoder
Chan, A., Tay, Y., Ong, Y.-S., and Zhang, A · 2020
Earlier work this paper cites.
MetaPoison: practical general-purpose clean-label data poisoning
Huang, W. R., Geiping, J., Fowl, L., Taylor, G., and Goldstein, T · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Kurita, K., Michel, P., and Neubig, G · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Concealed data poisoning attacks on NLP models
Wallace, E., Zhao, T. Z., Feng, S., and Singh, S · 2020
Cited alongside, same era.
Poisoning the unlabeled dataset of semi-supervised learning
Carlini, N · 2021
Cited alongside, same era.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., and Raffel, C · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
You autocomplete me: Poisoning vulnerabilities in neural code completion
Schuster, R., Song, C., Tromer, E., and Shmatikov, V · 2021
Cited alongside, same era.
Data poisoning attacks on federated machine learning
Sun, G., Cong, Y., Dong, J., Wang, Q., Lyu, L., and Liu, J · 2021
Deduplicating training data mitigates privacy risks in language models
Kandpal, N., Wallace, E., and Raffel, C · 2022
Later among the works it cites.
PoisonedEncoder: Poisoning the unlabeled pre-training data in contrastive learning
Liu, H., Jia, J., and Gong, N. Z · 2022
Later among the works it cites.
AI-powered coding assistant aims to help, not replace developers
Loten, A · 2022
Later among the works it cites.
MetaICL: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Analyzing dynamic adversarial training data in the limit
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in NLP models
Yang, W., Li, L., Zhang, Z., Ren, X., Sun, X., and He, B · 2021
Cited alongside, same era.
Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections
Zhong, R., Lee, K., Zhang, Z., and Klein, D · 2021
Cited alongside, same era.
Poisoning and backdooring contrastive learning
Carlini, N. and Terzis, A · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Wallace, E., Williams, A., Jia, R., and Kiela, D · 2022
Later among the works it cites.
Benchmarking generalization via in-context instructions on 1,600+ language tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Later among the works it cites.
Measuring forgetting of memorized training examples
Jagielski, M., Thakkar, O., Tramér, F., Ippolito, D., Lee, K., Carlini, N., Wallace, E., Song, S., Thakurta, A., Papernot, N., and Zhang, C · 2023
Closest in time.
Microsoft bets big on the creator of ChatGPT in race to dominate A.I
Metz, C. and Weise, K · 2023
Closest in time.