Fetching the paper…
Reading the bibliography…
Pretrained language models often generate outputs that are not in line with human preferences, such as harmful text or factually incorrect summaries.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 1904
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2005
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2009
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P · 2015
Earlier work this paper cites.
Dialogue learning with human-in-the-loop
Li, J., Miller, A. H., Chopra, S., Ranzato, M., and Weston, J · 2016
Earlier work this paper cites.
Dialog-based language learning
Weston, J. E · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S · 2017
Earlier work this paper cites.
Teaching machines to describe images with natural language feedback
Fidler, S. et al · 2017
Earlier work this paper cites.
Beating Atari with Natural Language Guided Reinforcement Learning, 2017
Kaplan, R., Sauer, C., and Sosa, A · 2017
Earlier work this paper cites.
TL;DR: Mining Reddit to learn automatic summarization
Völske, M., Potthast, M., Syed, S., and Stein, B · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Camburu, O.-M., Rocktäschel, T., Lukasiewicz, T., and Blunsom, P · 2018
Earlier work this paper cites.
Improving Language Understanding by Generative Pre-Training, 2018
Radford, A. and Narasimhan, K · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Guide me: Interacting with deep networks
Rupprecht, C., Laina, I., Navab, N., Hager, G. D., and Tombari, F · 2018
Earlier work this paper cites.
Using Natural Language for Reward Shaping in Reinforcement Learning, 2019
Goyal, P., Niekum, S., and Mooney, R. J · 2019
Earlier work this paper cites.
Learning from dialogue after deployment: Feed yourself, chatbot!
Hancock, B., Bordes, A., Mazare, P.-E., and Weston, J · 2019
Earlier work this paper cites.
A survey of reinforcement learning informed by natural language
Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktäschel, T · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Speak to your parser: Interactive text-to-SQL with natural language feedback
Elgohary, A., Hosseini, S., and Hassan Awadallah, A · 2020
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Cited alongside, same era.
Stanza: A Python Natural Language Processing Toolkit for Many Human Languages
Qi, P., Zhang, Y., Zhang, Y., Bolton, J., and Manning, C. D · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, 2020
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Scaling laws for reward model overoptimization, 2022
Gao, L., Schulman, J., and Hilton, J · 2022
Later among the works it cites.
Measuring goodhart’s law
Hilton, J. and Gao, L · 2022
Later among the works it cites.
Rl with kl penalties is better viewed as bayesian inference
Korbak, T., Perez, E., and Buckley, C. L · 2022
Later among the works it cites.
Can language models learn from explanations in context?
Lampinen, A. K., Dasgupta, I., Chan, S. C., Matthewson, K., Tessler, M. H., Creswell, A., McClelland, J. L., Wang, J. X., and Hill, F · 2022
Later among the works it cites.
Li, Z., Sharma, P., Lu, X. H., Cheung, J. C., and Reddy, S · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hase, P. and Bansal, M · 2021
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods, 2021
Lin, S., Hilton, J., and Evans, O · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2021
Cited alongside, same era.
Cut the carp: Fishing for zero-shot story evaluation
Matiana, S., Smith, J., Teehan, R., Castricato, L., Biderman, S., Gao, L., and Frazier, S · 2021
Cited alongside, same era.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Cited alongside, same era.
Interactive learning from activity description
Nguyen, K. X., Misra, D., Schapire, R., Dudík, M., and Shafto, P · 2021
Cited alongside, same era.
True few-shot learning with language models
Perez, E., Kiela, D., and Cho, K · 2021
Cited alongside, same era.
Later among the works it cites.
Inferring rewards from language in context
Lin, J., Fried, D., Klein, D., and Dragan, A · 2022
Later among the works it cites.
On improving summarization factual consistency from natural language feedback
Liu, Y., Deb, B., Teruel, M., Halfaker, A., Radev, D., and Awadallah, A. H · 2022
Later among the works it cites.
Text and Code Embeddings by Contrastive Pre-Training, 2022
Neelakantan, A., Xu, T., Puri, R., Radford, A., Han, J. M., Tworek, J., Yuan, Q., Tezak, N., Kim, J. W., Hallacy, C., Heidecke, J., Shyam, P., Power, B., Nekoul, T. E., Sastry, G., Krueger, G., Schnurr, D., Such, F. P., Hsu, K., Thompson, M., Khan, T., Sherbakov, T., Jang, J., Welinder, P., and Weng, L · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Later among the works it cites.
Training language models with language feedback
Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K., and Perez, E · 2022
Later among the works it cites.
Peer: A collaborative language model
Schick, T., Dwivedi-Yu, J., Jiang, Z., Petroni, F., Lewis, P., Izacard, G., You, Q., Nalmpantis, C., Grave, E., and Riedel, S · 2022
Later among the works it cites.
When life gives you lemons, make cherryade: Converting feedback from bad responses into good labels
Shi, W., Dinan, E., Shuster, K., Weston, J., and Xu, J · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D · 2022
Later among the works it cites.
Xu, J., Ung, M., Komeili, M., Arora, K., Boureau, Y.-L., and Weston, J · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Improving code generation by training with natural language feedback
Chen, A., Scheurer, J., Korbak, T., Campos, J. A., Chan, J. S., Bowman, S. R., Cho, K., and Perez, E · 2023
Closest in time.