Fetching the paper…
Reading the bibliography…
Making language models bigger does not inherently make them better at following a user's intent.
Learning from dialogue after deployment: Feed yourself, chatbot!
Hancock, B., Bordes, A., Mazare, P.-E., and Weston, J. (2019) · 1901
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M. (2019) · 1903
Earlier work this paper cites.
Yi, S., Goel, R., Khatri, C., Cervone, A., Chung, T., Hedayatnia, B., Venkatesh, A., Gabriel, R., and Hakkani-Tur, D. (2019) · 1904
Earlier work this paper cites.
Reducing gender bias in word-level language models with a gender-equalizing loss function
Qian, Y., Muaz, U., Zhang, B., and Hyun, J. W. (2019) · 1905
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2019) · 1905
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019) · 1907
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Dinan, E., Humeau, S., Chintagunta, B., and Weston, J. (2019b) · 1908
Earlier work this paper cites.
Release strategies and the social impacts of language models
Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., et al. (2019) · 1908
Earlier work this paper cites.
Better rewards yield better summaries: Learning to summarise without references
Böhm, F., Gao, Y., Meyer, C. M., Shapira, O., Dagan, I., and Gurevych, I. (2019) · 1909
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. (2019) · 1909
Earlier work this paper cites.
Finding generalizable evidence by learning to convince q&a models
Perez, E., Karamcheti, S., Fergus, R., Weston, J., Kiela, D., and Cho, K. (2019) · 1909
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N. (2019) · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 1909
Earlier work this paper cites.
Does gender matter? towards fairness in dialogue systems
Liu, H., Dacon, J., Fan, W., Liu, H., Liu, Z., and Tang, J. (2019) · 1910
Earlier work this paper cites.
Queens are powerful too: Mitigating gender bias in dialogue generation
Dinan, E., Fan, A., Williams, A., Urbanek, J., Kiela, D., and Weston, J. (2019a) · 1911
Earlier work this paper cites.
Reducing sentiment bias in language models via counterfactual evaluation
Huang, P.-S., Zhang, H., Jiang, R., Stanforth, R., Welbl, J., Rae, J., Maini, V., Yogatama, D., and Kohli, P. (2019) · 1911
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. (2019) · 1912
Earlier work this paper cites.
Zhou, W. and Xu, K. (2020) · 2002
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S. (2020) · 2004
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in nlp
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H. (2020) · 2005
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2005
Earlier work this paper cites.
Unifiedqa: Crossing format boundaries with a single qa system
Khashabi, D., Min, S., Khot, T., Sabharwal, A., Tafjord, O., Clark, P., and Hajishirzi, H. (2020) · 2005
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. (2020) · 2009
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. (2020) · 2009
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. (2020) · 2009
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. (2020) · 2010
Earlier work this paper cites.
Imitating interactive intelligence
Abramson, J., Ahuja, A., Barr, I., Brussee, A., Carnevale, F., Cassin, M., Chhaparia, R., Clark, S., Damoc, B., Dudzik, A., et al. (2020) · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. (2013) · 2013
Earlier work this paper cites.
Superintelligence
Bostrom, N. (2014) · 2014
Earlier work this paper cites.
Findings of the 2015 workshop on statistical machine translation
Bojar, O., Chatterjee, R., Federmann, C., Haddow, B., Huck, M., Hokamp, C., Koehn, P., Logacheva, V., Monz, C., Negri, M., Post, M., Scarton, C., Specia, L., and Turchi, M. (2015) · 2015
Earlier work this paper cites.
Corrigibility
Soares, N., Fallenstein, B., Armstrong, S., and Yudkowsky, E. (2015) · 2015
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A., and Bengio, Y. (2016) · 2016
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Nallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al. (2016) · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Thinking fast and slow with deep learning and tree search
Anthony, T., Tian, Z., and Barber, D. (2017) · 2017
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021) · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021) · 2021
Later among the works it cites.
Truth, lies, and automation
Buchanan, B., Lohn, A., Musser, M., and Sedova, K. (2021) · 2021
Later among the works it cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al. (2021) · 2021
Later among the works it cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A., Bryson, J. J., and Narayanan, A. (2017) · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Cited alongside, same era.
Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017) · 2017
Cited alongside, same era.
Tl; dr: Mining reddit to learn automatic summarization
Völske, M., Potthast, M., Syed, S., and Stein, B. (2017) · 2017
Cited alongside, same era.
Learning to understand goal specifications by modelling reward
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Eliciting latent knowledge: How to tell if your eyes deceive you
Christiano, P., Cotra, A., and Xu, M. (2021) · 2021
Later among the works it cites.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K.-W., and Gupta, R. (2021) · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N. (2021) · 2021
Later among the works it cites.
Kenton, Z., Everitt, T., Weidinger, L., Gabriel, I., Mikulik, V., and Irving, G. (2021) · 2021
Later among the works it cites.
How true is gpt-2? an empirical analysis of intersectional occupational biases
Kirk, H., Jun, Y., Iqbal, H., Benussi, E., Volpin, F., Dreyer, F. A., Shtedritski, A., and Asano, Y. M. (2021) · 2021
Later among the works it cites.
Towards understanding and mitigating social biases in language models
Liang, P. P., Wu, C., Morency, L.-P., and Salakhutdinov, R. (2021) · 2021
Later among the works it cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O. (2021) · 2021
Later among the works it cites.
Stereotype and skew: Quantifying gender bias in pre-trained and fine-tuned language models
Manela, D. d. V., Errington, D., Fisher, T., van Breugel, B., and Minervini, P. (2021) · 2021
Later among the works it cites.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H. (2021) · 2021
Later among the works it cites.
Training value-aligned reinforcement learning agents using a normative prior
Nahian, M. S. A., Frazier, S., Harrison, B., and Riedl, M. (2021) · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. (2021) · 2021
Later among the works it cites.
Mitigating harm in language models with conditional-likelihood filtration
Ngo, H., Raterink, C., Araújo, J. G., Zhang, I., Chen, C., Morisot, A., and Frosst, N. (2021) · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. (2021) · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. (2021) · 2021
Later among the works it cites.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Schick, T., Udupa, S., and Schütze, H. (2021) · 2021
Later among the works it cites.
Process for adapting language models to society (palms) with values-targeted datasets
Solaiman, I. and Dennison, C. (2021) · 2021
Later among the works it cites.
Understanding the capabilities, limitations, and societal impact of large language models
Tamkin, A., Brundage, M., Clark, J., and Ganguli, D. (2021) · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. (2021) · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., et al. (2021) · 2021
Later among the works it cites.
Challenges in detoxifying language models
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S. (2021) · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P. (2021) · 2021
Later among the works it cites.
Detoxifying language models risks marginalizing minority voices
Xu, A., Pathak, E., Wallace, E., Gururangan, S., Sap, M., and Klein, D. (2021) · 2021
Later among the works it cites.
On the evaluation of vision-and-language navigation instructions
Zhao, M., Anderson, P., Jain, V., Wang, S., Ku, A., Baldridge, J., and Ie, E. (2021) · 2021
Later among the works it cites.
Memory-assisted prompt editing to improve gpt-3 after deployment
Madaan, A., Tandon, N., Clark, P., and Yang, Y. (2022) · 2022
Closest in time.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. (2022) · 2022
Closest in time.