Fetching the paper…
Reading the bibliography…
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 1904
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2019
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 1909
Earlier work this paper cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
H. F. Song, A. Abdolmaleki, J. T. Springenberg, A. Clark, H. Soyer, J. W. Rae, S. Noury, A. Ahuja, S. Liu, D. Tirumala, N. Heess, D. Belov, M. Riedmiller, and M. M. Botvinick · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Logic and conversation
H. P. Grice · 1975
Earlier work this paper cites.
The Rating of Chessplayers, Past and Present
A. E. Elo · 1978
Earlier work this paper cites.
The case for motivated reasoning
Z. Kunda · 1990
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in nlp
S. L. Blodgett, S. Barocas, H. Daumé III, and H. Wallach · 2005
Earlier work this paper cites.
Selective question answering under domain shift
A. Kamath, R. Jia, and P. Liang · 2006
Earlier work this paper cites.
Bringing the people back in: Contesting benchmark machine learning datasets
E. Denton, A. Hanna, R. Amironesei, A. Smart, H. Nicole, and M. K. Scheuerman · 2007
Earlier work this paper cites.
Participation is not a design fix for machine learning
M. Sloane, E. Moss, O. Awomolo, and L. Forlano · 2007
Earlier work this paper cites.
The radicalization risks of GPT-3 and advanced neural language models
K. McGuffie and A. Newhouse · 2009
Earlier work this paper cites.
On the foundations of noise-free selective classification
R. El-Yaniv and Y. Wiener · 2010
Earlier work this paper cites.
Keeping women in the science pipeline
M. Goulden, M. A. Mason, and K. Frasch · 2011
Earlier work this paper cites.
Computing Krippendorff’s alpha-reliability, 2011
K. Krippendorff · 2011
Earlier work this paper cites.
Besting the Quiz Master: Crowdsourcing incremental classification games
J. Boyd-Graber, B. Satinoff, H. He, and H. Daumé III · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Deep learning with limited numerical precision
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan · 2015
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Harley, T. P. Lillicrap, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Selective classification for deep neural networks
Y. Geifman and R. El-Yaniv · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Earlier work this paper cites.
But who protects the moderators? the case of crowdsourced image moderation
B. Dang, M. J. Riedl, and M. Lease · 2018
Earlier work this paper cites.
G. Irving, P. Christiano, and D. Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
R. Rudinger, J. Naradowsky, B. Leonard, and B. Van Durme · 2018
Earlier work this paper cites.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
N. Shazeer and M. Stern · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
J. Zhao, T. Wang, M. Yatskar, V. Ordonez, and K.-W. Chang · 2018
Cited alongside, same era.
Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts
L. Breitfeller, E. Ahn, D. Jurgens, and Y. Tsvetkov · 2019
Cited alongside, same era.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
E. Dinan, S. Humeau, B. Chintagunta, and J. Weston · 2019
Cited alongside, same era.
Eli5: Long form question answering
A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli · 2019
Cited alongside, same era.
Selectivenet: A deep neural network with an integrated reject option
Y. Geifman and R. El-Yaniv · 2019
Cited alongside, same era.
How to never be wrong
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P.-S. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J.-B. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. Hechtman, L. Weidinger, I. Gabriel, W. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving · 2021
Later among the works it cites.
Cross-policy compliance detection via question answering
M. Saeidi, M. Yazdani, and A. Vlachos · 2021
Later among the works it cites.
The psychological well-being of content moderators: the emotional labor of commercial moderation and avenues for improving support
M. Steiger, T. J. Bharucha, S. Venkatagiri, M. J. Riedl, and M. Lease · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Gershman · 2019
Cited alongside, same era.
AI safety needs social scientists
G. Irving and A. Askell · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Cited alongside, same era.
WeBuildAI: Participatory framework for algorithmic governance
M. K. Lee, D. Kusbit, A. Kahng, J. T. Kim, X. Yuan, A. Chan, D. See, R. Noothigattu, S. Lee, A. Psomas, and A. D. Procaccia · 2019
Cited alongside, same era.
Finding generalizable evidence by learning to convince Q&A models
E. Perez, S. Karamcheti, R. Fergus, J. Weston, D. Kiela, and K. Cho · 2019
Cited alongside, same era.
Challenges and frontiers in abusive content detection
B. Vidgen, A. Harris, D. Nguyen, R. Tromble, S. Hale, and H. Margetts · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, C. Gulcehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wunsch, K. McKinney, O. Smith, T. Schaul, T. P. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Cited alongside, same era.
Fairness for unobserved characteristics: Insights from technological impacts on queer communities
N. Tomasev, K. R. McKee, J. Kay, and S. Mohamed · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Laura Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel · 2021
Later among the works it cites.
Challenges in detoxifying language models
J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Later among the works it cites.
Detoxifying language models risks marginalizing minority voices
A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein · 2021
Later among the works it cites.
Bot-adversarial dialogue for safe conversational agents
J. Xu, D. Ju, M. Li, Y.-L. Boureau, J. Weston, and E. Dinan · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan · 2022
Closest in time.
Power to the people? opportunities and challenges for participatory ai
A. Birhane, V. Prabhakaran, M. Diaz, I. Gabriel, M. C. Elish, S. Mohamed, and W. S. Isaac · 2022
Closest in time.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. van den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. de Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. W. Rae, E. Elsen, and L. Sifre · 2022
Closest in time.
Selection-inference: Exploiting large language models for interpretable logical reasoning
A. Creswell, M. Shanahan, and I. Higgins · 2022
Closest in time.
D. Dohan, W. Xu, A. Lewkowycz, J. Austin, D. Bieber, R. Gontijo Lopes, Y. Wu, H. Michalewski, R. A. Saurous, J. Sohl-dickstein, K. Murphy, and C. Sutton · 2022
Closest in time.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
N. Dziri, S. Milton, M. Yu, O. Zaiane, and S. Reddy · 2022
Closest in time.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. v. d. Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Closest in time.
In conversation with Artificial Intelligence: towards a theory of ideal speech for humans and language technologies
A. Kasirzadeh and I. Gabriel · 2022
Closest in time.
Policy compliance detection via expression tree inference
N. Kotonya, A. Vlachos, M. Yazdani, L. Mathias, and M. Saeidi · 2022
Closest in time.
Internet-augmented language models through few-shot prompting for open-domain question answering
A. Lazaridou, E. Gribovskaya, W. Stokowiec, and N. Grigorev · 2022
Closest in time.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra · 2022
Closest in time.
TruthfulQA: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Closest in time.
StreamingQA: A benchmark for adaptation to new knowledge over time in question answering models
A. Liška, T. Kočiský, E. Gribovskaya, T. Terzi, E. Sezener, D. Agrawal, C. de Masson d’Autume, T. Scholtes, M. Zaheer, S. Young, E. Gilsenan-McMahon, S. Austin, P. Blunsom, and A. Lazaridou · 2022
Closest in time.
Teaching language models to support answers with verified quotes
J. Menick, M. Trebacz, V. Mikulik, J. Aslanides, F. Song, M. Chadwick, M. Glaese, S. Young, L. Campbell-Gillingham, G. Irving, and N. McAleese · 2022
Closest in time.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Closest in time.
Single-turn debate does not help humans answer hard reading-comprehension questions
A. Parrish, H. Trivedi, E. Perez, A. Chen, N. Nangia, J. Phang, and S. R. Bowman · 2022
Closest in time.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Closest in time.
Explainable deep learning: A field guide for the uninitiated
G. Ras, N. Xie, M. van Gerven, and D. Doran · 2022
Closest in time.
Characteristics of harmful text: Towards rigorous benchmarking of language models
M. Rauh, J. Mellor, J. Uesato, P.-S. Huang, J. Welbl, L. Weidinger, S. Dathathri, A. Glaese, G. Irving, I. Gabriel, W. Isaac, and L. A. Hendricks · 2022
Closest in time.
Self-critiquing models for assisting human evaluators
W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike · 2022
Closest in time.
LaMDA: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, V. Zhao, Y. Zhou, C.-C. Chang, I. Krivokon, W. Rusch, M. Pickett, P. Srinivasan, L. Man, K. Meier-Hellstern, M. Ringel Morris, T. Doshi, R. Delos Santos, T. Duke, J. Soraker, B. Zevenbergen, V. Prabhakaran, M. Diaz, B. Hutchinson, K. Olson, A. Molina, E. Hoffman-John, J. Lee, L. Aroyo, R. Rajakumar, A. Butryna, M. Lamm, V. Kuzmina, J. Fenton, A. Cohen, R. Bernstein, R. Kurzweil, B. Aguera-Arcas, C. Cui, M. Croak, E. Chi, and Q. Le · 2022
Closest in time.
Conversational information seeking
H. Zamani, J. R. Trippas, J. Dalton, and F. Radlinski · 2022
Closest in time.