Fetching the paper…
Reading the bibliography…
Is it possible for machines to think like humans? And if it is, how should we go about teaching them to do so? As early as 1950, Alan Turing stated that we ought to teach machines in the way of teaching a child.
Programs with common sense
John McCarthy · 1960
Earlier work this paper cites.
Hannah arendt: From an interview, 1978
Hannah Arendt · 1978
Earlier work this paper cites.
Knowledge Acquisition, Knowledge Programming, and Knowledge Refinement
Frederick Hayes-Roth et al · 1980
Earlier work this paper cites.
Advice-taking and knowledge refinement: An iterative view of skill acquisition
Frederick Hayes-Roth, Philip Klahr, and David J Mostow · 1981
Earlier work this paper cites.
Information integrity: an emerging field and the state of knowledge
E. Geisler, P. Prabhaker, and M. Nayar · 2003
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Andrea Lockerd Thomaz, Cynthia Breazeal, et al · 2006
Earlier work this paper cites.
Methodological framework for analyzing social impact of technological innovations
Aistė Balžekienė, Eglė Butkevičienė, and Audronė Telešienė · 2008
Earlier work this paper cites.
Computing machinery and intelligence
Alan M Turing · 2009
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al · 2019
Earlier work this paper cites.
Artificial intelligence, values and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Better than our biases: Using psychological research to inform our approach to inclusive, effective feedback
Anne D Gordon · 2020
Earlier work this paper cites.
Autonomous automobilities: The social impacts of driverless vehicles
David Bissell, Thomas Birtchnell, Anthony Elliott, and Eric L Hsu · 2020
Earlier work this paper cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Earlier work this paper cites.
Evaluation of social impact measurement tools and techniques: a systematic review of the literature
Sally Kah and Temidayo Akenroye · 2020
Earlier work this paper cites.
Reinforcement learning with human advice: A survey
Anis Najar and Mohamed Chetouani · 2021
Cited alongside, same era.
The societal implications of deep reinforcement learning
Jess Whittlestone, Kai Arulkumaran, and Matthew Crosby · 2021
Cited alongside, same era.
This work is licensed by the State of Queensland Department of State Development, Infrastructure, Local Government and Planning under a Creative Commons Attribution (CC BY) 4.0 Australia licence. To view a copy of this license, visit creativecommons.org.au
Social impact evaluation guide: Business case development framework, release 3, June 2021 · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Cited alongside, same era.
Truth, lies, and automation
Ben Buchanan, Micah Musser, Andrew Lohn, and Katerina Sedova · 2021
Gpt-3 and instructgpt: technological dystopianism, utopianism, and “contextual” perspectives in ai ethics and industry
Anastasia Chan · 2022
Later among the works it cites.
Securing ai: How traditional vulnerability disclosure must adapt
Andrew J. Lohn and Wyatt Hoffman · 2022
Later among the works it cites.
A common language for responsible ai: Evolving and defining dod terms for implementation
Emelia S. Probasco · 2022
Later among the works it cites.
Improving multimodal interactive agents with reinforcement learning from human feedback
Josh Abramson, Arun Ahuja, Federico Carnevale, Petko Georgiev, Alex Goldin, Alden Hung, Jessica Landon, Jirka Lhotka, Timothy Lillicrap, Alistair Muldal, et al · 2022
Later among the works it cites.
The expertise problem: Learning from specialized feedback
Oliver Daniels-Koch and Rachel Freedman · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ai and the future of disinformation campaigns, part 2: A threat model
Katerina Sedova, Christine McNeill, Aurora Johnson, Aditi Joshi, and Ido Wulkan · 2021
Cited alongside, same era.
Alignment of language agents, 2021
Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving · 2021
Cited alongside, same era.
Reinforcement learning under moral uncertainty
Adrien Ecoffet and Joel Lehman · 2021
Cited alongside, same era.
Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle
Harini Suresh and John Guttag · 2021
Cited alongside, same era.
Ai for judges: A framework
James E. Baker, Laurie N. Hobart, and Matthew G. Mittelsteadt · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Aligning language models to follow instructions, 2022
OpenAI · 2022
Cited alongside, same era.
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Later among the works it cites.
Reinforcement learning from human feedback(rlhf)-chatgpt, 2023
Sthanikam Santhosh · 2023
Closest in time.
Illustrating reinforcement learning from human feedback (rlhf), 2023
Nazneen Rajani · 2023
Closest in time.
Introduction to reinforcement learning with human feedback, 2023
Edwin Chen · 2023
Closest in time.
The next generation of large language models, 2023
Rob Toews · 2023
Closest in time.
Social impact assessment
International Association for Impact Assessment · 2023
Closest in time.
The capacity for moral self-correction in large language models, 2023
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, Dawn Drain, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jackson Kernion, Jamie Kerr, Jared Mueller, Joshua Landau, Kamal Ndousse, Karina Nguyen, Liane Lovitt, Michael Sellitto, Nelson Elhage, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robert Lasenby, Robin Larson, Sam Ringer, Sandipan Kundu, Saurav Kadavath, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, Christopher Olah, Jack Clark, Samuel R. Bowman, and Jared Kaplan · 2023
Closest in time.
Reinforcement learning from human feedback, 2023
Oluwafemi Smith · 2023
Closest in time.
What is reinforcement learning from human feedback (rlhf)?, 2023
Ben Dickson · 2023
Closest in time.
An introduction to training llms using reinforcement learning from human feedback (rlhf), 2023
Ayush Thakur · 2023
Closest in time.