Fetching the paper…
Reading the bibliography…
Our intention is to provide a definitive reference on what it would take to safely make use of generative/predictive models in the absence of a solution to the Eliciting Latent Knowledge problem.
Risks from Learned Optimization in Advanced Machine Learning Systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulika, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Distributional Generalization: A New Kind of Generalization
Preetum Nakkiran and Yamini Bansal · 2009
Earlier work this paper cites.
An intuitive explanation of solomonoff induction, 2012
Alex Altair · 2012
Earlier work this paper cites.
An overview of 11 proposals for building safe advanced AI
Evan Hubinger · 2012
Earlier work this paper cites.
Tiling agents for self-modifying ai, and the löbian obstacle, 2013
Eliezer Yudkowsky and Marcello Herreshoff · 2013
Earlier work this paper cites.
Quantilizers: A Safer Alternative to Maximizers for Limited Optimization
Jessica Taylor · 2016
Earlier work this paper cites.
Good and safe uses of AI Oracles
Stuart Armstrong and Xavier O’Rorke · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Worst-case guarantees, 2019
Paul Christiano · 2019
Earlier work this paper cites.
Counterfactual oracles = online supervised learning with random selection of training episodes, 2019
Wei Dai · 2019
Earlier work this paper cites.
Relaxed adversarial training for inner alignment, 2019
Evan Hubinger · 2019
Earlier work this paper cites.
Current work in ai alignment, 2020
Paul Christiano · 2020
Earlier work this paper cites.
Homogeneity vs. heterogeneity in ai takeoff scenarios, 2020
Evan Hubinger · 2020
Earlier work this paper cites.
Multiple worlds, one universal wave function, 2020
Evan Hubinger · 2020
Earlier work this paper cites.
Causal Decision Theory
Paul Weirich · 2020
Earlier work this paper cites.
Decision Transformer: Reinforcement Learning via Sequence Modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Earlier work this paper cites.
Proposal for a regulation of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), 2021
European Commission · 2021
Earlier work this paper cites.
Why ai alignment could be hard with modern deep learning, 2021
Ajeya Cotra · 2021
Earlier work this paper cites.
Eliciting latent knowledge: How to tell if your eyes deceive you
Paul Christiano, Mark Xu, and Ajeya Cotra · 2021
Earlier work this paper cites.
How do we become confident in the safety of a machine learning system?
Evan Hubinger · 2021
Earlier work this paper cites.
Language models are multiverse generators
Janus · 2021
Cited alongside, same era.
Fun with +12 ooms of compute, 2021
Daniel Kokotajlo · 2021
Cited alongside, same era.
The Power of Scale for Parameter-Efficient Prompt Tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Lcdt, a myopic decision theory, 2021
Adam Shimi and Evan Hubinger · 2021
Cited alongside, same era.
Agents over cartesian world models, 2021
Mark Xu and Evan Hubinger · 2021
Cited alongside, same era.
Open problems with myopia, 2021
Mark Xu and Evan Hubinger · 2021
Cited alongside, same era.
Janus’ gpt wrangling, 2022
Scott Alexander · 2022
Conditioning, prompts, and fine-tuning, 2022
Adam Jermyn · 2022
Later among the works it cites.
Latent adversarial training, 2022
Adam Jermyn · 2022
Later among the works it cites.
Smoke without fire is scary, 2022
Adam Jermyn · 2022
Later among the works it cites.
Rl with kl penalties is better seen as bayesian inference, 2022
Tomek Korbak and Ethan Perez · 2022
Later among the works it cites.
A minimal viable product for alignment, 2022
Jan Leike · 2022
Later among the works it cites.
Attempts at forwarding speed priors, 2022
James Lucassen and Evan Hubinger · 2022
Later among the works it cites.
Factored cognition, 2022
James Lucassen and Evan Hubinger · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
You can still fetch the coffee today if you’re dead tomorrow, 2022
David A. Dalrymple · 2022
Cited alongside, same era.
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2022
Cited alongside, same era.
Path dependence in ml inductive biases, 2022
Vivek Hebbar and Evan Hubinger · 2022
Cited alongside, same era.
Acceptability verification: A research agenda, 2022
Evan Hubinger · 2022
Cited alongside, same era.
Later among the works it cites.
Strategy for conditioning generative models, 2022
James Lucassen and Evan Hubinger · 2022
Later among the works it cites.
Aligning Language Models to Follow Instructions, January 2022
Ryan Lowe and Jan Leike · 2022
Later among the works it cites.
In-context learning and induction heads, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Proper scoring rules don’t guarantee predicting fixed points, 2022
Johannes Treutlein, Rubi J. Hudson, and Caspar Oesterheld · 2022
Later among the works it cites.
Training goals for large language models, 2022
Johannes Treutlein · 2022
Later among the works it cites.
Verification is not easier than generation in general, 2022
John Wentworth · 2022
Later among the works it cites.
PromptChainer: Chaining Large Language Model Prompts through Visual Programming
Tongshuang Wu, Ellen Jiang, Aaron Donsbach, Jeff Gray, Alejandra Molina, Michael Terry, and Carrie J. Cai · 2022
Later among the works it cites.
Forecasting Future World Events with Neural Networks
Andy Zou, Tristan Xiao, Ryan Jia, Joe Kwon, Mantas Mazeika, Richard Li, Dawn Song, Jacob Steinhardt, Owain Evans, and Dan Hendrycks · 2022
Later among the works it cites.
Underspecification of oracle ai, 2023
Rubi J. Hudson, Adam Jermyn, and Johannes Treutlein · 2023
Closest in time.
Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2023
Closest in time.
Model index for researchers, 2023
OpenAI · 2023
Closest in time.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Closest in time.