Fetching the paper…
Reading the bibliography…
AI alignment research is the field of study dedicated to ensuring that artificial intelligence (AI) benefits humans.
“Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
“The structure of scientific revolutions”
Thomas Kuhn · 1970
Earlier work this paper cites.
“Let’s correct that small mistake”
Endel Põder · 2010
Earlier work this paper cites.
“Scikit-learn: Machine Learning in Python”
F. Pedregosa et al · 2011
Earlier work this paper cites.
“The elephant in the room: multi-authorship and the assessment of individual researchers”
George Lozano · 2013
Earlier work this paper cites.
“The AI alignment problem: why it is hard, and where to start”
Eliezer Yudkowsky · 2016
Earlier work this paper cites.
“Potential Risks from Advanced Artificial Intelligence”, 2016
Nick Beckstead and Luke Muehlhauser · 2016
Earlier work this paper cites.
“Deep reinforcement learning from human preferences”
Paul Christiano et al · 2017
Earlier work this paper cites.
“Superintelligence”
Nick Bostrom · 2017
Earlier work this paper cites.
“The Rocket Alignment Problem”
Eliezer Yudkowsky · 2018
Earlier work this paper cites.
“AI governance: a research agenda”
Allan Dafoe · 2018
Earlier work this paper cites.
“When will AI exceed human performance? Evidence from AI experts”
Katja Grace et al · 2018
Earlier work this paper cites.
“Alignment Forum”, 2018
Lightcone Infrastructure · 2018
Earlier work this paper cites.
“Announcing AlignmentForum.org Beta”, 2018
Raymond Arnold · 2018
Earlier work this paper cites.
“Umap: Uniform manifold approximation and projection for dimension reduction”
Leland McInnes, John Healy and James Melville · 2018
Earlier work this paper cites.
“Risks from learned optimization in advanced machine learning systems”
Evan Hubinger et al · 2019
Earlier work this paper cites.
“What failure looks like”
Paul Christiano · 2019
Earlier work this paper cites.
“The alignment problem: Machine learning and human values”
Brian Christian · 2020
Earlier work this paper cites.
“Artificial intelligence, values, and alignment”
Iason Gabriel · 2020
Cited alongside, same era.
“The precipice: Existential risk and the future of humanity”
Toby Ord · 2020
Cited alongside, same era.
“Current work in AI alignment”, 2020
Paul Christiano · 2020
Cited alongside, same era.
“Draft report on AI timelines”, 2020
Ayeja Cotra · 2020
Cited alongside, same era.
“AI Safety Papers”, 2020
Jess Riedel and Angelica Deibel · 2020
Cited alongside, same era.
“Toward trustworthy AI development: mechanisms for supporting verifiable claims”
Miles Brundage et al · 2020
Cited alongside, same era.
“Elicit: The AI research assistant”, 2021
Ought · 2021
Later among the works it cites.
“On the opportunities and risks of foundation models”
Rishi Bommasani et al · 2021
Later among the works it cites.
“Evaluating large language models trained on code”
Mark Chen et al · 2021
Later among the works it cites.
“A guide for many authors: Writing manuscripts in large collaborations”
Hannah Moshontz, Charles Ebersole, Sara Weston and Richard Klein · 2021
Later among the works it cites.
“Ethical and social risks of harm from Language Models”
Laura Weidinger et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arman Cohan et al · 2020
Cited alongside, same era.
“Some AI research areas and their relevance to existential safety”, 2020
Andrew Critch · 2020
Cited alongside, same era.
“Human-compatible artificial intelligence”
Stuart Russell · 2021
Cited alongside, same era.
“Alignment of language agents”
Zachary Kenton et al · 2021
Cited alongside, same era.
“Cooperative AI: machines must learn to find common ground”
Allan Dafoe et al · 2021
Cited alongside, same era.
“A General Language Assistant as a Laboratory for Alignment”
Amanda Askell et al · 2021
Cited alongside, same era.
Cornell University · 2021
Later among the works it cites.
“2021 AI Alignment Literature Review and Charity Comparison”, 2021
Larks · 2021
Later among the works it cites.
“seaborn: statistical data visualization”
Michael. Waskom · 2021
Later among the works it cites.
“Parametric UMAP Embeddings for Representation and Semisupervised Learning”
Tim Sainburg, Leland McInnes and Timothy Gentner · 2021
Later among the works it cites.
“Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX”, https://github.com/kingoflolz/mesh-transformer-jax , 2021
Ben Wang · 2021
Later among the works it cites.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, https://github.com/kingoflolz/mesh-transformer-jax , 2021
Ben Wang and Aran Komatsuzaki · 2021
Later among the works it cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Closest in time.
“Compute trends across three eras of machine learning”
Jaime Sevilla et al · 2022
Closest in time.
“Potential Risks from Advanced Artificial Intelligence”, 2022
FTX Foundation · 2022
Closest in time.
“How I failed to form views on AI safety”
Ada-Maaria Hyvärinen · 2022
Closest in time.
“Transcripts of interviews with AI researchers”, 2022
Vael Gates · 2022
Closest in time.
“AI safety resources”, 2022
Victoria Krakovna · 2022
Closest in time.
“Gini coefficient — Wikipedia, The Free Encyclopedia” [Online; accessed 27-May-2022], 2022
Wikipedia contributors · 2022
Closest in time.