Fetching the paper…
Reading the bibliography…
We aim to better understand the emergence of `situational awareness' in large language models (LLMs).
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Improving neural machine translation models with monolingual data, 2016
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Further advantages of data augmentation on convolutional neural networks
Alex Hernández-García and Peter König · 2018
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Data manipulation: Towards effective instance learning for neural dialogue generation via learning to augment and reweight
Hengyi Cai, Hongshen Chen, Yonghao Song, Cheng Zhang, Xiaofang Zhao, and Dawei Yin · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2020
Earlier work this paper cites.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment, 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Earlier work this paper cites.
The scaling hypothesis, 2021
Gwern Branwen · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Earlier work this paper cites.
Truthful ai: Developing and governing ai that does not lie
Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders · 2021
Earlier work this paper cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2021
Earlier work this paper cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Cited alongside, same era.
Is power-seeking ai an existential risk?
Joseph Carlsmith · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Cited alongside, same era.
Without specific countermeasures, the easiest path to transformative ai likely leads to ai takeover
Tasra: A taxonomy and analysis of societal-scale risks from ai
Andrew Critch and Stuart Russell · 2023
Closest in time.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Closest in time.
The pile
EleutherAI · 2023
Closest in time.
Studying large language model generalization with influence functions, 2023
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
Closest in time.
Methods for measuring, updating, and visualizing factual beliefs in language models
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ajeya Cotra · 2022
Cited alongside, same era.
Self-Locating Beliefs
Andy Egan and Michael G. Titelbaum · 2022
Cited alongside, same era.
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Nova DasSarma, T. J. Henighan, Andy Jones, Nicholas Joseph, John Kernion, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Nelson Elhage, Sheer El Showk, Stanislav Fort, Zac Hatfield-Dodds, Scott Johnston, Shauna Kravec, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Dario Amodei, Tom B. Brown, Jared Kaplan, Sam McCandlish, Christopher Olah, and Jack Clark · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
How likely is deceptive alignment?
Evan Hubringer · 2022
Cited alongside, same era.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2022
Cited alongside, same era.
An overview of catastrophic ai risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Closest in time.
Out-of-context meta-learning in large language models | openreview
Dmitrii Krasheninnikov, Egor Krasheninnikov, and David Krueger · 2023
Closest in time.
Introducing superalignment
OpenAI · 2023
Closest in time.
Openai api
OpenAI · 2023
Closest in time.
Early situational awareness and its implications: A story
Jacob Pfau · 2023
Closest in time.
Reinforcement learning from human feedback: Progress and challenges
John Schulman · 2023
Closest in time.
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al · 2023
Closest in time.
What will gpt-2030 look like?
Jacob Steinhardt · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Honesty is the best policy: Defining and mitigating ai deception
Francis Rhys Ward, Tom Everitt, Francesco Belardinelli, and Francesca Toni · 2023
Closest in time.
In-context instruction learning, 2023
Seonghyeon Ye, Hyeonbin Hwang, Sohee Yang, Hyeongu Yun, Yireun Kim, and Minjoon Seo · 2023
Closest in time.
Contextual augmentation: Data augmentation by words with paradigmatic relations
Sosuke Kobayashi · 2072
Closest in time.