Fetching the paper…
Reading the bibliography…
When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization.
Support vector method for novelty detection
Bernhard Schölkopf, Robert C Williamson, Alex Smola, John Shawe-Taylor, and John Platt · 1999
Earlier work this paper cites.
Anomaly detection: A survey
Varun Chandola, Arindam Banerjee, and Vipin Kumar · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
A survey on multi-view learning
Chang Xu, Dacheng Tao, and Chao Xu · 2013
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, Ian Molloy, and Biplav Srivastava · 2018
Earlier work this paper cites.
Deep learning for anomaly detection: A survey
Raghavendra Chalapathy and Sanjay Chawla · 2019
Earlier work this paper cites.
Specification gaming: the flip side of ai ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Earlier work this paper cites.
Neural unsupervised domain adaptation in nlp—a survey
Alan Ramponi and Barbara Plank · 2020
Earlier work this paper cites.
An empirical study on robustness to spurious correlations using pre-trained language models
Lifu Tu, Garima Lalwani, Spandana Gella, and He He · 2020
Earlier work this paper cites.
Consequences of misaligned ai
Simon Zhuang and Dylan Hadfield-Menell · 2020
Cited alongside, same era.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Cited alongside, same era.
Arc’s first technical report: Eliciting latent knowledge
Paul Christiano, Mark Xu, and Ajeya Cotra · 2021
Cited alongside, same era.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Tom Everitt, Marcus Hutter, Ramana Kumar, and Victoria Krakovna · 2021
Cited alongside, same era.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al · 2021
Cited alongside, same era.
Backdoor learning: A survey
Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2022
Later among the works it cites.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2022
Later among the works it cites.
Formal mathematics statement curriculum learning
Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever · 2022
Later among the works it cites.
Solving math word problems with process-and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Later among the works it cites.
Backdoorbench: A comprehensive benchmark of backdoor learning
Baoyuan Wu, Hongrui Chen, Mingda Zhang, Zihao Zhu, Shaokui Wei, Danni Yuan, and Chao Shen · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How ”discovering latent knowledge in language models without supervision” fits into a broader alignment scheme
Collin Burns · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
Mechanistic anomaly detection and elk
Paul Christiano · 2022
Cited alongside, same era.
Formalizing the presumption of independence
Paul Christiano, Eric Neyman, and Mark Xu · 2022
Cited alongside, same era.
Elk prize results
Paul Christiano and Mark Xu · 2022
Cited alongside, same era.
Last layer re-training is sufficient for robustness to spurious correlations
Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson · 2022
Cited alongside, same era.
Later among the works it cites.
Leace: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
Gpt-4 technical report
R OpenAI · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team et al · 2023
Closest in time.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Closest in time.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al · 2023
Closest in time.