Fetching the paper…
Reading the bibliography…
Neural networks are trained primarily based on their inputs and outputs, without regard for their internal mechanisms.
Root Mean Square Layer Normalization, October 2019
Biao Zhang and Rico Sennrich · 1910
Earlier work this paper cites.
The meta-pi network: Building distributed knowledge representations for robust multisource pattern recognition
A. Waibel and J. Hampshire II · 1939
Earlier work this paper cites.
The protection of information in computer systems
Jerome H Saltzer and Michael D Schroeder · 1975
Earlier work this paper cites.
Problems of Monetary Management: The UK Experience
C. A. E. Goodhart · 1984
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J Williams and David Zipser · 1989
Earlier work this paper cites.
Learning to coordinate behaviors
Pattie Maes and Rodney A Brooks · 1990
Earlier work this paper cites.
Markov decision processes
Martin L Puterman · 1990
Earlier work this paper cites.
Automatic programming of behavior-based robots using reinforcement learning
Sridhar Mahadevan and Jonathan Connell · 1992
Earlier work this paper cites.
Learning factorial codes by predictability minimization
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Transfer of learning by composing solutions of elemental sequential tasks
Satinder Pal Singh · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Access control: principle and practice
R.S. Sandhu and P. Samarati · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Access control: Policies, models, and mechanisms
Pierangela Samarati and Sabrina Capitani de Vimercati · 2001
Earlier work this paper cites.
The basic ai drives
Stephen M. Omohundro · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Introduction to Semi-Supervised Learning
Xiaojin Zhu, Andrew B. Goldberg, Ronald Brachman, and Thomas Dietterich · 2009
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Modeling skill dependence in probabilistic competence structures
D. de Chiusole and L. Stefanutti · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Censoring representations with an adversary
Harrison Edwards and Amos J. Storkey · 2015
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Yaroslav Ganin and Victor Lempitsky · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai · 2016
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Generalizing skills with semi-supervised reinforcement learning
Chelsea Finn, Tianhe Yu, Justin Fu, P. Abbeel, and Sergey Levine · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario March, and Victor Lempitsky · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning and Transfer of Modulated Locomotor Controllers
Nicolas Heess, Greg Wayne, Yuval Tassa, Timothy Lillicrap, Martin Riedmiller, and David Silver · 2016
Earlier work this paper cites.
Learning modular neural network policies for multi-task and multi-robot transfer
Coline Devin, Abhishek Gupta, Trevor Darrell, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Arun Mallya, Dillon Davis, and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2018
Earlier work this paper cites.
Intriguing properties of randomly weighted networks: Generalizing while learning next to nothing
Amir Rosenfeld and John K. Tsotsos · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Simple, scalable adaptation for neural machine translation
Ankur Bapna and Orhan Firat · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Disentangling disentanglement in variational autoencoders
Emile Mathieu, Tom Rainforth, N Siddharth, and Yee Whye Teh · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Intriguing Properties of Randomly Weighted Networks: Generalizing While Learning Next to Nothing
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models
Peter Henderson, Eric Mitchell, Christopher Manning, Dan Jurafsky, and Chelsea Finn · 2023
Later among the works it cites.
Backpack language models
John Hewitt, John Thickstun, Christopher D. Manning, and Percy Liang · 2023
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2023
Later among the works it cites.
Surgical fine-tuning improves adaptation to distribution shifts
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amir Rosenfeld and John K. Tsotsos · 2019
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2020
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Cited alongside, same era.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Mad-x: An adapter-based framework for multi-task cross-lingual transfer
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder · 2020
Cited alongside, same era.
Yoonho Lee, Annie S Chen, Fahim Tajwar, Ananya Kumar, Huaxiu Yao, Percy Liang, and Chelsea Finn · 2023
Later among the works it cites.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish · 2023
Later among the works it cites.
The hydra effect: Emergent self-repair in language model computations
Tom McGrath, Matthew Rahtz, János Kramár, Vladimir Mikulik, and Shane Legg · 2023
Later among the works it cites.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2023
Later among the works it cites.
Modular deep learning
Jonas Pfeiffer, Sebastian Ruder, Ivan Vulić, and Edoardo Ponti · 2023
Later among the works it cites.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Composing parameter-efficient modules with arithmetic operation
Jinghan Zhang, shiqi chen, Junteng Liu, and Junxian He · 2023
Later among the works it cites.
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, Colin Raffel, Shiyu Chang, Tatsunori Hashimoto, and William Yang Wang · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, Benjamin L. Edelman, Zhaowei Zhang, Mario Günther, Anton Korinek, Jose Hernandez-Orallo, Lewis Hammond, Eric J Bigelow, Alexander Pan, Lauro Langosco, Tomasz Korbak, Heidi Chenyu Zhang, Ruiqi Zhong, Sean O hEigeartaigh, Gabriel Recchia, Giulio Corsi, Alan Chan, Markus Anderljung, Lilian Edwards, Aleksandar Petrov, Christian Schroeder de Witt, Sumeet Ramesh Motwani, Yoshua Bengio, Danqi Chen, Philip Torr, Samuel Albanie, Tegan Maharaj, Jakob Nicolaus Foerster, Florian Tramèr, He He, Atoosa Kasirzadeh, Yejin Choi, and David Krueger · 2024
Closest in time.
John Beverley, David Limbaugh, Eric Merrell, Peter M. Koch, and Barry Smith · 2024
Closest in time.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu · 2024
Closest in time.
Robust unlearning via mechanistic localizations
Phillip Huang Guo, Aaquib Syed, Abhay Sheshadri, Aidan Ewart, and Gintare Karolina Dziugaite · 2024
Closest in time.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models
Chia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen, Chia-Mu Yu, and Chun ying Huang · 2024
Closest in time.
Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning
Tiansheng Huang, Gautam Bhattacharya, Pratik Joshi, Josh Kimball, and Ling Liu · 2024
Closest in time.
delphi: small language models training made easy, 2024
Jett Janiak, Jai Dhyani, Jannik Brinkmann, Gonçalo Paulo, Joshua Wendland, Víctor Abia Alonso, Siwei Li, Phan Anh Duong, and Alice Rigg · 2024
Closest in time.
Less is more: Selective layer finetuning with subtuning, 2024
Gal Kaplun, Andrey Gurevich, Tal Swisa, Mazor David, Shai Shalev-Shwartz, and eran malach · 2024
Closest in time.
Goodhart’s law in reinforcement learning
Jacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer, Charlie Griffin, and Joar Max Viktor Skalse · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al · 2024
Closest in time.
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Chris Liu, Hang Li, Kush R. Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu · 2024
Closest in time.
Unlearn efficient removal of knowledge in large language models, 2024
Tyler Lizzo and Larry Heck · 2024
Closest in time.
Large language models relearn removed concepts, 2024
Michelle Lo, Shay B. Cohen, and Fazl Barez · 2024
Closest in time.
An adversarial perspective on machine unlearning for ai safety, 2024
Jakub Łucki, Boyi Wei, Yangsibo Huang, Peter Henderson, Florian Tramèr, and Javier Rando · 2024
Closest in time.
Improve mathematical reasoning in language models by automated process supervision
Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, et al · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms, 2024
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell · 2024
Closest in time.
Transformer circuit evaluation metrics are not robust
Joseph Miller, Bilal Chughtai, and William Saunders · 2024
Closest in time.
Lottery ticket adaptation: Mitigating destructive interference in LLMs
Ashwinee Panda, Berivan Isik, Xiangyu Qi, Sanmi Koyejo, Tsachy Weissman, and Prateek Mittal · 2024
Closest in time.
The FineWeb datasets: Decanting the web for the finest text data at scale
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Closest in time.
Dissecting language models: Machine unlearning via selective pruning, 2024
Nicholas Pochinkov and Nandi Schoots · 2024
Closest in time.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner · 2024
Closest in time.
Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al · 2024
Closest in time.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer · 2024
Closest in time.
Tamper-resistant safeguards for open-weight llms, 2024
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, and Mantas Mazeika · 2024
Closest in time.
Disentangled representation learning
Xin Wang, Hong Chen, Si’ao Tang, Zihao Wu, and Wenwu Zhu · 2024
Closest in time.
A safety realignment framework via subspace-oriented model fusion for large language models
Xin Yi, Shunfan Zheng, Linlin Wang, Xiaoling Wang, and Liang He · 2024
Closest in time.
Instilling inductive biases with subnetworks, 2024
Enyan Zhang, Michael A. Lepori, and Ellie Pavlick · 2024
Closest in time.