Fetching the paper…
Reading the bibliography…
Sparse Autoencoders (SAEs) have shown promise in improving the interpretability of neural network activations, but can learn features that are not features of the input, limiting their effectiveness.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Avrim Blum and Tom Mitchell · 1998
Earlier work this paper cites.
Analyzing the effectiveness and applicability of co-training
Kamal Nigam, Andrew McCallum, Sebastian Thrun, and Tom Mitchell · 2000
Earlier work this paper cites.
Semi-supervised learning by disagreement
Zhi-Hua Zhou and Ming Li · 2005
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2009
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee · 2011
Earlier work this paper cites.
Unsupervised learning of sparse features for scalable audio classification
Mikael Henaff, Kevin Jarrett, Koray Kavukcuoglu, and Yann LeCun · 2011
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey · 2013
Earlier work this paper cites.
A deep learning based approach for traffic data imputation
Yanjie Duan, Yisheng Lv, Wenwen Kang, and Yulong Zhao · 2014
Earlier work this paper cites.
Anomaly detection using autoencoders with nonlinear dimensionality reduction
Mayu Sakurada and Takehisa Yairi · 2014
Earlier work this paper cites.
Detecting anomalous events in videos by learning deep representations of appearance and motion
Dan Xu, Yan Yan, Elisa Ricci, and Nicu Sebe · 2015
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2016
Earlier work this paper cites.
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2016
Earlier work this paper cites.
The temple university hospital eeg data corpus
Iyad Obeid and Joseph Picone · 2016
Cited alongside, same era.
Dual learning for machine translation
Yingce Xia, Di He, Tao Qin, Liwei Wang, Nenghai Yu, Tie-Yan Liu, and Wei-Ying Ma · 2016
Cited alongside, same era.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Cited alongside, same era.
Antti Tarvainen and Harri Valpola · 2017
Cited alongside, same era.
Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu · 2017
Cited alongside, same era.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Later among the works it cites.
Interpreting neural networks through the polytope lens
Sid Black, Lee Sharkey, Leo Grinsztajn, Eric Winsor, Dan Braun, Jacob Merizian, Kip Parker, Carlos Ramón Guevara, Beren Millidge, Gabriel Alfour, and Connor Leahy · 2022
Later among the works it cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Later among the works it cites.
Engineering monosemanticity in toy models
Adam S. Jermyn, Nicholas Schiefer, and Evan Hubinger · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Konrad Żołna, Devansh Arpit, Dendi Suhubdy, and Yoshua Bengio · 2017
Cited alongside, same era.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor W. Tsang, and Masashi Sugiyama · 2018
Cited alongside, same era.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Cited alongside, same era.
Deep co-training for semi-supervised image recognition
Siyuan Qiao, Wei Shen, Zhishuai Zhang, Bo Wang, and Alan Yuille · 2018
Cited alongside, same era.
Denoising sparse autoencoder-based ictal eeg classification
Yang Qiu, Weidong Zhou, Nana Yu, and Peidong Du · 2018
Cited alongside, same era.
Dual student: Breaking the limits of the teacher in semi-supervised learning
Zehao Ke, Di Qiu, Yihong Gong, and Dacheng Tao · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Deep sparse autoencoder and recursive neural network for eeg emotion recognition
Qi Li, Yunqing Liu, Yujie Shang, Qiong Zhang, and Fei Yan · 2022
Later among the works it cites.
Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and Beren Millidge · 2022
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Later among the works it cites.
Rigorously assessing natural language explanations of neurons
Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu, and Christopher Potts · 2023
Later among the works it cites.
Sparse autoencoders find composed features in small toy models
Evan Anders, Clement Neo, Jason Hoelscher-Obermaier, and Jessica Howard · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
Research report: Sparse autoencoders find only 9/180 board state features in othellogpt
Robert Huben · 2024
Closest in time.
Do sparse autoencoders find "true features"?
Demian Till · 2024
Closest in time.