Fetching the paper…
Reading the bibliography…
Representation learning, and interpreting learned representations, are key areas of focus in machine learning and neuroscience.
Analysis of orientation bias in cat retina
WR Levick and LN Thibos · 1982
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
David Marr · 1982
Earlier work this paper cites.
Neural representation and neural computation
Patricia Smith Churchland and Terrence J Sejnowski · 1990
Earlier work this paper cites.
Energy as a constraint on the coding and processing of sensory information
Simon B Laughlin · 2001
Earlier work this paper cites.
Colorbrewer. org: an online tool for selecting colour schemes for maps
Mark Harrower and Cynthia A Brewer · 2003
Earlier work this paper cites.
Representational similarity analysis-connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter A Bandettini · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
Optogenetic stimulation of a hippocampal engram activates fear memory recall
Xu Liu, Steve Ramirez, Petti T Pang, Corey B Puryear, Arvind Govindarajan, Karl Deisseroth, and Susumu Tonegawa · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Zipf’s word frequency law in natural language: A critical review and future directions
Steven T Piantadosi · 2014
Earlier work this paper cites.
Performance-optimized hierarchical models predict neural responses in higher visual cortex
Daniel LK Yamins, Ha Hong, Charles F Cadieu, Ethan A Solomon, Darren Seibert, and James J DiCarlo · 2014
Earlier work this paper cites.
Optogenetics: 10 years of microbial opsins in neuroscience
Karl Deisseroth · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Using goal-driven deep learning models to understand sensory cortex
Daniel LK Yamins and James J DiCarlo · 2016
Earlier work this paper cites.
Could a neuroscientist understand a microprocessor?
Eric Jonas and Konrad Paul Kording · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel · 2018
Earlier work this paper cites.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
Andrew K Lampinen and Surya Ganguli · 2018
Earlier work this paper cites.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Earlier work this paper cites.
Brain-score: Which artificial neural network for object recognition is most brain-like?
Martin Schrimpf, Jonas Kubilius, Ha Hong, Najib J Majaj, Rishi Rajalingham, Elias B Issa, Kohitij Kar, Pouya Bashivan, Jonathan Prescott-Roy, Franziska Geiger, et al · 2018
Earlier work this paper cites.
Representation in cognitive science
Nicholas Shea · 2018
Earlier work this paper cites.
On the learning dynamics of deep neural networks
Remi Tachet, Mohammad Pezeshki, Samira Shabanian, Aaron Courville, and Yoshua Bengio · 2018
Earlier work this paper cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q Camargo, and Ard A Louis · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Earlier work this paper cites.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2019
Earlier work this paper cites.
What do compressed deep neural networks forget?
Sara Hooker, Aaron Courville, Gregory Clark, Yann Dauphin, and Andrea Frome · 2019
Earlier work this paper cites.
Sarthak Jain and Byron C Wallace · 2019
Earlier work this paper cites.
SGD on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Peeling the onion of brain representations
Nikolaus Kriegeskorte and Jörn Diedrichsen · 2019
Earlier work this paper cites.
Universality and individuality in neural dynamics across large populations of recurrent networks
Niru Maheswaranathan, Alex Williams, Matthew Golub, Surya Ganguli, and David Sussillo · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen · 2019
Earlier work this paper cites.
Single-trial neural dynamics are dominated by richly varied movements
Simon Musall, Matthew T Kaufman, Ashley L Juavinett, Steven Gluf, and Anne K Churchland · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox · 2019
Cited alongside, same era.
Are disentangled representations helpful for abstract visual reasoning?
Sjoerd Van Steenkiste, Francesco Locatello, Jürgen Schmidhuber, and Olivier Bachem · 2019
Cited alongside, same era.
Welcome to the tidyverse
Hadley Wickham, Mara Averick, Jennifer Bryan, Winston Chang, Lucy D’Agostino McGowan, Romain François, Garrett Grolemund, Alex Hayes, Lionel Henry, Jim Hester, et al · 2019
Cited alongside, same era.
Problems with cosine as a measure of embedding similarity for high frequency words
Kaitlyn Zhou, Kawin Ethayarajh, Dallas Card, and Dan Jurafsky · 2022
Later among the works it cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarah Wiegreffe and Yuval Pinter · 2019
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani, Andrew M Saxe, and Haim Sompolinsky · 2020
Cited alongside, same era.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Cited alongside, same era.
What shapes feature representations? exploring datasets, architectures, and training
Katherine Hermann and Andrew Lampinen · 2020
Cited alongside, same era.
The origins and prevalence of texture bias in convolutional neural networks
Katherine Hermann, Ting Chen, and Simon Kornblith · 2020
Cited alongside, same era.
Rnns can generate bounded hierarchical languages with optimal memory
John Hewitt, Michael Hahn, Surya Ganguli, Percy Liang, and Christopher D Manning · 2020
Cited alongside, same era.
Does syntax need to grow on trees? sources of hierarchical inductive bias in sequence-to-sequence networks
R Thomas McCoy, Robert Frank, and Tal Linzen · 2020
Cited alongside, same era.
What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines?
Colin Conwell, Jacob S Prince, Kendrick N Kay, George A Alvarez, and Talia Konkle · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2023
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2023
Later among the works it cites.
On the foundations of shortcut learning
Katherine L Hermann, Hossein Mobahi, Thomas Fel, and Michael C Mozer · 2023
Later among the works it cites.
Soft matching distance: A metric on neural representations that captures single-neuron tuning
Meenakshi Khosla and Alex H Williams · 2023
Later among the works it cites.
Representations and computations in transformers that support generalization on structured tasks
Yuxuan Li and James McClelland · 2023
Later among the works it cites.
Benign oscillation of stochastic gradient descent with large learning rate
Miao Lu, Beining Wu, Xiaodong Yang, and Difan Zou · 2023
Later among the works it cites.
Mechanistic mode connectivity
Ekdeep Singh Lubana, Eric J Bigelow, Robert P Dick, David Krueger, and Hidenori Tanaka · 2023
Later among the works it cites.
Circuit component reuse across tasks in transformer language models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2023
Later among the works it cites.
Feature emergence via margin maximization: case studies in algebraic tasks
Depen Morwani, Benjamin L Edelman, Costin-Andrei Oncescu, Rosie Zhao, and Sham Kakade · 2023
Later among the works it cites.
Grokking of hierarchical structure in vanilla transformers
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D Manning · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Getting aligned on representational alignment
Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C Love, Erin Grant, Jascha Achterberg, Joshua B Tenenbaum, et al · 2023
Later among the works it cites.
Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions
Greta Tuckute, Jenelle Feather, Dana Boebinger, and Josh H McDermott · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
(un) interpretability of transformers: a case study with dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski · 2023
Later among the works it cites.
Benign overfitting and grokking in relu networks for xor cluster data
Zhiwei Xu, Yutong Wang, Spencer Frei, Gal Vardi, and Wei Hu · 2023
Later among the works it cites.
Which features are learnt by contrastive learning? On the role of simplicity bias in class collapse and feature suppression
Yihao Xue, Siddharth Joshi, Eric Gan, Pin-Yu Chen, and Baharan Mirzasoleiman · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Task structure and nonlinearity jointly determine learned representational geometry
Matteo Alleman, Jack W Lindsey, and Stefano Fusi · 2024
Closest in time.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, et al · 2024
Closest in time.
Understanding visual feature reliance through the lens of complexity
Thomas Fel, Louis Bethune, Andrew Kyle Lampinen, Thomas Serre, and Katherine Hermann · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah Goodman · 2024
Closest in time.
Artificial neural network language models predict human brain responses to language even after a developmentally realistic amount of training
Eghbal A Hosseini, Martin Schrimpf, Yian Zhang, Samuel Bowman, Noga Zaslavsky, and Evelina Fedorenko · 2024
Closest in time.
Asymmetric stimulus representations bias visual perceptual learning
Pooya Laamerad, Asmara Awada, Christopher C Pack, and Shahab Bakhtiari · 2024
Closest in time.
Aaron Mueller · 2024
Closest in time.
Improving neural network representations using human similarity judgments
Lukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A Vandermeulen, Katherine Hermann, Andrew Lampinen, and Simon Kornblith · 2024
Closest in time.
Linear explanations for individual neurons
Tuomas Oikarinen and Tsui-Wei Weng · 2024
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2024
Closest in time.
Complexity matters: Dynamics of feature learning in the presence of spurious correlations
GuanWen Qiu, Da Kuang, and Surbhi Goel · 2024
Closest in time.
Decomposing and editing predictions by modeling model computation
Harshay Shah, Andrew Ilyas, and Aleksander Madry · 2024
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman · 2024
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2024
Closest in time.