Fetching the paper…
Reading the bibliography…
Though modern neural networks have achieved impressive performance in both vision and language tasks, we know little about the functions that they implement.
Well-read students learn better: The impact of student initialization on knowledge distillation
Turc, I., Chang, M., Lee, K., and Toutanova, K · 1908
Earlier work this paper cites.
The compositionality papers
Fodor, J. A. and Lepore, E · 2002
Earlier work this paper cites.
The algebraic mind: Integrating connectionism and cognitive science
Marcus, G. F · 2003
Earlier work this paper cites.
Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language
Kirby, S., Cornish, H., and Smith, K · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep vs. shallow networks: An approximation theory perspective
Mhaskar, H. N. and Poggio, T · 2016
Earlier work this paper cites.
The logical primitives of thought: Empirical foundations for compositional cognitive models
Piantadosi, S. T., Tenenbaum, J. B., and Goodman, N. D · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Assessing composition in sentence vector representations
Ettinger, A., Elgohary, A., Phillips, C., and Resnik, P · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Lake, B. and Baroni, M · 2018
Earlier work this paper cites.
Learning sparse neural networks through l_0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2018
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
Saxton, D., Grefenstette, E., Hill, F., and Kohli, P · 2018
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Bau, A., Belinkov, Y., Sajjad, H., Durrani, N., Dalvi, F., and Glass, J · 2019
Earlier work this paper cites.
Compositional generalization through meta sequence-to-sequence learning
Lake, B. M · 2019
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Marvin, R. and Linzen, T · 2019
Earlier work this paper cites.
Compositional languages emerge in a neural iterated learning model
Ren, Y., Guo, S., Labeau, M., Cohen, S. B., and Kirby, S · 2019
Earlier work this paper cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J · 2019
Earlier work this paper cites.
Thread: Circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Cited alongside, same era.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
De Cao, N., Schlichtkrull, M. S., Aziz, W., and Titov, I · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Compositionality decomposed: How do neural networks generalise?
Hupkes, D., Dankers, V., Mul, M., and Bruni, E · 2020
Cited alongside, same era.
Cogs: A compositional generalization challenge based on semantic interpretation
Kim, N. and Linzen, T · 2020
Cited alongside, same era.
Concept bottleneck models
Can subnetwork structure be the key to out-of-distribution generalization?
Zhang, D., Ahuja, K., Xu, Y., Wang, Y., and Courville, A · 2021
Later among the works it cites.
Learning to generalize compositionally by transferring across semantic parsing tasks
Zhu, W., Shaw, P., Linzen, T., and Sha, F · 2021
Later among the works it cites.
Interpreting neural networks through the polytope lens
Black, S., Sharkey, L., Grinsztajn, L., Winsor, E., Braun, D., Merizian, J., Parker, K., Guevara, C. R., Millidge, B., Alfour, G., et al · 2022
Later among the works it cites.
The paradox of the compositionality of natural language: A neural machine translation case study
Dankers, V., Bruni, E., and Hupkes, D · 2022
Later among the works it cites.
Sparse interventions in language models with differentiable masking
De Cao, N., Schmid, L., Hupkes, D., and Titov, I · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P · 2020
Cited alongside, same era.
Learning compositional rules via neural program synthesis
Nye, M., Solar-Lezama, A., Tenenbaum, J., and Lake, B. M · 2020
Cited alongside, same era.
What’s hidden in a randomly weighted neural network?
Ramanujan, V., Wortsman, M., Kembhavi, A., Farhadi, A., and Rastegari, M · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y · 2020
Cited alongside, same era.
Winning the lottery with continuous sparsification
Savarese, P., Silva, H., and Maire, M · 2020
Cited alongside, same era.
Iterated learning for emergent systematicity in vqa
Vani, A., Schwarzer, M., Lu, Y., Dhekane, E., and Courville, A · 2020
Cited alongside, same era.
Supermasks in superposition
Wortsman, M., Ramanujan, V., Liu, R., Kembhavi, A., Rastegari, M., Yosinski, J., and Farhadi, A · 2020
Cited alongside, same era.
Later among the works it cites.
Quantifying local specialization in deep neural networks
Hod, S., Filan, D., Casper, S., Critch, A., and Russell, S · 2022
Later among the works it cites.
Kim, N., Linzen, T., and Smolensky, P · 2022
Later among the works it cites.
Tutorial 17: Self-supervised contrastive learning with simclr
Lippe, P · 2022
Later among the works it cites.
Unit testing for concepts in neural networks
Lovering, C. and Pavlick, E · 2022
Later among the works it cites.
Problems and mysteries of the many languages of thought
Mandelbaum, E., Dunham, Y., Feiman, R., Firestone, C., Green, E., Harris, D., Kibbe, M. M., Kurdi, B., Mylopoulos, M., Shepherd, J., et al · 2022
Later among the works it cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Olah, C · 2022
Later among the works it cites.
How deep sparse networks avoid the curse of dimensionality: Efficiently computable functions are compositionally sparse
Poggio, T · 2022
Later among the works it cites.
The best game in town: The re-emergence of the language of thought hypothesis across the cognitive sciences
Quilty-Dunn, J., Porot, N., and Mandelbaum, E · 2022
Later among the works it cites.
Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems
Smolensky, P., McCoy, R., Fernandez, R., Goldrick, M., and Gao, J · 2022
Later among the works it cites.
Causal distillation for language models
Wu, Z., Geiger, A., Rozner, J., Kreiss, E., Lu, H., Icard, T., Potts, C., and Goodman, N · 2022
Later among the works it cites.
A benchmark for compositional visual reasoning
Zerroug, A., Vaishnav, M., Colin, J., Musslick, S., and Serre, T · 2022
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations
Chughtai, B., Chan, L., and Nanda, N · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., et al · 2023
Closest in time.
Dreamcoder: growing generalizable, interpretable knowledge with wake–sleep bayesian program learning
Ellis, K., Wong, L., Nye, M., Sable-Meyer, M., Cary, L., Anaya Pozo, L., Hewitt, L., Solar-Lezama, A., and Tenenbaum, J. B · 2023
Closest in time.
Superposition, memorization, and double descent
Henighan, T., Carter, S., Humne, T., Elhage, N., Lasenby, R., Fort, S., Schiefer, N., and Olah, C · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
Merrill, W., Tsilivis, N., and Shukla, A · 2023
Closest in time.