Fetching the paper…
Reading the bibliography…
Despite the recent success of artificial neural networks on a variety of tasks, we have little knowledge or control over the exact solutions these models implement.
Exploring the origins and prevalence of texture bias in convolutional neural networks
Hermann, K. L. and Kornblith, S · 1911
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D · 1989
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches, 2014
Cho, K., van Merrienboer, B., Bahdanau, D., and Bengio, Y · 2014
Earlier work this paper cites.
The importance of shape in early lexical learning
Landau, B., Smith, L. B., and Jones, S. S · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation, 2016
Kim, Y. and Rush, A. M · 2016
Earlier work this paper cites.
A study and comparison of human and deep learning recognition performance under visual distortions, 2017
Dodge, S. and Karam, L · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks, 2019
Frankle, J. and Carbin, M · 2019
Earlier work this paper cites.
Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W · 2019
Earlier work this paper cites.
Doing more with less: meta-reasoning and meta-learning in humans and machines
Griffiths, T. L., Callaway, F., Chang, M. B., Grant, E., Krueger, P. M., and Lieder, F · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp, 2019
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Compositional generalization through meta sequence-to-sequence learning, 2019
Lake, B. M · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Good-enough compositional data augmentation
Andreas, J · 2020
Cited alongside, same era.
Reminder of the first paper on transfer learning in neural networks, 1976
Bozinovski, S · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Learning task-general representations with generative neuro-symbolic modeling
Feinman, R. and Lake, B. M · 2020
Cited alongside, same era.
Branch specialization
Voss, C., Goh, G., Cammarata, N., Petrov, M., Schubert, L., and Olah, C · 2021
Later among the works it cites.
Graphical clusterability and local specialization in deep neural networks
Casper, S., Hod, S., Filan, D., Wild, C., Critch, A., and Russell, S · 2022
Later among the works it cites.
Pruning for interpretable, feature-preserving circuits in cnns
Hamblin, C., Konkle, T., and Alvarez, G · 2022
Later among the works it cites.
Using natural language and program abstractions to instill human inductive biases in machines
Kumar, S., Correa, C. G., Dasgupta, I., Marjieh, R., Hu, M. Y., Hawkins, R., Cohen, J. D., Narasimhan, K., Griffiths, T., et al · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T. J., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T. B., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, W., Liu, H., and Bowman, S · 2020
Cited alongside, same era.
Does data augmentation improve generalization in nlp?
Jha, R., Lovering, C., and Pavlick, E · 2020
Cited alongside, same era.
More bang for your buck: Natural perturbation for robust question answering
Khashabi, D., Khot, T., and Sabharwal, A · 2020
Cited alongside, same era.
Meta-learning of structured task distributions in humans and machines
Kumar, S., Dasgupta, I., Cohen, J., Daw, N., and Griffiths, T · 2020
Cited alongside, same era.
Universal linguistic inductive biases via meta-learning, 2020
McCoy, R. T., Grant, E., Smolensky, P., Griffiths, T. L., and Linzen, T · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R · 2020
Cited alongside, same era.
What is being transferred in transfer learning?
Neyshabur, B., Sedghi, H., and Zhang, C · 2020
Cited alongside, same era.
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision, 2022
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Later among the works it cites.
Improving systematic generalization through modularity and augmentation
Ruis, L. and Lake, B. M · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations
Chughtai, B., Chan, L., and Nanda, N · 2023
Closest in time.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model, 2023
Hanna, M., Liu, O., and Variengien, A · 2023
Closest in time.
Segment anything, 2023
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., and Girshick, R · 2023
Closest in time.
Break it down: Evidence for structural compositionality in neural networks, 2023
Lepori, M. A., Serre, T., and Pavlick, E · 2023
Closest in time.
Embers of autoregression: Understanding large language models through the problem they are trained to solve, 2023
McCoy, R. T., Yao, S., Friedman, D., Hardy, M., and Griffiths, T. L · 2023
Closest in time.
Language models implement simple word2vec-style vector arithmetic, 2023
Merullo, J., Eickhoff, C., and Pavlick, E · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Closest in time.