Fetching the paper…
Reading the bibliography…
Central to the success of artificial neural networks is their ability to generalize.
Meta-dataset: A dataset of datasets for learning to learn from few examples
Triantafillou, E., Zhu, T., Dumoulin, V., Lamblin, P., Evci, U., Xu, K., Goroshin, R., Gelada, C., Swersky, K., Manzagol, P.-A., et al. (2019) · 1903
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
Saxton, D., Grefenstette, E., Hill, F., and Kohli, P. (2019) · 1904
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. (2019) · 1911
Earlier work this paper cites.
Measuring compositional generalization: A comprehensive method on realistic data
Keysers, D., Schärli, N., Scales, N., Buisman, H., Furrer, D., Kashubin, S., Momchev, N., Sinopalnikov, D., Stafiniak, L., Tihon, T., et al. (2019) · 1912
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A. and Pylyshyn, Z. W. (1988) · 1988
Earlier work this paper cites.
Solving raven’s progressive matrices with neural networks
Zhuo, T. and Kankanhalli, M. (2020) · 2002
Earlier work this paper cites.
Learning compositional rules via neural program synthesis
Nye, M. I., Solar-Lezama, A., Tenenbaum, J. B., and Lake, B. M. (2020) · 2003
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2005
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C. (2020) · 2008
Earlier work this paper cites.
Learning explanations that are hard to vary
Parascandolo, G., Neitz, A., Orvieto, A., Gresele, L., and Schölkopf, B. (2020) · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
One shot learning of simple visual concepts
Lake, B., Salakhutdinov, R., Gross, J., and Tenenbaum, J. (2011) · 2011
Earlier work this paper cites.
The bilingual brain: Flexibility and control in the human cortex
Buchweitz, A. and Prat, C. (2013) · 2013
Earlier work this paper cites.
Indirection and symbol-like processing in the prefrontal cortex and basal ganglia
Kriete, T., Noelle, D. C., Cohen, J. D., and O’Reilly, R. C. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Analysis of boolean functions
O’Donnell, R. (2014) · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
Hierarchical error representation: a computational model of anterior cingulate and dorsolateral prefrontal cortex
Alexander, W. H. and Brown, J. W. (2015) · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
A model for structured information representation in neural networks
Müller, M. G., Papadimitriou, C. H., Maass, W., and Legenstein, R. (2016) · 2016
Cited alongside, same era.
Making neural programming architectures generalize via recursion
Cai, J., Shin, R., and Song, D. (2017) · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R. (2017) · 2017
Cited alongside, same era.
Indirection explains flexible tuning of neurons in prefrontal cortex
Noelle, D. C. (2017) · 2017
Cited alongside, same era.
Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition
Winkler, J. K., Fink, C., Toberer, F., Enk, A., Deinlein, T., Hofmann-Wellenhof, R., Thomas, L., Lallas, A., Blum, A., Stolz, W., et al. (2019) · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J. (2019) · 2019
Later among the works it cites.
Compositional generalization via neural-symbolic stack machines
Chen, X., Liang, C., Yu, A. W., Song, D., and Zhou, D. (2020) · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020) · 2020
Later among the works it cites.
Streaming object detection for 3-d point clouds
Han, W., Zhang, Z., Caine, B., Yang, B., Sprunk, C., Alsharif, O., Ngiam, J., Vasudevan, V., Shlens, J., and Chen, Z. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oh, J., Singh, S., Lee, H., and Kohli, P. (2017) · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R. S. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Cited alongside, same era.
Measuring abstract reasoning in neural networks
Barrett, D., Hill, F., Santoro, A., Morcos, A., and Lillicrap, T. (2018) · 2018
Cited alongside, same era.
Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate
Belkin, M., Hsu, D., and Mitra, P. (2018) · 2018
Cited alongside, same era.
How thalamic relays might orchestrate supervised deep training and symbolic computation in the brain
Hayworth, K. J. and Marblestone, A. H. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Measuring compositional generalization: A comprehensive method on realistic data
Keysers, D., Schärli, N., Scales, N., Buisman, H., Furrer, D., Kashubin, S., Momchev, N., Sinopalnikov, D., Stafiniak, L., Tihon, T., Tsarkov, D., Wang, X., van Zee, M., and Bousquet, O. (2020) · 2020
Later among the works it cites.
Bongard-logo: A new benchmark for human-level concept learning and reasoning
Nie, W., Yu, Z., Mao, L., Patel, A. B., Zhu, Y., and Anandkumar, A. (2020) · 2020
Later among the works it cites.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Oakden-Rayner, L., Dunnmon, J., Carneiro, G., and Ré, C. (2020) · 2020
Later among the works it cites.
A benchmark for systematic generalization in grounded language understanding
Ruis, L., Andreas, J., Baroni, M., Bouchacourt, D., and Lake, B. M. (2020) · 2020
Later among the works it cites.
Deep learning needs a prefrontal cortex
Russin, J., O’Reilly, R. C., and Bengio, Y. (2020) · 2020
Later among the works it cites.
Belkin, M. (2021) · 2021
Closest in time.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al. (2021) · 2021
Closest in time.
Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset
Ettinger, S., Cheng, S., Caine, B., Liu, C., Zhao, H., Pradhan, S., Chai, Y., Sapp, B., Qi, C., Zhou, Y., et al. (2021) · 2021
Closest in time.
Highly accurate protein structure prediction with alphafold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al. (2021) · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Closest in time.
Compositional processing emerges in neural networks solving math problems
Russin, J., Fernandez, R., Palangi, H., Rosen, E., Jojic, N., Smolensky, P., and Gao, J. (2021) · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Keysers, D., Uszkoreit, J., Lucic, M., et al. (2021) · 2021
Closest in time.