Fetching the paper…
Reading the bibliography…
Biological systems perceive the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc.
Certain topics in telegraph transmission theory
Nyquist, H · 1928
Earlier work this paper cites.
Gestalt psychology
Köhler, W · 1967
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K · 1980
Earlier work this paper cites.
Spatiotemporal energy models for the perception of motion
Adelson, E. H. and Bergen, J. R · 1985
Earlier work this paper cites.
Distributed hierarchical processing in the primate cerebral cortex
Felleman, D. J. and Essen, D. C. V · 1991
Earlier work this paper cites.
A neurobiological model of visual attention and invariant pattern recognition based on dynamic routing of information
Olshausen, B. A., Anderson, C. H., and Van Essen, D. C · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
A model of neuronal responses in visual area MT
Simoncelli, E. P. and Heeger, D. J · 1998
Earlier work this paper cites.
Competition for consciousness among visual events: The psychophysics of reentrant visual processes
Lollo, V. D., Enns, J. T., and Rensink, R. A · 2000
Earlier work this paper cites.
Normalized cuts and image segmentation
Shi, J. and Malik, J · 2000
Earlier work this paper cites.
Combining top-down and bottom-up segmentation
Borenstein, E., Sharon, E., and Ullman, S · 2004
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2006
Earlier work this paper cites.
Why don’t we see changes? the role of attentional bottlenecks and limited visual memory
Wolfe, J. M., Reinecke, A., and Brawn, P · 2006
Earlier work this paper cites.
Compositional pattern producing networks: A novel abstraction of development
Stanley, K. O · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
High-performance neural networks for visual object classification
Cireşan, D. C., Meier, U., Masci, J., Gambardella, L. M., and Schmidhuber, J · 2011
Earlier work this paper cites.
Principles of Neural Science
Kandel, E., Schwartz, J., Jessell, T., Siegelbaum, S., and Hudspeth, A · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A., and Hinton, G · 2013
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3D convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M · 2015
Earlier work this paper cites.
Memory networks
Weston, J., Chopra, S., and Bordes, A · 2015
Earlier work this paper cites.
3D shapenets: A deep representation for volumetric shapes
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Human pose estimation with iterative error feedback
Carreira, J., Agrawal, P., Fragkiadaki, K., and Malik, J · 2016
Earlier work this paper cites.
Group equivariant convolutional networks
Cohen, T. and Welling, M · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Bottom-up and top-down reasoning with hierarchical rectified gaussians
Hu, P. and Ramanan, D · 2016
Cited alongside, same era.
Geometric deep learning: Going beyond Euclidean data
Bronstein, M., Bruna, J., LeCun, Y., Szlam, A., and Vandergheynst, P · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Cited alongside, same era.
Audio Set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Cited alongside, same era.
Kaiser, L., Gomez, A. N., Shazeer, N., Vaswani, A., Parmar, N., Jones, L., and Uszkoreit, J · 2017
Cited alongside, same era.
PointNet++: Deep hierarchical feature learning on point sets in a metric space
The DeepMind JAX Ecosystem, 2020
Babuschkin, I., Baumli, K., Bell, A., Bhupatiraju, S., Bruce, J., Buchlovsky, P., Budden, D., Cai, T., Clark, A., Danihelka, I., Fantacci, C., Godwin, J., Jones, C., Hennigan, T., Hessel, M., Kapturowski, S., Keck, T., Kemaev, I., King, M., Martens, L., Mikulik, V., Norman, T., Quan, J., Papamakarios, G., Ring, R., Ruiz, F., Sanchez, A., Schneider, R., Sezener, E., Spencer, S., Srinivasan, S., Stokowiec, W., and Viola, F · 2020
Later among the works it cites.
Longformer: The long-document Transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Later among the works it cites.
End-to-end object detection with Transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qi, C. R., Yi, L., Su, H., and Guibas, L. J · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Objects that sound
Arandjelovic, R. and Zisserman, A · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Learning SO(3) equivariant representations with spherical CNNs
Esteves, C., Allen-Blanchette, C., Makadia, A., and Daniilidis, K · 2018
Cited alongside, same era.
Audio set classification with attention model: A probabilistic perspective
Kong, Q., Xu, Y., Wang, W., and Plumbley, M. D · 2018
Cited alongside, same era.
Image Transformer
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., and Tran, D · 2018
Cited alongside, same era.
Large scale audiovisual learning of sounds with weakly labeled data
Fayek, H. M. and Kumar, A · 2020
Later among the works it cites.
Guo, M.-H., Cai, J.-X., Liu, Z.-N., Mu, T.-J., Martin, R. R., and Hu, S.-M · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Transformers are RNNs: Fast autoregressive Transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Reformer: The efficient Transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Later among the works it cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D · 2020
Later among the works it cites.
Albert: A lite BERT for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Later among the works it cites.
Making sense of vision and touch: Learning multimodal representations for contact-rich tasks
Lee, M. A., Zhu, Y., Zachares, P., Tan, M., Srinivasan, K., Savarese, S., Fei-Fei, L., Garg, A., and Bohg, J · 2020
Later among the works it cites.
Context-gated convolution
Lin, X., Ma, L., Liu, W., and Chang, S.-F · 2020
Later among the works it cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Later among the works it cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorth, R., and Ng, R · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Later among the works it cites.
Efficient content-based sparse attention with routing Transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2020
Later among the works it cites.
Fourier features let networks learn high frequency functions in low dimensional domains
Tancik, M., Srinivasan, P. P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J. T., and Ng, R · 2020
Later among the works it cites.
Efficient Transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2020
Later among the works it cites.
Training data-efficient image Transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Later among the works it cites.
Visual Transformers: Token-based image representation and processing for computer vision
Wu, B., Xu, C., Dai, X., Wan, A., Zhang, P., Yan, Z., Tomizuka, M., Gonzalez, J., Keutzer, K., and Vajda, P · 2020
Later among the works it cites.
Audiovisual slowfast networks for video recognition
Xiao, F., Lee, Y. J., Grauman, K., Malik, J., and Feichtenhofer, C · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Later among the works it cites.
Exploring self-attention for image recognition
Zhao, H., Jia, J., and Koltun, V · 2020
Later among the works it cites.
Towards robust image classification using sequential attention models
Zoran, D., Chrzanowski, M., Huang, P.-S., Gowal, S., Mott, A., and Kohli, P · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization
Brock, A., De, S., Smith, S. L., and Simonyan, K · 2021
Closest in time.
Rethinking attention with Performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L., and Weller, A · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Closest in time.
Coordination among neural modules through a shared global workspace
Goyal, A., Didolkar, A., Lamb, A., Badola, K., Ke, N. R., Rahaman, N., Binas, J., Blundell, C., Mozer, M., and Bengio, Y · 2021
Closest in time.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L · 2021
Closest in time.
Bottleneck Transformers for visual recognition
Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A · 2021
Closest in time.
Max-deeplab: End-to-end panoptic segmentation with mask transformers
Wang, H., Zhu, Y., Adam, H., Yuille, A., and Chen, L.-C · 2021
Closest in time.
Nyströmformer: A Nyström-based algorithm for approximating self-attention
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., and Singh, V · 2021
Closest in time.