Fetching the paper…
Reading the bibliography…
As an essential ingredient of modern deep learning, attention mechanism, especially self-attention, plays a vital role in the global correlation discovery.
Generalized low rank models
Madeleine Udell, Corinne Horn, Reza Zadeh, and Stephen Boyd · 1935
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos et al · 1990
Earlier work this paper cites.
Quantization
R. M. Gray and D. L. Neuhoff · 1998
Earlier work this paper cites.
Learning the parts of objects by non-negative matrix factorization
Daniel D Lee and H Sebastian Seung · 1999
Earlier work this paper cites.
Concept decompositions for large sparse text data using clustering
Inderjit S. Dhillon and Dharmendra S. Modha · 2001
Earlier work this paper cites.
Algorithms for non-negative matrix factorization
Daniel D. Lee and H. Sebastian Seung · 2001
Earlier work this paper cites.
On spectral clustering: Analysis and an algorithm
Andrew Ng, Michael Jordan, and Yair Weiss · 2002
Earlier work this paper cites.
Laplacian eigenmaps for dimensionality reduction and data representation
Mikhail Belkin and Partha Niyogi · 2003
Earlier work this paper cites.
Think globally, fit locally: unsupervised learning of low dimensional manifolds
Lawrence K Saul and Sam T Roweis · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Tensor decompositions and applications
Tamara G. Kolda and Brett W. Bader · 2009
Earlier work this paper cites.
The augmented lagrange multiplier method for exact recovery of corrupted low-rank matrices
Zhouchen Lin, Minming Chen, Leqin Wu, and Yi Ma · 2009
Earlier work this paper cites.
Robust principal component analysis: Exact recovery of corrupted low-rank matrices via convex optimization
John Wright, Arvind Ganesh, Shankar Rao, Yigang Peng, and Yi Ma · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
Online learning for matrix factorization and sparse coding
Julien Mairal, Francis Bach, Jean Ponce, and Guillermo Sapiro · 2010
Earlier work this paper cites.
Tilt: Transform invariant low-rank textures
Zhengdong Zhang, Arvind Ganesh, Xiao Liang, and Yi Ma · 2012
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
T. Sainath, Brian Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran · 2013
Earlier work this paper cites.
Generalized nonconvex nonsmooth low-rank minimization
Canyi Lu, Jinhui Tang, Shuicheng Yan, and Zhouchen Lin · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, et al · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Group equivariant convolutional networks
Taco S Cohen and Max Welling · 2016
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
S. M. Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, koray kavukcuoglu, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
Deep dictionary learning
Snigdha Tariyal, A. Majumdar, R. Singh, and Mayank Vatsa · 2016
Earlier work this paper cites.
Optnet: Differentiable optimization as a layer in neural networks
Brandon Amos and J. Z. Kolter · 2017
Earlier work this paper cites.
Yoshua Bengio · 2017
Cited alongside, same era.
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei · 2017
Cited alongside, same era.
Neural expectation maximization
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2017
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Cited alongside, same era.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Cited alongside, same era.
Low rank factorization for compact multi-head self-attention
Sneha Mehta, H. Rangwala, and N. Ramakrishnan · 2019
Later among the works it cites.
Stand-alone self-attention in vision models
Niki Parmar, Prajit Ramachandran, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Meta-learning with implicit gradients
A. Rajeswaran, Chelsea Finn, S. Kakade, and Sergey Levine · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A Smith · 2019
Later among the works it cites.
Truncated back-propagation for bilevel optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Pyramid scene parsing network
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia · 2017
Cited alongside, same era.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Cited alongside, same era.
Matrix capsules with EM routing
Geoffrey E Hinton, Sara Sabour, and Nicholas Frosst · 2018
Cited alongside, same era.
On the generalization of equivariance and convolution in neural networks to the action of compact groups
Risi Kondor and Shubhendu Trivedi · 2018
Cited alongside, same era.
Sequential attend, infer, repeat: Generative modelling of moving objects
Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner · 2018
Cited alongside, same era.
Amirreza Shaban, Ching-An Cheng, Nathan Hatch, and Byron Boots · 2019
Later among the works it cites.
General E(2)-equivariant steerable CNNs
Maurice Weiler and Gabriele Cesa · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Later among the works it cites.
Logan: Latent optimisation for generative adversarial networks
Yan Wu, Jeff Donahue, David Balduzzi, Karen Simonyan, and Timothy P. Lillicrap · 2019
Later among the works it cites.
Ada-tucker: Compressing deep neural networks via adaptive dimension adjustment tucker decomposition
Zhisheng Zhong, Fangyin Wei, Zhouchen Lin, and Chao Zhang · 2019
Later among the works it cites.
Asymmetric non-local neural networks for semantic segmentation
Zhen Zhu, Mengde Xu, Song Bai, Tengteng Huang, and Xiang Bai · 2019
Later among the works it cites.
Tesa: Tensor element self-attention via matricization
Francesca Babiloni, Ioannis Marras, Gregory Slabaugh, and Stefanos Zafeiriou · 2020
Later among the works it cites.
Your local gan: Designing two dimensional local attention mechanisms for generative models
Giannis Daras, Augustus Odena, Han Zhang, and A. Dimakis · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, M. Dehghani, Matthias Minderer, G. Heigold, S. Gelly, Jakob Uszkoreit, and N. Houlsby · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Franccois Fleuret · 2020
Later among the works it cites.
Spatial pyramid based graph reasoning for semantic segmentation
Xia Li, Y. Yang, Qijie Zhao, Tian cheng Shen, Zhouchen Lin, and Hong-Cheu Liu · 2020
Later among the works it cites.
Optimizing Millions of Hyperparameters by Implicit Differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Later among the works it cites.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, et al · 2020
Later among the works it cites.
Attentive group equivariant convolutional networks
David Romero, Erik Bekkers, Jakub Tomczak, and Mark Hoogendoorn · 2020
Later among the works it cites.
PDO-eConvs: Partial differential operator based equivariant convolutions
Zhengyang Shen, Lingshen He, Zhouchen Lin, and Jinwen Ma · 2020
Later among the works it cites.
Kyungwoo Song, Yohan Jung, Dong-Jun Kim, and I. Moon · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Object-contextual representations for semantic segmentation
Yuhui Yuan, Xilin Chen, and Jingdong Wang · 2020
Later among the works it cites.
Consistency regularization for generative adversarial networks
Han Zhang, Zizhao Zhang, Augustus Odena, and Honglak Lee · 2020
Later among the works it cites.
Point transformer, 2020
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun · 2020
Later among the works it cites.
Squeeze-and-attention networks for semantic segmentation
Zilong Zhong, Zhong Qiu Lin, Rene Bidart, Xiaodan Hu, Ibrahim Ben Daya, Zhifeng Li, Wei-Shi Zheng, Jonathan Li, and Alexander Wong · 2020
Later among the works it cites.
Vivit: A video vision transformer
A. Arnab, M. Dehghani, G. Heigold, Chen Sun, Mario Lucic, and C. Schmid · 2021
Closest in time.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Closest in time.
Vision transformers with patch diversification, 2021
Chengyue Gong, Dilin Wang, Meng Li, Vikas Chandra, and Qiang Liu · 2021
Closest in time.
Neural architecture search without training
Joe Mellor, Jack Turner, Amos Storkey, and Elliot J Crowley · 2021
Closest in time.
Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing
Vishal Monga, Yuelong Li, and Yonina C Eldar · 2021
Closest in time.
Daniel Neimark, O. Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.
Group equivariant stand-alone self-attention for vision
David W Romero and Jean-Baptiste Cordonnier · 2021
Closest in time.