Fetching the paper…
Reading the bibliography…
\textit{Attention} computes the dependency between representations, and it encourages the model to focus on the important selective features.
Geodesic distance estimation with spherelets
Li, D.; and Dunson, D. B. 2019 · 1907
Earlier work this paper cites.
Distribution functions in dimensions and their margins
Sklar, M. 1959 · 1959
Earlier work this paper cites.
II: Fourier Analysis, Self-Adjointness , volume 2
Reed, M.; and Simon, B. 1975 · 1975
Earlier work this paper cites.
Correlation Theory of Stationary and Related Random Functions
Yaglom, A. M. 1987 · 1987
Earlier work this paper cites.
Link-based classification
Geetor, L.; and Lu, Q. 2003 · 2003
Earlier work this paper cites.
Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X.; Davis, J.; Sarlos, T.; Belanger, D.; Colwell, L.; and Weller, A. 2020a · 2006
Earlier work this paper cites.
Approximate Inference for Spectral Mixture Kernel
Jung, Y.; Song, K.; and Park, J. 2020 · 2006
Earlier work this paper cites.
An introduction to copulas
Nelsen, R. B. 2007 · 2007
Earlier work this paper cites.
Visualizing data using t-SNE
Maaten, L. v. d.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A.; and Recht, B. 2008 · 2008
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlos, T.; Hawkins, P.; Davis, J.; Mohiuddin, A.; Kaiser, L.; et al. 2020b · 2009
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Maas, A. L.; Hannun, A. Y.; and Ng, A. Y. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Convolutional Neural Networks for Sentence Classification
Kim, Y. 2014 · 2014
Cited alongside, same era.
Auto-Encoding Variational Bayes
Kingma, D. P.; and Welling, M. 2014 · 2014
Cited alongside, same era.
Deepwalk: Online learning of social representations
Perozzi, B.; Al-Rfou, R.; and Skiena, S. 2014 · 2014
Cited alongside, same era.
Copula variational inference
Tran, D.; Blei, D.; and Airoldi, E. M. 2015 · 2015
Cited alongside, same era.
Semi-Supervised Classification with Graph Convolutional Networks
Kipf, T. N.; and Welling, M. 2017 · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Later among the works it cites.
Metrics for deep generative models
Chen, N.; Klushyn, A.; Kurle, R.; Jiang, X.; Bayer, J.; and Smagt, P. 2018 · 2018
Later among the works it cites.
Latent alignment and variational attention
Deng, Y.; Kim, Y.; Chiu, J.; Guo, D.; and Rush, A. 2018 · 2018
Later among the works it cites.
Classical Structured Prediction Losses for Sequence to Sequence Learning
Edunov, S.; Ott, M.; Auli, M.; Grangier, D.; and Ranzato, M. 2018 · 2018
Later among the works it cites.
Spatial mapping with Gaussian processes and nonstationary Fourier features
Ton, J.-F.; Flaxman, S.; Sejdinovic, D.; and Bhatt, S. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convolutional neural networks on graphs with fast localized spectral filtering
Defferrard, M.; Bresson, X.; and Vandergheynst, P. 2016 · 2016
Cited alongside, same era.
Implicit Kernel Learning
Li, C.-L.; Chang, W.-C.; Mroueh, Y.; Yang, Y.; and Poczos, B. 2019 · 2016
Cited alongside, same era.
Sequence-to-Sequence Learning as Beam-Search Optimization
Wiseman, S.; and Rush, A. M. 2016 · 2016
Cited alongside, same era.
An Actor-Critic Algorithm for Sequence Prediction
Bahdanau, D.; Brakel, P.; Xu, K.; Goyal, A.; Lowe, R.; Pineau, J.; Courville, A. C.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
Categorical Reparametrization with Gumble-Softmax
Jang, E.; Gu, S.; and Poole, B. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Graph Attention Networks
Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; and Bengio, Y. 2018 · 2018
Later among the works it cites.
Black Box Quantiles for Kernel Learning
Tompkins, A.; Senanayake, R.; Morere, P.; and Ramos, F. 2019 · 2019
Later among the works it cites.
Transformer Dissection: An Unified Understanding for Transformer’s Attention via the Lens of Kernel
Tsai, Y.-H. H.; Bai, S.; Yamada, M.; Morency, L.-P.; and Salakhutdinov, R. 2019 · 2019
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020 · 2020
Closest in time.
Encoding word order in complex embeddings
Wang, B.; Zhao, D.; Lioma, C.; Li, Q.; Zhang, P.; and Simonsen, J. G. 2020 · 2020
Closest in time.