Fetching the paper…
Reading the bibliography…
We address the problem of learning on sets of features, motivated by the need of performing pooling operations in long biological sequences of varying sizes, with long-range dependencies, and possibly few labeled data.
Concerning nonnegative matrices and doubly stochastic matrices
Richard Sinkhorn and Paul Knopp · 1967
Earlier work this paper cites.
The earth mover’s distance as a metric for image retrieval
Yossi Rubner, Carlo Tomasi, and Leonidas J. Guibad · 2000
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf and Alexander J. Smola · 2001
Earlier work this paper cites.
Using the nyström method to speed up kernel machines
Christopher K.I. Williams and Matthias Seeger · 2001
Earlier work this paper cites.
Mercer kernels for object recognition with local features
Siwei Lyu · 2004
Earlier work this paper cites.
Optimal Transport: Old and New
Cédric Villani · 2008
Earlier work this paper cites.
Improved nyström low-rank approximation and error analysis
Kai Zhang, Ivor W. Tsang, and James T. Kwok · 2008
Earlier work this paper cites.
Non-Local Means Denoising
Antoni Buades, Bartomeu Coll, and Jean-Michel Morel · 2011
Earlier work this paper cites.
Wasserstein barycenter and its application to texture mixing
Julien Rabin, Gabriel Peyré, Julie Delon, and Marc Bernot · 2011
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Fast computation of wasserstein barycenters
Marco Cuturi and Arnaud Doucet · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
To aggregate or not to aggregate: Selective match kernels for image search
Giorgos Tolias, Yannis Avrithis, and Hervé Jégou · 2013
Earlier work this paper cites.
Convolutional kernel networks
Julien Mairal, Piotr Koniusz, Zaid Harchaoui, and Cordelia Schmid · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzimitry Bahdanau, Kyunghyun Cho, and Joshua Bengio · 2015
Earlier work this paper cites.
Predicting effects of noncoding variants with deep learning–based sequence model
Jian Zhou and Olga G Troyanskaya · 2015
Cited alongside, same era.
Sliced wasserstein kernels for probabilit distributions
Soheil Kolouri, Yang Zou, and Gustavo K. Rohde · 2016
Cited alongside, same era.
End-to-end kernel learning with supervised convolutional kernel networks
Julien Mairal · 2016
Cited alongside, same era.
Multilevel clustering via wasserstein means
Nhat Ho, XuanLong Nguyen, Mikhail Yurochkin, Hung Hai Bui, Viet Huynh, and Dinh Phung · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Nikki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross B. Girshick, Abhinav Gupta, and Kaiming He · 2017
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Later among the works it cites.
Computational optimal transport
Gabriel Peyré and Marco Cuturi · 2019
Later among the works it cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Siddharth Goyal, Joshua Meier, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus · 2019
Later among the works it cites.
Transformer dissection: A unified understanding of transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbhakhsh, Barnabás Póczos, Ruslan Salakhutdinov, and Alexander J. Smola · 2017
Cited alongside, same era.
On the definiteness of earth mover’s distance and its relation to set intersection
Andrew Gardner, Christian A. Duncan, Jinko Kanno, and Rastko R. Selmic · 2018
Cited alongside, same era.
Learning generative models with sinkhorn divergences
Aude Genevay, Gabriel Peyré, and Marco Cuturi · 2018
Cited alongside, same era.
Convolution, attention and structure embedding
Jean-Marc Andreoli · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Glue: a multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 2019
Later among the works it cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, David Belanger, Lucy Colwell, and Adrian Weller · 2020
Closest in time.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2020
Closest in time.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Closest in time.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Tiedemann Jörg · 2020
Closest in time.
Rep the set: Neural networks for learning set representations
Konstantinos Skianis, Giannis Nikolentzos, Stratis Limnios, and Michalis Vazirgiannis · 2020
Closest in time.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Closest in time.
Hard-coded gaussian attention for neural machine translation
You Weiqiu, Simeng Sun, and Mohit Iyyer · 2020
Closest in time.