Fetching the paper…
Reading the bibliography…
Regularized optimal transport (OT) is now increasingly used as a loss or as a matching layer in neural networks.
Concerning nonnegative matrices and doubly stochastic matrices
Richard Sinkhorn and Paul Knopp · 1967
Earlier work this paper cites.
Algorithms for network programming
Jeff L Kennington and Richard V Helgason · 1980
Earlier work this paper cites.
A finite algorithm for finding the projection of a point onto the canonical simplex of ℝ n \mathbb{R}^{n}
Christian Michelot · 1986
Earlier work this paper cites.
Network flows
Ravindra K Ahuja, Thomas L Magnanti, and James B Orlin · 1988
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Convex analysis and minimization algorithms II , volume 305
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal · 1993
Earlier work this paper cites.
A note on constrained k k -means algorithms
Michael K Ng · 2000
Earlier work this paper cites.
Automated colour grading using colour distribution transfer
Francois Pitié, Anil C Kokaram, and Rozenn Dahyot · 2007
Earlier work this paper cites.
Efficient projections onto the ℓ 1 \ell_{1} -ball for learning in high dimensions
John Duchi, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Variational analysis , volume 317
R Tyrrell Rockafellar and Roger J-B Wets · 2009
Earlier work this paper cites.
A multiscale approach to optimal transport
Quentin Mérigot · 2011
Earlier work this paper cites.
Sparse prediction with the k k -support norm
Andreas Argyriou, Rina Foygel, and Nathan Srebro · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Sparse projections onto the simplex
Anastasios Kyrillidis, Stephen Becker, Volkan Cevher, and Christoph Koch · 2013
Earlier work this paper cites.
Proximal alternating linearized minimization for nonconvex and nonsmooth problems
Jérôme Bolte, Shoham Sabach, and Marc Teboulle · 2014
Earlier work this paper cites.
Wasserstein propagation for semi-supervised learning
Justin Solomon, Raif Rustamov, Guibas Leonidas, and Adrian Butscher · 2014
Earlier work this paper cites.
Numerical methods for matching for teams and Wasserstein barycenters
Guillaume Carlier, Adam Oberman, and Edouard Oudet · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Cited alongside, same era.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger · 2015
Cited alongside, same era.
Top-k multiclass SVM
Maksim Lapin, Matthias Hein, and Bernt Schiele · 2015
Cited alongside, same era.
On the minimization over sparse symmetric sets: projections, optimality conditions, and algorithms
Amir Beck and Nadav Hallak · 2016
Cited alongside, same era.
New perspectives on k-support and cluster norms
Andrew M McDonald, Massimiliano Pontil, and Dimitris Stamos · 2016
Cited alongside, same era.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Cited alongside, same era.
SuperGlue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich · 2020
Later among the works it cites.
An image is worth 16 × \times 16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Wasserstein k k -means with sparse simplex projection
Takumi Fukunaga and Hiroyuki Kasai · 2021
Later among the works it cites.
Unbiased gradient estimation with balanced assignments for mixtures of experts
Wouter Kool, Chris J. Maddison, and Andriy Mnih · 2021
Later among the works it cites.
GShard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lucas Roberts, Leo Razoumov, Lin Su, and Yuyang Wang · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Cited alongside, same era.
Smooth and sparse optimal transport
Mathieu Blondel, Vivien Seguy, and Antoine Rolet · 2018
Cited alongside, same era.
Regularized optimal transport and the rot mover’s distance
Arnaud Dessein, Nicolas Papadakis, and Jean-Luc Rouas · 2018
Cited alongside, same era.
A smoother way to train structured prediction models
Venkata Krishna Pillutla, Vincent Roulet, Sham M Kakade, and Zaid Harchaoui · 2018
Cited alongside, same era.
BASE layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Later among the works it cites.
Quadratically regularized optimal transport
Dirk A Lorenz, Paul Manns, and Christian Meyer · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Later among the works it cites.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, and Jason E Weston · 2021
Later among the works it cites.
Brandon Amos, Samuel Cohen, Giulia Luise, and Ievgen Redko · 2022
Closest in time.
Efficient optimal transport algorithm by accelerated gradient descent
Dongsheng An, Na Lei, Xiaoyin Xu, and Xianfeng Gu · 2022
Closest in time.
TPU-KNN: K nearest neighbor search at peak FLOP/s
Felix Chern, Blake Hechtman, Andy Davis, Ruiqi Guo, David Majnemer, and Sanjiv Kumar · 2022
Closest in time.
Unified scaling laws for routed language models
Aidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Closest in time.
Multimodal contrastive learning with LIMoE: the language-image mixture of experts
Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby · 2022
Closest in time.
On the adversarial robustness of mixture of experts
Joan Puigcerver, Rodolphe Jenatton, Carlos Riquelme, Pranjal Awasthi, and Srinadh Bhojanapalli · 2022
Closest in time.
Sinkformers: Transformers with doubly stochastic attention
Michael E Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyré · 2022
Closest in time.
SpeechMoE2: Mixture-of-experts model with improved routing
Zhao You, Shulin Feng, Dan Su, and Dong Yu · 2022
Closest in time.