Fetching the paper…
Reading the bibliography…
Wasserstein Discriminant Analysis (WDA) is a new supervised method that can improve classification of high-dimensional data by computing a suitable linear map onto a lower dimensional subspace.
Bonnans JF, Shapiro A (1998) Optimization problems with perturbations: A guided tour. SIAM review 40(2):228–264
1998
Earlier work this paper cites.
Friedman J, Hastie T, Tibshirani R (2001) The elements of statistical learning. Springer series in statistics Springer, Berlin
2001
Earlier work this paper cites.
Chapelle O, Vapnik V, Bousquet O, Mukherjee S (2002) Choosing multiple parameters for support vector machines. Machine learning 46(1-3):131–159
2002
Earlier work this paper cites.
Schölkopf B, Smola AJ (2002) Learning with kernels: Support vector machines, regularization, optimization, and beyond. MIT press
2002
Earlier work this paper cites.
Fern XZ, Brodley CE (2003) Random projection for high dimensional data clustering: A cluster ensemble approach. In: ICML, vol 3, pp 186–193
2003
Earlier work this paper cites.
Xing EP, Ng AY, Jordan MI, Russell S (2003) Distance metric learning with application to clustering with side-information. Advances in neural information processing systems 15:505–512
2003
Earlier work this paper cites.
Bach FR, Lanckriet GR, Jordan MI (2004) Multiple kernel learning, conic duality, and the smo algorithm. In: Proceedings of the twenty-first international conference on Machine learning, ACM, p 6
2004
Earlier work this paper cites.
Colson B, Marcotte P, Savard G (2007) An overview of bilevel optimization. Annals of operations research 153(1):235–256
2007
Earlier work this paper cites.
Griffin G, Holub A, Perona P (2007) Caltech-256 Object Category Dataset. Tech. Rep. CNS-TR-2007-001, California Institute of Technology
2007
Earlier work this paper cites.
Sugiyama M (2007) Dimensionality reduction of multimodal labeled data by local fisher discriminant analysis. The Journal of Machine Learning Research 8:1027–1061
2007
Earlier work this paper cites.
Knight PA (2008) The sinkhorn–knopp algorithm: convergence and applications. SIAM Journal on Matrix Analysis and Applications 30(1):261–275
2008
Earlier work this paper cites.
Van der Maaten L, Hinton G (2008) Visualizing data using t-sne. Journal of Machine Learning Research 9(2579-2605):85
2008
Earlier work this paper cites.
Petersen KB, Pedersen MS, et al (2008) The matrix cookbook. Technical University of Denmark 7:15
2008
Earlier work this paper cites.
Schmidt M (2008) Minconf-projection methods for optimization with simple constraints in matlab
2008
Earlier work this paper cites.
Villani C (2008) Optimal transport: old and new, vol 338. Springer Science & Business Media
2008
Cited alongside, same era.
Absil PA, Mahony R, Sepulchre R (2009) Optimization algorithms on matrix manifolds. Princeton University Press
2009
Cited alongside, same era.
Bengio Y (2009) Learning deep architectures for ai. Foundations and trends® in Machine Learning 2(1):1–127
2009
Cited alongside, same era.
Van Der Maaten L, Postma E, Van den Herik J (2009) Dimensionality reduction: a comparative review. Journal of Machine Learning Research 10:66–71
2009
Cited alongside, same era.
Weinberger KQ, Saul LK (2009) Distance metric learning for large margin nearest neighbor classification. The Journal of Machine Learning Research 10:207–244
2009
Cited alongside, same era.
Solomon J, Rustamov R, Leonidas G, Butscher A (2014) Wasserstein propagation for semi-supervised learning. In: ICML, pp 306–314
2014
Later among the works it cites.
Benamou JD, Carlier G, Cuturi M, Nenna L, Peyré G (2015) Iterative bregman projections for regularized transportation problems. SIAM Journal on Scientific Computing 37(2):A1111–A1138
2015
Later among the works it cites.
Emigh M, Kriminger E, Prîncipe JC (2015) Linear discriminant analysis with an information divergence criterion. In: 2015 International Joint Conference on Neural Networks (IJCNN), IEEE, pp 1–6
2015
Later among the works it cites.
Mueller J, Jaakkola T (2015) Principal differences analysis: Interpretable characterization of differences between distributions. In: NIPS, pp 1693–1701
2015
Later among the works it cites.
Seguy V, Cuturi M (2015) Principal geodesic analysis for probability measures under the optimal transport metric. In: NIPS, pp 3294–3302
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Burges CJ (2010) Dimension reduction: A guided tour. Now Publishers
2010
Cited alongside, same era.
2010
Cited alongside, same era.
Cuturi M (2013) Sinkhorn distances: Lightspeed computation of optimal transport. In: NIPS, pp 2292–2300
2013
Cited alongside, same era.
Giraldo LGS, Principe JC (2013) Information theoretic learning with infinitely divisible kernels. In: Proceedings of the first International Conference on Representation Learning (ICLR), pp 1–8
2013
Cited alongside, same era.
Lichman M (2013) UCI machine learning repository. URL http://archive.ics.uci.edu/ml
2013
Cited alongside, same era.
Suzuki T, Sugiyama M (2013) Sufficient dimension reduction via squared-loss mutual information estimation. Neural computation 25(3):725–758
2013
Cited alongside, same era.
Boumal N, Mishra B, Absil PA, Sepulchre R (2014) Manopt, a matlab toolbox for optimization on manifolds. The Journal of Machine Learning Research 15(1):1455–1459
2014
Cited alongside, same era.
2015
Later among the works it cites.
Tangkaratt V, Sasaki H, Sugiyama M (2015) Direct estimation of the derivative of quadratic mutual information with application in supervised dimension reduction. arXiv preprint arXiv:150801019
2015
Later among the works it cites.
Bonneel N, Peyré G, Cuturi M (2016) Wasserstein barycentric coordinates: Histogram regression using optimal transport. ACM Transactions on Graphics 35(4)
2016
Closest in time.
Courty N, Flamary R, Tuia D, Rakotomamonjy A (2016) Optimal transport for domain adaptation. Pattern Analysis and Machine Intelligence, IEEE Transactions on
2016
Closest in time.
Huang G, Guo C, Kusner MJ, Sun Y, Sha F, Weinberger KQ (2016) Supervised word mover’s distance. In: Advances in Neural Information Processing Systems, pp 4862–4870
2016
Closest in time.
Koep N, Weichwald S (2016) Pymanopt: A python toolbox for optimization on manifolds using automatic differentiation. Journal of Machine Learning Research 17:1–5
2016
Closest in time.
Flamary R, Courty N (2017) Pot python optimal transport library
2017
Closest in time.
Peyré G, Cuturi M (2018) Computational Optimal Transport. To be published in Foundations and Trends in Computer Science, URL https://optimaltransport.github.io
2018
Closest in time.
Frogner C, Zhang C, Mobahi H, Araya M, Poggio T (2015) Learning with a wasserstein loss. In: NIPS, pp 2044–2052
2052
Closest in time.