Fetching the paper…
Reading the bibliography…
We introduce a differentiable clustering method based on stochastic perturbations of minimum-weight spanning forests.
Differentiable deep clustering with cluster size constraints
Genevay, A., Dulac-Arnold, G., and Vert, J.-P. (2019) · 1910
Earlier work this paper cites.
Differentiation of blackbox combinatorial solvers
Vlastelica, M., Paulus, A., Musil, V., Martius, G., and Rolínek, M. (2019) · 1912
Earlier work this paper cites.
Statistical Theory of Extreme Values and some Practical Applications: A Series of Lectures
Gumbel, E. J. (1954) · 1954
Earlier work this paper cites.
On the shortest spanning subtree of a graph and the traveling salesman problem
Kruskal, J. B. (1956) · 1956
Earlier work this paper cites.
Minimum spanning trees and single linkage cluster analysis
Gower, J. C. and Ross, G. J. (1969) · 1969
Earlier work this paper cites.
Reducibility among combinatorial problems
Karp, R. M. (1972) · 1972
Earlier work this paper cites.
Convex analysis
Rockafellar, R. T. (1997) · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Clustering large graphs via the singular value decomposition
Drineas, P., Frieze, A., Kannan, R., Vempala, S., and Vinay, V. (2004) · 2004
Earlier work this paper cites.
K-means++: the advantages of careful seeding
Arthur, D. and Vassilvitskii, S. (2007) · 2007
Earlier work this paper cites.
Structured prediction models via the matrix-tree theorem
Koo, T., Globerson, A., Carreras Pérez, X., and Collins, M. (2007) · 2007
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Earlier work this paper cites.
Mnist handwritten digit database
LeCun, Y., Cortes, C., and Burges, C. (2010) · 2010
Earlier work this paper cites.
Perturb-and-MAP random fields: Using discrete optimization to learn and sample from energy models
Papandreou, G. and Yuille, A. L. (2011) · 2011
Earlier work this paper cites.
How the initialization affects the stability of the k-means algorithm
Bubeck, S., Meilă, M., and von Luxburg, U. (2012) · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Cuturi, M. (2013) · 2013
Earlier work this paper cites.
On sampling from the gibbs distribution with random maximum a-posteriori perturbations
Hazan, T., Maji, S., and Jaakkola, T. (2013) · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S. (2014) · 2014
Cited alongside, same era.
Learning with maximum a-posteriori perturbation models
Gane, A., Hazan, T., and Jaakkola, T. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Tensorflow: a system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. (2016) · 2016
Cited alongside, same era.
Perturbation techniques in online learning and optimization
Abernethy, J., Lee, C., and Tewari, A. (2016) · 2016
Cited alongside, same era.
Perturbations, Optimization, and Statistics
Hazan, T., Papandreou, G., and Tarlow, D. (2016) · 2016
Structured prediction with partial labelling through the infimum loss
Cabannes, V., Rudi, A., and Bach, F. (2020) · 2020
Later among the works it cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. (2020) · 2020
Later among the works it cites.
Gradient estimation with stochastic softmax tricks
Paulus, M., Choi, D., Tarlow, D., Krause, A., and Maddison, C. J. (2020) · 2020
Later among the works it cites.
Self-supervised learning of audio representations from permutations with differentiable ranking
Carr, A. N., Berthet, Q., Blondel, M., Teboul, O., and Zeghidour, N. (2021) · 2021
Later among the works it cites.
Differentiable patch selection for image recognition
Cordonnier, J.-B., Mahendran, A., Dosovitskiy, A., Weissenborn, D., Uszkoreit, J., and Unterthiner, T. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W. (2016) · 2016
Cited alongside, same era.
Unsupervised deep embedding for clustering analysis
Xie, J., Girshick, R., and Farhadi, A. (2016) · 2016
Cited alongside, same era.
Soft-DTW: a differentiable loss function for time-series
Cuturi, M. and Blondel, M. (2017) · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B. (2017) · 2017
Cited alongside, same era.
Towards k-means-friendly spaces: Simultaneous deep learning and clustering
Yang, B., Fu, X., Sidiropoulos, N. D., and Hong, M. (2017) · 2017
Cited alongside, same era.
Kumar, A., Brazil, G., and Liu, X. (2021) · 2021
Later among the works it cites.
Differentiable rendering with perturbed optimizers
Le Lidec, Q., Laptev, I., Schmid, C., and Carpentier, J. (2021) · 2021
Later among the works it cites.
Leveraging recursive gumbel-max trick for approximate inference in combinatorial spaces
Struminsky, K., Gadetsky, A., Rakitin, D., Karpushkin, D., and Vetrov, D. P. (2021) · 2021
Later among the works it cites.
Efficient computation of expectations under spanning tree distributions
Zmigrod, R., Vieira, T., and Cotterell, R. (2021) · 2021
Later among the works it cites.
Efficient and modular implicit differentiation
Blondel, M., Berthet, Q., Cuturi, M., Frostig, R., Hoyer, S., Llinares-López, F., Pedregosa, F., and Vert, J.-P. (2022) · 2022
Later among the works it cites.
Introduction to Algorithms
Cormen, T. H., Leiserson, C. E., Rivest, R. L., and Stein, C. (2022) · 2022
Later among the works it cites.
Fast stochastic composite minimization and an accelerated frank-wolfe algorithm under parallelization
Dubois-Taine, B., Bach, F., Berthet, Q., and Taylor, A. (2022) · 2022
Later among the works it cites.
A review of the gumbel-max trick and its extensions for discrete stochasticity in machine learning
Huijben, I. A., Kool, W., Paulus, M. B., and Van Sloun, R. J. (2022) · 2022
Later among the works it cites.
Deepconsensus improves the accuracy of sequences with a gap-aware sequence transformer
Baid, G., Cook, D. E., Shafin, K., Yun, T., Llinares-López, F., Berthet, Q., Belyaeva, A., Töpfer, A., Wenger, A. M., Rowell, W. J., et al. (2023) · 2023
Closest in time.
Deep embedding and alignment of protein sequences
Llinares-López, F., Berthet, Q., Blondel, M., Teboul, O., and Vert, J.-P. (2023) · 2023
Closest in time.
Fast, differentiable and sparse top-k: a convex analysis perspective
Sander, M. E., Puigcerver, J., Djolonga, J., Peyré, G., and Blondel, M. (2023) · 2023
Closest in time.
Learning by sorting: Self-supervised learning with group ordering constraints
Shvetsova, N., Petersen, F., Kukleva, A., Schiele, B., and Kuehne, H. (2023) · 2023
Closest in time.