Fetching the paper…
Reading the bibliography…
The top-k operator returns a sparse vector, where the non-zero values correspond to the k largest values of the input.
The theory of max-min, with applications
Danskin, J. M · 1966
Earlier work this paper cites.
Diagonal equivalence to matrices with prescribed row and column sums
Sinkhorn, R · 1967
Earlier work this paper cites.
Permutation polyhedra
Bowman, V · 1972
Earlier work this paper cites.
A method for finding projections onto the intersection of convex sets in hilbert spaces
Boyle, J. P. and Dykstra, R. L · 1986
Earlier work this paper cites.
Minimizing separable convex functions subject to simple chain constraints
Best, M. J., Chakravarti, N., and Ubhaya, V. A · 2000
Earlier work this paper cites.
Randomized online pca algorithms with regret bounds that are logarithmic in the dimension
Warmuth, M. K. and Kuzmin, D · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Gradient descent optimization of smoothed information retrieval metrics
Chapelle, O. and Wu, M · 2010
Earlier work this paper cites.
Ranking via sinkhorn propagation
Adams, R. P. and Zemel, R. S · 2011
Earlier work this paper cites.
Proximal splitting methods in signal processing
Combettes, P. L. and Pesquet, J.-C · 2011
Earlier work this paper cites.
Sparse prediction with the k k -support norm
Argyriou, A., Foygel, R., and Srebro, N · 2012
Earlier work this paper cites.
How to project onto the monotone nonnegative cone using pool adjacent violators type algorithms
Németh, A. and Németh, S · 2012
Earlier work this paper cites.
Lectures on polytopes , volume 152
Ziegler, G. M · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Cuturi, M · 2013
Earlier work this paper cites.
Reflection methods for user-friendly submodular optimization
Jegelka, S., Bach, F., and Sra, S · 2013
Earlier work this paper cites.
Sparse projections onto the simplex
Kyrillidis, A., Becker, S., Cevher, V., and Koch, C · 2013
Earlier work this paper cites.
Memory bounded deep convolutional networks
Collins, M. D. and Kohli, P · 2014
Earlier work this paper cites.
Spectral k-support norm regularization
McDonald, A. M., Pontil, M., and Stamos, D · 2014
Earlier work this paper cites.
The ordered weighted l 1 l_{1} norm: Atomic formulation, projections, and algorithms
Zeng, X. and Figueiredo, M. A · 2014
Earlier work this paper cites.
The k-support norm and convex envelopes of cardinality and rank
Eriksson, A., Thanh Pham, T., Chin, T.-J., and Reid, I · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Cited alongside, same era.
Top-k multiclass svm
Lapin, M., Hein, M., and Schiele, B · 2015
Cited alongside, same era.
Loss functions for top-k error: Analysis and insights
Lapin, M., Hein, M., and Schiele, B · 2016
Cited alongside, same era.
Efficient bregman projections onto the permutahedron and related polytopes
Lim, C. H. and Wright, S. J · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Wiseman, S. and Rush, A. M · 2016
Cited alongside, same era.
First-order methods in optimization
Beck, A · 2017
Cited alongside, same era.
Learning with differentiable pertubed optimizers
Berthet, Q., Blondel, M., Teboul, O., Cuturi, M., Vert, J.-P., and Bach, F · 2020
Later among the works it cites.
What is the state of neural network pruning?
Blalock, D., Gonzalez Ortiz, J. J., Frankle, J., and Guttag, J · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Softsort: A continuous relaxation for the argsort operator
Prillo, S. and Eisenschlos, J · 2020
Later among the works it cites.
Optimizing rank-based metrics with blackbox differentiation
Rolínek, M., Musil, V., Paulus, A., Vlastelica, M., Michaelis, C., and Martius, G · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ir evaluation methods for retrieving highly relevant documents
Järvelin, K. and Kekäläinen, J · 2017
Cited alongside, same era.
Learning what’s easy: Fully differentiable neural easy-first taggers
Martins, A. F. and Kreutzer, J · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A · 2017
Cited alongside, same era.
Dykstra’s algorithm, admm, and coordinate descent: Connections, insights, and extensions
Tibshirani, R. J · 2017
Cited alongside, same era.
Smooth loss functions for deep top-k classification
Berrada, L., Zisserman, A., and Kumar, M. P · 2018
Cited alongside, same era.
Xie, Y., Dai, H., Chen, M., Dai, B., Zhao, T., Zha, H., Wei, W., and Pfister, T · 2020
Later among the works it cites.
Efficient and modular implicit differentiation
Blondel, M., Berthet, Q., Cuturi, M., Frostig, R., Hoyer, S., Llinares-López, F., Pedregosa, F., and Vert, J.-P · 2021
Later among the works it cites.
Differentiable patch selection for image recognition
Cordonnier, J.-B., Mahendran, A., Dosovitskiy, A., Weissenborn, D., Uszkoreit, J., and Unterthiner, T · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
Hazimeh, H., Zhao, Z., Chowdhery, A., Sathiamoorthy, M., Chen, Y., Mazumder, R., Hong, L., and Chi, E · 2021
Later among the works it cites.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Later among the works it cites.
Implicit mle: backpropagating through discrete exponential family distributions
Niepert, M., Minervini, P., and Franceschi, L · 2021
Later among the works it cites.
Differentiable sorting networks for scalable sorting and ranking supervision
Petersen, F., Borgelt, C., Kuehne, H., and Deussen, O · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N · 2021
Later among the works it cites.
A review of sparse expert models in deep learning
Fedus, W., Dean, J., and Zoph, B · 2022
Later among the works it cites.
Sparsity-constrained optimal transport
Liu, T., Puigcerver, J., and Blondel, M · 2022
Later among the works it cites.
Differentiable top-k classification learning
Petersen, F., Kuehne, H., Borgelt, C., and Deussen, O · 2022
Later among the works it cites.
Multi-vector retrieval as sparse alignment
Qian, Y., Lee, J., Duddu, S. M. K., Dai, Z., Brahma, S., Naim, I., Lei, T., and Zhao, V. Y · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A., Chen, Z., Le, Q., and Laudon, J · 2022
Later among the works it cites.