Fetching the paper…
Reading the bibliography…
The top-k operation, i.e., finding the k largest or smallest elements from a collection of scores, is an important model component, which is widely used in information retrieval, machine learning, and data mining.
Mathematical methods of organizing and planning production
Kantorovich, L. V · 1960
Earlier work this paper cites.
Algorithm 65: Find
Hoare, C. A · 1961
Earlier work this paper cites.
Concerning nonnegative matrices and doubly stochastic matrices
Sinkhorn, R. and Knopp, P · 1967
Earlier work this paper cites.
Speech understanding systems: A summary of results of the five-year research effort. department of computer science, 1977
Reddy, D. R. et al · 1977
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Efficient projections onto the l 1-ball for learning in high dimensions
Duchi, J., Shalev-Shwartz, S., Singer, Y., and Chandra, T · 2008
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation , volume 105
Griewank, A. and Walther, A · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T., Tibshirani, R., and Friedman, J · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Cuturi, M · 2013
Earlier work this paper cites.
Neural codes for image retrieval
Babenko, A., Slesarev, A., Chigorin, A., and Lempitsky, V · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Edinburgh’s phrase-based machine translation systems for wmt-14
Durrani, N., Haddow, B., Koehn, P., and Heafield, K · 2014
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Jean, S., Cho, K., Memisevic, R., and Bengio, Y · 2014
Cited alongside, same era.
Addressing the rare word problem in neural machine translation
Luong, M.-T., Sutskever, I., Le, Q. V., Vinyals, O., and Zaremba, W · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Cited alongside, same era.
Iterative bregman projections for regularized transportation problems
Benamou, J.-D., Carlier, G., Cuturi, M., Nenna, L., and Peyré, G · 2015
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Klein, G., Kim, Y., Deng, Y., Senellart, J., and Rush, A. M · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Later among the works it cites.
A continuous relaxation of beam search for end-to-end training of neural sequence models
Goyal, K., Neubig, G., Dyer, C., and Berg-Kirkpatrick, T · 2018
Later among the works it cites.
Differential properties of sinkhorn approximation for learning with wasserstein distance
Luise, G., Rudi, A., Pontil, M., and Ciliberto, C · 2018
Later among the works it cites.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Papernot, N. and McDaniel, P · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Cited alongside, same era.
Deep image retrieval: Learning global representations for image search
Gordo, A., Almazán, J., Revaud, J., and Larlus, D · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
Cnn image retrieval learns from bow: Unsupervised fine-tuning with hard examples
Radenović, F., Tolias, G., and Chum, O · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Wiseman, S. and Rush, A. M · 2016
Cited alongside, same era.
Plötz, T. and Roth, S · 2018
Later among the works it cites.
Surprisingly easy hard-attention for sequence to sequence learning
Shankar, S., Garg, S., and Sarawagi, S · 2018
Later among the works it cites.
Fine-grained video categorization with redundancy reduction attention
Zhu, C., Tan, X., Zhou, F., Liu, X., Yue, K., Ding, E., and Ma, Y · 2018
Later among the works it cites.
Differentiable ranking and sorting using optimal transport
Cuturi, M., Teboul, O., and Vert, J.-P · 2019
Later among the works it cites.
Stochastic optimization of sorting networks via continuous relaxations
Grover, A., Wang, E., Zweig, A., and Ermon, S · 2019
Later among the works it cites.
Kool, W., Van Hoof, H., and Welling, M · 2019
Later among the works it cites.
Attention gated networks: Learning to leverage salient regions in medical images
Schlemper, J., Oktay, O., Schaap, M., Heinrich, M., Kainz, B., Glocker, B., and Rueckert, D · 2019
Later among the works it cites.
Reparameterizable subset sampling via continuous relaxations
Xie, S. M. and Ermon, S · 2019
Later among the works it cites.