An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Later among the works it cites.
Neural combinatorial optimization with reinforcement learning
Irwan Bello, Hieu Pham, Quoc V. Le, Mohammad Norouzi, and Samy Bengio · 2017
Later among the works it cites.
Max-sum diversification, monotone submodular functions, and dynamic updates
Allan Borodin, Aadhar Jain, Hyun Chul Lee, and Yuli Ye · 2017
Later among the works it cites.
Unbiased learning-to-rank with biased feedback
Thorsten Joachims, Adith Swaminathan, and Tobias Schnabel · 2017
Later among the works it cites.
Neural models for information retrieval
Original
Bhaskar Mitra and Nick Craswell · 2017
Later among the works it cites.
Deep choice model using pointer networks for airline itinerary prediction
Alejandro Mottini and Rodrigo Acuna-Agost · 2017
Later among the works it cites.
Maximizing subset accuracy with recurrent neural networks in multi-label classification
Jinseok Nam, Eneldo Loza Mencía, Hyunwoo J Kim, and Johannes Fürnkranz · 2017
Later among the works it cites.
Deeprank: A new deep architecture for relevance ranking in information retrieval
Liang Pang, Yanyan Lan, Jiafeng Guo, Jun Xu, Jingfang Xu, and Xueqi Cheng · 2017
Later among the works it cites.
Visual permutation learning
Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, and Stephen Gould · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
An attention-based deep net for learning to rank
Original
Baiyang Wang and Diego Klabjan · 2017
Later among the works it cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Original
Victor Zhong, Caiming Xiong, and Richard Socher · 2017
Later among the works it cites.
Learning permutations with sinkhorn policy gradient
Patrick Emami and Sanjay Ranka · 2018
Closest in time.
Deep learning with logged bandit feedback
T. Joachims, A. Swaminathan, and M. de Rijke · 2018
Closest in time.
Reparameterizing the birkhoff polytope for variational permutation inference
Scott Linderman, Gonzalo Mena, Hal Cooper, Liam Paninski, and John Cunningham · 2018
Closest in time.
Learning latent permutations with gumbel-sinkhorn networks
Gonzalo Mena, David Belanger, Scott Linderman, and Jasper Snoek · 2018
Closest in time.
Top-k off-policy correction for a reinforce recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed Chi · 2019
Closest in time.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2019
Closest in time.