Fetching the paper…
Reading the bibliography…
Search, recommendation, and online advertising are the three most important information-providing mechanisms on the web.
Model-based reinforcement learning for whole-chain recommendations
Zhao, X · 1902
Earlier work this paper cites.
Optimal control of markov processes with incomplete state information
Åström, K. J · 1965
Earlier work this paper cites.
The optimal control of partially observable markov processes over a finite horizon
Smallwood, R. D · 1973
Earlier work this paper cites.
The optimal control of partially observable markov processes over the infinite horizon: Discounted costs
Sondik, E. J · 1978
Earlier work this paper cites.
Multi-armed bandit problems and resource sharing systems
Varaiya, P · 1983
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, A. W · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A · 1994
Earlier work this paper cites.
Learning from delayed rewards
Kröse, B. J. A · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P · 1996
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Sutton, R. S · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R · 1999
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Bowling, M. H · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G · 2003
Earlier work this paper cites.
Multi-agent reinforcement learning: a critical survey
Shoham, Y · 2003
Earlier work this paper cites.
Natural actor-critic
Peters, J · 2005
Earlier work this paper cites.
An mdp-based recommender system
Shani, G · 2005
Earlier work this paper cites.
A survey of web information extraction systems
Chang, C · 2006
Earlier work this paper cites.
Incremental natural actor-critic algorithms
Bhatnagar, S · 2007
Earlier work this paper cites.
Adarank: a boosting algorithm for information retrieval
Xu, J · 2007
Earlier work this paper cites.
A support vector method for optimizing average precision
Yue, Y · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L · 2008
Earlier work this paper cites.
Natural actor-critic
Peters, J · 2008
Earlier work this paper cites.
Directly optimizing evaluation measures in learning to rank
Xu, J · 2008
Earlier work this paper cites.
Natural actor-critic algorithms
Bhatnagar, S · 2009
Cited alongside, same era.
Adaptive relevance feedback in information retrieval
Lv, Y · 2009
Cited alongside, same era.
Reinforcement learning and dynamic programming using function approximators
Busoniu, L · 2010
Cited alongside, same era.
Query representation and understanding workshop
Croft, W. B · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Li, L · 2010
Cited alongside, same era.
Contextual multi-armed bandits
Lu, T · 2010
Cited alongside, same era.
PILCO: A model-based and data-efficient approach to policy search
Customer acquisition via display advertising using multi-armed bandit experiments
Schwartz, E. M · 2017
Later among the works it cites.
Reinforcement mechanism design
Tang, P · 2017
Later among the works it cites.
Efficient ordered combinatorial semi-bandits for whole-page recommendation
Wang, Y · 2017
Later among the works it cites.
Returning is believing: Optimizing long-term user engagement in recommender systems
Wu, Q · 2017
Later among the works it cites.
Adapting markov decision process for search result diversification
Xia, L · 2017
Later among the works it cites.
Directly optimize diversity evaluation measures: A new approach to search result diversification
Xu, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deisenroth, M. P · 2011
Cited alongside, same era.
Information seeking: convergence of search, recommendations, and advertising
Garcia-Molina, H · 2011
Cited alongside, same era.
A unified optimization framework for auction and guaranteed delivery in online advertising
Salomatin, K · 2012
Cited alongside, same era.
Dynamic programming
Bellman, R · 2013
Cited alongside, same era.
Automatic ad format selection via contextual bandits
Tang, L · 2013
Cited alongside, same era.
Estimation bias in multi-armed bandit algorithms for search advertising
Xu, M · 2013
Cited alongside, same era.
Reinforcement learning to rank with markov decision process
Zeng, W · 2017
Later among the works it cites.
Deep reinforcement learning for list-wise recommendations
Zhao, X · 2017
Later among the works it cites.
Large-scale interactive recommendation with tree-structured policy gradient
Chen, H · 2018
Closest in time.
Stabilizing reinforcement learning in dynamic environment with application to online recommendation
Chen, S · 2018
Closest in time.
Reinforcement learning based recommender system using biclustering technique
Choi, S · 2018
Closest in time.
Learning to collaborate: Multi-scenario ranking via multi-agent reinforcement learning
Feng, J · 2018
Closest in time.
From greedy selection to exploratory decision-making: Diverse ranking with policy-value networks
Feng, Y · 2018
Closest in time.
Targeting optimization for internet advertising by learning from logged bandit feedback
Gasparini, M · 2018
Closest in time.
Reinforcement learning to rank in e-commerce search engine: Formalization, analysis, and application
Hu, Y · 2018
Closest in time.
Real-time bidding with multi-agent reinforcement learning in display advertising
Jin, J · 2018
Closest in time.
Balanced news using constrained bandit-based personalization
Kapoor, S · 2018
Closest in time.
A change-detection based framework for piecewise-stationary multi-armed bandit problem
Liu, F · 2018
Closest in time.
Deep reinforcement learning based recommendation with explicit user-item interactions modeling
Liu, F · 2018
Closest in time.
Learning to coordinate multiple reinforcement learning agents for diverse query reformulation
Nogueira, R · 2018
Closest in time.
A combinatorial-bandit algorithm for the online joint bid/budget optimization of pay-per-click advertising campaigns
Nuara, A · 2018
Closest in time.
Rohde, D · 2018
Closest in time.
Learning to advertise with adaptive exposure via constrained two-level reinforcement learning
Wang, W · 2018
Closest in time.
Optimizing whole-page presentation for web search
Wang, Y · 2018
Closest in time.
A multi-agent reinforcement learning method for impression allocation in online display advertising
Wu, D · 2018
Closest in time.
Budget constrained bidding by model-free reinforcement learning in display advertising
Wu, D · 2018
Closest in time.
Learning contextual bandits in a non-stationary environment
Wu, Q · 2018
Closest in time.
Multi page search with reinforcement learning to rank
Zeng, W · 2018
Closest in time.
Deep reinforcement learning for sponsored search real-time bidding
Zhao, J · 2018
Closest in time.
DRN: A deep reinforcement learning framework for news recommendation
Zheng, G · 2018
Closest in time.
Reinforcement learning to diversify recommendations
Zou, L · 2019
Closest in time.