Fetching the paper…
Reading the bibliography…
Real-time advertising allows advertisers to bid for each impression for a visiting user.
Asynchronous methods for deep reinforcement learning. In ICML . 1928–1937
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Equilibrium in a stochastic n n -person game
Arlington M Fink et al · 1964
Earlier work this paper cites.
Game Theory Cambridge MA
Drew Fudenberg and Jean Tirole. 1991 · 1991
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents. In 10th ICML . 330–337
Ming Tan. 1993 · 1993
Earlier work this paper cites.
Multiagent reinforcement learning: theoretical framework and an algorithm.. In ICML , Vol. 98. Citeseer, 242–250
Junling Hu, Michael P Wellman, et al · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction . Vol. 1
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games. In Proceedings of the Sixteenth conference on Uncertainty in artificial intelligence . Morgan Kaufmann Publishers Inc., 541–548
S. Singh, M. Kearns, and Y. Mansour. 2000 · 2000
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games. In JMLR . ., 1039–1069
Junling Hu and Michael P. Wellman. 2003 · 2003
Earlier work this paper cites.
Reinforcement learning to play an optimal Nash equilibrium in team Markov games. In NIPS . 1603–1610
Xiaofeng Wang and Tuomas Sandholm. 2003 · 2003
Earlier work this paper cites.
Internet advertising and the generalized second-price auction: Selling billions of dollars worth of keywords
Benjamin Edelman, Michael Ostrovsky, and Michael Schwarz. 2007 · 2007
Earlier work this paper cites.
Algorithmic game theory . Vol. 1
Noam Nisan, Tim Roughgarden, Eva Tardos, and Vijay V Vazirani. 2007 · 2007
Earlier work this paper cites.
The online advertising industry: Economics, evolution, and privacy
David S Evans. 2009 · 2009
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation. In 19th WWW
Lihong Li, Wei Chu, John Langford, and Robert E Schapire. 2010 · 2010
Cited alongside, same era.
Online display advertising: Targeting and obtrusiveness
Avi Goldfarb and Catherine Tucker. 2011 · 2011
Cited alongside, same era.
Model-free reinforcement learning with continuous action in practice. In ACC, 2012 . IEEE
Thomas Degris, Patrick M Pilarski, and Richard S Sutton. 2012 · 2012
Cited alongside, same era.
Bid optimizing and inventory scoring in targeted online advertising. In 18th SIGKDD . ACM, 804–812
Claudia Perlich, Brian Dalessandro, Rod Hook, Ori Stitelman, Troy Raeder, and Foster Provost. 2012 · 2012
Cited alongside, same era.
Smart pacing for effective online ad campaign optimization. In 21st SIGKDD . ACM, 2217–2226
Jian Xu, Kuang-chih Lee, Wentong Li, Hang Qi, and Quan Lu. 2015 · 2015
Later among the works it cites.
Learning to communicate with deep multi-agent reinforcement learning. In NIPS . 2137–2145
J. Foerster, I. A. Assael, N. de Freitas, and S. Whiteson. 2016 · 2016
Later among the works it cites.
Display advertising with real-time bidding (RTB) and behavioural targeting
Jun Wang, Weinan Zhang, and Shuai Yuan. 2016 · 2016
Later among the works it cites.
Real-Time Bidding by Reinforcement Learning in Display Advertising. In 10th WSDM . ACM, 661–670
Han Cai, Kan Ren, Weinan Zhang, Kleanthis Malialis, Jun Wang, Yong Yu, and Defeng Guo. 2017 · 2017
Later among the works it cites.
Improving Real-Time Bidding Using a Constrained Markov Decision Process. In ICADMA . Springer, 711–726
Manxing Du, Redouane Sassioui, Georgios Varisteas, Mats Brorsson, Omar Cherkaoui, et al · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kuang-Chih Lee, Ali Jalali, and Ali Dasdan. 2013 · 2013
Cited alongside, same era.
Real-time bidding for online advertising: measurement and analysis. In 7th ADKDD . ACM, 3
Shuai Yuan, Jun Wang, and Xiaoxue Zhao. 2013 · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms. In ICML
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Cited alongside, same era.
An empirical study of reserve price optimisation in real-time bidding. In 20th SIGKDD
Shuai Yuan, Jun Wang, Bowei Chen, Peter Mason, and Sam Seljan. 2014 · 2014
Cited alongside, same era.
Optimal real-time bidding for display advertising. In 20th SIGKDD . 1077–1086
Weinan Zhang, Shuai Yuan, and Jun Wang. 2014 · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
Cooperative multi-agent control using deep reinforcement learning. In AAMAS . Springer
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer. 2017 · 2017
Later among the works it cites.
Multi-agent Reinforcement Learning in Sequential Social Dilemmas. In 16th AAMAS . 464–473
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel. 2017 · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments. In NIPS . 6382–6393
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel. 2017 · 2017
Later among the works it cites.
LADDER: A Human-Level Bidding Agent for Large-Scale Real-Time Online Auctions
Yu Wang, Jiayi Liu, Yuxiang Liu, Jun Hao, Yang He, Jinghe Hu, Weipeng Yan, and Mantian Li. 2017 · 2017
Later among the works it cites.
Optimized cost per click in taobao display advertising. In 23rd SIGKDD . ACM
Han Zhu, Junqi Jin, Chang Tan, Fei Pan, Yifan Zeng, Han Li, and Kun Gai. 2017 · 2017
Later among the works it cites.