Fetching the paper…
Reading the bibliography…
Successful applications of reinforcement learning in real-world problems often require dealing with partially observable states.
Reinforcement Learning for Robots using Neural Networks
Lin, Long-Ji · 1993
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Tesauro, Gerald · 1995
Earlier work this paper cites.
Customer lifetime valuation to support marketing decision making
Dwyer, F. Robert · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, Leslie Pack, Littman, Michael L., and Cassandra, Anthony R · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bakker, Bram · 2002
Earlier work this paper cites.
Predictive representations of state
Littman, Michael L., Sutton, Richard S., and Singh, Satinder P · 2002
Earlier work this paper cites.
Sequential cost-sensitive decision-making with reinforcement learning
Pednault, Edwin P. D., Abe, Naoki, and Zadrozny, Bianca · 2002
Cited alongside, same era.
Point-based value iteration: An anytime algorithm for POMDPs
Pineau, Joelle, Gordon, Geoffrey J., and Thrun, Sebastian · 2003
Cited alongside, same era.
Data Mining Techniques: For Marketing, Sales, and Customer Relationship Management
Berry, Michael J.A. and Linoff, Gordon S · 2004
Cited alongside, same era.
Partially observable Markov decision processes for spoken dialog systems
Williams, Jason D. and Young, Steve J · 2007
Cited alongside, same era.
A hidden Markov model of customer relationship dynamics
Netzer, Oded, Lattin, James M., and Srinivasan, V · 2008
Cited alongside, same era.
Closing the learning-planning loop with predictive state representations
Concurrent reinforcement learning from customer interactions
Silver, David, Newnham, Leonard, Barker, David, Weller, Suzanne, and McFall, Jason · 2013
Later among the works it cites.
Deep learning: Methods and applications
Deng, Li and Yu, Dong · 2014
Later among the works it cites.
Deep recurrent Q-learning for partially observable MDPs, 2015
Hausknecht, Matthew and Stone, Peter · 2015
Closest in time.
Improved Empirical Methods in Reinforcement-Learning Evaluation
Marivate, Vukosi N · 2015
Closest in time.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boots, Byron, Siddiqi, Sajid M., and Gordon, Geoffrey J · 2011
Cited alongside, same era.
Customer Relationship Management: Concept, Strategy, and Tools
Kumar, V. and Reinartz, Werner · 2012
Cited alongside, same era.
Recent advances in deep learning for speech research at Microsoft
Deng, Li, Li, Jinyu, Huang, Jui-Ting, Yao, Kaisheng, Yu, Dong, Seide, Frank, Seltzer, Michael, Zweig, Geoff, He, Xiaodong, Williams, Jason, Gong, Yifan, and Acero, Alex · 2013
Cited alongside, same era.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, Karthik, Kulkarni, Tejas, and Barzilay, Regina · 2015
Closest in time.
Personalized ad recommendation systems for life-time value optimization with guarantees
Theocharous, Georgios, Thomas, Philip S., and Ghavamzadeh, Mohammad · 2015
Closest in time.
Tkachenko, Yegor · 2015
Closest in time.