Fetching the paper…
Reading the bibliography…
We give an overview of recent exciting achievements of deep reinforcement learning (RL).
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Neuronlike elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W. (1983) · 1983
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. (1990) · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton, R. S. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G. (1994) · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P., Littman, M. L., and Moore, A. (1996) · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Investment Science
Luenberger, D. G. (1997) · 1997
Earlier work this paper cites.
Enhancing q-learning for optimal asset allocation
Neuneier, R. (1997) · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G. (2000) · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. and Russell, S. (2000) · 2000
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
Stone, P. and Veloso, M. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Valuing American options by simulation: a simple least-squares approach
Longstaff, F. A. and Schwartz, E. S. (2001) · 2001
Earlier work this paper cites.
Learning to trade via direct reinforcement
Moody, J. and Saffell, M. (2001) · 2001
Earlier work this paper cites.
Regression methods for pricing complex American-style options
Tsitsiklis, J. N. and Van Roy, B. (2001) · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. (2002) · 2002
Earlier work this paper cites.
Computer go
Müller, M. (2002) · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S. (2003) · 2003
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
Least squares policy evaluation algorithms with linear function approximation
Nedić, A. and Bertsekas, D. P. (2003) · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Convex Optimization
Boyd, S. and Vandenberghe, L. (2004) · 2004
Earlier work this paper cites.
The Adaptive Markets Hypothesis: Market efficiency from an evolutionary perspective
Lo, A. W. (2004) · 2004
Earlier work this paper cites.
A simulation approach to dynamic portfolio choice with an application to learning about return predictability
Brandt, M. W., Goyal, A., Santa-Clara, P., and Stroud, J. R. (2005) · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L. (2005) · 2005
Earlier work this paper cites.
Cognitive radio: brain-empowered wireless communications
Haykin, S. (2005) · 2005
Earlier work this paper cites.
Markov decision processes : discrete stochastic dynamic programming
Puterman, M. L. (2005) · 2005
Earlier work this paper cites.
Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. (2005) · 2005
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning
Ghavamzadeh, M., Mahadevan, S., and Makar, R. (2006) · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R. (2006) · 2006
Earlier work this paper cites.
Combining online and offline knowledge in uct
Gelly, S. and Silver, D. (2007) · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B. (2007) · 2007
Earlier work this paper cites.
If multi-agent learning is the answer, what is the question?
Shoham, Y., Powers, R., and Grenager, T. (2007) · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L., Babuska, R., and Schutter, B. D. (2008) · 2008
Earlier work this paper cites.
Neural Networks and Learning Machines (third edition)
Haykin, S. (2008) · 2008
Earlier work this paper cites.
Essentials of Game Theory: A Concise, Multidisciplinary Introduction
Leyton-Brown, K. and Shoham, Y. (2008) · 2008
Earlier work this paper cites.
Introduction to Information Retrieval
Manning, C. D., Raghavan, P., and Schütze, H. (2008) · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B. (2009) · 2009
Earlier work this paper cites.
Learning deep architectures for ai
Bengio, Y. (2009) · 2009
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Hastie, T., Tibshirani, R., and Friedman, J. (2009) · 2009
Earlier work this paper cites.
Learning exercise policies for American options
Li, Y., Szepesvári, C., and Schuurmans, D. (2009) · 2009
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach (3rd edition)
Russell, S. and Norvig, P. (2009) · 2009
Earlier work this paper cites.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Shoham, Y. and Leyton-Brown, K. (2009) · 2009
Earlier work this paper cites.
RL-Glue : Language-independent software for reinforcement-learning experiments
Tanner, B. and White, A. (2009) · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P. (2009) · 2009
Earlier work this paper cites.
A general projection property for distribution families
Yu, Y.-L., Li, Y., Szepesvári, C., and Schuurmans, D. (2009) · 2009
Earlier work this paper cites.
Introduction to semi-supervised learning
Zhu, X. and Goldberg, A. B. (2009) · 2009
Earlier work this paper cites.
Consumer credit-risk models via machine-learning algorithms
Khandani, A. E., Kim, A. J., and Lo, A. W. (2010) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q. (2010) · 2010
Earlier work this paper cites.
A Comprehensive Survey of Data Mining-based Fraud Detection Research
Phua, C., Lee, V., Smith, K., and Gayler, R. (2010) · 2010
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Szepesvári, C. (2010) · 2010
Earlier work this paper cites.
Adaptive stochastic control for the smart grid
Anderson, R. N., Boulanger, A., Powell, W. B., and Scott, W. (2011) · 2011
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Bishop, C. (2011) · 2011
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M. P. and Rasmussen, C. E. (2011) · 2011
Earlier work this paper cites.
Approximate Dynamic Programming: Solving the curses of dimensionality (2nd Edition)
Powell, W. B. (2011) · 2011
Earlier work this paper cites.
Informing sequential clinical decision-making through reinforcement learning: an empirical study
Shortreed, S. M., Laber, E., Lizotte, D. J., Stroup, T. S., Pineau, J., and Murphy, S. A. (2011) · 2011
Earlier work this paper cites.
Semi-supervised recursive autoencoders for predicting sentiment distributions
Socher, R., Pennington, J., Huang, E. H., Ng, A. Y., and Manning, C. D. (2011) · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction, , proc. of 10th
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Earlier work this paper cites.
Dynamic programming and optimal control (Vol. II, 4th Edition: Approximate Dynamic Programming)
Bertsekas, D. P. (2012) · 2012
Earlier work this paper cites.
A survey of Monte Carlo tree search methods
Browne, C., Powley, E., Whitehouse, D., Lucas, S., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S. (2012) · 2012
Earlier work this paper cites.
A few useful things to know about machine learning
Domingos, P. (2012) · 2012
Earlier work this paper cites.
Smart grid - the new and improved power grid: A survey
Fang, X., Misra, S., Xue, G., and Yang, D. (2012) · 2012
Earlier work this paper cites.
The grand challenge of computer go: Monte carlo tree search and extensions
Gelly, S., Schoenauer, M., Sebag, M., Teytaud, O., Kocsis, L., Silver, D., and Szepesvári, C. (2012) · 2012
Earlier work this paper cites.
Q-learning with censored data
Goldberg, Y. and Kosorok, M. R. (2012) · 2012
Earlier work this paper cites.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
Grondman, I., Busoniu, L., Lopes, G. A., and Babuška, R. (2012) · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
Hinton, G., Deng, L., Yu, D., Dahl, G. E., rahman Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., , and Kingsbury, B. (2012) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Le, Q. V., Ranzato, M., Monga, R., Devin, M., Chen, K., Corrado, G. S., Dean, J., and Ng, A. Y. (2012) · 2012
Earlier work this paper cites.
Sentiment Analysis and Opinion Mining
Liu, B. (2012) · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Murphy, K. P. (2012) · 2012
Earlier work this paper cites.
Reinforcement Learning: State-of-the-Art (edited book)
Wiering, M. and van Otterlo, M. (2012) · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., and Peters, J. (2013) · 2013
Earlier work this paper cites.
Machine learning paradigms for speech recognition: An overview
Deng, L. and Li, X. (2013) · 2013
Earlier work this paper cites.
Multiagent reinforcement learning for integrated network of adaptive traffic signal controllers (marlin-atsc): methodology and large-scale application on downtown toronto
El-Tantawy, S., Abdulhai, B., and Abdelgawad, H. (2013) · 2013
Earlier work this paper cites.
Learning and reasoning in cognitive radio networks
Gavrilovska, L., Atanasovski, V., Macaluso, I., and DaSilva, L. A. (2013) · 2013
Earlier work this paper cites.
A tutorial on linear function approximators for dynamic programming and reinforcement learning
Geramifard, A., Walsh, T. J., Tellex, S., Chowdhary, G., Roy, N., and How, J. P. (2013) · 2013
Earlier work this paper cites.
Speech-centric information processing: An optimization-oriented approach
He, X. and Deng, L. (2013) · 2013
Earlier work this paper cites.
An Introduction to Statistical Learning with Applications in R
James, G., Witten, D., Hastie, T., and Tibshirani, R. (2013) · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Kalchbrenner, N. and Blunsom, P. (2013) · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Earlier work this paper cites.
Applied Predictive Modeling
Kuhn, M. and Johnson, K. (2013) · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Earlier work this paper cites.
A survey of real-time strategy game ai research and competition in starcraft
Ontañón, S., Synnaeve, G., Uriarte, A., Richoux, F., Churchill, D., and Preuss, M. (2013) · 2013
Earlier work this paper cites.
Data Science for Business
Provost, F. and Fawcett, T. (2013) · 2013
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R. (2013) · 2013
Earlier work this paper cites.
Concurrent reinforcement learning from customer interactions
Silver, D., Newnham, L., Barker, D., Weller, S., and McFall, J. (2013) · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment tree- bank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C., Ng, A., and Potts, C. (2013) · 2013
Earlier work this paper cites.
POMDP-based statistical spoken dialogue systems: a review
Young, S., Gašić, M., Thomson, B., and Williams, J. D. (2013) · 2013
Earlier work this paper cites.
Multiple object recognition with visual attention
Ba, J., Mnih, V., and Kavukcuoglu, K. (2014) · 2014
Earlier work this paper cites.
Introduction to Intelligent Systems in Traffic and Transportation
Bazzan, A. L. and Klügl, F. (2014) · 2014
Earlier work this paper cites.
TORCS, The Open Racing Car Simulator
Bernhard Wymann, E. E., Guionneau, C., Dimitrakakis, C., and Rémi Coulom, A. S. (2014) · 2014
Earlier work this paper cites.
Dynamic treatment regimes
Chakraborty, B. and Murphy, S. A. (2014) · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Game-theoretic security patrolling with dynamic execution uncertainty and a case study on a real transit system
Delle Fave, F. M., Jiang, A. X., Yin, Z., Zhang, C., Tambe, M., Kraus, S., and Sullivan, J. P. (2014) · 2014
Earlier work this paper cites.
Deep Learning: Methods and Applications
Deng, L. and Dong, Y. (2014) · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma, M. W. (2014) · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., , and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Neural Turing Machines
Graves, A., Wayne, G., and Danihelka, I. (2014) · 2014
Earlier work this paper cites.
Options, Futures and Other Derivatives (9th edition)
Hull, J. C. (2014) · 2014
Earlier work this paper cites.
Learning from limited demonstrations
Kim, B., massoud Farahmand, A., Pineau, J., and Precup, D. (2014) · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Kingma, D. P., Rezende, D. J., Mohamed, S., and Welling, M. (2014) · 2014
Earlier work this paper cites.
Trading off scientific knowledge and user learning with multi-armed bandits
Liu, Y.-E., Mandel, T., Brunskill, E., and Popović, Z. (2014) · 2014
Earlier work this paper cites.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A. R., van Hasselt, H., and Sutton, R. S. (2014) · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
Mandel, T., Liu, Y. E., Levine, S., Brunskill, E., and Popović, Z. (2014) · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Mnih, V., Heess, N., Graves, A., and Kavukcuoglu, K. (2014) · 2014
Earlier work this paper cites.
A $3 trillion challenge to computational scientists: Transforming healthcare delivery
Saria, S. (2014) · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Earlier work this paper cites.
Internet of things in industries: A survey
Xu, L. D., He, W., and Li, S. (2014) · 2014
Earlier work this paper cites.
Universal option models
Yao, H., Szepesvari, C., Sutton, R. S., Modayil, J., and Bhatnagar, S. (2014) · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014) · 2014
Earlier work this paper cites.
Maximum entropy semi-supervised inverse reinforcement learning
Audiffren, J., Valko, M., Lazaric, A., and Ghavamzadeh, M. (2015) · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O. (2015) · 2015
Earlier work this paper cites.
Active object localization with deep reinforcement learning
Caicedo, J. C. and Lazebnik, S. (2015) · 2015
Earlier work this paper cites.
Natural Language Understanding with Distributed Representation
Cho, K. (2015) · 2015
Earlier work this paper cites.
Multi-task learning for multiple language translation
Dong, D., Wu, H., He, W., Yu, D., and Wang, H. (2015) · 2015
Earlier work this paper cites.
Deep Reinforcement Learning in Large Discrete Action Spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B. (2015) · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcìa, J. and Fernàndez, F. (2015) · 2015
Earlier work this paper cites.
Rlpy: A value-function-based reinforcement learning framework for education and research
Geramifard, A., Dann, C., Klein, R. H., Dabney, W., and How, J. P. (2015) · 2015
Earlier work this paper cites.
Bayesian reinforcement learning: a survey
Ghavamzadeh, M., Mannor, S., Pineau, J., and Tamar, A. (2015) · 2015
Earlier work this paper cites.
Fast R-CNN
Girshick, R. (2015) · 2015
Earlier work this paper cites.
Draw: A recurrent neural network for image generation
Gregor, K., Danihelka, I., Graves, A., Rezende, D., and Wierstra, D. (2015) · 2015
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable MDPs
Hausknecht, M. and Stone, P. (2015) · 2015
Earlier work this paper cites.
Advances in natural language processing
Hirschberg, J. and Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, C. S. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Spatial transformer networks
Jaderberg, M., Simonyan, K., Zisserman, A., and Kavukcuoglu, K. (2015) · 2015
Earlier work this paper cites.
Machine learning: Trends, perspectives, and prospects
Jordan, M. I. and Mitchell, T. (2015) · 2015
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Koch, G., Zemel, R., and Salakhutdinov, R. (2015) · 2015
Earlier work this paper cites.
Adaptive Treatment Strategies in Practice: Planning Trials and Analyzing Data for Personalized Medicine
Kosorok, M. R. and Moodie, E. E. M. (2015) · 2015
Earlier work this paper cites.
Deep convolutional inverse graphics network
Kulkarni, T. D., Whitney, W., Kohli, P., and Tenenbaum, J. B. (2015) · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B. (2015) · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Earlier work this paper cites.
Recurrent Reinforcement Learning: A Hybrid Approach
Li, X., Li, L., Gao, J., He, X., Chen, J., Deng, L., and He, J. (2015) · 2015
Earlier work this paper cites.
Reinforcement learning improves behaviour from evaluative feedback
Littman, M. L. (2015) · 2015
Earlier work this paper cites.
Learning transferable features with deep adaptation networks
Long, M., Cao, Y., Wang, J., and Jordan, M. I. (2015) · 2015
Earlier work this paper cites.
Using recurrent neural networks for slot filling in spoken language understanding
Mesnil, G., Dauphin, Y., Yao, K., Bengio, Y., Deng, L., He, X., Heck, L., Tur, G., Hakkani-Tür, D., Yu, D., and Zweig, G. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., De Maria, A., Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., Legg, S., Mnih, V., Kavukcuoglu, K., and Silver, D. (2015) · 2015
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, K., Kulkarni, T., and Barzilay, R. (2015) · 2015
Earlier work this paper cites.
Big data in manufacturing: a systematic mapping study
O’Donovan, P., Leahy, K., Bruton, K., and O’Sullivan, D. T. J. (2015) · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R., and Singh, S. (2015) · 2015
Earlier work this paper cites.
Is object localization for free? – weakly-supervised learning with convolutional neural networks
Oquab, M., Bottou, L., Laptev, I., and Sivic, J. (2015) · 2015
Earlier work this paper cites.
Policy search: Methods and applications
Peters, J. and Neumann, G. (2015) · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J. (2015) · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Schmidhuber, J. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Earlier work this paper cites.
A survey of available corpora for building data-driven dialogue systems
Serban, I. V., Lowe, R., Charlin, L., and Pineau, J. (2015) · 2015
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Weston, J., and Fergus, R. (2015) · 2015
Earlier work this paper cites.
Personalized ad recommendation systems for life-time value optimization with guarantees
Theocharous, G., Thomas, P. S., and Ghavamzadeh, M. (2015) · 2015
Earlier work this paper cites.
Pointer networks
Vinyals, O., Fortunato, M., and Jaitly, N. (2015) · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M. (2015) · 2015
Earlier work this paper cites.
Memory networks
Weston, J., Chopra, S., and Bordes, A. (2015) · 2015
Earlier work this paper cites.
Galileo: Perceiving physical object properties by integrating a physics engine with deep learning
Wu, J., Yildirim, I., Lim, J. J., Freeman, B., and Tenenbaum, J. (2015) · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J. L., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R. S., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Stacked Attention Networks for Image Question Answering
Yang, Z., He, X., Gao, J., Deng, L., and Smola, A. (2015) · 2015
Earlier work this paper cites.
Reinforcement Learning Neural Turing Machines - Revised
Zaremba, W. and Sutskever, I. (2015) · 2015
Earlier work this paper cites.
Object detectors emerge in deep scene CNNs
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A. (2015) · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L. (2016) · 2016
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, P., Nair, A., Abbeel, P., Malik, J., and Levine, S. (2016) · 2016
Earlier work this paper cites.
Concrete Problems in AI Safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D. (2016) · 2016
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Colmenarejo, S. G., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and de Freitas, N. (2016) · 2016
Earlier work this paper cites.
A sequence-to-sequence model for user simulation in spoken dialogue systems
Asri, L. E., He, J., and Suleman, K. (2016) · 2016
Earlier work this paper cites.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C. (2016) · 2016
Earlier work this paper cites.
Differentially private policy evaluation
Balle, B., Gomrokchi, M., and Precup, D. (2016) · 2016
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Battaglia, P. W., Pascanu, R., Lai, M., Rezende, D., and Kavukcuoglu, K. (2016) · 2016
Earlier work this paper cites.
DeepMind Lab
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S. (2016) · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Schaul, T., Srinivasan, S., Saxton, D., Ostrovski, G., and Munos, R. (2016) · 2016
Earlier work this paper cites.
Neural Combinatorial Optimization with Reinforcement Learning
Bello, I., Pham, H., Le, Q. V., Norouzi, M., and Bengio, S. (2016) · 2016
Earlier work this paper cites.
Playing Doom with SLAM-Augmented Deep Reinforcement Learning
Bhatti, S., Desmaison, A., Miksik, O., Nardelli, N., Siddharth, N., and Torr, P. H. S. (2016) · 2016
Cited alongside, same era.
End to End Learning for Self-Driving Cars
Bojarski, M., Testa, D. D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L. D., Monfort, M., Muller, U., Zhang, J., Zhang, X., Zhao, J., and Zieba, K. (2016) · 2016
Cited alongside, same era.
Path integral guided policy search
Chebotar, Y., Kalakrishnan, M., Yahya, A., Li, A., Schaal, S., and Levine, S. (2016) · 2016
Cited alongside, same era.
Unsupervised Learning of Predictors from Unpaired Input-Output Samples
Chen, J., Huang, P.-S., He, X., Gao, J., and Deng, L. (2016) · 2016
Cited alongside, same era.
Lifelong Machine Learning
Chen, Z. and Liu, B. (2016) · 2016
Cited alongside, same era.
Semi-supervised learning for neural machine translation
Cheng, Y., Xu, W., He, Z., He, W., Wu, H., Sun, M., and Liu, Y. (2016) · 2016
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
He, F. S., Liu, Y., Schwing, A. G., and Peng, J. (2017) · 2017
Closest in time.
Mask R-CNN
He, K., Gkioxari, G., Dollár, P., and Girshick, R. (2017) · 2017
Closest in time.
Deep semantic role labeling: What works and what’s next
He, L., Lee, K., Lewis, M., and Zettlemoyer, L. (2017) · 2017
Closest in time.
Emergence of Locomotion Behaviours in Rich Environments
Heess, N., TB, D., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, A., Riedmiller, M., and Silver, D. (2017) · 2017
Closest in time.
A benchmark environment motivated by industrial control problems
Hein, D., Depeweg, S., Tokic, M., Udluft, S., Hentschel, A., Runkler, T. A., and Sterzing, V. (2017) · 2017
Closest in time.
Automatic Goal Generation for Reinforcement Learning Agents
Held, D., Geng, X., Florensa, C., and Abbeel, P. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement Learning Using Quantum Boltzmann Machines
Crawford, D., Levit, A., Ghadermarzy, N., Oberoi, J. S., and Ronagh, P. (2016) · 2016
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y. (2016) · 2016
Cited alongside, same era.
Associative long short-term memory
Danihelka, I., Wayne, G., Uria, B., Kalchbrenner, N., and Graves, A. (2016) · 2016
Cited alongside, same era.
Deep direct reinforcement learning for financial signal representation and trading
Deng, Y., Bao, F., Kong, Y., Ren, Z., and Dai, Q. (2016) · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
RL 2 : Fast Reinforcement Learning via Slow Reinforcement Learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Closest in time.
Model-Based Planning in Discrete Action Spaces
Henaff, M., Whitney, W. F., and LeCun, Y. (2017) · 2017
Closest in time.
Intrinsically motivated model learning for developing curious robots
Hester, T. and Stone, P. (2017) · 2017
Closest in time.
β \beta -VAE: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A. (2017) · 2017
Closest in time.
Vain: Attentional multi-agent predictive modeling
Hoshen, Y. (2017) · 2017
Closest in time.
On Unifying Deep Generative Models
Hu, Z., Yang, Z., Salakhutdinov, R., and Xing, E. P. (2017) · 2017
Closest in time.
Densely connected convolutional networks
Huang, G., Liu, Z., Weinberger, K. Q., and van der Maaten, L. (2017) · 2017
Closest in time.
Adversarial Attacks on Neural Network Policies
Huang, S., Papernot, N., Goodfellow, I., Duan, Y., and Abbeel, P. (2017) · 2017
Closest in time.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Islam, R., Henderson, P., Gomrokchi, M., and Precup, D. (2017) · 2017
Closest in time.
Population Based Training of Neural Networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K. (2017) · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Closest in time.
Tuning recurrent neural networks with reinforcement learning
Jaques, N., Gu, S., Turner, R. E., and Eck, D. (2017) · 2017
Closest in time.
Bag of tricks for efficient text classification
Joulin, A., Grave, E., Bojanowski, P., and Mikolov, T. (2017) · 2017
Closest in time.
Speech and Language Processing (3rd ed. draft)
Jurafsky, D. and Martin, J. H. (2017) · 2017
Closest in time.
Deep Learning for Video Game Playing
Justesen, N., Bontrager, P., Togelius, J., and Risi, S. (2017) · 2017
Closest in time.
Learning macromanagement in starcraft from replays using deep learning
Justesen, N. and Risi, S. (2017) · 2017
Closest in time.
Batch policy gradient methods for improving neural conversation models
Kandasamy, K., Bachrach, Y., Tomioka, R., Tarlow, D., and Carter, D. (2017) · 2017
Closest in time.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Kansky, K., Silver, T., Mély, D. A., Eldawy, M., Lázaro-Gredilla, M., Lou, X., Dorfman, N., Sidor, S., Phoenix, S., and George, D. (2017) · 2017
Closest in time.
A new softmax operator for reinforcement learning
Kavosh and Littman, M. L. (2017) · 2017
Closest in time.
Generalization in Deep Learning
Kawaguchi, K., Pack Kaelbling, L., and Bengio, Y. (2017) · 2017
Closest in time.
Robust and efficient transfer learning with hidden-parameter markov decision processes
Killian, T., Daulton, S., Konidaris, G., and Doshi-Velez, F. (2017) · 2017
Closest in time.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R. (2017) · 2017
Closest in time.
Self-Normalizing Neural Networks
Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S. (2017) · 2017
Closest in time.
OpenNMT: Open-Source Toolkit for Neural Machine Translation
Klein, G., Kim, Y., Deng, Y., Senellart, J., and Rush, A. M. (2017) · 2017
Closest in time.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P. (2017) · 2017
Closest in time.
Continual curiosity-driven skill acquisition from high-dimensional video inputs for humanoid robots
Kompella, V. R., Stollenga, M., Luciw, M., and Schmidhuber, J. (2017) · 2017
Closest in time.
Collaborative deep reinforcement learning for joint object search
Kong, X., Xin, B., Wang, Y., and Hua, G. (2017) · 2017
Closest in time.
Natural language does not emerge ’naturally’ in multi-agent dialog
Kottur, S., Moura, J. M., Lee, S., and Batra, D. (2017) · 2017
Closest in time.
Poseagent: Budget-constrained 6d object pose estimation via reinforcement learning
Krull, A., Brachmann, E., Nowozin, S., Michel, F., Shotton, J., and Rother, C. (2017) · 2017
Closest in time.
Playing FPS games with deep reinforcement learning
Lample, G. and Chaplot, D. S. (2017) · 2017
Closest in time.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Perolat, J., Silver, D., and Graepel, T. (2017) · 2017
Closest in time.
Learning visual servoing with deep features and trust region fitted Q-iteration
Lee, A. X., Levine, S., and Abbeel, P. (2017) · 2017
Closest in time.
Safe Mutations for Deep and Recurrent Neural Networks through Output Gradients
Lehman, J., Chen, J., Clune, J., and Stanley, K. O. (2017) · 2017
Closest in time.
Multi-agent reinforcement learning in sequential social dilemmas
Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., and Graepel, T. (2017) · 2017
Closest in time.
Deal or no deal? end-to-end learning for negotiation dialogues
Lewis, M., Yarats, D., Dauphin, Y. N., Parikh, D., and Batra, D. (2017) · 2017
Closest in time.
Learning to optimize
Li, K. and Malik, J. (2017) · 2017
Closest in time.
Learning to Optimize Neural Nets
Li, K. and Malik, J. (2017) · 2017
Closest in time.
Infogail: Interpretable imitation learning from visual demonstrations
Li, Y., Song, J., and Ermon, S. (2017) · 2017
Closest in time.
Ray rllib: A composable and scalable reinforcement learning library
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Gonzalez, J., Goldberg, K., and Stoica, I. (2017c) · 2017
Closest in time.
Stardata: A starcraft ai research dataset
Lin, Z., Gehring, J., Khalidov, V., and Synnaeve, G. (2017) · 2017
Closest in time.
Diagnostic inferencing via improving clinical concept extraction with deep reinforcement learning: A preliminary study
Ling, Y., Hasan, S. A., Datla, V., Qadir, A., Lee, K., Liu, J., and Farri, O. (2017) · 2017
Closest in time.
Designing the robot behavior for safe human robot interactions, in Trends in Control and Decision-Making for Human-Robot Collaboration Systems (Y. Wang and F. Zhang (Eds.))
Liu, C. and Tomizuka, M. (2017) · 2017
Closest in time.
Progressive Neural Architecture Search
Liu, C., Zoph, B., Shlens, J., Hua, W., Li, L.-J., Fei-Fei, L., Yuille, A., Huang, J., and Murphy, K. (2017) · 2017
Closest in time.
3DCNN-DQN-RNN: A deep reinforcement learning framework for semantic parsing of large-scale 3d point clouds
Liu, F., Li, S., Zhang, L., Zhou, C., Ye, R., Wang, Y., and Lu, J. (2017) · 2017
Closest in time.
Hierarchical Representations for Efficient Architecture Search
Liu, H., Simonyan, K., Vinyals, O., Fernando, C., and Kavukcuoglu, K. (2017) · 2017
Closest in time.
A hierarchical framework of cloud resource allocation and power management using deep reinforcement learning
Liu, N., Li, Z., Xu, Z., Xu, J., Lin, S., Qiu, Q., Tang, J., and Wang, Y. (2017) · 2017
Closest in time.
Unsupervised Sequence Classification using Sequential Output Statistics
Liu, Y., Chen, J., and Deng, L. (2017) · 2017
Closest in time.
Learning multiple tasks with multilinear relationship networks
Long, M., Cao, Z., Wang, J., and Yu, P. S. (2017) · 2017
Closest in time.
Deep Network Guided Proof Search
Loos, S., Irving, G., Szegedy, C., and Kaliszyk, C. (2017) · 2017
Closest in time.
Gradient Episodic Memory for Continuum Learning
Lopez-Paz, D. and Ranzato, M. (2017) · 2017
Closest in time.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I. (2017) · 2017
Closest in time.
A Laplacian framework for option discovery in reinforcement learning
Machado, M. C., Bellemare, M. G., and Bowling, M. (2017) · 2017
Closest in time.
Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2017) · 2017
Closest in time.
Towards Deep Learning Models Resistant to Adversarial Attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2017) · 2017
Closest in time.
Dex-Net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics
Mahler, J., Liang, J., Niyaz, S., Laskey, M., Doan, R., Liu, X., Aparicio Ojea, J., and Goldberg, K. (2017) · 2017
Closest in time.
Teacher-Student Curriculum Learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2017) · 2017
Closest in time.
Data-efficient reinforcement learning in continuous-state POMDPs
McAllister, R. and Rasmussen, C. E. (2017) · 2017
Closest in time.
Learned in Translation: Contextualized Word Vectors
McCann, B., Bradbury, J., Xiong, C., and Socher, R. (2017) · 2017
Closest in time.
On the State of the Art of Evaluation in Neural Language Models
Melis, G., Dyer, C., and Blunsom, P. (2017) · 2017
Closest in time.
Learning human behaviors from motion capture by adversarial imitation
Merel, J., Tassa, Y., TB, D., Srinivasan, S., Lemmon, J., Wang, Z., Wayne, G., and Heess, N. (2017) · 2017
Closest in time.
Dynamic safe interruptibility for decentralized multi-agent reinforcement learning
Mhamdi, E. M. E., Guerraoui, R., Hendrikx, H., and Maurer, A. (2017) · 2017
Closest in time.
Advances in Pre-Training Distributed Word Representations
Mikolov, T., Grave, E., Bojanowski, P., Puhrsch, C., and Joulin, A. (2017) · 2017
Closest in time.
Explanation in Artificial Intelligence: Insights from the Social Sciences
Miller, T. (2017) · 2017
Closest in time.
Deep learning for healthcare: review, opportunities and challenges
Miotto, R., Wang, F., Wang, S., Jiang, X., and Dudley, J. T. (2017) · 2017
Closest in time.
Device placement optimization with reinforcement learning
Mirhoseini, A., Pham, H., Le, Q. V., Steiner, B., Larsen, R., Zhou, Y., Kumar, N., and Mohammad Norouzi, Samy Bengio, J. D. (2017) · 2017
Closest in time.
Learning to navigate in complex environments
Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., Kumaran, D., and Hadsell, R. (2017) · 2017
Closest in time.
Neural Models for Information Retrieval
Mitra, B. and Craswell, N. (2017) · 2017
Closest in time.
Deep learning takes on translation
Monroe, D. (2017) · 2017
Closest in time.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisý, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M. (2017) · 2017
Closest in time.
Improving policy gradient by exploring under-appreciated rewards
Nachum, O., Norouzi, M., and Schuurmans, D. (2017) · 2017
Closest in time.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Closest in time.
Geometry of Optimization and Implicit Regularization in Deep Learning
Neyshabur, B., Tomioka, R., Salakhutdinov, R., and Srebro, N. (2017) · 2017
Closest in time.
Task-Oriented Query Reformulation with Reinforcement Learning
Nogueira, R. and Cho, K. (2017) · 2017
Closest in time.
PGQ: Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V. (2017) · 2017
Closest in time.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Closest in time.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Omidshafiei, S., Pazis, J., Amato, C., How, J. P., and Vian, J. (2017) · 2017
Closest in time.
Count-Based Exploration with Neural Density Models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R. (2017) · 2017
Closest in time.
Semi-supervised knowledge transfer for deep learning from private training data
Papernot, N., Abadi, M., Erlingsson, Ú., Goodfellow, I., and Talwar, K. (2017) · 2017
Closest in time.
Neuro-symbolic program synthesis
Parisotto, E., rahman Mohamed, A., Singh, R., Li, L., Zhou, D., and Kohli, P. (2017) · 2017
Closest in time.
Reinforced video captioning with entailment rewards
Pasunuru, R. and Bansal, M. (2017) · 2017
Closest in time.
A Deep Reinforced Model for Abstractive Summarization
Paulus, R., Xiong, C., and Socher, R. (2017) · 2017
Closest in time.
DeepXplore: Automated Whitebox Testing of Deep Learning Systems
Pei, K., Cao, Y., Yang, J., and Jana, S. (2017) · 2017
Closest in time.
C-learn: Learning geometric constraints from demonstrations for multi-step manipulation in shared autonomy
Pérez-D’Arpino, C. and Shah, J. A. (2017) · 2017
Closest in time.
A multi-agent reinforcement learning model of common-pool resource appropriation
Perolat, J., Leibo, J. Z., Zambaldi, V., Beattie, C., Tuyls, K., and Graepel, T. (2017) · 2017
Closest in time.
Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
Petroski Such, F., Madhavan, V., Conti, E., Lehman, J., Stanley, K. O., and Clune, J. (2017) · 2017
Closest in time.
Data-efficient Deep Reinforcement Learning for Dexterous Manipulation
Popov, I., Heess, N., Lillicrap, T., Hafner, R., Barth-Maron, G., Vecerik, M., Lampe, T., Tassa, Y., Erez, T., and Riedmiller, M. (2017) · 2017
Closest in time.
The intelligent industry of the future: A survey on emerging trends, research challenges and opportunities in industry 4.0
Preuveneers, D. and Ilie-Zudor, E. (2017) · 2017
Closest in time.
Neural Episodic Control
Pritzel, A., Uria, B., Srinivasan, S., Puigdomènech, A., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C. (2017) · 2017
Closest in time.
Learning to Generate Reviews and Discovering Sentiment
Radford, A., Jozefowicz, R., and Sutskever, I. (2017) · 2017
Closest in time.
Attend, adapt and transfer: Attentive deep architecture for adaptive transfer from multiple sources in the same domain
Rajendran, J., Lakshminarayanan, A., Khapra, M. M., P, P., and Ravindran, B. (2017) · 2017
Closest in time.
Attention-aware deep reinforcement learning for video face recognition
Rao, Y., Lu, J., and Zhou, J. (2017) · 2017
Closest in time.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H. (2017) · 2017
Closest in time.
Deep reinforcement learning-based image captioning with embedding reward
Ren, Z., Wang, X., Zhang, N., Lv, X., and Li, L.-J. (2017) · 2017
Closest in time.
Self-critical sequence training for image captioning
Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V. (2017) · 2017
Closest in time.
First-person activity forecasting with online inverse reinforcement learning
Rhinehart, N. and Kitani, K. M. (2017) · 2017
Closest in time.
End-to-end Differentiable Proving
Rocktäschel, T. and Riedel, S. (2017) · 2017
Closest in time.
An Overview of Multi-Task Learning in Deep Neural Networks
Ruder, S. (2017) · 2017
Closest in time.
Dynamic routing between capsules
Sabour, S., Frosst, N., and Hinton, G. E. (2017) · 2017
Closest in time.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Salimans, T., Ho, J., Chen, X., and Sutskever, I. (2017) · 2017
Closest in time.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D. G. T., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T. (2017) · 2017
Closest in time.
Equivalence Between Policy Gradients and Soft Q-Learning
Schulman, J., Abbeel, P., and Chen, X. (2017) · 2017
Closest in time.
Learning to Plan Chemical Syntheses
Segler, M. H. S., Preuss, M., and Waller, M. P. (2017) · 2017
Closest in time.
A Deep Reinforcement Learning Chatbot
Serban, I. V., Sankar, C., Germain, M., Zhang, S., Lin, Z., Subramanian, S., Kim, T., Pieper, M., Chandar, S., Ke, N. R., Mudumba, S., de Brebisson, A., Sotelo, J. M. R., Suhubdy, D., Michalski, V., Nguyen, A., Pineau, J., and Bengio, Y. (2017) · 2017
Closest in time.
Failures of gradient-based deep learning
Shalev-Shwartz, S., Shamir, O., and Shammah, S. (2017) · 2017
Closest in time.
Learning to repeat: Fine grained action repetition for deep reinforcement learning
Sharma, S., Lakshminarayanan, A. S., and Ravindran, B. (2017) · 2017
Closest in time.
Interactive learning for acquisition of grounded verb semantics towards human-robot communication
She, L. and Chai, J. (2017) · 2017
Closest in time.
Reasonet: Learning to stop reading in machine comprehension
Shen, Y., Huang, P.-S., Gao, J., and Chen, W. (2017) · 2017
Closest in time.
Learning from simulated and unsupervised images through adversarial training
Shrivastava, A., Pfister, T., Tuzel, O., Susskind, J., Wang, W., and Webb, R. (2017) · 2017
Closest in time.
Opening the Black Box of Deep Neural Networks via Information
Shwartz-Ziv, R. and Tishby, N. (2017) · 2017
Closest in time.
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2017) · 2017
Closest in time.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. (2017) · 2017
Closest in time.
A Brief Introduction to Machine Learning for Engineers
Simeone, O. (2017) · 2017
Closest in time.
Best Practices for Applying Deep Learning to Novel Applications
Smith, L. N. (2017) · 2017
Closest in time.
Federated multi-task learning
Smith, V., Chiang, C.-K., Sanjabi, M., and Talwalkar, A. (2017) · 2017
Closest in time.
Prototypical Networks for Few-shot Learning
Snell, J., Swersky, K., and Zemel, R. S. (2017) · 2017
Closest in time.
Machine Learning with World Knowledge: The Position and Survey
Song, Y. and Roth, D. (2017) · 2017
Closest in time.
Scalable and sustainable deep learning via randomized hashing
Spring, R. and Shrivastava, A. (2017) · 2017
Closest in time.
Third person imitation learning
Stadie, B. C., Abbeel, P., and Sutskever, I. (2017) · 2017
Closest in time.
A berkeley view of systems challenges for AI
Stoica, I., Song, D., Popa, R. A., Patterson, D. A., Mahoney, M. W., Katz, R. H., Joseph, A. D., Jordan, M., Hellerstein, J. M., Gonzalez, J., Goldberg, K., Ghodsi, A., Culler, D. E., and Abbeel, P. (2017) · 2017
Closest in time.
End-to-end optimization of goal-driven and visually grounded dialogue systems
Strub, F., de Vries, H., Mary, J., Piot, B., Courville, A., and Pietquin, O. (2017) · 2017
Closest in time.
Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning
Supančič, III, J. and Ramanan, D. (2017) · 2017
Closest in time.
Efficient Processing of Deep Neural Networks: A Tutorial and Survey
Sze, V., Chen, Y.-H., Yang, T.-J., and Emer, J. (2017) · 2017
Closest in time.
Exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P. (2017) · 2017
Closest in time.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S. (2017) · 2017
Closest in time.
ELF: An Extensive, Lightweight and Flexible Research Platform for Real-time Strategy Games
Tian, Y., Gong, Q., Shang, W., Wu, Y., and Zitnick, L. (2017) · 2017
Closest in time.
Ensemble Adversarial Training: Attacks and Defenses
Tramèr, F., Kurakin, A., Papernot, N., Boneh, D., and McDaniel, P. (2017) · 2017
Closest in time.
Deep probabilistic programming
Tran, D., Hoffman, M. D., Saurous, R. A., Brevdo, E., Murphy, K., and Blei, D. M. (2017) · 2017
Closest in time.
Episodic exploration for deep deterministic policies: An application to StarCraft micromanagement tasks
Usunier, N., Synnaeve, G., Lin, Z., and Chintala, S. (2017) · 2017
Closest in time.
Coordinated deep reinforcement learners for traffic light control
van der Pol, E. and Oliehoek, F. A. (2017) · 2017
Closest in time.
Hybrid reward architecture for reinforcement learning
van Seijen, H., Fatemi, M., Romoff, J., Laroche, R., Barnes, T., and Tsang, J. (2017) · 2017
Closest in time.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Closest in time.
Predictive-state decoders: Encoding the future into recurrent networks
Venkatraman, A., Rhinehart, N., Sun, W., Pinto, L., Hebert, M., Boots, B., Kitani, K. M., and Bagnell, J. A. (2017) · 2017
Closest in time.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Večerík, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M. (2017) · 2017
Closest in time.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Closest in time.
On the Origin of Deep Learning
Wang, H. and Raj, B. (2017) · 2017
Closest in time.
Robust Imitation of Diverse Behaviors
Wang, Z., Merel, J., Reed, S., Wayne, G., de Freitas, N., and Heess, N. (2017) · 2017
Closest in time.
Visual interaction networks: Learning a physics simulator from video
Watters, N., Tacchetti, A., Weber, T., Pascanu, R., Battaglia, P., and Zoran, D. (2017) · 2017
Closest in time.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Jimenez Rezende, D., Puigdomènech Badia, A., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P., Silver, D., and Wierstra, D. (2017) · 2017
Closest in time.
Sequence-to-Sequence Models Can Directly Transcribe Foreign Speech
Weiss, R. J., Chorowski, J., Jaitly, N., Wu, Y., and Chen, Z. (2017) · 2017
Closest in time.
Saliency-based sequential image attention with multiset prediction
Welleck, S., Mao, J., Cho, K., and Zhang, Z. (2017) · 2017
Closest in time.
A network-based end-to-end trainable task-oriented dialogue system
Wen, T.-H., Vandyke, D., Mrksic, N., Gasic, M., Rojas-Barahona, L. M., Su, P.-H., Ultes, S., and Young, S. (2017) · 2017
Closest in time.
Distral: Robust multitask reinforcement learning
Whye Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R. (2017) · 2017
Closest in time.
Hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning
Williams, J. D., Asadi, K., and Zweig, G. (2017) · 2017
Closest in time.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B. (2017) · 2017
Closest in time.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Liao, S., Grosse, R., and Ba, J. (2017) · 2017
Closest in time.
Training agent for first-person shooter game with actor-critic curriculum learning
Wu, Y. and Tian, Y. (2017) · 2017
Closest in time.
The Microsoft 2017 Conversational Speech Recognition System
Xiong, W., Wu, L., Alleva, F., Droppo, J., Huang, X., and Stolcke, A. (2017) · 2017
Closest in time.
Neural Task Programming: Learning to Generalize Across Hierarchical Tasks
Xu, D., Nair, S., Zhu, Y., Gao, J., Garg, A., Fei-Fei, L., and Savarese, S. (2017) · 2017
Closest in time.
Leveraging knowledge bases in lstms for improving machine reading
Yang, B. and Mitchell, T. (2017) · 2017
Closest in time.
Semi-supervised qa with generative domain-adaptive nets
Yang, Z., Hu, J., Salakhutdinov, R., and Cohen, W. W. (2017) · 2017
Closest in time.
Dualgan: Unsupervised dual learning for image-to-image translation
Yi, Z., Zhang, H., Tan, P., and Gong, M. (2017) · 2017
Closest in time.
Learning to compose words into sentences with reinforcement learning
Yogatama, D., Blunsom, P., Dyer, C., Grefenstette, E., and Ling, W. (2017) · 2017
Closest in time.
Recent Trends in Deep Learning Based Natural Language Processing
Young, T., Hazarika, D., Poria, S., and Cambria, E. (2017) · 2017
Closest in time.
Seqgan: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y. (2017) · 2017
Closest in time.
Action-decision networks for visual tracking with deep reinforcement learning
Yun, S., Choi, J., Yoo, Y., Yun, K., and Young Choi, J. (2017) · 2017
Closest in time.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Zagoruyko, S. and Komodakis, N. (2017) · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Closest in time.
Sentence simplification with deep reinforcement learning
Zhang, X. and Lapata, M. (2017) · 2017
Closest in time.
Practical Network Blocks Design with Q-Learning
Zhong, Z., Yan, J., and Liu, C.-L. (2017) · 2017
Closest in time.
Emotional Chatting Machine: Emotional Conversation Generation with Internal and External Memory
Zhou, H., Huang, M., Zhang, T., Zhu, X., and Liu, B. (2017) · 2017
Closest in time.
VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
Zhou, Y. and Tuzel, O. (2017) · 2017
Closest in time.
Deep forest: Towards an alternative to deep neural networks
Zhou, Z.-H. and Feng, J. (2017) · 2017
Closest in time.
Rules of Machine Learning: Best Practices for ML Engineering
Zinkevich, M. (2017) · 2017
Closest in time.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V. (2017) · 2017
Closest in time.
Learning Transferable Architectures for Scalable Image Recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V. (2017) · 2017
Closest in time.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Albrechta, S. V. and Stone, P. (2018) · 2018
Closest in time.
Machine learning in wireless sensor networks: Algorithms, strategies, and applications
Alsheikh, M. A., Lin, S., Niyato, D., and Tan, H.-P. (2014) · 2018
Closest in time.
Teaching a machine to read maps with deep reinforcement learning
Brunner, G., Richter, O., Wang, Y., and Wattenhofer, R. (2018) · 2018
Closest in time.
Multi-step reinforcement learning: A unifying algorithm
De Asis, K., Hernandez-Garcia, J. F., Zacharias Holland, G., and Sutton, R. S. (2018) · 2018
Closest in time.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S. (2018) · 2018
Closest in time.
Learning with options that terminate off-policy
Harutyunyan, A., Vrancx, P., Bacon, P.-L., Precup, D., and Nowe, A. (2018) · 2018
Closest in time.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018) · 2018
Closest in time.
Deep Q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Dulac-Arnold, G., Osband, I., Agapiou, J., Leibo, J. Z., and Gruslys, A. (2018) · 2018
Closest in time.
Psychlab: A Psychology Laboratory for Deep Reinforcement Learning Agents
Leibo, J. Z., de Masson d’Autume, C., Zoran, D., Amos, D., Beattie, C., Anderson, K., García Castañeda, A., Sanchez, M., Green, S., Gruslys, A., Legg, S., Hassabis, D., and Botvinick, M. M. (2018) · 2018
Closest in time.
Theoretical Impediments to Machine Learning With Seven Sparks from the Causal Revolution
Pearl, J. (2018) · 2018
Closest in time.
Reinforcement Learning: An Introduction (2nd Edition, in preparation)
Sutton, R. S. and Barto, A. G. (2018) · 2018
Closest in time.
DeepMind Control Suite
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., and Riedmiller, M. (2018) · 2018
Closest in time.
Artificial Intelligence and Games
Yannakakis, G. N. and Togelius, J. (2018) · 2018
Closest in time.
Deep Learning for Sentiment Analysis : A Survey
Zhang, L., Wang, S., and Liu, B. (2018) · 2018
Closest in time.
Visual interpretability for deep learning: a survey
Zhang, Q. and Zhu, S.-C. (2018) · 2018
Closest in time.