Fetching the paper…
Reading the bibliography…
We discuss deep reinforcement learning in an overview style.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
The transfer of learning
Ellis, H. C. (1965) · 1965
Earlier work this paper cites.
An Introduction to Artificial Intelligence
Bellman, R. (1978) · 1978
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W. (1983) · 1983
Earlier work this paper cites.
A theory of the learnable
Valiant, L. (1984) · 1984
Earlier work this paper cites.
Introduction to Artificial Intelligence Programming
Charniak, E. and McDermott, D. (1985) · 1985
Earlier work this paper cites.
Evolutionary principles in self-referential learning
Schmidhuber, J. (1987) · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Artificial Intelligence: The Very Idea
Haugeland, J. (1989) · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H. (1989) · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. (1990) · 1990
Earlier work this paper cites.
Learning a synaptic learning rule
Bengio, Y., Bengio, S., and Cloutier, J. (1991) · 1991
Earlier work this paper cites.
Artificial Intelligence
Rich, E. and Knight, K. (1991) · 1991
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J. (1991) · 1991
Earlier work this paper cites.
The Age of Intelligent Machines
Kurzweil, R. (1992) · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton, R. S. (1992) · 1992
Earlier work this paper cites.
Reinforcement learning is direct adaptive optimal control
Sutton, R. S., Barto, A. G., and Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Artificial Intelligence
Winston, P. H. (1992) · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P. (1993) · 1993
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E. (1993) · 1993
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, A. and Atkeson, C. (1993) · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M. (1993) · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning?
Littman, M. L. (1994) · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist sytems
Rummery, G. A. and Niranjan, M. (1994) · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G. (1994) · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Multitask learning
Caruana, R. (1997) · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Investment Science
Luenberger, D. G. (1997) · 1997
Earlier work this paper cites.
Machine Learning
Mitchell, T. (1997) · 1997
Earlier work this paper cites.
Enhancing Q-learning for optimal asset allocation
Neuneier, R. (1997) · 1997
Earlier work this paper cites.
One Jump Ahead
Schaeffer, J. (1997) · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Artificial Intelligence: A New Synthesis
Nilsson, N. J. (1998) · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. (1998) · 1998
Earlier work this paper cites.
Computational Intelligence: A Logical Approach
Poole, D., Mackworth, A., and Goebel, R. (1998) · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Learning to Learn
Thrun, S. and Pratt, L., editors (1998) · 1998
Earlier work this paper cites.
Statistical Learning Theory
Vapnik, V. N. (1998) · 1998
Earlier work this paper cites.
How machines have learned to play othello
Buro, M. (1999) · 1999
Earlier work this paper cites.
Technical Analysis of the Financial Markets: A Comprehensive Guide to Trading Methods and Applications
Murphy, J. J. (1999) · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G. (2000) · 2000
Earlier work this paper cites.
Foundations of technical analysis: Computational algorithms, statistical inference, and empirical implementation
Lo, A. W., Mamaysky, H., and Wang, J. (2000) · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. and Russell, S. (2000) · 2000
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J. (2000) · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, S., Jaakkola, T., Littman, M. L., and Szepesvári, C. (2000) · 2000
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
Stone, P. and Veloso, M. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Relational reinforcement learning
Džeroski, S., Raedt, L. D., and Driessens, K. (2001) · 2001
Earlier work this paper cites.
GIB: Imperfect information in a computationally challenging game
Ginsberg, M. L. (2001) · 2001
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A. (2001) · 2001
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R. (2001) · 2001
Earlier work this paper cites.
Predictive representations of state
Littman, M. L., Sutton, R. S., and Singh, S. (2001) · 2001
Earlier work this paper cites.
Valuing American options by simulation: a simple least-squares approach
Longstaff, F. A. and Schwartz, E. S. (2001) · 2001
Earlier work this paper cites.
Learning to trade via direct reinforcement
Moody, J. and Saffell, M. (2001) · 2001
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S. (2001) · 2001
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond
Schölkopf, B. and Smola, A. J. (2001) · 2001
Earlier work this paper cites.
Regression methods for pricing complex American-style options
Tsitsiklis, J. N. and Van Roy, B. (2001) · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. (2002) · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E. (2002) · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Bowling, M. and Veloso, M. (2002) · 2002
Earlier work this paper cites.
R-MAX - a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2002) · 2002
Earlier work this paper cites.
Strategic Asset Allocation Portfolio Choice for Long-Term Investors
Campbell, J. Y. and Viceira, L. M. (2002) · 2002
Earlier work this paper cites.
Deep blue
Campbell, M., Hoane, A. J., and Hsu, F. (2002) · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. (2002) · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Computer go
Müller, M. (2002) · 2002
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D. and Sen, Ś. (2002) · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S. (2003) · 2003
Earlier work this paper cites.
Generalizing plans to new environments in relational MDPs
Guestrin, C., Koller, D., Gearhart, C., and Kanodia, N. (2003) · 2003
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Hu, J. and Wellman, M. P. (2003) · 2003
Earlier work this paper cites.
On actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2003) · 2003
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
Least squares policy evaluation algorithms with linear function approximation
Nedić, A. and Bertsekas, D. P. (2003) · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Convex Optimization
Boyd, S. and Vandenberghe, L. (2004) · 2004
Earlier work this paper cites.
Knowledge Representation and Reasoning
Brachman, R. and Levesque, H. (2004) · 2004
Earlier work this paper cites.
Monte Carlo Methods in Financial Engineering
Glasserman, P. (2004) · 2004
Earlier work this paper cites.
A survey of multi-agent organizational paradigms
Horling, B. and Lesser, V. (2004) · 2004
Earlier work this paper cites.
The Adaptive Markets Hypothesis: Market efficiency from an evolutionary perspective
Lo, A. W. (2004) · 2004
Earlier work this paper cites.
Autonomous helicopter flight via reinforcement learning
Ng, A. Y., Kim, H. J., Jordan, M. I., and Sastry, S. (2004) · 2004
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Singh, S., James, M., and Rudary, M. (2004) · 2004
Earlier work this paper cites.
Temporal-difference networks
Sutton, R. S. and Tanner, B. (2004) · 2004
Earlier work this paper cites.
Relational reinforcement learning: An overview
Tadepalli, P., Givan, R., and Driessens, K. (2004) · 2004
Earlier work this paper cites.
A simulation approach to dynamic portfolio choice with an application to learning about return predictability
Brandt, M. W., Goyal, A., Santa-Clara, P., and Stroud, J. R. (2005) · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L. (2005) · 2005
Earlier work this paper cites.
Cognitive radio: brain-empowered wireless communications
Haykin, S. (2005) · 2005
Earlier work this paper cites.
Universal Artificial Intelligence: Sequential Decisions Based on Algorithmic Probability
Hutter, M. (2005) · 2005
Earlier work this paper cites.
Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. (2005) · 2005
Earlier work this paper cites.
Temporal abstraction in temporal-difference networks
Sutton, R. S., Rafols, E. J., and Koop, A. (2005) · 2005
Earlier work this paper cites.
Model compression
Bucila, C., Caruana, R., and Niculescu-Mizil, A. (2006) · 2006
Earlier work this paper cites.
Settling the complexity of two-player nash equilibrium
Chen, X. and Deng, X. (2006) · 2006
Earlier work this paper cites.
Hierarchical multi-agent reinforcement learning
Ghavamzadeh, M., Mahadevan, S., and Makar, R. (2006) · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R. (2006) · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Rasmussen, C. E. and Williams, C. K. I. (2006) · 2006
Earlier work this paper cites.
Handbook of constraint programming
Rossi, F., Beek, P. V., and Walsh, T., editors (2006) · 2006
Earlier work this paper cites.
Dynamic movement primitives -a framework for motor control in humans and humanoid robotics
Schaal, S. (2006) · 2006
Earlier work this paper cites.
Combining online and offline knowledge in UCT
Gelly, S. and Silver, D. (2007) · 2007
Earlier work this paper cites.
Introduction to Statistical Relational Learning
Getoor, L. and Taskar, B., editors (2007) · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J. and Zhang, T. (2007) · 2007
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Mahadevan, S. and Maggioni, M. (2007) · 2007
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, P.-Y. and Kaplan, F. (2007) · 2007
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S. (2007) · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B. (2007) · 2007
Earlier work this paper cites.
Checkers is solved
Schaeffer, J., Burch, N., Björnsson, Y., Kishimoto, A., Müller, M., Lake, R., Lu, P., and Sutphen, S. (2007) · 2007
Earlier work this paper cites.
If multi-agent learning is the answer, what is the question?
Shoham, Y., Powers, R., and Grenager, T. (2007) · 2007
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E. (2007) · 2007
Earlier work this paper cites.
Interactive storytelling: A player modelling approach
Thue, D., Bulitko, V., Spetch, M., and Wasylishen, E. (2007) · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L., Babuska, R., and Schutter, B. D. (2008) · 2008
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Diuk, C., Cohen, A., and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Essentials of Game Theory: A Concise, Multidisciplinary Introduction
Leyton-Brown, K. and Shoham, Y. (2008) · 2008
Earlier work this paper cites.
An Introduction to Kolmogorov Complexity and Its Applications (3rd edition)
Li, M. and Vitányi, P. (2008) · 2008
Earlier work this paper cites.
Introduction to Information Retrieval
Manning, C. D., Raghavan, P., and Schütze, H. (2008) · 2008
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
Oliehoek, F. A., Spaan, M. T. J., and Vlassis, N. (2008) · 2008
Earlier work this paper cites.
Probabilistic inductive logic programming: theory and applications
Raedt, L. D., Frasconi, P., Kersting, K., and Muggleton, S., editors (2008) · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Apprenticeship learning using linear programming
Syed, U., Bowling, M., and Schapire, R. E. (2008) · 2008
Earlier work this paper cites.
Visualizing data using t-SNE
van der Maaten, L. and Hinton, G. (2008) · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J., and Dey, A. K. (2008) · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B. (2009) · 2009
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
Metalearning: Applications to Data Mining
Brazdil, P., Carrier, C. G., Soares, C., and Vilalta, R. (2009) · 2009
Earlier work this paper cites.
Anomaly detection : A survey
Chandola, V., Banerjee, A., and Kumar, V. (2009) · 2009
Earlier work this paper cites.
Search-based structured prediction
Daumé, III, H., Langford, J., and Marcu, D. (2009) · 2009
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Hastie, T., Tibshirani, R., and Friedman, J. (2009) · 2009
Earlier work this paper cites.
Probabilistic Graphical Models: Principles and Techniques
Koller, D. and Friedman, N. (2009) · 2009
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, J. Z. and Ng, A. Y. (2009) · 2009
Earlier work this paper cites.
Learning exercise policies for American options
Li, Y., Szepesvári, C., and Schuurmans, D. (2009) · 2009
Earlier work this paper cites.
Predictive systems: Living with imperfect predictors
Pastor, L. and Stambaugh, R. F. (2009) · 2009
Earlier work this paper cites.
Causality
Pearl, J. (2009) · 2009
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach (3rd edition)
Russell, S. and Norvig, P. (2009) · 2009
Earlier work this paper cites.
The graph neural network model
Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., and Monfardini, G. (2009) · 2009
Earlier work this paper cites.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Shoham, Y. and Leyton-Brown, K. (2009) · 2009
Earlier work this paper cites.
Where do rewards come from?
Singh, S., Lewis, R. L., and Barto, A. G. (2009) · 2009
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Strehl, A. L., Li, L., and Littman, M. L. (2009) · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P. (2009) · 2009
Earlier work this paper cites.
A theoretical and empirical analysis of expected sarsa
van Seijen, H., van Hasselt, H., Whiteson, S., and Wiering, M. (2009) · 2009
Earlier work this paper cites.
A general projection property for distribution families
Yu, Y.-L., Li, Y., Szepesvári, C., and Schuurmans, D. (2009) · 2009
Earlier work this paper cites.
Introduction to semi-supervised learning
Zhu, X. and Goldberg, A. B. (2009) · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, Y. A. (2010) · 2010
Earlier work this paper cites.
Consumer credit-risk models via machine-learning algorithms
Khandani, A. E., Kim, A. J., and Lo, A. W. (2010) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q. (2010) · 2010
Earlier work this paper cites.
A Comprehensive Survey of Data Mining-based Fraud Detection Research
Phua, C., Lee, V., Smith, K., and Gayler, R. (2010) · 2010
Earlier work this paper cites.
Merging AI and OR to solve high-dimensional stochastic optimization problems using approximate dynamic programming
Powell, W. B. (2010) · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G. J., and Bagnell, J. A. (2010) · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990-2010)
Schmidhuber, J. (2010) · 2010
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Singh, S., Lewis, R., Barto, A., and Sorg, J. (2010) · 2010
Earlier work this paper cites.
Outside the closed world: On using machine learning for network intrusion detection
Sommer, R. and Paxson, V. (2010) · 2010
Earlier work this paper cites.
A reduction from apprenticeship learning to classification
Syed, U. and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Szepesvári, C. (2010) · 2010
Earlier work this paper cites.
Double Q-learning
van Hasselt, H. (2010) · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Ziebart, B. D., Bagnell, J. A., and Dey, A. K. (2010) · 2010
Earlier work this paper cites.
Adaptive stochastic control for the smart grid
Anderson, R. N., Boulanger, A., Powell, W. B., and Scott, W. (2011) · 2011
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Bishop, C. (2011) · 2011
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M. P. and Rasmussen, C. E. (2011) · 2011
Earlier work this paper cites.
Data Mining: Concepts and Techniques (3rd edition)
Han, J., Kamber, M., and Pei, J. (2011) · 2011
Earlier work this paper cites.
Thinking, Fast and Slow
Kahneman, D. (2011) · 2011
Earlier work this paper cites.
Knows what it knows: a framework for self-aware learning
Li, L., Littman, M. L., Walsh, T. J., and Strehl, A. L. (2011) · 2011
Earlier work this paper cites.
Approximate Dynamic Programming: Solving the curses of dimensionality (2nd Edition)
Powell, W. B. (2011) · 2011
Earlier work this paper cites.
Informing sequential clinical decision-making through reinforcement learning: an empirical study
Shortreed, S. M., Laber, E., Lizotte, D. J., Stroup, T. S., Pineau, J., and Murphy, S. A. (2011) · 2011
Earlier work this paper cites.
Semi-supervised recursive autoencoders for predicting sentiment distributions
Socher, R., Pennington, J., Huang, E. H., Ng, A. Y., and Manning, C. D. (2011) · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction, , proc. of 10th
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Earlier work this paper cites.
Dynamic programming and optimal control (Vol. II, 4th Edition: Approximate Dynamic Programming)
Bertsekas, D. P. (2012) · 2012
Earlier work this paper cites.
Learning to win by reading manuals in a monte-carlo framework
Branavan, S. R. K., Silver, D., and Barzilay, R. (2012) · 2012
Earlier work this paper cites.
A survey of Monte Carlo tree search methods
Browne, C., Powley, E., Whitehouse, D., Lucas, S., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S. (2012) · 2012
Earlier work this paper cites.
Tractable objectives for robust policy optimization
Chen, K. and Bowling, M. (2012) · 2012
Earlier work this paper cites.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S. (2012) · 2012
Earlier work this paper cites.
A few useful things to know about machine learning
Domingos, P. (2012) · 2012
Earlier work this paper cites.
Smart grid - the new and improved power grid: A survey
Fang, X., Misra, S., Xue, G., and Yang, D. (2012) · 2012
Earlier work this paper cites.
The grand challenge of computer go: Monte carlo tree search and extensions
Gelly, S., Schoenauer, M., Sebag, M., Teytaud, O., Kocsis, L., Silver, D., and Szepesvári, C. (2012) · 2012
Earlier work this paper cites.
Q-learning with censored data
Goldberg, Y. and Kosorok, M. R. (2012) · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
Hinton, G., Deng, L., Yu, D., Dahl, G. E., rahman Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., , and Kingsbury, B. (2012) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Le, Q. V., Ranzato, M., Monga, R., Devin, M., Chen, K., Corrado, G. S., Dean, J., and Ng, A. Y. (2012) · 2012
Earlier work this paper cites.
Sample complexity bounds of exploration
Li, L. (2012) · 2012
Earlier work this paper cites.
Sentiment Analysis and Opinion Mining
Liu, B. (2012) · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Murphy, K. P. (2012) · 2012
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, K., Toussaint, M., and Vijayakumar, S. (2012) · 2012
Earlier work this paper cites.
Temporal-difference search in computer Go
Silver, D., Sutton, R. S., and Müller, M. (2012) · 2012
Earlier work this paper cites.
Artist agent: A reinforcement learning approach to automatic stroke generation in oriental ink painting
Xie, N., Hachiya, H., and Sugiyama, M. (2012) · 2012
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Barto, A. (2013) · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P. (2013) · 2013
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., and Peters, J. (2013) · 2013
Earlier work this paper cites.
Machine learning paradigms for speech recognition: An overview
Deng, L. and Li, X. (2013) · 2013
Earlier work this paper cites.
Multiagent reinforcement learning for integrated network of adaptive traffic signal controllers (marlin-atsc): methodology and large-scale application on downtown toronto
El-Tantawy, S., Abdulhai, B., and Abdelgawad, H. (2013) · 2013
Earlier work this paper cites.
Learning and reasoning in cognitive radio networks
Gavrilovska, L., Atanasovski, V., Macaluso, I., and DaSilva, L. A. (2013) · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., rahman Mohamed, A., and Hinton, G. (2013) · 2013
Earlier work this paper cites.
Speech-centric information processing: An optimization-oriented approach
He, X. and Deng, L. (2013) · 2013
Earlier work this paper cites.
Mohex 2.0: A pattern-based mcts hex player
Huang, S.-C., Arneson, B., Hayward, R. B., and Müller, M. (2013) · 2013
Earlier work this paper cites.
An Introduction to Statistical Learning with Applications in R
James, G., Witten, D., Hastie, T., and Tibshirani, R. (2013) · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Kalchbrenner, N. and Blunsom, P. (2013) · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Earlier work this paper cites.
Applied Predictive Modeling
Kuhn, M. and Johnson, K. (2013) · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Earlier work this paper cites.
A survey of real-time strategy game ai research and competition in starcraft
Ontañón, S., Synnaeve, G., Uriarte, A., Richoux, F., Churchill, D., and Preuss, M. (2013) · 2013
Earlier work this paper cites.
Data Science for Business
Provost, F. and Fawcett, T. (2013) · 2013
Earlier work this paper cites.
Concurrent reinforcement learning from customer interactions
Silver, D., Newnham, L., Barker, D., Weller, S., and McFall, J. (2013) · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment tree- bank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C., Ng, A., and Potts, C. (2013) · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2013) · 2013
Earlier work this paper cites.
POMDP-based statistical spoken dialogue systems: a review
Young, S., Gašić, M., Thomson, B., and Williams, J. D. (2013) · 2013
Earlier work this paper cites.
Multiple object recognition with visual attention
Ba, J., Mnih, V., and Kavukcuoglu, K. (2014) · 2014
Earlier work this paper cites.
Introduction to Intelligent Systems in Traffic and Transportation
Bazzan, A. L. and Klügl, F. (2014) · 2014
Earlier work this paper cites.
From machine learning to machine reasoning
Bottou, L. (2014) · 2014
Earlier work this paper cites.
Dynamic treatment regimes
Chakraborty, B. and Murphy, S. A. (2014) · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma, M. W. (2014) · 2014
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., and Darrell, T. (2014) · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., , and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Neural Turing Machines
Graves, A., Wayne, G., and Danihelka, I. (2014) · 2014
Earlier work this paper cites.
Bayes-adaptive simulation-based search with value function approximation
Guez, A., Heess, N., Silver, D., and Dayan, P. (2014) · 2014
Earlier work this paper cites.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X. (2014) · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2014) · 2014
Earlier work this paper cites.
Options, Futures and Other Derivatives (9th edition)
Hull, J. C. (2014) · 2014
Earlier work this paper cites.
Learning from limited demonstrations
Kim, B., Farahmand, A.-m., Pineau, J., and Precup, D. (2014) · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Kingma, D. P., Rezende, D. J., Mohamed, S., and Welling, M. (2014) · 2014
Earlier work this paper cites.
Learning complex neural network policies with trajectory optimization
Levine, S. and Koltun, V. (2014) · 2014
Earlier work this paper cites.
Trading off scientific knowledge and user learning with multi-armed bandits
Liu, Y.-E., Mandel, T., Brunskill, E., and Popović, Z. (2014) · 2014
Earlier work this paper cites.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A. R., van Hasselt, H., and Sutton, R. S. (2014) · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
Mandel, T., Liu, Y. E., Levine, S., Brunskill, E., and Popović, Z. (2014) · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Mnih, V., Heess, N., Graves, A., and Kavukcuoglu, K. (2014) · 2014
Earlier work this paper cites.
A $3 trillion challenge to computational scientists: Transforming healthcare delivery
Saria, S. (2014) · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Earlier work this paper cites.
TORCS, The Open Racing Car Simulator
Wymann, B., Espié, E., Guionneau, C., Dimitrakakis, C., and Rémi Coulom, A. S. (2014) · 2014
Earlier work this paper cites.
Internet of things in industries: A survey
Xu, L. D., He, W., and Li, S. (2014) · 2014
Earlier work this paper cites.
Universal option models
Yao, H., Szepesvari, C., Sutton, R. S., Modayil, J., and Bhatnagar, S. (2014) · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014) · 2014
Earlier work this paper cites.
Model-based relative entropy stochastic search
Abdolmaleki, A., Lioutikov, R., Lau, N., Reis, L. P., Peters, J., and Neumann, G. (2015) · 2015
Earlier work this paper cites.
Maximum entropy semi-supervised inverse reinforcement learning
Audiffren, J., Valko, M., Lazaric, A., and Ghavamzadeh, M. (2015) · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O. (2015) · 2015
Earlier work this paper cites.
Active object localization with deep reinforcement learning
Caicedo, J. C. and Lazebnik, S. (2015) · 2015
Earlier work this paper cites.
Natural Language Understanding with Distributed Representation
Cho, K. (2015) · 2015
Earlier work this paper cites.
The Master Algorithm: How the Quest for the Ultimate Learning Machine Will Remake Our World
Domingos, P. (2015) · 2015
Earlier work this paper cites.
Multi-task learning for multiple language translation
Dong, D., Wu, H., He, W., Yu, D., and Wang, H. (2015) · 2015
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S. (2015) · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcìa, J. and Fernàndez, F. (2015) · 2015
Earlier work this paper cites.
Bayesian reinforcement learning: a survey
Ghavamzadeh, M., Mannor, S., Pineau, J., and Tamar, A. (2015) · 2015
Earlier work this paper cites.
Fast R-CNN
Girshick, R. (2015) · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Draw: A recurrent neural network for image generation
Gregor, K., Danihelka, I., Graves, A., Rezende, D., and Wierstra, D. (2015) · 2015
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable MDPs
Hausknecht, M. and Stone, P. (2015) · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Tassa, Y., and Erez, T. (2015) · 2015
Earlier work this paper cites.
Advances in natural language processing
Hirschberg, J. and Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Spatial transformer networks
Jaderberg, M., Simonyan, K., Zisserman, A., and Kavukcuoglu, K. (2015) · 2015
Earlier work this paper cites.
Machine learning: Trends, perspectives, and prospects
Jordan, M. I. and Mitchell, T. (2015) · 2015
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Koch, G., Zemel, R., and Salakhutdinov, R. (2015) · 2015
Earlier work this paper cites.
Adaptive Treatment Strategies in Practice: Planning Trials and Analyzing Data for Personalized Medicine
Kosorok, M. R. and Moodie, E. E. M. (2015) · 2015
Earlier work this paper cites.
Deep convolutional inverse graphics network
Kulkarni, T. D., Whitney, W., Kohli, P., and Tenenbaum, J. B. (2015) · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B. (2015) · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Earlier work this paper cites.
DeepMPC: Learning deep latent features for model predictive control
Lenz, I., Knepper, R., and Saxena, A. (2015) · 2015
Earlier work this paper cites.
Reinforcement learning improves behaviour from evaluative feedback
Littman, M. L. (2015) · 2015
Earlier work this paper cites.
Learning transferable features with deep adaptation networks
Long, M., Cao, Y., Wang, J., and Jordan, M. I. (2015) · 2015
Earlier work this paper cites.
Using recurrent neural networks for slot filling in spoken language understanding
Mesnil, G., Dauphin, Y., Yao, K., Bengio, Y., Deng, L., He, X., Heck, L., Tur, G., Hakkani-Tür, D., Yu, D., and Zweig, G. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., De Maria, A., Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., Legg, S., Mnih, V., Kavukcuoglu, K., and Silver, D. (2015) · 2015
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, K., Kulkarni, T., and Barzilay, R. (2015) · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R., and Singh, S. (2015) · 2015
Earlier work this paper cites.
Is object localization for free? – weakly-supervised learning with convolutional neural networks
Oquab, M., Bottou, L., Laptev, I., and Sivic, J. (2015) · 2015
Earlier work this paper cites.
Economic reasoning and artificial intelligence
Parkes, D. C. and Wellman, M. P. (2015) · 2015
Earlier work this paper cites.
Policy search: Methods and applications
Peters, J. and Neumann, G. (2015) · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J. (2015) · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Schmidhuber, J. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Earlier work this paper cites.
UCL reinforcement learning course
Silver, D. (2015) · 2015
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Weston, J., and Fergus, R. (2015) · 2015
Earlier work this paper cites.
Personalized ad recommendation systems for life-time value optimization with guarantees
Theocharous, G., Thomas, P. S., and Ghavamzadeh, M. (2015) · 2015
Earlier work this paper cites.
Pointer networks
Vinyals, O., Fortunato, M., and Jaitly, N. (2015) · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M. (2015) · 2015
Earlier work this paper cites.
Optimal demand response using device-based reinforcement learning
Wen, Z., O’Neill, D., and Maei, H. (2015) · 2015
Earlier work this paper cites.
Memory networks
Weston, J., Chopra, S., and Bordes, A. (2015) · 2015
Earlier work this paper cites.
Galileo: Perceiving physical object properties by integrating a physics engine with deep learning
Wu, J., Yildirim, I., Lim, J. J., Freeman, B., and Tenenbaum, J. (2015) · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J. L., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R. S., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Reinforcement Learning Neural Turing Machines - Revised
Zaremba, W. and Sutskever, I. (2015) · 2015
Earlier work this paper cites.
Object detectors emerge in deep scene CNNs
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A. (2015) · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, P., Nair, A., Abbeel, P., Malik, J., and Levine, S. (2016) · 2016
Cited alongside, same era.
Concrete Problems in AI Safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D. (2016) · 2016
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Colmenarejo, S. G., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and de Freitas, N. (2016) · 2016
Cited alongside, same era.
A sequence-to-sequence model for user simulation in spoken dialogue systems
Asri, L. E., He, J., and Suleman, K. (2016) · 2016
Cited alongside, same era.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C. (2016) · 2016
Cited alongside, same era.
Layer Normalization
Ba, J., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Transferring context-dependent test inputs
Reichstaller, A. and Knapp, A. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning-based image captioning with embedding reward
Ren, Z., Wang, X., Zhang, N., Lv, X., and Li, L.-J. (2017) · 2017
Later among the works it cites.
Self-critical sequence training for image captioning
Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V. (2017) · 2017
Later among the works it cites.
First-person activity forecasting with online inverse reinforcement learning
Rhinehart, N. and Kitani, K. M. (2017) · 2017
Later among the works it cites.
End-to-end Differentiable Proving
Rocktäschel, T. and Riedel, S. (2017) · 2017
Later among the works it cites.
Learning a health knowledge graph from electronic medical records
Rotmensch, M., Halpern, Y., Tlimat, A., Horng, S., and Sontag, D. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Interaction networks for learning about objects, relations and physics
Battaglia, P. W., Pascanu, R., Lai, M., Rezende, D., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
DeepMind Lab
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S. (2016) · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Schaul, T., Srinivasan, S., Saxton, D., Ostrovski, G., and Munos, R. (2016) · 2016
Cited alongside, same era.
Neural Combinatorial Optimization with Reinforcement Learning
Bello, I., Pham, H., Le, Q. V., Norouzi, M., and Bengio, S. (2016) · 2016
Cited alongside, same era.
Playing Doom with SLAM-Augmented Deep Reinforcement Learning
Bhatti, S., Desmaison, A., Miksik, O., Nardelli, N., Siddharth, N., and Torr, P. H. S. (2016) · 2016
Cited alongside, same era.
End to End Learning for Self-Driving Cars
Bojarski, M., Testa, D. D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L. D., Monfort, M., Muller, U., Zhang, J., Zhang, X., Zhao, J., and Zieba, K. (2016) · 2016
Cited alongside, same era.
An Overview of Multi-Task Learning in Deep Neural Networks
Ruder, S. (2017) · 2017
Later among the works it cites.
Sim-to-real robot learning from pixels with progressive nets
Rusu, A. A., Vecerik, M., Rothörl, T., Heess, N., Pascanu, R., and Hadsell, R. (2017) · 2017
Later among the works it cites.
Dynamic routing between capsules
Sabour, S., Frosst, N., and Hinton, G. E. (2017) · 2017
Later among the works it cites.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Salimans, T., Ho, J., Chen, X., and Sutskever, I. (2017) · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D. G. T., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T. (2017) · 2017
Later among the works it cites.
The nuts and bolts of deep reinforcement learning research
Schulman, J. (2017) · 2017
Later among the works it cites.
Learning to repeat: Fine grained action repetition for deep reinforcement learning
Sharma, S., Lakshminarayanan, A. S., and Ravindran, B. (2017) · 2017
Later among the works it cites.
Interactive learning for acquisition of grounded verb semantics towards human-robot communication
She, L. and Chai, J. (2017) · 2017
Later among the works it cites.
Reasonet: Learning to stop reading in machine comprehension
Shen, Y., Huang, P.-S., Gao, J., and Chen, W. (2017) · 2017
Later among the works it cites.
Learning from simulated and unsupervised images through adversarial training
Shrivastava, A., Pfister, T., Tuzel, O., Susskind, J., Wang, W., and Webb, R. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning with subgoals
Silver, D. (2017) · 2017
Later among the works it cites.
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2017) · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. (2017) · 2017
Later among the works it cites.
Steps towards continual learning
Singh, S. (2017) · 2017
Later among the works it cites.
Federated multi-task learning
Smith, V., Chiang, C.-K., Sanjabi, M., and Talwalkar, A. (2017) · 2017
Later among the works it cites.
Prototypical Networks for Few-shot Learning
Snell, J., Swersky, K., and Zemel, R. S. (2017) · 2017
Later among the works it cites.
Third person imitation learning
Stadie, B. C., Abbeel, P., and Sutskever, I. (2017) · 2017
Later among the works it cites.
A berkeley view of systems challenges for AI
Stoica, I., Song, D., Popa, R. A., Patterson, D. A., Mahoney, M. W., Katz, R. H., Joseph, A. D., Jordan, M., Hellerstein, J. M., Gonzalez, J., Goldberg, K., Ghodsi, A., Culler, D. E., and Abbeel, P. (2017) · 2017
Later among the works it cites.
End-to-end optimization of goal-driven and visually grounded dialogue systems
Strub, F., de Vries, H., Mary, J., Piot, B., Courville, A., and Pietquin, O. (2017) · 2017
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T. (2017) · 2017
Later among the works it cites.
Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning
Supančič, III, J. and Ramanan, D. (2017) · 2017
Later among the works it cites.
Exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P. (2017) · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S. (2017) · 2017
Later among the works it cites.
ELF: An extensive, lightweight and flexible research platform for real-time strategy games
Tian, Y., Gong, Q., Shang, W., Wu, Y., and Zitnick, L. (2017) · 2017
Later among the works it cites.
Episodic exploration for deep deterministic policies: An application to StarCraft micromanagement tasks
Usunier, N., Synnaeve, G., Lin, Z., and Chintala, S. (2017) · 2017
Later among the works it cites.
Coordinated deep reinforcement learners for traffic light control
van der Pol, E. and Oliehoek, F. A. (2017) · 2017
Later among the works it cites.
Hybrid reward architecture for reinforcement learning
van Seijen, H., Fatemi, M., Romoff, J., Laroche, R., Barnes, T., and Tsang, J. (2017) · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Predictive-state decoders: Encoding the future into recurrent networks
Venkatraman, A., Rhinehart, N., Sun, W., Pinto, L., Hebert, M., Boots, B., Kitani, K. M., and Bagnell, J. A. (2017) · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Večerík, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M. (2017) · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
StarCraft II: A New Challenge for Reinforcement Learning
Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Sasha Vezhnevets, A., Yeo, M., Makhzani, A., Küttler, H., Agapiou, J., Schrittwieser, J., Quan, J., Gaffney, S., Petersen, S., Simonyan, K., Schaul, T., van Hasselt, H., Silver, D., Lillicrap, T., Calderone, K., Keet, P., Brunasso, A., Lawrence, D., Ekermo, A., Repp, J., and Tsing, R. (2017) · 2017
Later among the works it cites.
Robust Imitation of Diverse Behaviors
Wang, Z., Merel, J., Reed, S., Wayne, G., de Freitas, N., and Heess, N. (2017) · 2017
Later among the works it cites.
Visual interaction networks: Learning a physics simulator from video
Watters, N., Tacchetti, A., Weber, T., Pascanu, R., Battaglia, P., and Zoran, D. (2017) · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Jimenez Rezende, D., Puigdomènech Badia, A., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P., Silver, D., and Wierstra, D. (2017) · 2017
Later among the works it cites.
Saliency-based sequential image attention with multiset prediction
Welleck, S., Mao, J., Cho, K., and Zhang, Z. (2017) · 2017
Later among the works it cites.
A network-based end-to-end trainable task-oriented dialogue system
Wen, T.-H., Vandyke, D., Mrksic, N., Gasic, M., Rojas-Barahona, L. M., Su, P.-H., Ultes, S., and Young, S. (2017) · 2017
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Whye Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R. (2017) · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
Wichrowska, O., Maheswaranathan, N., Hoffman, M. W., Gomez Colmenarejo, S., Denil, M., de Freitas, N., and Sohl-Dickstein, J. (2017) · 2017
Later among the works it cites.
Hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning
Williams, J. D., Asadi, K., and Zweig, G. (2017) · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Liao, S., Grosse, R., and Ba, J. (2017) · 2017
Later among the works it cites.
Training agent for first-person shooter game with actor-critic curriculum learning
Wu, Y. and Tian, Y. (2017) · 2017
Later among the works it cites.
The Microsoft 2017 Conversational Speech Recognition System
Xiong, W., Wu, L., Alleva, F., Droppo, J., Huang, X., and Stolcke, A. (2017) · 2017
Later among the works it cites.
Collective robot reinforcement learning with distributed asynchronous guided policy search
Yahya, A., Li, A., Kalakrishnan, M., Chebotar, Y., and Levine, S. (2017) · 2017
Later among the works it cites.
Leveraging knowledge bases in LSTMs for improving machine reading
Yang, B. and Mitchell, T. (2017) · 2017
Later among the works it cites.
Semi-supervised qa with generative domain-adaptive nets
Yang, Z., Hu, J., Salakhutdinov, R., and Cohen, W. W. (2017) · 2017
Later among the works it cites.
Dualgan: Unsupervised dual learning for image-to-image translation
Yi, Z., Zhang, H., Tan, P., and Gong, M. (2017) · 2017
Later among the works it cites.
Learning to compose words into sentences with reinforcement learning
Yogatama, D., Blunsom, P., Dyer, C., Grefenstette, E., and Ling, W. (2017) · 2017
Later among the works it cites.
Recent Trends in Deep Learning Based Natural Language Processing
Young, T., Hazarika, D., Poria, S., and Cambria, E. (2017) · 2017
Later among the works it cites.
SeqGAN: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y. (2017) · 2017
Later among the works it cites.
Adversarial Examples: Attacks and Defenses for Deep Learning
Yuan, X., He, P., Zhu, Q., and Li, X. (2017) · 2017
Later among the works it cites.
Action-decision networks for visual tracking with deep reinforcement learning
Yun, S., Choi, J., Yoo, Y., Yun, K., and Young Choi, J. (2017) · 2017
Later among the works it cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Zagoruyko, S. and Komodakis, N. (2017) · 2017
Later among the works it cites.
THUMT: An Open Source Toolkit for Neural Machine Translation
Zhang, J., Ding, Y., Shen, S., Cheng, Y., Sun, M., Luan, H., and Liu, Y. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning with successor features for navigation across similar environments
Zhang, J., Springenberg, J. T., Boedecker, J., and Burgard, W. (2017) · 2017
Later among the works it cites.
Deep Learning based Recommender System: A Survey and New Perspectives
Zhang, S., Yao, L., Sun, A., and Tay, Y. (2017) · 2017
Later among the works it cites.
Sentence simplification with deep reinforcement learning
Zhang, X. and Lapata, M. (2017) · 2017
Later among the works it cites.
Practical Network Blocks Design with Q-Learning
Zhong, Z., Yan, J., and Liu, C.-L. (2017) · 2017
Later among the works it cites.
Deep forest: Towards an alternative to deep neural networks
Zhou, Z.-H. and Feng, J. (2017) · 2017
Later among the works it cites.
Rules of Machine Learning: Best Practices for ML Engineering
Zinkevich, M. (2017) · 2017
Later among the works it cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V. (2017) · 2017
Later among the works it cites.
Learning Transferable Architectures for Scalable Image Recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V. (2017) · 2017
Later among the works it cites.
End to-end goal-oriented question answering systems
Agarwal, D., Chen, B.-C., He, Q., Obukhov?, M., Yang, J., and Zhang, L. (2018) · 2018
Closest in time.
Prediction Machines: The Simple Economics of Artificial Intelligence
Agrawal, A., Gans, J., and Goldfarb, A. (2018) · 2018
Closest in time.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P. (2018) · 2018
Closest in time.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Albrechta, S. V. and Stone, P. (2018) · 2018
Closest in time.
Machine learning in wireless sensor networks: Algorithms, strategies, and applications
Alsheikh, M. A., Lin, S., Niyato, D., and Tan, H.-P. (2014) · 2018
Closest in time.
Differentiable MPC for end-to-end planning and control
Amos, B., Jimenez, I., Sacks, J., Boots, B., and Kolter, J. Z. (2018) · 2018
Closest in time.
Learning to Evade Static PE Machine Learning Malware Models via Reinforcement Learning
Anderson, H. S., Kharkar, A., Filar, B., Evans, D., and Roth, P. (2018) · 2018
Closest in time.
Connecting language and vision to actions
Anderson, P., Das, A., and Wu, Q. (2018) · 2018
Closest in time.
Toward theoretical understanding of deep learning
Arora, S. (2018) · 2018
Closest in time.
Unsupervised neural machine translation
Artetxe, M., Labaka, G., Agirre, E., and Cho, K. (2018) · 2018
Closest in time.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A., Carlini, N., and Wagner, D. (2018) · 2018
Closest in time.
Playing hard exploration games by watching YouTube
Aytar, Y., Pfaff, T., Budden, D., Paine, T., Wang, Z., and de Freitas, N. (2018) · 2018
Closest in time.
An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
Bai, S., Zico Kolter, J., and Koltun, V. (2018) · 2018
Closest in time.
Vector-based navigation using grid-like representations in artificial agents
Banino, A., Barry, C., Uria, B., Blundell, C., Lillicrap, T., Mirowski, P., Pritzel, A., Chadwick, M. J., Degris, T., Modayil, J., Wayne, G., Soyer, H., Viola, F., Zhang, B., Goroshin, R., Rabinowitz, N., Pascanu, R., Beattie, C., Petersen, S., Sadik, A., Gaffney, S., King, H., Kavukcuoglu, K., Hassabis, D., Hadsell, R., and Kumaran, D. (2018) · 2018
Closest in time.
Emergent complexity via multi-agent competition
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I. (2018) · 2018
Closest in time.
Distributed distributional deterministic policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., and Lillicrap, T. (2018) · 2018
Closest in time.
A history of reinforcement learning
Barto, A. (2018) · 2018
Closest in time.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., Gulcehre, C., Song, F., Ballard, A., Gilmer, J., Dahl, G., Vaswani, A., Allen, K., Nash, C., Langston, V., Dyer, C., Heess, N., Wierstra, D., Kohli, P., Botvinick, M., Vinyals, O., Li, Y., and Pascanu, R. (2018) · 2018
Closest in time.
Dopamine
Bellemare, M. G., Castro, P. S., Gelada, C., Kumar, S., and Moitra, S. (2018) · 2018
Closest in time.
Expert level control of ramp metering based on multi-task deep reinforcement learning
Belletti, F., Haziza, D., Gomes, G., and Bayen, A. M. (2018) · 2018
Closest in time.
From deep learning of disentangled representations to higher-level cognition
Bengio, Y. (2018) · 2018
Closest in time.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J. (2018) · 2018
Closest in time.
Deep Learning Techniques for Music Generation
Briot, J.-P., Hadjeres, G., and Pachet, F. (2018) · 2018
Closest in time.
Large Scale GAN Training for High Fidelity Natural Image Synthesis
Brock, A., Donahue, J., and Simonyan, K. (2018) · 2018
Closest in time.
Teaching a machine to read maps with deep reinforcement learning
Brunner, G., Richter, O., Wang, Y., and Wattenhofer, R. (2018) · 2018
Closest in time.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H. (2018) · 2018
Closest in time.
Everybody Dance Now
Chan, C., Ginosar, S., Zhou, T., and Efros, A. A. (2018) · 2018
Closest in time.
Neural Ordinary Differential Equations
Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. (2018) · 2018
Closest in time.
A Lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Ghavamzadeh, M., and Duenez-Guzman, E. (2018) · 2018
Closest in time.
Data-efficient model-based reinforcement learning with deep probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S. (2018) · 2018
Closest in time.
AutoAugment: Learning Augmentation Policies from Data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. (2018) · 2018
Closest in time.
Emergence of grid-like representations by training recurrent neural networks to perform spatial localization
Cueva, C. J. and Wei, X.-X. (2018) · 2018
Closest in time.
Implicit quantile networks for distributional reinforcement learning
Dabney, W., Ostrovski, G., Silver, D., and Munos, R. (2018) · 2018
Closest in time.
Human-level intelligence or animal-like abilities?
Darwiche, A. (2018) · 2018
Closest in time.
Multi-step reinforcement learning: A unifying algorithm
De Asis, K., Hernandez-Garcia, J. F., Zacharias Holland, G., and Sutton, R. S. (2018) · 2018
Closest in time.
End-to-end differentiable physics for learning and control
de Avila Belbute-Peres, F., Smith, K., Allen, K., Tenenbaum, J., and Kolter, J. Z. (2018) · 2018
Closest in time.
Universal Transformers
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł. (2018) · 2018
Closest in time.
Deep Learning in Natural Language Processing
Deng, L. and Liu, Y., editors (2018) · 2018
Closest in time.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Closest in time.
Deep learning of aftershock patterns following large earthquakes
DeVries, P. M. R., Viégas, F., Wattenberg, M., and Meade, B. J. (2018) · 2018
Closest in time.
The challenge of realistic music generation: modelling raw audio at scale
Dieleman, S., van den Oord, A., and Simonyan, K. (2018) · 2018
Closest in time.
Reflections on innateness in machine learning
Dietterich, T. G. (2018) · 2018
Closest in time.
Scalable coordinated exploration in concurrent reinforcement learning
Dimakopoulou, M., Osband, I., and Roy, B. V. (2018) · 2018
Closest in time.
GAN Q-learning
Doan, T., Mazoure, B., and Lyle, C. (2018) · 2018
Closest in time.
An information-theoretic analysis of Thompson sampling for large action spaces
Dong, S. and Roy, B. V. (2018) · 2018
Closest in time.
Neural Architecture Search: A Survey
Elsken, T., Hendrik Metzen, J., and Hutter, F. (2018) · 2018
Closest in time.
Neural scene representation and rendering
Eslami, S. M. A., Rezende, D. J., Besse, F., Viola, F., Morcos, A. S., Garnelo, M., Ruderman, A., Rusu, A. A., Danihelka, I., Gregor, K., Reichert, D. P., Buesing, L., Weber, T., Vinyals, O., Rosenbaum, D., Rabinowitz, N., King, H., Hillier, C., Botvinick, M., Wierstra, D., Kavukcuoglu, K., and Hassabis, D. (2018) · 2018
Closest in time.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K. (2018) · 2018
Closest in time.
Learning explanatory rules from noisy data
Evans, R. and Grefenstette, E. (2018) · 2018
Closest in time.
Robust physical-world attacks on deep learning models
Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., and Song, D. (2018) · 2018
Closest in time.
Iterative value-aware model learning
Farahmand, A.-m. (2018) · 2018
Closest in time.
Generalization and Regularization in DQN
Farebrother, J., Machado, M. C., and Bowling, M. (2018) · 2018
Closest in time.
TreeQN and ATreeC: Differentiable tree-structured models for deep reinforcement learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S. (2018) · 2018
Closest in time.
Resilient Computing with Reinforcement Learning on a Dynamical System: Case Study in Sorting
Faust, A., Aimone, J. B., James, C. D., and Tapia, L. (2018) · 2018
Closest in time.
Clinically applicable deep learning for diagnosis and referral in retinal disease
Fauw, J. D., Ledsam, J. R., Romera-Paredes, B., Nikolov, S., Tomasev, N., Blackwell, S., Askham, H., Glorot, X., O’Donoghue, B., Visentin, D., van den Driessche, G., Lakshminarayanan, B., Meyer, C., Mackinder, F., Bouton, S., Ayoub, K., Chopra, R., King, D., Karthikesalingam, A., Hughes, C. O., Raine, R., Hughes, J., Sim, D. A., Egan, C., Tufail, A., Montgomery, H., Hassabis, D., Rees, G., Back, T., Khaw, P. T., Suleyman, M., Cornebise, J., Keane, P. A., and Ronneberger, O. (2018) · 2018
Closest in time.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Finn, C. and Levine, S. (2018) · 2018
Closest in time.
Probabilistic model-agnostic meta-learning
Finn, C., Xu, K., and Levine, S. (2018) · 2018
Closest in time.
Noisy networks for exploration
Fortunato, M., Gheshlaghi Azar, M., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S. (2018) · 2018
Closest in time.
DeepTraffic: Driving Fast through Dense Traffic with Deep Reinforcement Learning
Fridman, L., Jenik, B., and Terwilliger, J. (2018) · 2018
Closest in time.
Neural approaches to Conversational AI
Gao, J., Galley, M., and Li, L. (2018b) · 2018
Closest in time.
Model-free, model-based, and general intelligence
Geffner, H. (2018) · 2018
Closest in time.
The successor representation: Its computational logic and neural substrates
Gershman, S. J. (2018) · 2018
Closest in time.
Unsupervised video object segmentation for deep reinforcement learning
Goel, V., Weng, J., and Poupart, P. (2018) · 2018
Closest in time.
Recasting gradient-based meta-learning as hierarchical bayes
Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T. (2018) · 2018
Closest in time.
Learning to search with MCTSnets
Guez, A., Weber, T., Antonoglou, I., Simonyan, K., Vinyals, O., Wierstra, D., Munos, R., and Silver, D. (2018) · 2018
Closest in time.
A Survey of Learning Causality with Data: Problems and Methods
Guo, R., Cheng, L., Li, J., Hahn, P. R., and Liu, H. (2018) · 2018
Closest in time.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S. (2018) · 2018
Closest in time.
A neural representation of sketch drawings
Ha, D. and Eck, D. (2018) · 2018
Closest in time.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J. (2018) · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Closest in time.
Learning to play with intrinsically-motivated, self-aware agents
Haber, N., Mrowca, D., Wang, S., Fei-Fei, L., and Yamins, D. (2018) · 2018
Closest in time.
Learning with options that terminate off-policy
Harutyunyan, A., Vrancx, P., Bacon, P.-L., Precup, D., and Nowe, A. (2018) · 2018
Closest in time.
Online robust policy learning in the presence of unknown adversaries
Havens, A., Jiang, Z., and Sarkar, S. (2018) · 2018
Closest in time.
Amc: Automl for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S. (2018) · 2018
Closest in time.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2018) · 2018
Closest in time.
IBM has a Watson dilemma
Hernandez, D. and Greenwald, T. (2018) · 2018
Closest in time.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018) · 2018
Closest in time.
Deep Q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Dulac-Arnold, G., Osband, I., Agapiou, J., Leibo, J. Z., and Gruslys, A. (2018) · 2018
Closest in time.
Matrix capsules with EM routing
Hinton, G. E., Sabour, S., and Frosst, N. (2018) · 2018
Closest in time.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D. (2018) · 2018
Closest in time.
Evolved policy gradients
Houthooft, R., Chen, Y., Isola, P., Stadie, B., Wolski, F., Ho, J., and Abbeel, P. (2018) · 2018
Closest in time.
Unsupervised Learning via Meta-Learning
Hsu, K., Levine, S., and Finn, C. (2018) · 2018
Closest in time.
Learning safe policies with expert guidance
Huang, J., Wu, F., Precup, D., and Cai, Y. (2018) · 2018
Closest in time.
Compositional attention networks for machine reasoning
Hudson, D. A. and Manning, C. D. (2018) · 2018
Closest in time.
Inequity aversion improves cooperation in intertemporal social dilemmas
Hughes, E., Leibo, J., Phillips, M., karl Tuyls, Dueñez-Guzman, E., Castañeda, A. G., Dunning, I., Zhu, T., McKee, K., Koster, R., Roff, H., and Graepel, T. (2018) · 2018
Closest in time.
Basic instincts
Hutson, M. (2018) · 2018
Closest in time.
Human-level performance in first-person multiplayer games with population-based deep reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Garcia Castaneda, A., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., Sonnerat, N., Green, T., Deason, L., Leibo, J. Z., Silver, D., Hassabis, D., Kavukcuoglu, K., and Graepel, T. (2018) · 2018
Closest in time.
Learning to look around: Intelligently exploring unseen environments for unknown tasks
Jayaraman, D. and Grauman, K. (2018) · 2018
Closest in time.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. (2018) · 2018
Closest in time.
Efficient Neural Architecture Search with Network Morphism
Jin, H., Song, Q., and Hu, X. (2018) · 2018
Closest in time.
Artificial intelligence???the revolution hasn?t happened yet
Jordan, M. (2018) · 2018
Closest in time.
Adversarial attacks on stochastic bandits
Jun, K.-S., Li, L., Ma, Y., and Zhu, X. (2018) · 2018
Closest in time.
Confounding-robust policy improvement
Kallus, N. and Zhou, A. (2018) · 2018
Closest in time.
Neural architecture search with Bayesian optimisation and optimal transport
Kandasamy, K., Neiswanger, W., Schneider, J., Poczos, B., and Xing, E. (2018) · 2018
Closest in time.
Strategic Object Oriented Reinforcement Learning
Keramati, R., Whang, J., Cho, P., and Brunskill, E. (2018) · 2018
Closest in time.
Evolutionary reinforcement learning
Khadka, S. and Tumer, K. (2018) · 2018
Closest in time.
Re-evaluate: Reproducibility in evaluating reinforcement learning algorithms
Khetarpal, K., Ahmed, Z., Cianflone, A., Islam, R., and Pineau, J. (2018) · 2018
Closest in time.
Do Better ImageNet Models Transfer Better?
Kornblith, S., Shlens, J., and Le, Q. V. (2018) · 2018
Closest in time.
The case for learned index structures
Kraska, T., Beutel, A., Chi, E. H., Dean, J., and Polyzotis, N. (2018) · 2018
Closest in time.
Cognitive computational neuroscience
Kriegeskorte, N. and Douglas, P. K. (2018) · 2018
Closest in time.
Learning to Optimize Join Queries With Deep Reinforcement Learning
Krishnan, S., Yang, Z., Goldberg, K., Hellerstein, J., and Stoica, I. (2018) · 2018
Closest in time.
Context-dependent upper-confidence bounds for directed exploration
Kumaraswamy, R., Schlegel, M., White, A., and White, M. (2018) · 2018
Closest in time.
The GAN Landscape: Losses, Architectures, Regularization, and Normalization
Kurach, K., Lucic, M., Zhai, X., Michalski, M., and Gelly, S. (2018) · 2018
Closest in time.
Human-in-the-loop interpretability prior
Lage, I., Ross, A., Gershman, S. J., Kim, B., and Doshi-Velez, F. (2018) · 2018
Closest in time.
Actor-critic policy optimization in partially observable multiagent environments
Lanctot, M., Srinivasan, S., Zambaldi, V., Perolat, J., karl Tuyls, Munos, R., and Bowling, M. (2018) · 2018
Closest in time.
Toprank: A practical algorithm for online stochastic ranking
Lattimore, T., Kveton, B., Li, S., and Szepesvári, C. (2018) · 2018
Closest in time.
Bandit Algorithms
Lattimore, T. and Szepesvári, C. (2018) · 2018
Closest in time.
Data center cooling using model-predictive control
Lazic, N., Boutilier, C., Lu, T., Wong, E., Roy, B., Ryu, M., and Imwalle, G. (2018) · 2018
Closest in time.
Hierarchical imitation and reinforcement learning
Le, H. M., Jiang, N., Agarwal, A., Dudík, M., Yue, Y., and Daumé, III, H. (2018) · 2018
Closest in time.
Learning world models: The next step towards AI
LeCun, Y. (2018) · 2018
Closest in time.
What innate priors should we build into the architecture of deep learning systems?
LeCun, Y. and Manning, C. (2018) · 2018
Closest in time.
AI Superpowers: China, Silicon Valley, and the New World Order
Lee, K.-F. (2018) · 2018
Closest in time.
Gated path planning networks
Lee, L., Parisotto, E., Chaplot, D. S., Xing, E., and Salakhutdinov, R. (2018) · 2018
Closest in time.
Reward learning from human preferences and demonstrations in Atari
Leike, J., Ibarz, B., Amodei, D., Irving, G., and Legg, S. (2018) · 2018
Closest in time.
CS 294: Deep reinforcement learning
Levine, S. (2018) · 2018
Closest in time.
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Levine, S. (2018) · 2018
Closest in time.
About AI conferences
Li, Y. (2018) · 2018
Closest in time.
Memory augmented policy optimization for program synthesis with generalization
Liang, C., Norouzi, M., Berant, J., Le, Q. V., and Lao, N. (2018) · 2018
Closest in time.
BBQ-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Lipton, Z., Li, X., Gao, J., Li, L., Ahmed, F., and Deng, L. (2018) · 2018
Closest in time.
The mythos of model interpretability
Lipton, Z. C. (2018) · 2018
Closest in time.
Troubling trends in machine learning scholarship
Lipton, Z. C. and Steinhardt, J. (2018) · 2018
Closest in time.
Non-delusional Q-learning and value-iteration
Lu, T., Boutilier, C., and Schuurmans, D. (2018) · 2018
Closest in time.
Are GANs created equal? a large-scale study
Lucic, M., Kurach, K., Michalski, M., Gelly, S., and Bousquet, O. (2018) · 2018
Closest in time.
Neural architecture optimization
Luo, R., Tian, F., Qin, T., Chen, E., and Liu, T. (2018) · 2018
Closest in time.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
Madhavan, V., Such, F. P., Clune, J., Stanley, K., and Lehman, J. (2018) · 2018
Closest in time.
Imagination machines
Mahadevan, S. (2018a) · 2018
Closest in time.
Exploring the Limits of Weakly Supervised Pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and van der Maaten, L. (2018) · 2018
Closest in time.
IJCAI Research Excellence Award talk
Malik, J. (2018) · 2018
Closest in time.
Learning Visual Question Answering by Bootstrapping Hard Attention
Malinowski, M., Doersch, C., Santoro, A., and Battaglia, P. (2018) · 2018
Closest in time.
DeepProbLog: Neural probabilistic logic programming
Manhaeve, R., Dumancic, S., Kimmig, A., Demeester, T., and Raedt, L. D. (2018) · 2018
Closest in time.
Deep Learning: A Critical Appraisal
Marcus, G. (2018) · 2018
Closest in time.
Solving the Rubik’s Cube Without Human Knowledge
McAleer, S., Agostinelli, F., Shmakov, A., and Baldi, P. (2018) · 2018
Closest in time.
The Natural Language Decathlon: Multitask Learning as Question Answering
McCann, B., Shirish Keskar, N., Xiong, C., and Socher, R. (2018) · 2018
Closest in time.
Towards robust interpretability with self-explaining neural networks
Melis, D. A. and Jaakkola, T. (2018) · 2018
Closest in time.
On the state of the art of evaluation in neural language models
Melis, G., Dyer, C., and Blunsom, P. (2018) · 2018
Closest in time.
When Recurrent Models Don’t Need To Be Recurrent
Miller, J. and Hardt, M. (2018) · 2018
Closest in time.
Learning to navigate in cities without a map
Mirowski, P., Grimes, M., Malinowski, M., Hermann, K. M., Anderson, K., Teplyashin, D., Simonyan, K., koray kavukcuoglu, Zisserman, A., and Hadsell, R. (2018) · 2018
Closest in time.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P. (2018) · 2018
Closest in time.
Personalizing a dialogue system with transfer learning
Mo, K., Li, S., Zhang, Y., Li, J., and Yang, Q. (2018) · 2018
Closest in time.
A flexible neural representation for physics prediction
Mrowca, D., Zhuang, C., Wang, E., Haber, N., Fei-Fei, L., Tenenbaum, J., and Yamins, D. (2018) · 2018
Closest in time.
Toward a smart cloud: A review of fault-tolerance methods in cloud systems
Mukwevho, M. A. and Celik, T. (2018) · 2018
Closest in time.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S. (2018) · 2018
Closest in time.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2018) · 2018
Closest in time.
Visual goal-conditioned reinforcement learning by representation learning
Nair, A., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S. (2018) · 2018
Closest in time.
Reinforcement learning for solving the vehicle routing problem
Nazari, M., Oroojlooy, A., Snyder, L., and Takáč, M. (2018) · 2018
Closest in time.
Learning beam search policies via imitation learning
Negrinho, R., Gormley, M., and Gordon, G. (2018) · 2018
Closest in time.
Machine Learning Yearning (draft)
Ng, A. (2018) · 2018
Closest in time.
Scalable end-to-end autonomous vehicle testing via rare-event simulation
O’Kelly, M., Sinha, A., Namkoong, H., Tedrake, R., and Duchi, J. C. (2018) · 2018
Closest in time.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A. (2018) · 2018
Closest in time.
Learning Dexterous In-Hand Manipulation
OpenAI (2018) · 2018
Closest in time.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J. S., and Cassirer, A. (2018) · 2018
Closest in time.
Recurrent relational networks
Palm, R., Paquet, U., and Winther, O. (2018) · 2018
Closest in time.
Reinforcement learning with function-valued action spaces for partial differential equation control
Pan, Y., Farahmand, A.-m., White, M., Nabi, S., Grover, P., and Nikovski, D. (2018) · 2018
Closest in time.
On Reinforcement Learning for Full-length Game of StarCraft
Pang, Z.-J., Liu, R.-Z., Meng, Z.-Y., Zhang, Y., Yu, Y., and Lu, T. (2018) · 2018
Closest in time.
The seven pillars of causal reasoning with reflections on machine learning
Pearl, J. (2018) · 2018
Closest in time.
The Book of Why: The New Science of Cause and Effect
Pearl, J. and Mackenzie, D. (2018) · 2018
Closest in time.
Temporal difference models: Model-free deep rl for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S. (2018) · 2018
Closest in time.
Deep reinforcement learning for de novo drug design
Popova, M., Isayev, O., and Tropsha, A. (2018) · 2018
Closest in time.
Temporal abstraction
Precup, D. (2018) · 2018
Closest in time.
Machine theory of mind
Rabinowitz, N. C., Perbet, F., Song, H. F., Zhang, C., Eslami, S. A., and Botvinick, M. (2018) · 2018
Closest in time.
Scalable and accurate deep learning for electronic health records
Rajkomar, A., Oren, E., Chen, K., Dai, A. M., Hajaj, N., Liu, P. J., Liu, X., Sun, M., Sundberg, P., Yee, H., Zhang, K., Duggan, G. E., Flores, G., Hardt, M., Irvine, J., Le, Q., Litsch, K., Marcus, J., Mossin, A., Tansuwan, J., Wang, D., Wexler, J., Wilson, J., Ludwig, D., Volchenboum, S. L., Chou, K., Pearson, M., Madabushi, S., Shah, N. H., Butte, A. J., Howell, M., Cui, C., Corrado, G., and Dean, J. (2018) · 2018
Closest in time.
Mura: Large dataset for abnormality detection in musculoskeletal radiographs
Rajpurkar, P., Irvin, J., Bagul, A., Ding, D., Duan, T., Mehta, H., Yang, B., Zhu, K., Laird, D., Ball, R. L., Langlotz, C., Shpanskaya, K., Lungren, M. P., and Ng, A. Y. (2018) · 2018
Closest in time.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S. (2018) · 2018
Closest in time.
Regularized Evolution for Image Classifier Architecture Search
Real, E., Aggarwal, A., Huang, Y., and Le, Q. V. (2018) · 2018
Closest in time.
Glider soaring via reinforcement learning in the field
Reddy, G., Wong-Ng, J., Celani, A., Sejnowski, T. J., and Vergassola, M. (2018) · 2018
Closest in time.
Inverse reinforcement learning for computer vision
Rhinehart, N., Kitani, K., and Vernaza, P. (2018) · 2018
Closest in time.
Learning abstract options
Riemer, M., Liu, M., and Tesauro, G. (2018) · 2018
Closest in time.
An analysis of categorical distributional reinforcement learning
Rowland, M., Bellemare, M. G., Dabney, W., Munos, R., and Teh, Y. W. (2018) · 2018
Closest in time.
Sim2real view invariant visual servoing by recurrent control
Sadeghi, F., Toshev, A., Jang, E., and Levine, S. (2018) · 2018
Closest in time.
Relational recurrent neural networks
Santoro, A., Faulkner, R., Raposo, D., Rae, J., Chrzanowski, M., Weber, T., Wierstra, D., Vinyals, O., Pascanu, R., and Lillicrap, T. (2018) · 2018
Closest in time.
How does batch normalization help optimization? (no, it is not about internal covariate shift)
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A. (2018) · 2018
Closest in time.
Winner’s curse? on pace, progress, and empirical rigor
Sculley, D., Snoek, J., Wiltschko, A., and Rahimi, A. (2018) · 2018
Closest in time.
Planning chemical syntheses with deep neural networks and symbolic AI
Segler, M. H. S., Preuss, M., and Waller, M. P. (2018) · 2018
Closest in time.
Multi-task learning as multi-objective optimization
Sener, O., Sener, O., and Koltun, V. (2018) · 2018
Closest in time.
A survey of available corpora for building data-driven dialogue systems: The journal version
Serban, I. V., Lowe, R., Charlin, L., and Pineau, J. (2018) · 2018
Closest in time.
Accelerating learning in constructive predictive frameworks with the successor representation
Sherstan, C., Machado, M. C., and Pilarski, P. M. (2018) · 2018
Closest in time.
Virtual-Taobao: Virtualizing Real-world Online Retail Environment for Reinforcement Learning
Shi, J.-C., Yu, Y., Da, Q., Chen, S.-Y., and Zeng, A.-X. (2018) · 2018
Closest in time.
Principles of deep rl
Silver, D. (2018) · 2018
Closest in time.
Multitask reinforcement learning for zero-shot generalization with subtask dependencies
Sohn, S., Oh, J., and Lee, H. (2018) · 2018
Closest in time.
AI and security: Lessons, challenges and future directions
Song, D. (2018) · 2018
Closest in time.
Multi-agent generative adversarial imitation learning
Song, J., Ren, H., Sadigh, D., and Ermon, S. (2018) · 2018
Closest in time.
Universal planning networks
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C. (2018) · 2018
Closest in time.
The importance of sampling in meta-reinforcement learning
Stadie, B., Yang, G., Abbeel, P., Wu, Y., Duan, Y., Chen, X., Houthooft, R., and Sutskever, I. (2018) · 2018
Closest in time.
TStarBots: Defeating the Cheating Level Builtin AI in StarCraft II in the Full Game
Sun, P., Sun, X., Han, L., Xiong, J., Wang, Q., Li, B., Zheng, Y., Liu, J., Liu, Y., Liu, H., and Zhang, T. (2018) · 2018
Closest in time.
Dual policy iteration
Sun, W., Gordon, G., Boots, B., and Bagnell, J. (2018) · 2018
Closest in time.
The next big step in AI: Planning with a learned model
Sutton, R. (2018) · 2018
Closest in time.
Reinforcement Learning: An Introduction (2nd Edition)
Sutton, R. S. and Barto, A. G. (2018) · 2018
Closest in time.
Learning plannable representations with causal InfoGAN
Tamar, A., Abbeel, P., Yang, G., Kurutach, T., and Russell, S. (2018) · 2018
Closest in time.
DeepMind Control Suite
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., and Riedmiller, M. (2018) · 2018
Closest in time.
Building machines that learn & think like people
Tenenbaum, J. (2018) · 2018
Closest in time.
Deep reinforcement learning of marked temporal point processes
Upadhyay, U., De, A., and Rodriguez, M. G. (2018) · 2018
Closest in time.
Reinforcement learning of theorem proving
Urban, J., Kaliszyk, C., Michalewski, H., and Olšák, M. (2018) · 2018
Closest in time.
Representation Learning with Contrastive Predictive Coding
van den Oord, A., Li, Y., and Vinyals, O. (2018) · 2018
Closest in time.
Multi-agent reinforcement learning via double averaging primal-dual optimization
Wai, H.-T., Wang, P. Z., Yang, Z., and Hong, M. (2018) · 2018
Closest in time.
Deep reinforcement learning for NLP
Wang, W. Y., Li, J., and He, X. (2018d) · 2018
Closest in time.
Machine learning in compiler optimization
Wang, Z. and O’Boyle, M. (2018) · 2018
Closest in time.
Unsupervised Predictive Memory in a Goal-Directed Agent
Wayne, G., Hung, C.-C., Amos, D., Mirza, M., Ahuja, A., Grabska-Barwinska, A., Rae, J., Mirowski, P., Leibo, J. Z., Santoro, A., Gemici, M., Reynolds, M., Harley, T., Abramson, J., Mohamed, S., Rezende, D., Saxton, D., Cain, A., Hillier, C., Silver, D., Kavukcuoglu, K., Botvinick, M., Hassabis, D., and Lillicrap, T. (2018) · 2018
Closest in time.
Constrained cross-entropy method for safe reinforcement learning
Wen, M. and Topcu, U. (2018) · 2018
Closest in time.
Transfer learning with neural AutoML
Wong, C., Houlsby, N., Lu, Y., and Gesmundo, A. (2018) · 2018
Closest in time.
Reinforced co-training
Wu, J., Li, L., and Wang, W. Y. (2018) · 2018
Closest in time.
Model-level dual learning
Xia, Y., Tan, X., Tian, F., Qin, T., Yu, N., and Liu, T.-Y. (2018) · 2018
Closest in time.
Memory-augmented monte carlo tree search
Xiao, C., JinchengMei, and Müller, M. (2018) · 2018
Closest in time.
Reinforced continual learning
Xu, J. and Zhu, Z. (2018) · 2018
Closest in time.
Meta-Gradient Reinforcement Learning
Xu, Z., van Hasselt, H., and Silver, D. (2018) · 2018
Closest in time.
Peorl: Integrating symbolic planning and hierarchical reinforcement learning for robust decision-making
Yang, F., Lyu, D., Liu, B., and Gustafson, S. (2018) · 2018
Closest in time.
Artificial Intelligence and Games
Yannakakis, G. N. and Togelius, J. (2018) · 2018
Closest in time.
Neural-Symbolic VQA: Disentangling reasoning from vision and language understanding
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., and Tenenbaum, J. (2018) · 2018
Closest in time.
Bayesian model-agnostic meta-learning
Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S. (2018) · 2018
Closest in time.
BDD100K: A Diverse Driving Video Database with Scalable Annotation Tooling
Yu, F., Xian, W., Chen, Y., Liu, F., Liao, M., Madhavan, V., and Darrell, T. (2018) · 2018
Closest in time.
One-shot imitation from observing humans via domain-adaptive meta-learning
Yu, T., Finn, C., Xie, A., Dasari, S., Zhang, T., Abbeel, P., and Levine, S. (2018) · 2018
Closest in time.
Imitation learning
Yue, Y. and Le, H. M. (2018) · 2018
Closest in time.
Learn what not to learn: Action elimination with deep reinforcement learning
Zahavy, T., Harush, M., Merlis, N., Mankowitz, D. J., and Mannor, S. (2018) · 2018
Closest in time.
Relational Deep Reinforcement Learning
Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., Lockhart, E., Shanahan, M., Langston, V., Pascanu, R., Botvinick, M., Vinyals, O., and Battaglia, P. (2018) · 2018
Closest in time.
Deep Learning in Mobile and Wireless Networking: A Survey
Zhang, C., Patras, P., and Haddadi, H. (2018) · 2018
Closest in time.
Neural guided constraint logic programming for program synthesis
Zhang, L., Rosenblatt, G., Fetaya, E., Liao, R., Byrd, W., Might, M., Urtasun, R., and Zemel, R. (2018) · 2018
Closest in time.
Deep Learning for Sentiment Analysis : A Survey
Zhang, L., Wang, S., and Liu, B. (2018) · 2018
Closest in time.
Visual interpretability for deep learning: a survey
Zhang, Q. and Zhu, S.-C. (2018) · 2018
Closest in time.
Emotional chatting machine: Emotional conversation generation with internal and external memory
Zhou, H., Huang, M., Zhang, T., Zhu, X., and Liu, B. (2018) · 2018
Closest in time.
Voxelnet: End-to-end learning for point cloud based 3d object detection
Zhou, Y. and Tuzel, O. (2018) · 2018
Closest in time.
Multi-agent online learning with asynchronous feedback loss
Zhou, Z., Mertikopoulos, P., Athey, S., Bambos, N., Glynn, P. W., and Ye, Y. (2018) · 2018
Closest in time.
XiaoIce Band: A melody and arrangement generation framework for pop music
Zhu, H., Liu, Q., Yuan, N. J., Qin, C., Li, J., Zhang, K., Zhou, G., Wei, F., Xu, Y., and Chen, E. (2018) · 2018
Closest in time.
Applications of deep learning and reinforcement learning to biological data
Mahmud, M., Kaiser, M. S., Hussain, A., and Vassanelli, S. (2018) · 2079
Closest in time.