Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL), bolstered by the expressive capabilities of Deep Neural Networks (DNNs) for function approximation, has demonstrated considerable success in numerous applications.
Language Models are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 1901
Earlier work this paper cites.
MONet: Unsupervised Scene Decomposition and Representation
Burgess, C., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A. (2019) · 1901
Earlier work this paper cites.
Mesh-based Tools to Analyze Deep Reinforcement Learning Policies for Underactuated Biped Locomotion
Talele, N., and Byl, K. (2019) · 1903
Earlier work this paper cites.
Skill Transfer in Deep Reinforcement Learning Under Morphological Heterogeneity
Hu, Y., and Montana, G. (2019) · 1908
Earlier work this paper cites.
Reinforcement Learning With Structured Hierarchical Grammar Representations of Actions
Christodoulou, P., Lange, R., Shafti, A., and Faisal, A. (2019) · 1910
Earlier work this paper cites.
Faster and Safer Training by Embedding High-level Knowledge Into Deep Reinforcement Learning
Zhang, H., Gao, Z., Zhou, Y., Zhang, H., Wu, K., and Lin, F. (2019a) · 1910
Earlier work this paper cites.
Faster and Safer Training by Embedding High-level Knowledge Into Deep Reinforcement Learning
Zhang, H., Gao, Z., Zhou, Y., Zhang, H., Wu, K., and Lin, F. (2019b) · 1910
Earlier work this paper cites.
Some Applications of the Theory of Dynamic Programming - A Review
Bellman, R. (1954) · 1954
Earlier work this paper cites.
Learning to Predict by the Methods of Temporal Differences
Sutton, R. (1988) · 1988
Earlier work this paper cites.
Simple Statistical Gradient-following Algorithms for Connectionist Reinforcement Learning
Williams, R. (1992) · 1992
Earlier work this paper cites.
Improving Generalization for Temporal Difference Learning: The Successor Representation
Dayan, P. (1993) · 1993
Earlier work this paper cites.
A Framework for Behavioural Cloning
Bain, M., and Sammut, C. (1995) · 1995
Earlier work this paper cites.
Exploiting Structure in Policy Construction
Boutilier, C., Dearden, R., and Goldszmidt, M. (1995) · 1995
Earlier work this paper cites.
Reinforcement Learning of Non-markov Decision Processes
Whitehead, S., and Lin, L. (1995) · 1995
Earlier work this paper cites.
Reinforcement Learning With Hierarchies of Machines
Parr, R., and Russell, S. (1997) · 1997
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Efficient Reinforcement Learning in Factored MDPs
Kearns, M., and Koller, D. (1999) · 1999
Earlier work this paper cites.
Computing Factored Value Functions for Policies in Structured MDPs
Koller, D., and Parr, R. (1999) · 1999
Earlier work this paper cites.
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
Ng, A., Harada, D., and Russell, S. (1999) · 1999
Earlier work this paper cites.
Stochastic Dynamic Programming With Factored Representations
Boutilier, C., Dearden, R., and Goldszmidt, M. (2000) · 2000
Earlier work this paper cites.
Relational Reinforcement Learning
Dzeroski, S., Raedt, L., and Driessens, K. (2001) · 2001
Earlier work this paper cites.
Dynamic Bayesian Networks: A State of the Art
Mihajlovic, V., and Petkovic, M. (2001) · 2001
Earlier work this paper cites.
Solving Factored MDPs With Large Action Space Using Algebraic Decision Diagrams
Kim, K., and Dean, T. (2002) · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Kakade, S. (2003) · 2003
Earlier work this paper cites.
Solving Relational MDPs With First-order Machine Learning
Mausam, and Weld, D. (2003) · 2003
Earlier work this paper cites.
Payani, A., and Fekri, F. (2020) · 2003
Earlier work this paper cites.
Q-Decomposition for Reinforcement Learning Agents
Russell, S., and Zimdars, A. (2003) · 2003
Earlier work this paper cites.
Structural Abstraction Experiments in Reinforcement Learning
Fitch, R., Hengst, B., Suc, D., Calbert, G., and Scholz, J. (2005) · 2005
Earlier work this paper cites.
Approximate Linear Programming for First-order MDPs
Sanner, S., and Boutilier, C. (2005) · 2005
Earlier work this paper cites.
Approximate Policy Iteration With a Policy Language Bias: Solving Relational Markov Decision Processes
Fern, A., Yoon, S., and Givan, R. (2006) · 2006
Earlier work this paper cites.
Towards a Unified Theory of State Abstraction for MDPs.
Li, L., Walsh, T., and Littman, M. (2006) · 2006
Earlier work this paper cites.
Safe Reinforcement Learning With Mixture Density Network: A Case Study in Autonomous Highway Driving
Baheri, A. (2020) · 2007
Earlier work this paper cites.
Proto-value Functions: A Laplacian Framework for Learning Representation and Control in Markov Decision Processes
Mahadevan, S., and Maggioni, M. (2007) · 2007
Earlier work this paper cites.
Multi-task reinforcement learning as a hidden-parameter block MDP
Zhang, A., Sodhani, S., Khetarpal, K., and Pineau, J. (2020) · 2007
Earlier work this paper cites.
An Object-oriented Representation for Efficient Reinforcement Learning
Diuk, C., Cohen, A., and Littman, M. (2008) · 2008
Earlier work this paper cites.
Model-based Bayesian Reinforcement Learning in Large Structured Domains
Ross, S., and Pineau, J. (2008) · 2008
Earlier work this paper cites.
Simple Local Models for Complex Dynamical Systems
Talvitie, E., and Singh, S. (2008) · 2008
Earlier work this paper cites.
Symbolic Relational Deep Reinforcement Learning Based on Graph Neural Networks
Janisch, J., Pevný, T., and Lisý, V. (2020) · 2009
Earlier work this paper cites.
Reinforcement Learning in Finite MDPs: PAC Analysis
Strehl, A., Li, L., and Littman, M. (2009) · 2009
Earlier work this paper cites.
Probabilistic Relational Planning With First Order Decision Diagrams
Joshi, S., and Khardon, R. (2011) · 2011
Earlier work this paper cites.
Reinforcement Learning With Subspaces Using Free Energy Paradigm
Ghorbani, M., Hosseini, R., Shariatpanahi, S., and Ahmadabadi, M. (2020) · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. (2014) · 2014
Earlier work this paper cites.
Optimal Behavioral Hierarchy
Solway, A., Diuk, C., Córdova, N., Yee, D., Barto, A., Niv, Y., and Botvinick, M. (2014) · 2014
Earlier work this paper cites.
Factored MDPs for Detecting Topics of User Sessions
Tavakol, M., and Brefeld, U. (2014) · 2014
Earlier work this paper cites.
Goal-Based Action Priors
Abel, D., Hershkowitz, D., Barth-Maron, G., Brawner, S., O’Farrell, K., MacGlashan, J., and Tellex, S. (2015) · 2015
Earlier work this paper cites.
A Comprehensive Survey on Safe Reinforcement Learning
Garcia, J., and Fernandez, F. (2015) · 2015
Earlier work this paper cites.
Contextual Markov Decision Processes
Hallak, A., Castro, D. D., and Mannor, S. (2015) · 2015
Earlier work this paper cites.
Patterns for Learning with Side Information
Jonschkowski, R., Höfer, S., and Brock, O. (2015) · 2015
Earlier work this paper cites.
Mankowitz, D., Mann, T., and Mannor, S. (2015) · 2015
Earlier work this paper cites.
Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning
Mohamed, S., and Rezende, D. (2015) · 2015
Earlier work this paper cites.
Universal Value Function Approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Earlier work this paper cites.
Learning Shared Representations in Multi-task Reinforcement Learning
Borsa, D., Graepel, T., and Shawe-Taylor, J. (2016) · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Taming the Noise in Reinforcement Learning via Soft Updates
Fox, R., Pakman, A., and Tishby, N. (2016) · 2016
Earlier work this paper cites.
Towards Deep Symbolic Reinforcement Learning
Garnelo, M., Arulkumaran, K., and Shanahan, M. (2016) · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D., and Wierstra, D. (2016) · 2016
Earlier work this paper cites.
Learning and Transfer of Modulated Locomotor Controllers
Heess, N., Wayne, G., Tassa, Y., Lillicrap, T., Riedmiller, M., and Silver, D. (2016) · 2016
Earlier work this paper cites.
Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
Kulkarni, T., Narasimhan, K., Saeedi, A., and Tenenbaum, J. (2016) · 2016
Earlier work this paper cites.
Source Task Creation for Curriculum Learning
Narvekar, S., Sinapov, J., Leonetti, M., and Stone, P. (2016) · 2016
Earlier work this paper cites.
Causal Inference by Using Invariant Prediction: Identification and Confidence Intervals
Peters, J., Buhlmann, P., and Meinshausen, N. (2016) · 2016
Earlier work this paper cites.
Policy Distillation
Rusu, A., Colmenarejo, S., Gülçehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2016) · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C., Guez, A., Sifre, L., Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Earlier work this paper cites.
Integrating Symmetry of Environment by Designing Special Basis Functions for Value Function Approximation in Reinforcement Learning
Wang, G., Fang, Z., Li, B., and Li, P. (2016) · 2016
Earlier work this paper cites.
Active Exploration for Learning Symbolic Representations
Andersen, G., and Konidaris, G. (2017) · 2017
Earlier work this paper cites.
Reinforcement Learning in Rich-observation MDPs Using Spectral Methods
Azizzadenesheli, K., Lazaric, A., and Anandkumar, A. (2017) · 2017
Earlier work this paper cites.
The Option-critic Architecture
Bacon, P., Harb, J., and Precup, D. (2017) · 2017
Earlier work this paper cites.
Successor Features for Transfer in Reinforcement Learning
Barreto, A., Dabney, W., Munos, R., Hunt, J., Schaul, T., van Hasselt, H., and Silver, D. (2017) · 2017
Earlier work this paper cites.
Learning Modular Neural Network Policies for Multi-task and Multi-robot Transfer
Devin, C., Gupta, A., Darrell, T., Abbeel, P., and Levine, S. (2017) · 2017
Earlier work this paper cites.
Stochastic Neural Networks for Hierarchical Reinforcement Learning
Florensa, C., Duan, Y., and Abbeel, P. (2017) · 2017
Earlier work this paper cites.
Learning Invariant Feature Spaces to Transfer Skills With Reinforcement Learning
Gupta, A., Devin, C., Liu, Y., Abbeel, P., and Levine, S. (2017) · 2017
Earlier work this paper cites.
DARLA: Improving Zero-shot Transfer in Reinforcement Learning
Higgins, I., Pal, A., Rusu, A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A. (2017) · 2017
Earlier work this paper cites.
On Decomposability in Robot Reinforcement Learning
Höfer, S. (2017) · 2017
Earlier work this paper cites.
Active Exploration and Parameterized Reinforcement Learning Applied to a Simulated Human-robot Interaction Task
Khamassi, M., Velentzas, G., Tsitsimis, T., and Tzafestas, C. (2017) · 2017
Earlier work this paper cites.
Symmetry Learning for Function Approximation in Reinforcement Learning
Mahajan, A., and Tulabandhula, T. (2017) · 2017
Earlier work this paper cites.
Relational Reinforcement Learning With Guided Demonstrations
Martinez, D., Alenya, G., and Torras, C. (2017) · 2017
Earlier work this paper cites.
Discrete Sequential Prediction of Continuous Actions for Deep RL
Metz, L., Ibarz, J., Jaitly, N., and Davidson, J. (2017) · 2017
Earlier work this paper cites.
Curiosity-driven Exploration by Self-supervised Prediction
Pathak, D., Agrawal, P., Efros, A., and Darrell, T. (2017) · 2017
Earlier work this paper cites.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Salimans, T., Ho, J., Chen, X., and Sutskever, I. (2017) · 2017
Earlier work this paper cites.
Hierarchy Through Composition With Multitask LMDPs
Saxe, A., Earle, A., and Rosman, B. (2017) · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A., Kaiser, L., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Optimizing Chemical Reactions With Deep Reinforcement Learning
Zhou, Z., Li, X., and Zare, R. (2017) · 2017
Earlier work this paper cites.
Unpaired Image-to-image Translation Using Cycle-consistent Adversarial Networks
Zhu, J., Park, T., Isola, P., and Efros, A. (2017) · 2017
Earlier work this paper cites.
Symbolic Relation Networks for Reinforcement Learning
Adjodah, D., Klinger, T., and Joseph, J. (2018) · 2018
Earlier work this paper cites.
Learning With Latent Language
Andreas, J., Klein, D., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement
Barreto, A., Borsa, D., Quan, J., Schaul, T., Silver, D., Hessel, M., Mankowitz, D., Zidek, A., and Munos, R. (2018) · 2018
Earlier work this paper cites.
Planning and Learning With Stochastic Action Sets
Boutilier, C., Cohen, A., Hassidim, A., Mansour, Y., Meshi, O., Mladenov, M., and Schuurmans, D. (2018) · 2018
Earlier work this paper cites.
Meta-Reinforcement Learning of Structured Exploration Strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Composable Deep Reinforcement Learning for Robotic Manipulation
Haarnoja, T., Pong, V., Zhou, A., Dalal, M., Abbeel, P., and Levine, S. (2018b) · 2018
Earlier work this paper cites.
Learning an Embedding Space for Transferable Robot Skills
Hausman, K., Springenberg, J., Wang, Z., Heess, N., and Riedmiller, M. (2018) · 2018
Earlier work this paper cites.
Learning to Interrupt: A Hierarchical Deep Reinforcement Learning Framework for Efficient Exploration
Li, T., Pan, J., Zhu, D., and Meng, M. (2018) · 2018
Earlier work this paper cites.
The Mythos of Model Interpretability
Lipton, Z. (2018) · 2018
Earlier work this paper cites.
Robot Representation and Reasoning With Knowledge From Reinforcement Learning
Lu, K., Zhang, S., Stone, P., and Chen, X. (2018) · 2018
Earlier work this paper cites.
Data-efficient Hierarchical Reinforcement Learning
Nachum, O., Gu, S., Lee, H., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Exploration in Structured Reinforcement Learning
Ok, J., Proutière, A., and Tranos, D. (2018) · 2018
Earlier work this paper cites.
Hierarchical and Interpretable Skill Acquisition in Multi-task Reinforcement Learning
Shu, T., Xiong, C., and Socher, R. (2018) · 2018
Earlier work this paper cites.
Hierarchical Reinforcement Learning for Zero-shot Generalization With Subtask Dependencies
Sohn, S., Oh, J., and Lee, H. (2018) · 2018
Earlier work this paper cites.
Structured Control Nets for Deep Reinforcement Learning
Srouji, M., Zhang, J., and Salakhutdinov, R. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R., and Barto, A. (2018) · 2018
Cited alongside, same era.
Action Branching Architectures for Deep Reinforcement Learning
Tavakoli, A., Pardo, F., and Kormushev, P. (2018) · 2018
Cited alongside, same era.
Programmatically Interpretable Reinforcement Learning
Verma, A., Murali, V., Singh, R., Kohli, P., and Chaudhuri, S. (2018) · 2018
Cited alongside, same era.
Nervenet: Learning Structured Policy With Graph Neural Networks
Wang, T., Liao, R., Ba, J., and Fidler, S. (2018) · 2018
Cited alongside, same era.
Variance Reduction for Policy Gradient With Action-dependent Factorized Baselines
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A., Kakade, S., Mordatch, I., and Abbeel, P. (2018) · 2018
Cited alongside, same era.
PEORL: Integrating Symbolic Planning and Hierarchical Reinforcement Learning for Robust Decision-making
Solving Compositional Reinforcement Learning Problems via Task Reduction
Li, Y., Wu, Y., Xu, H., Wang, X., and Wu, Y. (2021) · 2021
Later among the works it cites.
Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning
Liao, L., Fu, Z., Yang, Z., Wang, Y., Kolar, M., and Wang, Z. (2021) · 2021
Later among the works it cites.
Reinforcement Learning in Factored Action Spaces Using Tensor Decompositions
Mahajan, A., Samvelyan, M., Mao, L., Makoviychuk, V., Garg, A., Kossaifi, J., Whiteson, S., Zhu, Y., and Anandkumar, A. (2021) · 2021
Later among the works it cites.
Task-agnostic Exploration via Policy Gradient of a Non-parametric State Entropy Estimate
Mutti, M., Pratissoli, L., and Restelli, M. (2021) · 2021
Later among the works it cites.
Reinforcement Learning in Linear MDPs: Constant Regret and Representation Selection
Papini, M., Tirinzoni, A., Pacchiano, A., Restelli, M., Lazaric, A., and Pirotta, M. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, F., Lyu, D., Liu, B., and Gustafson, S. (2018) · 2018
Cited alongside, same era.
Structured Agents for Physical Construction
Bapst, V., Sanchez-Gonzalez, A., Doersch, C., Stachenfeld, K., Kohli, P., Battaglia, P., and Hamrick, J. (2019) · 2019
Cited alongside, same era.
The Option Keyboard: Combining Skills in Reinforcement Learning
Barreto, A., Borsa, D., Hou, S., Comanici, G., Aygün, E., Hamel, P., Toyama, D., Hunt, J., Mourad, S., Silver, D., and Precup, D. (2019) · 2019
Cited alongside, same era.
Dot-to-dot: Explainable Hierarchical Reinforcement Learning for Robotic Manipulation
Beyret, B., Shafti, A., and Faisal, A. (2019) · 2019
Cited alongside, same era.
Universal Successor Features Approximators
Borsa, D., Barreto, A., Quan, J., Mankowitz, D., van Hasselt, H., Munos, R., Silver, D., and Schaul, T. (2019) · 2019
Cited alongside, same era.
Decomposition methods with deep corrections for reinforcement learning
Bouton, M., Julian, K., Nakhaei, A., Fujimura, K., and Kochenderfer, M. (2019) · 2019
Cited alongside, same era.
Computation of Weighted Sums of Rewards for Concurrent MDPs
Buchholz, P., and Scheftelowitsch, D. (2019) · 2019
Cited alongside, same era.
Modular Networks Prevent Catastrophic Interference in Model-based Multi-task Reinforcement Learning
Schiewer, R., and Wiskott, L. (2021) · 2021
Later among the works it cites.
Causal Influence Detection for Improving Efficiency in Reinforcement Learning
Seitzer, M., Schölkopf, B., and Martius, G. (2021) · 2021
Later among the works it cites.
AlwaysSafe: Reinforcement Learning Without Safety Constraint Violations During Training
Simao, T., Jansen, N., and Spaan, M. (2021) · 2021
Later among the works it cites.
Structured World Belief for Reinforcement Learning in POMDPs
Singh, G., Peri, S., Kim, J., Kim, H., and Ahn, S. (2021) · 2021
Later among the works it cites.
Multi-task Reinforcement Learning With Context-based Representations
Sodhani, S., Zhang, A., and Pineau, J. (2021) · 2021
Later among the works it cites.
Factored Policy Gradients: Leveraging Structure for Efficient Learning in MOMDPs
Spooner, T., Vadori, N., and Ganesh, S. (2021) · 2021
Later among the works it cites.
Unsupervised Learning for Reinforcement Learning.
Srinivas, A., and Abbeel, P. (2021) · 2021
Later among the works it cites.
TempLe: Learning Template of Transitions for Sample Efficient Multi-task RL
Sun, Y., Yin, X., and Huang, F. (2021) · 2021
Later among the works it cites.
Human-level Reinforcement Learning Through Theory-based Modeling, Exploration, and Planning
Tsividis, P., Loula, J., Burga, J., Foss, N., Campero, A., Pouncy, T., Gershman, S., and Tenenbaum, J. (2021) · 2021
Later among the works it cites.
A Novel Approach to Curiosity and Explainable Reinforcement Learning via Interpretable Sub-goals
van Rossum, C., Feinberg, C., Shumays, A., Baxter, K., and Bartha, B. (2021) · 2021
Later among the works it cites.
Alchemy: A Benchmark and Analysis Toolkit for Meta-reinforcement Learning Agents
Wang, J., King, M., Porcel, N., Kurth-Nelson, Z., Zhu, T., Deck, C., Choy, P., Cassin, M., Reynolds, M., Song, H., Buttimore, G., Reichert, D., Rabinowitz, N., Matthey, L., Hassabis, D., Lerchner, A., and Botvinick, M. (2021) · 2021
Later among the works it cites.
Interpretable Model-based Hierarchical Reinforcement Learning Using Inductive Logic Programming
Xu, D., and Fekri, F. (2021) · 2021
Later among the works it cites.
Reinforcement Learning With Prototypical Representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2021) · 2021
Later among the works it cites.
Learning Invariant Representations for Reinforcement Learning Without Reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S. (2021) · 2021
Later among the works it cites.
Domain Knowledge Guided Offline Q Learning
Zhang, X., Zhang, S., and Yu, Y. (2021) · 2021
Later among the works it cites.
Context-specific Representation Abstraction for Deep Option Learning
Abdulhai, M., Kim, D., Riemer, M., Liu, M., Tesauro, G., and How, J. (2022) · 2022
Later among the works it cites.
Automated Dynamic Algorithm Configuration
Adriaensen, S., Biedenkapp, A., Shala, G., Awad, N., Eimer, T., Lindauer, M., and Hutter, F. (2022) · 2022
Later among the works it cites.
Experiential Explanations for Reinforcement Learning
Alabdulkarim, A., Singh, M., Mansi, G., Hall, K., and Riedl, M. (2022) · 2022
Later among the works it cites.
Interpretable Preference-based Reinforcement Learning With Tree-structured Reward Functions
Bewley, T., and Lecune, F. (2022) · 2022
Later among the works it cites.
Deep Surrogate Assisted Generation of Environments
Bhatt, V., Tjanaka, B., Fontaine, M., and Nikolaidis, S. (2022) · 2022
Later among the works it cites.
Nuclear Norm Maximization Based Curiosity-driven Learning
Chen, C., Gao, Z., Xu, K., Yang, S., Li, Y., Ding, B., Feng, D., and Wang, H. (2022) · 2022
Later among the works it cites.
Uncertainty Estimation Based Intrinsic Reward For Efficient Reinforcement Learning
Chen, C., Wan, T., Shi, P., Ding, B., Gao, Z., and Feng, D. (2022) · 2022
Later among the works it cites.
Generalizing Goal-conditioned Reinforcement Learning With Variational Causal Reasoning
Ding, W., Lin, H., Li, B., and Zhao, D. (2022) · 2022
Later among the works it cites.
Multi-agent Deep Reinforcement Learning: a Survey
Gronauer, S., and Diepold, K. (2022) · 2022
Later among the works it cites.
A Relational Intervention Approach for Unsupervised Dynamics Generalization in Model-based Reinforcement Learning
Guo, J., Gong, M., and Tao, D. (2022) · 2022
Later among the works it cites.
Bisimulation Makes Analogies in Goal-conditioned Reinforcement Learning
Hansen-Estruch, P., Zhang, A., Nair, A., Yin, P., and Levine, S. (2022) · 2022
Later among the works it cites.
Exploration via Elliptical Episodic Bonuses
Henaff, M., Raileanu, R., Jiang, M., and Rocktäschel, T. (2022) · 2022
Later among the works it cites.
Bilinear Value Networks
Hong, Z., Yang, G., and Agrawal, P. (2022) · 2022
Later among the works it cites.
Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
Icarte, R., Klassen, T., Valenzano, R., and McIlraith, S. (2022) · 2022
Later among the works it cites.
Discrete Factorial Representations as an Abstraction for Goal Conditioned Reinforcement Learning
Islam, R., Zang, H., Goyal, A., Lamb, A., Kawaguchi, K., Li, X., Laroche, R., Bengio, Y., and Combes, R. (2022) · 2022
Later among the works it cites.
Unsupervised Skill Discovery via Recurrent Skill Training
Jiang, Z., Gao, J., and Chen, J. (2022) · 2022
Later among the works it cites.
Relational Abstractions for Generalized Reinforcement Learning on Symbolic Problems
Karia, R., and Srivastava, S. (2022) · 2022
Later among the works it cites.
Disentangled (Un)Controllable Features
Kooi, J., Hoogendoorn, M., and François-Lavet, V. (2022) · 2022
Later among the works it cites.
Using Natural Language and Program Abstractions to Instill Human Inductive Biases in Machines
Kumar, S., Correa, C., Dasgupta, I., Marjieh, R., Hu, M., Hawkins, R., Daw, N., Cohen, J., Narasimhan, K., and Griffiths, T. (2022) · 2022
Later among the works it cites.
Tell me Why! Explanations Support Learning Relational and Causal Structure
Lampinen, A., Roy, N., Dasgupta, I., Chan, S., Tam, A., Mcclelland, J., Yan, C., Santoro, A., Rabinowitz, N., Wang, J., and Hill, F. (2022) · 2022
Later among the works it cites.
Reinforcement Learning Algorithm Selection
Laroche, R., and Feraud, R. (2022) · 2022
Later among the works it cites.
Recursive Constraints to Prevent Instability in Constrained Reinforcement Learning
Lee, J., Sedwards, S., and Czarnecki, K. (2022) · 2022
Later among the works it cites.
Discovered Policy Optimisation
Lu, C., Kuba, J., Letcher, A., Metz, L., de Witt, C., and Foerster, J. (2022) · 2022
Later among the works it cites.
Multi-objective Evolution for Generalizable Policy Gradient Algorithms
Luis, J., Miao, Y., Co-Reyes, J., Parisi, A., Tan, J., Real, E., and Faust, A. (2022) · 2022
Later among the works it cites.
Compositional Multi-object Reinforcement Learning With Linear Relation Networks
Mambelli, D., Träuble, F., Bauer, S., Schölkopf, B., and Locatello, F. (2022) · 2022
Later among the works it cites.
Task Factorization in Curriculum Learning
Mirsky, R., Shperberg, S., Zhang, Y., Xu, Z., Jiang, Y., Cui, J., and Stone, P. (2022) · 2022
Later among the works it cites.
Unsupervised Reinforcement Learning in Multiple Environments
Mutti, M., Mancassola, M., and Restelli, M. (2022) · 2022
Later among the works it cites.
Skill-based Meta-reinforcement Learning
Nam, T., Sun, S., Pertsch, K., Hwang, S., and Lim, J. (2022) · 2022
Later among the works it cites.
Graph Neural Networks for Relational Inductive Bias in Vision-based Deep Reinforcement Learning of Robot Control
Oliva, M., Banik, S., Josifovski, J., and Knoll, A. (2022) · 2022
Later among the works it cites.
Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., and Lindauer, M. (2022) · 2022
Later among the works it cites.
Hierarchical Reinforcement Learning: A Comprehensive Survey
Pateria, S., Subagdja, B., Tan, A., and Quek, C. (2022) · 2022
Later among the works it cites.
Towards an Interpretable Hierarchical Agent Framework Using Semantic Goals
Prakash, B., Waytowich, N., Oates, T., and Mohsenin, T. (2022) · 2022
Later among the works it cites.
SymNet 2.0: Effectively Handling Non-fluents and Actions in Generalized Neural Policies for RDDL Relational MDPs
Sharma, V., Arora, D., Geisser, F., Mausam, and Singla, P. (2022) · 2022
Later among the works it cites.
Hierarchical Representation Learning for Markov Decision Processes
Steccanella, L., Totaro, S., and Jonsson, A. (2022) · 2022
Later among the works it cites.
Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in Healthcare
Tang, S., Makar, M., Sjoding, M., Doshi-Velez, F., and Wiens, J. (2022) · 2022
Later among the works it cites.
Bayesian Generational Population-based Training
Wan, X., Lu, C., Parker-Holder, J., Ball, P., Nguyen, V., Ru, B., and Osborne, M. (2022) · 2022
Later among the works it cites.
Model-based Meta Reinforcement Learning Using Graph Structured Surrogate Models and Amortized Policy Search
Wang, Q., and van Hoof, H. (2022) · 2022
Later among the works it cites.
Denoised MDPs: Learning World Models Better Than the World Itself
Wang, T., Du, S., Torralba, A., Isola, P., Zhang, A., and Tian, Y. (2022) · 2022
Later among the works it cites.
Structure Learning-based Task Decomposition for Reinforcement Learning in Non-stationary Environments
Woo, H., Yoo, G., and Yoo, M. (2022) · 2022
Later among the works it cites.
Training a Resilient Q-network Against Observational Interference
Yang, C., Hung, I., Ouyang, Y., and Chen, P. (2022) · 2022
Later among the works it cites.
Reachability Constrained Reinforcement Learning
Yu, D., Ma, H., Li, S., and Chen, J. (2022) · 2022
Later among the works it cites.
APD: Learning Diverse Behaviors for Reinforcement Learning Through Unsupervised Active Pre-training
Zeng, K., Zhang, Q., Chen, B., Liang, B., and Yang, J. (2022) · 2022
Later among the works it cites.
A Survey of Knowledge-based Sequential Decision-making Under Uncertainty
Zhang, S., and Sridharan, M. (2022) · 2022
Later among the works it cites.
Policy Architectures for Compositional Generalization in Control
Zhou, A., Kumar, V., Finn, C., and Rajeswaran, A. (2022) · 2022
Later among the works it cites.
Human-timescale Adaptation in an Open-ended Task Space
Bauer, J., Baumli, K., Baveja, S., Behbahani, F., Bhoopchand, A., Bradley-Schmieg, N., Chang, M., Clay, N., Collister, A., Dasagi, V., Gonzalez, L., Gregor, K., Hughes, E., Kashem, S., Loks-Thompson, M., Openshaw, H., Parker-Holder, J., Pathak, S., Nieves, N., Rakicevic, N., Rocktäschel, T., Schroecker, Y., Sygnowski, J., Tuyls, K., York, S., Zacherl, A., and Zhang, L. (2023) · 2023
Closest in time.
A Survey of Meta-reinforcement Learning
Beck, J., Vuorio, R., Liu, E., Xiong, Z., Zintgraf, L., Finn, C., and Whiteson, S. (2023) · 2023
Closest in time.
Contextualize Me - The Case for Context in Reinforcement Learning
Benjamins, C., Eimer, T., Schubert, F., Mohan, A., Döhler, S., Biedenkapp, A., Rosenhahn, B., Hutter, F., and Lindauer, M. (2023) · 2023
Closest in time.
A Kernel Perspective on Behavioural Metrics for Markov Decision Processes
Castro, P., Kastner, T., Panangaden, P., and Rowland, M. (2023) · 2023
Closest in time.
Meta-reinforcement Learning via Exploratory Task Clustering
Chu, Z., and Wang, H. (2023) · 2023
Closest in time.
State and Action Abstraction for Search and Reinforcement Learning Algorithms
Dockhorn, A., and Kruse, R. (2023) · 2023
Closest in time.
Hyperparameters in Reinforcement Learning and How To Tune Them
Eimer, T., Lindauer, M., and Raileanu, R. (2023) · 2023
Closest in time.
Guided Reinforcement Learning: A Review and Evaluation for Efficient and Effective Real-World Robotics [Survey]
Eßer, J., Bach, N., Jestel, C., Urbann, O., and Kerner, S. (2023) · 2023
Closest in time.
Cell-free Latent Go-explore
Gallouedec, Q., and Dellandrea, E. (2023) · 2023
Closest in time.
Mastering Diverse Domains Through World Models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. (2023) · 2023
Closest in time.
The Big World Hypothesis and its Ramifications on Reinforcement Learning.
Javed, K. (2023) · 2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A., Lo, W., Dollár, P., and Girshick, R. (2023) · 2023
Closest in time.
A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T. (2023) · 2023
Closest in time.
Deep Laplacian-based Options for Temporally-extended Exploration
Klissarov, M., and Machado, M. (2023) · 2023
Closest in time.
Revisiting Bisimulation: A Sampling-based State Similarity Pseudo-metric
Lan, C., and Agarwal, R. (2023) · 2023
Closest in time.
Learning to Optimize for Reinforcement Learning
Lan, Q., Mahmood, A., Yan, S., and Xu, Z. (2023) · 2023
Closest in time.
Scaling Up Q-learning via Exploiting State-action Equivalence
lyu, Y., Côme, A., Zhang, Y., and Talebi, M. (2023) · 2023
Closest in time.
Towards Deployable RL – what’s Broken With RL Research and a Potential Fix
Mannor, S., and Tamar, A. (2023) · 2023
Closest in time.
Feudal Graph Reinforcement Learning
Marzi, T., Khehra, A., Cini, A., and Alippi, C. (2023) · 2023
Closest in time.
Model-based Reinforcement Learning: A Survey
Moerland, T., Broekens, J., Plaat, A., and Jonker, C. (2023) · 2023
Closest in time.
AutoRL Hyperparameter Landscapes
Mohan, A., Benjamins, C., Wienecke, K., Dockhorn, A., and Lindauer, M. (2023) · 2023
Closest in time.
OpenAI (2023) · 2023
Closest in time.
A Survey on Offline Reinforcement Learning: Taxonomy, Review, and Open Problems
Prudencio, R., Maximo, M., and Colombini, E. (2023) · 2023
Closest in time.
Ensemble Reinforcement Learning: A Survey
Song, Y., Suganthan, P., Pedrycz, W., Ou, J., He, Y., and Chen, Y. (2023) · 2023
Closest in time.
SMART: Self-supervised Multi-task pretrAining With contRol Transformers
Sun, Y., Ma, S., Madaan, R., Bonatti, R., Huang, F., and Kapoor, A. (2023) · 2023
Closest in time.
Reinforcement Learning With Exogenous States and Rewards
Trimponias, G., and Dietterich, T. (2023) · 2023
Closest in time.
Optimal Goal-reaching Reinforcement Learning via Quasimetric Learning
Wang, T., Torralba, A., Isola, P., and Zhang, A. (2023) · 2023
Closest in time.
Augmented Modular Reinforcement Learning Based on Heterogeneous Knowledge
Wolf, L., and Musolesi, M. (2023) · 2023
Closest in time.
Sample Efficient Deep Reinforcement Learning via Local Planning
Yin, D., Thiagarajan, S., Lazic, N., Rajaraman, N., Hao, B., and Szepesvári, C. (2023) · 2023
Closest in time.
The Benefits of Model-based Generalization in Reinforcement Learning
Young, K., Ramesh, A., Kirsch, L., and Schmidhuber, J. (2023) · 2023
Closest in time.
Latent State Marginalization as a Low-cost Approach for Improving Exploration
Zhang, D., Courville, A., Bengio, Y., Zheng, Q., Zhang, A., and Chen, R. (2023) · 2023
Closest in time.
Decision Transformer is a Robust Contender for Offline Reinforcement Learning
Bhargava, P., Chitnis, R., Geramifard, A., Sodhani, S., and Zhang, A. (2024) · 2024
Closest in time.
Motif: Intrinsic Motivation from Artificial Intelligence Feedback
Klissarov, M., D’Oro, P., Sodhani, S., Raileanu, R., Bacon, P., Vincent, P., Zhang, A., and Henaff, M. (2024) · 2024
Closest in time.