Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (DRL) and Deep Multi-agent Reinforcement Learning (MARL) have achieved significant successes across a wide range of domains, including game AI, autonomous vehicles, robotics, and so on.
W. R. Thompson, “On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,” Biometrika , vol. 25, no. 3/4, pp. 285–294, 1933
1933
Earlier work this paper cites.
T. L. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in applied mathematics , vol. 6, no. 1, pp. 4–22, 1985
1985
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, pp. 229–256, 1992
1992
Earlier work this paper cites.
P. Dayan, “Improving generalization for temporal difference learning: The successor representation,” Neural Computation , vol. 5, no. 4, pp. 613–624, 1993
1993
Earlier work this paper cites.
D. A. Nix and A. S. Weigend, “Estimating the mean and variance of the target probability distribution,” in ICNN , vol. 1, 1994, pp. 55–60
1994
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning - an introduction , ser. Adaptive computation and machine learning, 1998
1998
Earlier work this paper cites.
R. Dearden, N. Friedman, and S. J. Russell, “Bayesian q-learning,” in AAAI , 1998
1998
Earlier work this paper cites.
D. Carmel and S. Markovitch, “Exploration strategies for model-based learning in multi-agent systems: Exploration strategies,” JAAMAS , vol. 2, no. 2, pp. 141–172, 1999
1999
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML , 1999
1999
Earlier work this paper cites.
R. M. Ryan and E. L. Deci, “Intrinsic and extrinsic motivations: Classic definitions and new directions,” Contemporary educational psychology , vol. 25, no. 1, pp. 54–67, 2000
2000
Earlier work this paper cites.
M. J. A. Strens, “A bayesian framework for reinforcement learning,” in ICML , 2000
2000
Earlier work this paper cites.
P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,” Machine learning , vol. 47, no. 2-3, pp. 235–256, 2002
2002
Earlier work this paper cites.
P. Auer, “Using confidence bounds for exploitation-exploration trade-offs,” JMLR , vol. 3, pp. 397–422, 2002
2002
Earlier work this paper cites.
G. Chalkiadakis and C. Boutilier, “Coordination in multiagent reinforcement learning: a bayesian approach,” in AAMAS , 2003
2003
Earlier work this paper cites.
L. Itti and P. Baldi, “Bayesian surprise attracts human attention,” in NeurIPS , 2005
2005
Earlier work this paper cites.
K. Verbeeck, A. Nowé, M. Peeters, and K. Tuyls, “Multi-agent reinforcement learning in stochastic single and multi-stage games,” in ALAMAS , 2005
2005
Earlier work this paper cites.
L. Busoniu, R. Babuska, and B. De Schutter, “A comprehensive survey of multiagent reinforcement learning,” TSMC , vol. 38, no. 2, pp. 156–172, 2008
2008
Earlier work this paper cites.
P. Auer, T. Jaksch, and R. Ortner, “Near-optimal regret bounds for reinforcement learning,” in NeurIPS , 2009
2009
Earlier work this paper cites.
N. Srinivas, A. Krause, S. M. Kakade, and M. W. Seeger, “Gaussian process optimization in the bandit setting: No regret and experimental design,” in ICML , 2010
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
M. Agueh and G. Carlier, “Barycenters in the wasserstein space,” SIAM J. Math. Anal. , vol. 43, no. 2, pp. 904–924, 2011
2011
Earlier work this paper cites.
A. Graves, “Practical variational inference for neural networks,” in NeurIPS , 2011
2011
Earlier work this paper cites.
T. Bolander and M. B. Andersen, “Epistemic planning for single and multi-agent systems,” JANCL , vol. 21, no. 1, pp. 9–34, 2011
2011
Earlier work this paper cites.
E. Kaufmann, O. Cappé, and A. Garivier, “On bayesian upper confidence bounds for bandit problems,” in AISTATS , 2012
2012
Earlier work this paper cites.
A. G. Barto, “Intrinsic motivation and reinforcement learning,” in Intrinsically motivated learning in natural and artificial systems , 2013
2013
Earlier work this paper cites.
I. Osband, D. Russo, and B. V. Roy, “(more) efficient reinforcement learning via posterior sampling,” in NeurIPS , 2013
2013
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. A. Riedmiller, “Deterministic policy gradient algorithms,” in ICML , 2014
2014
Earlier work this paper cites.
C. Salge, C. Glackin, and D. Polani, “Empowerment–an introduction,” in Guided Self-Organization: Inception , 2014
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. P. Singh, “Action-conditional video prediction using deep networks in atari games,” in NeurIPS , 2015
2015
Earlier work this paper cites.
J. García and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , vol. 16, pp. 1437–1480, 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization,” in ICML , 2015
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in ICLR , 2016
2016
Earlier work this paper cites.
S. S. Mousavi, M. Schukat, and E. Howley, “Deep reinforcement learning: an overview,” in SAI Intelligent Systems Conference , 2016
2016
Earlier work this paper cites.
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” in NeurIPS , 2016
2016
Earlier work this paper cites.
H. van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in AAAI , 2016
2016
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas, “Dueling network architectures for deep reinforcement learning,” in ICML , 2016
2016
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” in ICLR , 2016
2016
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML , 2016
2016
Earlier work this paper cites.
M. J. Hausknecht and P. Stone, “Deep reinforcement learning in parameterized action space,” in ICLR , 2016
2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, “Openai gym,” 2016
2016
Earlier work this paper cites.
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaskowski, “Vizdoom: A doom-based AI research platform for visual reinforcement learning,” in CIG , 2016
2016
Earlier work this paper cites.
I. Osband, B. V. Roy, and Z. Wen, “Generalization and exploration via randomized value functions,” in ICML , 2016
2016
Earlier work this paper cites.
I. Osband, C. Blundell, A. Pritzel, and B. V. Roy, “Deep exploration via bootstrapped DQN,” in NeurIPS , 2016
2016
Earlier work this paper cites.
K. Azizzadenesheli, A. Lazaric, and A. Anandkumar, “Reinforcement learning of pomdps using spectral methods,” in CoLT , 2016
2016
Earlier work this paper cites.
Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in ICML , 2016
2016
Earlier work this paper cites.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. D. Turck, and P. Abbeel, “VIME: variational information maximizing exploration,” in NeurIPS , 2016
2016
Earlier work this paper cites.
A. van den Oord, N. Kalchbrenner, L. Espeholt, K. Kavukcuoglu, O. Vinyals, and A. Graves, “Conditional image generation with pixelcnn decoders,” in NeurIPS , 2016
2016
Earlier work this paper cites.
M. Turchetta, F. Berkenkamp, and A. Krause, “Safe exploration in finite markov decision processes with gaussian processes,” in NeurIPS , 2016, pp. 4305–4313
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in ICML , 2017
2017
Earlier work this paper cites.
Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274 , 2017
2017
Earlier work this paper cites.
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Processing Magazine , vol. 34, no. 6, pp. 26–38, 2017
2017
Earlier work this paper cites.
M. G. Bellemare, W. Dabney, and R. Munos, “A distributional perspective on reinforcement learning,” in ICML , 2017
2017
Earlier work this paper cites.
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in NeurIPS , 2017
2017
Earlier work this paper cites.
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor, “A deep hierarchical approach to lifelong learning in minecraft,” in AAAI , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, D. Silver, and H. van Hasselt, “Successor features for transfer in reinforcement learning,” in NeurIPS , 2017
2017
Earlier work this paper cites.
H. Tang, R. Houthooft, D. Foote, A. Stooke, X. Chen, Y. Duan, J. Schulman, F. D. Turck, and P. Abbeel, “#exploration: A study of count-based exploration for deep reinforcement learning,” in NeurIPS , 2017
2017
Earlier work this paper cites.
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos, “Count-based exploration with neural density models,” in ICML , 2017
2017
Earlier work this paper cites.
J. Martin, S. N. Sasikumar, T. Everitt, and M. Hutter, “Count-based exploration in feature space for reinforcement learning,” in IJCAI , 2017
2017
Earlier work this paper cites.
J. Fu, J. D. Co-Reyes, and S. Levine, “EX2: exploration with exemplar models for deep reinforcement learning,” in NeurIPS , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, “Deep variational information bottleneck,” in ICLR , 2017
2017
Earlier work this paper cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in ICML , vol. 70. PMLR, 2017, pp. 22–31
2017
Earlier work this paper cites.
M. Chakraborty, K. Y. P. Chua, S. Das, and B. Juba, “Coordinated versus decentralized exploration in multi-agent multi-armed bandits,” in IJCAI , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
G. Kalweit and J. Boedecker, “Uncertainty-driven imagination for continuous deep reinforcement learning,” in CoRL , 2017, pp. 195–206
2017
Earlier work this paper cites.
S. Agrawal and R. Jia, “Posterior sampling for reinforcement learning: worst-case regret bounds,” in Advances in Neural Information Processing Systems , 2017, pp. 1184–1194
2017
Earlier work this paper cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba, “Hindsight experience replay,” in NeurIPS , 2017
2017
Earlier work this paper cites.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in ICML , ser. Proceedings of Machine Learning Research, vol. 70, 2017, pp. 1352–1361
2017
Cited alongside, same era.
Y. Liu, P. Ramachandran, Q. Liu, and J. Peng, “Stein variational policy gradient,” in UAI , 2017
2017
Cited alongside, same era.
T. Rashid, M. Samvelyan, C. S. de Witt, G. Farquhar, J. N. Foerster, and S. Whiteson, “QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning,” in ICML , 2018
2018
Cited alongside, same era.
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos, “Distributional reinforcement learning with quantile regression,” in AAAI , 2018
2018
Cited alongside, same era.
S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in ICML , 2018
M. Fang, C. Zhou, B. Shi, B. Gong, J. Xu, and T. Zhang, “Dher: Hindsight experience replay for dynamic goals,” in ICLR , 2019
2019
Later among the works it cites.
H. Liu, A. Trott, R. Socher, and C. Xiong, “Competitive experience replay,” in ICLR , 2019
2019
Later among the works it cites.
R. Zhao, X. Sun, and V. Tresp, “Maximum entropy-regularized multi-goal reinforcement learning,” in ICML , 2019
2019
Later among the works it cites.
M. Fang, T. Zhou, Y. Du, L. Han, and Z. Zhang, “Curriculum-guided hindsight experience replay,” in NeurIPS , 2019
2019
Later among the works it cites.
Z. Ren, K. Dong, Y. Zhou, Q. Liu, and J. Peng, “Exploration via hindsight goal generation,” in NeurIPS , 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
P. I. Frazier, “A tutorial on bayesian optimization,” arXiv preprint arXiv:1807.02811 , 2018
2018
Cited alongside, same era.
I. Osband, J. Aslanides, and A. Cassirer, “Randomized prior functions for deep reinforcement learning,” in NeurIPS , 2018
2018
Cited alongside, same era.
K. Azizzadenesheli, E. Brunskill, and A. Anandkumar, “Efficient exploration through bayesian deep q-networks,” in Information Theory and Applications Workshop , 2018
2018
Cited alongside, same era.
B. O’Donoghue, I. Osband, R. Munos, and V. Mnih, “The uncertainty bellman equation and exploration,” in ICML , 2018
2018
Cited alongside, same era.
J. Kirschner and A. Krause, “Information directed sampling and bandits with heteroscedastic noise,” in CoLT , 2018
2018
Cited alongside, same era.
L. Fox, L. Choshen, and Y. Loewenstein, “DORA the explorer: Directed outreaching reinforcement action-selection,” in ICLR , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine, “Diversity is all you need: Learning skills without a reward function,” in ICLR , 2019
2019
Later among the works it cites.
T. Gangwani, Q. Liu, and J. Peng, “Learning self-imitating diverse policies,” in ICLR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Later among the works it cites.
H.-n. Wang, N. Liu, Y.-y. Zhang, D.-w. Feng, F. Huang, D.-s. Li, and Y.-m. Zhang, “Deep reinforcement learning: a survey,” FRONT INFORM TECH EL , pp. 1–19, 2020
2020
Later among the works it cites.
C. Dann, “Strategic exploration in reinforcement learning - new algorithms and learning guarantees,” Ph.D. dissertation, School of Computer Science, Carnegie Mellon University, 2020
2020
Later among the works it cites.
T. Lattimore and C. Szepesvári, Bandit Algorithms . Cambridge University Press, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
T. Wang, J. Wang, Y. Wu, and C. Zhang, “Influence-based multi-agent exploration,” in ICLR , 2020
2020
Later among the works it cites.
F. Zhou, J. Wang, and X. Feng, “Non-crossing quantile regression for distributional reinforcement learning,” in NeurIPS , 2020
2020
Later among the works it cites.
A. Zanette, D. Brandfonbrener, E. Brunskill, M. Pirotta, and A. Lazaric, “Frequentist regret bounds for randomized least-squares value iteration,” in AISTATS , 2020
2020
Later among the works it cites.
M. C. Machado, M. G. Bellemare, and M. Bowling, “Count-based exploration with the successor representation,” in AAAI , 2020
2020
Later among the works it cites.
Y. Song, Y. Chen, Y. Hu, and C. Fan, “Exploring unknown states with action balance,” in CIG , 2020
2020
Later among the works it cites.
R. Y. Tao, V. François-Lavet, and J. Pineau, “Novelty search in representational space for sample efficient exploration,” in NeurIPS , 2020
2020
Later among the works it cites.
A. Stooke, J. Achiam, and P. Abbeel, “Responsive safety in reinforcement learning by PID lagrangian methods,” in ICML , vol. 119, 2020, pp. 9133–9143
2020
Later among the works it cites.
B. Thananjeyan, A. Balakrishna, U. Rosolia, F. Li, R. McAllister, J. E. Gonzalez, S. Levine, F. Borrelli, and K. Goldberg, “Safety augmented value estimation from demonstrations (SAVED): safe deep model-based RL for sparse cost robotic tasks,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 3612–3619, 2020
2020
Later among the works it cites.
T. G. Karimpanal, S. Rana, S. Gupta, T. Tran, and S. Venkatesh, “Learning transferable domain priors for safe exploration in reinforcement learning,” in IJCNN , 2020, pp. 1–10
2020
Later among the works it cites.
Y. Guo, J. Choi, M. Moczulski, S. Feng, S. Bengio, M. Norouzi, and H. Lee, “Memory based trajectory-conditioned policies for learning from sparse rewards,” in NeurIPS , 2020
2020
Later among the works it cites.
E. Zhao, S. Deng, Y. Zang, Y. Kang, K. Li, and J. Xing, “Potential driven reinforcement learning for hard exploration tasks,” in IJCAI , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
X. Lyu and C. Amato, “Likelihood quantile networks for coordinating multi-agent reinforcement learning,” in AAMAS , 2020
2020
Later among the works it cites.
Z. Zheng, J. Oh, M. Hessel, Z. Xu, M. Kroiss, H. van Hasselt, D. Silver, and S. Singh, “What can learned intrinsic rewards capture?” in ICML , 2020
2020
Later among the works it cites.
R. Chitnis, S. Tulsiani, S. Gupta, and A. Gupta, “Intrinsic motivation for encouraging synergistic behavior,” in ICLR , 2020
2020
Later among the works it cites.
G. Chen, “A new framework for multi-agent reinforcement learning - centralized training and exploration with decentralized execution via policy distillation,” in AAMAS , 2020
2020
Later among the works it cites.
A. A. Taiga, W. Fedus, M. C. Machado, A. Courville, and M. G. Bellemare, “On bonus based exploration methods in the arcade learning environment,” in ICLR , 2020
2020
Later among the works it cites.
G. Farquhar, L. Gustafson, Z. Lin, S. Whiteson, N. Usunier, and G. Synnaeve, “Growing action spaces,” in ICML , 2020
2020
Later among the works it cites.
M. Laskin, A. Srinivas, and P. Abbeel, “CURL: contrastive unsupervised representations for reinforcement learning,” in ICML , ser. Proceedings of Machine Learning Research, vol. 119, 2020, pp. 5639–5650
2020
Later among the works it cites.
K. H. Kim, M. Sano, J. De Freitas, N. Haber, and D. Yamins, “Active world model learning in agent-rich environments with progress curiosity,” in ICML , 2020
2020
Later among the works it cites.
Q. Cai, Z. Yang, C. Jin, and Z. Wang, “Provably efficient exploration in policy optimization,” in ICML , 2020
2020
Later among the works it cites.
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan, “Provably efficient reinforcement learning with linear function approximation,” in CoLT , 2020
2020
Later among the works it cites.
S. Thudumu, P. Branch, J. Jin, and J. J. Singh, “A comprehensive survey of anomaly detection techniques for high dimensional big data,” Journal of Big Data , vol. 7, no. 1, pp. 1–30, 2020
2020
Later among the works it cites.
T. T. Nguyen, N. D. Nguyen, P. Vamplew, S. Nahavandi, R. Dazeley, and C. P. Lim, “A multi-objective deep reinforcement learning framework,” EAAI , vol. 96, p. 103915, 2020
2020
Later among the works it cites.
K. Ciosek, V. Fortuin, R. Tomioka, K. Hofmann, and R. E. Turner, “Conservative uncertainty estimation by fitting prior networks,” in ICLR , 2020
2020
Later among the works it cites.
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak, “Planning to explore via self-supervised world models,” in ICML , 2020
2020
Later among the works it cites.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” in ICLR , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Curi, F. Berkenkamp, and A. Krause, “Efficient model-based reinforcement learning through optimistic policy search and planning,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine, “Skew-fit: State-covering self-supervised reinforcement learning,” in ICML , 2020
2020
Later among the works it cites.
A. Sharma, S. Gu, S. Levine, V. Kumar, and K. Hausman, “Dynamics-aware unsupervised discovery of skills,” in ICLR , 2020
2020
Later among the works it cites.
V. Campos, A. Trott, C. Xiong, R. Socher, X. Giro-i Nieto, and J. Torres, “Explore, discover and learn: Unsupervised discovery of state-covering skills,” in ICML , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Tang, “Self-imitation learning via generalized lower bound q-learning,” in NeurIPS , 2020
2020
Later among the works it cites.
M. Li, Z. Cao, and Z. Li, “A reinforcement learning-based vehicle platoon control strategy for reducing energy consumption in traffic oscillations,” IEEE Transactions on Neural Networks and Learning Systems , vol. 32, no. 12, pp. 5309–5322, 2021
2021
Closest in time.
L. Kong, W. He, W. Yang, Q. Li, and O. Kaynak, “Fuzzy approximation-based finite-time control for a robot with actuator saturation under time-varying constraints of work space,” IEEE Transactions on Cybernetics , vol. 51, no. 10, pp. 4873–4884, 2021
2021
Closest in time.
C. Sun, W. Liu, and L. Dong, “Reinforcement learning with task decomposition for cooperative multiagent systems,” IEEE Transactions on Neural Networks and Learning Systems , vol. 32, no. 5, pp. 2054–2065, 2021
2021
Closest in time.
J. Fu, X. Teng, C. Cao, Z. Ju, and P. Lou, “Robot motor skill transfer with alternate learning in two spaces,” IEEE Transactions on Neural Networks and Learning Systems , vol. 32, no. 10, pp. 4553–4564, 2021
2021
Closest in time.
2021
Closest in time.
K. Lee, M. Laskin, A. Srinivas, and P. Abbeel, “SUNRISE: A simple unified framework for ensemble learning in deep reinforcement learning,” in ICML , 2021
2021
Closest in time.
C. Bai, L. Wang, L. Han, J. Hao, A. Garg, P. Liu, and Z. Wang, “Principled exploration via optimistic bootstrapping and backward induction,” in ICML , 2021
2021
Closest in time.
C. Bai, P. Liu, K. Liu, L. Wang, Y. Zhao, L. Han, and Z. Wang, “Variational dynamic for self-supervised exploration in deep reinforcement learning,” IEEE Transactions on Neural Networks and Learning Systems , 2021
2021
Closest in time.
2021
Closest in time.
T. Zhang, H. Xu, X. Wang, Y. Wu, K. Keutzer, J. E. Gonzalez, and Y. Tian, “Noveld: A simple yet effective exploration criterion,” in Advances in Neural Information Processing Systems , 2021
2021
Closest in time.
H. Bharadhwaj, A. Kumar, N. Rhinehart, S. Levine, F. Shkurti, and A. Garg, “Conservative safety critics for exploration,” in ICLR , 2021
2021
Closest in time.
N. Hunt, N. Fulton, S. Magliacane, T. N. Hoang, S. Das, and A. Solar-Lezama, “Verifiably safe exploration for end-to-end reinforcement learning,” in HSCC , 2021, pp. 14:1–14:11
2021
Closest in time.
G. Thomas, Y. Luo, and T. Ma, “Safe reinforcement learning by imagining the near future,” in NeurIPS , 2021, pp. 13 859–13 869
2021
Closest in time.
E. Adrien, H. Joost, L. Joel, S. K. O, and C. Jeff, “First return, then explore,” Nature , vol. 590, no. 7847, pp. 580–586, 2021
2021
Closest in time.
W. Sun, C. Lee, and C. Lee, “DFAC framework: Factorizing the value function via quantile mixture for multi-agent distributional q-learning,” in ICML , 2021
2021
Closest in time.
I.-J. Liu, U. Jain, R. A. Yeh, and A. Schwing, “Cooperative exploration for multi-agent deep reinforcement learning,” in International Conference on Machine Learning , 2021, pp. 6826–6836
2021
Closest in time.
L. Zheng, J. Chen, J. Wang, J. He, Y. Hu, Y. Chen, C. Fan, Y. Gao, and C. Zhang, “Episodic multi-agent reinforcement learning with curiosity-driven exploration,” in NeurIPS , M. Ranzato, A. Beygelzimer, Y. N. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021, pp. 3757–3769
2021
Closest in time.
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto, “Reinforcement learning with prototypical representations,” in ICML , ser. Proceedings of Machine Learning Research, vol. 139, 2021, pp. 11 920–11 931
2021
Closest in time.
C. Bai, L. Wang, L. Han, A. Garg, J. Hao, P. Liu, and Z. Wang, “Dynamic bottleneck for robust self-supervised exploration,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Closest in time.
2021
Closest in time.
J. Ferret, O. Pietquin, and M. Geist, “Self-imitation advantage learning,” in AAMAS , 2021, pp. 501–509
2021
Closest in time.
2021
Closest in time.
P. Ladosz, L. Weng, M. Kim, and H. Oh, “Exploration in deep reinforcement learning: A survey,” Information Fusion , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
W. Saunders, G. Sastry, A. Stuhlmüller, and O. Evans, “Trial without error: Towards safe reinforcement learning via human intervention,” in AAMAS , 2018, pp. 2067–2069
2069
Closest in time.