Fetching the paper…
Reading the bibliography…
Action advising endeavors to leverage supplementary guidance from expert teachers to alleviate the issue of sampling inefficiency in Deep Reinforcement Learning (DRL).
2013
Earlier work this paper cites.
L. Torrey and M. E. Taylor, “Teaching on a budget: agents advising agents in reinforcement learning,” in International conference on Autonomous Agents and Multi-Agent Systems , 2013
2013
Earlier work this paper cites.
S. Griffith, K. Subramanian, J. Scholz, C. L. Isbell, and A. L. Thomaz, “Policy shaping: Integrating human feedback with reinforcement learning,” Conference on Neural Information Processing Systems , 2013
2013
Earlier work this paper cites.
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy, “Deep exploration via bootstrapped dqn,” in Conference on Neural Information Processing Systems , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2016, pp. 770–778
2016
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning , 2016
2016
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Conference on Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
B. Sangiovanni, A. Rendiniello, G. P. Incremona, A. Ferrara, and M. Piastra, “Deep reinforcement learning for collision avoidance of robotic manipulators,” in European Control Conference , 2018
2018
Earlier work this paper cites.
R. Toro Icarte, T. Q. Klassen, R. A. Valenzano, and S. A. McIlraith, “Advice-based exploration in model-based reinforcement learning,” in Canadian Conference on Artificial Intelligence , 2018
2018
Earlier work this paper cites.
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in AAAI Conference on Artificial Intelligence , 2018
2018
Earlier work this paper cites.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” Nature , 2019
2019
Earlier work this paper cites.
J. Chen, B. Yuan, and M. Tomizuka, “Model-free deep reinforcement learning for urban autonomous driving,” in IEEE intelligent transportation systems conference , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
E. Ilhan, J. Gow, and D. Perez-Liebana, “Teaching on a budget in multi-agent deep reinforcement learning,” in IEEE Conference on Games , 2019
2019
Earlier work this paper cites.
M. Ye, X. Zhang, P. C. Yuen, and S.-F. Chang, “Unsupervised embedding learning via invariant and spreading instance feature,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2019, pp. 6210–6219
2019
Cited alongside, same era.
Y. Burda, H. Edwards, A. Storkey, and O. Klimov, “Exploration by random network distillation,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
A. G. Barto, R. S. Sutton, and C. W. Anderson, “Looking back on the actor–critic architecture,” IEEE Transactions on Systems, Man, and Cybernetics: Systems , 2020
2020
Cited alongside, same era.
D. Ye, Z. Liu, M. Sun, B. Shi, P. Zhao, H. Wu, H. Yu, S. Yang, X. Wu, Q. Guo et al. , “Mastering complex control in moba games with deep reinforcement learning,” in AAAI Conference on Artificial Intelligence , 2020
2020
Cited alongside, same era.
D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus, “Improving sample efficiency in model-free reinforcement learning from images,” in AAAI Conference on Artificial Intelligence , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
L. Zhou and K. Small, “Inverse reinforcement learning with natural language goals,” in AAAI Conference on Artificial Intelligence , 2021
2021
Later among the works it cites.
E. Ilhan, J. Gow, and D. Perez, “Student-initiated action advising via advice novelty,” IEEE Transactions on Games , 2021
2021
Later among the works it cites.
S. Arora and P. Doshi, “A survey of inverse reinforcement learning: Challenges, methods and progress,” Artificial Intelligence , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray et al. , “Learning dexterous in-hand manipulation,” The International Journal of Robotics Research , 2020
2020
Cited alongside, same era.
K. Zhou, S. Song, A. Xue, K. You, and H. Wu, “Smart train operation algorithms based on expert knowledge and reinforcement learning,” IEEE Transactions on Systems, Man, and Cybernetics: Systems , 2020
2020
Cited alongside, same era.
L. Yang, Q. Sun, N. Zhang, and Z. Liu, “Optimal energy operation strategy for we-energy of energy internet based on hybrid reinforcement learning with human-in-the-loop,” IEEE Transactions on Systems, Man, and Cybernetics: Systems , 2020
2020
Cited alongside, same era.
F. L. D. Silva, P. Hernandez-Leal, B. Kartal, and M. E. Taylor, “Uncertainty-aware action advising for deep reinforcement learning agents,” in AAAI Conference on Artificial Intelligence , 2020
2020
Cited alongside, same era.
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” Annual Conference on Neural Information Processing Systems , 2020
2020
Cited alongside, same era.
Y. Tian, D. Krishnan, and P. Isola, “Contrastive multiview coding,” in European Conference on Computer Vision , 2020
2020
Cited alongside, same era.
K. He, H. Fan, Y. Wu, S. Xie, and R. B. Girshick, “Momentum contrast for unsupervised visual representation learning,” in IEEE Conference on Computer Vision and Pattern Recognition , 2020
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. E. Hinton, “A simple framework for contrastive learning of visual representations,” in International Conference on Machine Learning , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Bignold, F. Cruz, R. Dazeley, P. Vamplew, and C. Foale, “Persistent rule-based interactive reinforcement learning,” Neural Computing and Applications , 2021
2021
Later among the works it cites.
E. Ilhan, J. Gow, and D. P. Liebana, “Action advising with advice imitation in deep reinforcement learning,” in International Conference on Autonomous Agents and Multiagent Systems , 2021
2021
Later among the works it cites.
Z. Ye, Y. Chen, X. Jiang, G. Song, B. Yang, and S. Fan, “Improving sample efficiency in multi-agent actor-critic methods,” Applied Intelligence , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
D. Harnack, J. Pivin-Bachler, and N. Navarro-Guerrero, “Quantifying the effect of feedback frequency in interactive reinforcement learning for robotic tasks,” Neural Computing and Applications , 2022
2022
Later among the works it cites.
S. Liu, K. Chen, N. Yu, J. Song, Z. Feng, and M. Song, “Ask-ac: An initiative advisor-in-the-loop actor–critic framework,” IEEE Transactions on Systems, Man, and Cybernetics: Systems , pp. 1–12, 2023
2023
Closest in time.
OpenAI, “GPT-4 technical report,” arXiv preprint arXiv.2303.08774 , 2023
2023
Closest in time.
A. Bignold, F. Cruz, M. E. Taylor, T. Brys, R. Dazeley, P. Vamplew, and C. Foale, “A conceptual framework for externally-influenced agents: An assisted reinforcement learning review,” Journal of Ambient Intelligence and Humanized Computing , 2023
2023
Closest in time.
M. Towers, J. K. Terry, A. Kwiatkowski, J. U. Balis, G. d. Cola, T. Deleu, M. Goulão, A. Kallinteris, A. KG, M. Krimmel, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, A. T. J. Shen, and O. G. Younis, “Gymnasium,” Mar. 2023. [Online]. Available: https://zenodo.org/record/8127025
2023
Closest in time.