Fetching the paper…
Reading the bibliography…
The MineRL competition is designed for the development of reinforcement learning and imitation learning algorithms that can efficiently leverage human demonstrations to drastically reduce the number of environment interactions needed to solve the complex \emph{ObtainDiamond} task with sparse rewards.
MacQueen, J., et al.: Some methods for classification and analysis of multivariate observations. In: Proceedings of the fifth Berkeley symposium on mathematical statistics and probability. vol. 1, pp. 281–297. Oakland, CA, USA (1967)
1967
Earlier work this paper cites.
Lloyd, S.: Least squares quantization in pcm. IEEE transactions on information theory 28
1982
Earlier work this paper cites.
Likas, A., Vlassis, N., Verbeek, J.J.: The global k-means clustering algorithm. Pattern recognition 36
2003
Earlier work this paper cites.
Blei, D.M., Griffiths, T.L., Jordan, M.I.: The nested chinese restaurant process and bayesian nonparametric inference of topic hierarchies. Journal of the ACM (JACM) 57
2010
Earlier work this paper cites.
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., Riedmiller, M.: Deterministic policy gradient algorithms. In: ICML. pp. 387–395. PMLR (2014)
2014
Earlier work this paper cites.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., et al.: Human-level control through deep reinforcement learning. nature 518
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Identity mappings in deep residual networks. In: European conference on computer vision. pp. 630–645. Springer (2016)
2016
Earlier work this paper cites.
Ho, J., Ermon, S.: Generative adversarial imitation learning. Advances in neural information processing systems 29
2016
Earlier work this paper cites.
Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., Wierstra, D.: Continuous control with deep reinforcement learning. In: ICLR (2016)
2016
Earlier work this paper cites.
Gu, S., Holly, E., Lillicrap, T., Levine, S.: Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In: 2017 IEEE international conference on robotics and automation (ICRA). pp. 3389–3396. IEEE (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al.: Deep q-learning from demonstrations. In: Thirty-second AAAI conference on artificial intelligence (2018)
2018
Cited alongside, same era.
Kang, B., Jie, Z., Feng, J.: Policy optimization with demonstrations. In: International Conference on Machine Learning. pp. 2469–2478. PMLR (2018)
2018
Cited alongside, same era.
Osa, T., Pajarinen, J., Neumann, G., Bagnell, J.A., Abbeel, P., Peters, J., et al.: An algorithmic perspective on imitation learning. Foundations and Trends in Robotics 7
2018
Cited alongside, same era.
Reddy, S., Dragan, A.D., Levine, S.: Sqil: Imitation learning via reinforcement learning with sparse rewards. In: ICLR (2019)
2019
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Mao, H., Liu, W., Hao, J., Luo, J., Li, D., Zhang, Z., Wang, J., Xiao, Z.: Neighborhood cognition consistent multi-agent reinforcement learning. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 7219–7226 (2020)
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R.S., Barto, A.G.: Reinforcement learning: An introduction. MIT press (2018)
2018
Cited alongside, same era.
Fujimoto, S., Meger, D., Precup, D.: Off-policy deep reinforcement learning without exploration. In: International Conference on Machine Learning. pp. 2052–2062. PMLR (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Guss, W.H., Codel, C., Hofmann, K., Houghton, B., Kuno, N., Milani, S., Mohanty, S., Perez Liebana, D., Salakhutdinov, R., Topin, N., et al.: The minerl competition on sample efficient reinforcement learning using human priors. arXiv e-prints (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Mao, H., Zhang, Z., Xiao, Z., Gong, Z.: Modelling the dynamic joint policy of teammates with attention multi-agent ddpg. In: Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems (2019)
2019
Cited alongside, same era.
Mao, H., Zhang, Z., Xiao, Z., Gong, Z., Ni, Y.: Learning agent communication under limited bandwidth by message pruning. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 5142–5149 (2020)
2020
Later among the works it cites.
Mao, H., Zhang, Z., Xiao, Z., Gong, Z., Ni, Y.: Learning multi-agent communication with double attentional deep reinforcement learning. Autonomous Agents and Multi-Agent Systems 34
2020
Later among the works it cites.
Milani, S., Topin, N., Houghton, B., Guss, W.H., Mohanty, S.P., Nakata, K., Vinyals, O., Kuno, N.S.: Retrospective analysis of the 2019 minerl competition on sample efficient reinforcement learning. In: NeurIPS 2019 Competition and Demonstration Track. pp. 203–214. PMLR (2020)
2020
Later among the works it cites.
Scheller, C., Schraner, Y., Vogel, M.: Sample efficient reinforcement learning through learning from demonstrations in minecraft. In: NeurIPS 2019 Competition and Demonstration Track. pp. 67–76. PMLR (2020)
2020
Later among the works it cites.
2021
Closest in time.
Guss, W.H., Milani, S., Topin, N., Houghton, B., Mohanty, S., Melnik, A., Harter, A., Buschmaas, B., Jaster, B., Berganski, C., et al.: Towards robust and domain agnostic reinforcement learning competitions: Minerl 2020. In: NeurIPS 2020 Competition and Demonstration Track. pp. 233–252. PMLR (2021)
2021
Closest in time.
Skrynnik, A., Staroverov, A., Aitygulov, E., Aksenov, K., Davydov, V., Panov, A.I.: Hierarchical deep q-network from imperfect demonstrations in minecraft. Cognitive Systems Research 65
2021
Closest in time.