Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) tasks are challenging to implement, execute and test due to algorithmic instability, hyper-parameter sensitivity, and heterogeneous distributed communication patterns.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Earlier work this paper cites.
keras-rl
Plappert, M · 2016
Earlier work this paper cites.
Cntk: Microsoft’s open-source deep-learning toolkit
Seide, F. and Agarwal, A · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Tf.learn: Tensorflow’s high-level module for distributed machine learning
Tang, Y · 2016
Earlier work this paper cites.
A computational model for tensorflow: An introduction
Abadi, M., Isard, M., and Murray, D. G · 2017
Earlier work this paper cites.
Reinforcement learning coach, December 2017
Caspi, I., Leibovich, G., and Novik, G · 2017
Earlier work this paper cites.
Sonnet: TensorFlow-based neural network library
DeepMind · 2017
Earlier work this paper cites.
ONNX- Open Neural Network Exchange Format
Facebook Inc · 2017
Earlier work this paper cites.
Tensorflow agents: Efficient batched reinforcement learning in tensorflow
Hafner, D., Davidson, J., and Vanhoucke, V · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Cited alongside, same era.
Ray: A distributed framework for emerging AI applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Paul, W., Jordan, M. I., and Stoica, I · 2017
Cited alongside, same era.
Weld: A common runtime for high performance data analytics
Palkar, S., Thomas, J. J., Shanbhag, A., Narayanan, D., Pirk, H., Schwarzkopf, M., Amarasinghe, S., Zaharia, M., and InfoLab, S · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D · 2018
Closest in time.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2018
Closest in time.
Rllib: Abstractions for distributed reinforcement learning
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Goldberg, K., Gonzalez, J., Jordan, M., and Stoica, I · 2018
Closest in time.
Simple random search provides a competitive approach to reinforcement learning
Mania, H., Guy, A., and Recht, B · 2018
Closest in time.
Hierarchical planning for device placement
Mirhoseini, A., Goldie, A., Pham, H., Steiner, B., Le, Q. V., and Dean, J · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Openai baselines
Sidor, S. and Schulman, J · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Domain randomization and generative models for robotic grasping
Tobin, J., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
Dopamine
Bellemare, M. G., Castro, P. S. C., Gelada, C., Kumar, S., and Moitra, S · 2018
Cited alongside, same era.
Tvm: An automated end-to-end optimizing compiler for deep learning
Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Cowan, M., Shen, H., Wang, L., Hu, Y., Ceze, L., Guestrin, C., and Krishnamurthy, A · 2018
Cited alongside, same era.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Horizon: Facebook’s open source applied reinforcement learning platform
Gauci, J., Conti, E., Liang, Y., Virochsiri, K., He, Y., Kaden, Z., Narayanan, V., and Ye, X · 2018
Cited alongside, same era.
Autograph: Imperative-style coding with graph-based performance
Moldovan, D., Decker, J. M., Wang, F., Johnson, A. A., Lee, B. K., Nado, Z., Sculley, D., Rompf, T., and Wiltschko, A. B · 2018
Closest in time.
The impact of nondeterminism on reproducibility in deep reinforcement learning
Nagarajan, P., Warnell, G., and Stone, P · 2018
Closest in time.
OpenAI Five DOTA
OpenAI · 2018
Closest in time.
Gluon - A clear, concise, simple yet powerful and efficient API for deep learning
Rochel, S. et al · 2018
Closest in time.
Lift: Reinforcement learning in computer systems by learning from demonstrations
Schaarschmidt, M., Kuhnle, A., Ellis, B., Fricke, K., Gessert, F., and Yoneki, E · 2018
Closest in time.
Horovod: fast and easy distributed deep learning in tensorflow
Sergeev, A. and Balso, M. D · 2018
Closest in time.
Unsupervised predictive memory in a goal-directed agent
Wayne, G., Hung, C., Amos, D., Mirza, M., Ahuja, A., Grabska-Barwinska, A., Rae, J. W., Mirowski, P., Leibo, J. Z., Santoro, A., Gemici, M., Reynolds, M., Harley, T., Abramson, J., Mohamed, S., Rezende, D. J., Saxton, D., Cain, A., Hillier, C., Silver, D., Kavukcuoglu, K., Botvinick, M., Hassabis, D., and Lillicrap, T. P · 2018
Closest in time.
Dynamic control flow in large-scale machine learning
Yu, Y., Abadi, M., Barham, P., Brevdo, E., Burrows, M., Davis, A., Dean, J., Ghemawat, S., Harley, T., Hawkins, P., Isard, M., Kudlur, M., Monga, R., Murray, D., and Zheng, X · 2018
Closest in time.