Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms involve the deep nesting of highly irregular computation patterns, each of which typically exhibits opportunities for distributed computation.
Encapsulation of parallelism and architecture-independence in extensible database query execution
Graefe, G. and Davison, D. L · 1993
Earlier work this paper cites.
A high-performance, portable implementation of the MPI message passing interface standard
Gropp, W., Lusk, E., Doss, N., and Skjellum, A · 1996
Earlier work this paper cites.
Bigtable: A distributed storage system for structured data
Chang, F., Dean, J., Ghemawat, S., Hsieh, W. C., Wallach, D. A., Burrows, M., Chandra, T., Fikes, A., and Gruber, R. E · 2008
Earlier work this paper cites.
MapReduce: simplified data processing on large clusters
Dean, J. and Ghemawat, S · 2008
Earlier work this paper cites.
Composing parallel software efficiently with Lithe
Pan, H., Hindman, B., and Asanović, K · 2010
Earlier work this paper cites.
Spark: Cluster computing with working sets
Zaharia, M., Chowdhury, N. M., Franklin, M., Shenker, S., and Stoica, I · 2010
Earlier work this paper cites.
Scientific computing with EC2 spot instances
Amazon · 2011
Earlier work this paper cites.
The datacenter as a computer: An introduction to the design of warehouse-scale machines
Barroso, L. A., Clidaras, J., and Hölzle, U · 2013
Earlier work this paper cites.
The tail at scale
Dean, J. and Barroso, L. A · 2013
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A., and Ahmed, A · 2014
Earlier work this paper cites.
Preemptible virtual machines
Google · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., et al · 2016
Cited alongside, same era.
OpenAI gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Hyperband: Bandit-based configuration evaluation for hyperparameter optimization
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A · 2016
Cited alongside, same era.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al · 2017
Closest in time.
PyTorch implementation of advantage actor critic (A2C), proximal policy optimization (PPO) and scalable trust-region method for deep reinforcement learning
Kostrikov, I · 2017
Closest in time.
ONNX: Open neural network exchange format
Microsoft · 2017
Closest in time.
Ray: A distributed framework for emerging AI applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Paul, W., Jordan, M. I., and Stoica, I · 2017
Closest in time.
Evolution Strategies Starter Agent
OpenAI · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Universe Starter Agent
OpenAI · 2016
Cited alongside, same era.
Amazon EC2 pricing
Amazon · 2017
Cited alongside, same era.
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
Reinforcement learning coach by Intel
Caspi, I · 2017
Cited alongside, same era.
NNVM compiler: Open compiler for AI frameworks
DMLC · 2017
Cited alongside, same era.
The Gluon API specification
Gluon · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Closest in time.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., and Sutskever, I · 2017
Closest in time.
TensorForce: A TensorFlow library for applied reinforcement learning
Schaarschmidt, M., Kuhnle, A., and Fricke, K · 2017
Closest in time.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Closest in time.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Closest in time.
ELF: an extensive, lightweight and flexible research platform for real-time strategy games
Tian, Y., Gong, Q., Shang, W., Wu, Y., and Zitnick, L · 2017
Closest in time.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Closest in time.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D · 2018
Closest in time.