Fetching the paper…
Reading the bibliography…
Low-precision training has become a popular approach to reduce compute requirements, memory footprint, and energy consumption in supervised learning.
Further remarks on reducing truncation errors, commun
Kahan, W · 1965
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Ieee standard 754 for binary floating-point arithmetic
Kahan, W · 1996
Earlier work this paper cites.
Adaptive vector quantization for reinforcement learning
Lau, H., Mak, K., and Lee, I · 2002
Earlier work this paper cites.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2004
Earlier work this paper cites.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2004
Earlier work this paper cites.
Precision & performance: Floating point and ieee 754 compliance for nvidia gpus
Whitehead, N. and Fit-Florea, A · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Reinforcement learning through asynchronous advantage actor-critic on a gpu
Babaeizadeh, M., Frosio, I., Tyree, S., Clemons, J., and Kautz, J · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Lin, D., Talathi, S., and Annapureddy, S · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Understanding and optimizing asynchronous low-precision stochastic gradient descent
De Sa, C., Feldman, M., Ré, C., and Olukotun, K · 2017
Earlier work this paper cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Earlier work this paper cites.
Rllib: Abstractions for distributed reinforcement learning. arxiv e-prints, art
Liang, E., Liaw, R., Moritz, P., Nishihara, R., Fox, R., Goldberg, K., Gonzalez, J. E., Jordan, M. I., and Stoica, I · 2017
Earlier work this paper cites.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al · 2017
Earlier work this paper cites.
Song, Z., Liu, Z., and Wang, D · 2017
Earlier work this paper cites.
Mixed precision training of convolutional neural networks using integer operations
Das, D., Mellempudi, N., Mudigere, D., Kalamkar, D., Avancha, S., Banerjee, K., Sridharan, S., Vaidyanathan, K., Kaul, B., Georganas, E., et al · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.
Highly efficient 8-bit low precision inference of convolutional neural networks with intelcaffe
Gong, J., Shen, H., Zhang, G., Liu, X., Li, S., Jin, G., Maheshwari, N., Fomenko, E., and Segal, E · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., Van Hasselt, H., and Silver, D · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2018
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2018
rlpyt: A research code base for deep reinforcement learning in pytorch
Stooke, A. and Abbeel, P · 2019
Later among the works it cites.
Battery charge scheduling in long-life autonomous mobile robots
Tomy, M., Lacerda, B., Hawes, N., and Wyatt, J. L · 2019
Later among the works it cites.
Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames
Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., and Batra, D · 2019
Later among the works it cites.
Swalp: Stochastic weight averaging in low precision training
Yang, G., Zhang, T., Kirichenko, P., Bai, J., Wilson, A. G., and De Sa, C · 2019
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ray: A distributed framework for emerging { \{ AI } \} applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Training deep neural networks with 8-bit floating point numbers
Wang, N., Choi, J., Brand, D., Chen, C.-Y., and Gopalakrishnan, K · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Accelerating reinforcement learning through gpu atari emulation
Dalton, S., Garland, M., and Frosio, I · 2019
Cited alongside, same era.
Seed rl: Scalable and efficient deep-rl with accelerated central inference
Espeholt, L., Marinier, R., Stanczyk, P., Wang, K., and Michalski, M · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Cited alongside, same era.
Qpytorch: A low-precision arithmetic simulation framework, 2019
Zhang, T., Lin, Z., Yang, G., and Sa, C. D · 2019
Later among the works it cites.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z · 2020
Later among the works it cites.
Automatic mixed precision package - torch.cuda.amp
Carilli, M · 2020
Later among the works it cites.
A statistical framework for low-bitwidth training of deep neural networks
Chen, J., Gai, Y., Yao, Z., Mahoney, M. W., and Gonzalez, J. E · 2020
Later among the works it cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Later among the works it cites.
Post-training piecewise linear quantization for deep neural networks
Fang, J., Shafiee, A., Abdel-Aziz, H., Thorsley, D., Georgiadis, G., and Hassoun, J. H · 2020
Later among the works it cites.
Stochastic rounding and reduced-precision fixed-point arithmetic for solving neural ordinary differential equations
Hopkins, M., Mikaitis, M., Lester, D. R., and Furber, S · 2020
Later among the works it cites.
Evaluating the performance of reinforcement learning algorithms
Jordan, S., Chandak, Y., Cohen, D., Zhang, M., and Thomas, P · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2020
Later among the works it cites.
Deep reinforcement learning-based beam tracking for low-latency services in vehicular networks
Liu, Y., Jiang, Z., Zhang, S., and Xu, S · 2020
Later among the works it cites.
Wrapnet: Neural net inference with ultra-low-resolution arithmetic
Ni, R., Chu, H.-m., Castañeda, O., Chiang, P.-y., Studer, C., and Goldstein, T · 2020
Later among the works it cites.
Sample factory: Egocentric 3d control from pixels at 100000 fps with asynchronous reinforcement learning
Petrenko, A., Huang, Z., Kumar, T., Sukhatme, G., and Koltun, V · 2020
Later among the works it cites.
dm control: Software and tasks for continuous control, 2020
Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., and Heess, N · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
Yarats, D. and Kostrikov, I · 2020
Later among the works it cites.
Zamirai, P., Zhang, J., Aberger, C. R., and De Sa, C · 2020
Later among the works it cites.
A framework for efficient robotic manipulation
Zhan, A., Zhao, P., Pinto, L., Abbeel, P., and Laskin, M · 2020
Later among the works it cites.
The ingredients of real-world robotic reinforcement learning
Zhu, H., Yu, J., Gupta, A., Shah, D., Hartikainen, K., Singh, A., Kumar, V., and Levine, S · 2020
Later among the works it cites.