Fetching the paper…
Reading the bibliography…
Stateful optimizers maintain gradient statistics over time, e.g., the exponentially smoothed sum (SGD with momentum) or squared sum (Adam) of past gradient values.
Computing extremely accurate quantiles using t-digests
Dunning, T. and Ertl, O. (2019) · 1902
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. (2019) · 1904
Earlier work this paper cites.
Mixed precision training with 8-bit floating point
Mellempudi, N., Srinivasan, S., Das, D., and Kaul, B. (2019) · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. (2019) · 1909
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. (2019) · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2019) · 1910
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N. (1999) · 1999
Earlier work this paper cites.
Quantile and histogram estimation
Chen, E. J. and Kelton, W. D. (2001) · 2001
Earlier work this paper cites.
Space-efficient online computation of quantile summaries
Greenwald, M. and Khanna, S. (2001) · 2001
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Training large neural networks with constant memory using a new execution algorithm
Pudipeddi, B., Mesmakhosroshahi, M., Xi, J., and Bharadwaj, S. (2020) · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K. (2020b) · 2003
Earlier work this paper cites.
Binary neural networks: A survey
Qin, H., Gong, R., Liu, X., Bai, X., Song, J., and Sebe, N. (2020) · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2005
Earlier work this paper cites.
Fast and approximate stream mining of quantiles and frequencies using graphics processors
Govindaraju, N. K., Raghuvanshi, N., and Manocha, D. (2005) · 2005
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. (2020) · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al. (2020) · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Courbariaux, M., Bengio, Y., and David, J.-P. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
Results of the wmt14 metrics shared task
Macháček, M. and Bojar, O. (2014) · 2014
Cited alongside, same era.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J. (2015) · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A. (2015) · 2015
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M. (2018) · 2018
Later among the works it cites.
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V. (2018) · 2018
Later among the works it cites.
Training deep neural networks with 8-bit floating point numbers
Wang, N., Choi, J., Brand, D., Chen, C., and Gopalakrishnan, K. (2018b) · 2018
Later among the works it cites.
Memory efficient adaptive optimization
Anil, R., Gupta, V., Koren, T., and Singer, Y. (2019) · 2019
Later among the works it cites.
Accurate and efficient 2-bit quantized neural networks
Choi, J., Venkataramani, S., Srinivasan, V., Gopalakrishnan, K., Wang, Z., and Chuang, P. (2019) · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Cited alongside, same era.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016) · 2016
Cited alongside, same era.
Binarynet: Training deep neural networks with weights and activations constrained to +1 or -1
Courbariaux, M. and Bengio, Y. (2016) · 2016
Cited alongside, same era.
8-bit approximations for parallelism in deep learning
Dettmers, T. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A. (2016) · 2016
Cited alongside, same era.
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019) · 2019
Later among the works it cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V. (2019) · 2019
Later among the works it cites.
Differentiable soft quantization: Bridging full-precision and low-bit neural networks
Gong, R., Liu, X., Jiang, S., Li, T., Hu, P., Lin, J., Yu, F., and Yan, J. (2019) · 2019
Later among the works it cites.
Fully quantized network for object detection
Li, R., Wang, Y., Liang, F., Qin, H., Yan, J., and Fan, R. (2019) · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Later among the works it cites.
Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks
Sun, X., Choi, J., Chen, C., Wang, N., Venkataramani, S., Srinivasan, V., Cui, X., Zhang, W., and Gopalakrishnan, K. (2019) · 2019
Later among the works it cites.
Shifted and squeezed 8-bit floating point format for low-precision training of deep neural networks
Cambier, L., Bhiwandiwalla, A., Gong, T., Elibol, O. H., Nekuii, M., and Tang, H. (2020) · 2020
Later among the works it cites.
Extreme tensoring for low-memory preconditioning
Chen, X., Agarwal, N., Hazan, E., Zhang, C., and Zhang, Y. (2020a) · 2020
Later among the works it cites.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. (2020) · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. (2020) · 2020
Later among the works it cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, E. (2020) · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N. (2021) · 2021
Closest in time.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L. (2021) · 2021
Closest in time.
End-to-end quantized training via log-barrier extensions
Li, J. B., Qu, S., Li, X., Strubell, E., and Metze, F. (2021) · 2021
Closest in time.
Xilinx/brevitas
Pappalardo, A. (2021) · 2021
Closest in time.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y. (2021) · 2021
Closest in time.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021) · 2021
Closest in time.