Fetching the paper…
Reading the bibliography…
The rapid scaling of language models is motivating research using low-bitwidth quantization.
Selective brain damage: Measuring the disparate impact of model pruning
Hooker, S., Courville, A. C., Dauphin, Y. N., and Frome, A · 1911
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Minimum Bayes-risk decoding for statistical machine translation
Kumar, S. and Byrne, W · 2004
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., et al · 2016
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Earlier work this paper cites.
Results of the WMT17 metrics shared task
Bojar, O., Graham, Y., and Kamran, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm
Liu, Z., Wu, B., Luo, W., Yang, X., Liu, W., and Cheng, K.-T · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Post, M · 2018
Earlier work this paper cites.
Neural machine translation with 4-bit precision and beyond
Aji, A. F. and Heafield, K · 2019
Earlier work this paper cites.
Findings of the 2019 conference on machine translation (WMT19)
Barrault, L., Bojar, O., Costa-jussà, M. R., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Huck, M., Koehn, P., Malmasi, S., Monz, C., Müller, M., Pal, S., Post, M., and Zampieri, M · 2019
Cited alongside, same era.
Efficient 8-bit quantization of transformer neural machine language translation model
Bhandare, A., Sripathi, V., Karkada, D., Menon, V., Choi, S., Datta, K., and Saletore, V · 2019
Cited alongside, same era.
Fully quantized transformer for machine translation
Prato, G., Charlaix, E., and Rezagholizadeh, M · 2019
Cited alongside, same era.
Q8bert: Quantized 8bit bert
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Cited alongside, same era.
Binarybert: Pushing the limit of bert quantization
Bai, H., Zhang, W., Hou, L., Shang, L., Jin, J., Jiang, X., Liu, Q., Lyu, M., and King, I · 2020
{GS}hard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Later among the works it cites.
How do adam and training strategies help bnns optimization
Liu, Z., Shen, Z., Li, S., Helwegen, K., Huang, D., and Cheng, K.-T · 2021
Later among the works it cites.
Scaling language model training to a trillion parameters using megatron, 2021
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., and Catanzaro, B · 2021
Later among the works it cites.
An evaluation of edge tpu accelerators for convolutional neural networks
Seshadri, K., Akin, B., Laudon, J., Narayanaswami, R., and Yazdanbakhsh, A · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BLEU might be guilty but references are not innocent
Freitag, M., Grangier, D., and Caswell, I · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A · 2020
Cited alongside, same era.
Findings of the 2021 conference on machine translation (WMT21)
Akhbardeh, F., Arkhangorodsky, A., Biesialska, M., Bojar, O., Chatterjee, R., Chaudhary, V., Costa-jussa, M. R., España-Bonet, C., Fan, A., Federmann, C., Freitag, M., Graham, Y., Grundkiewicz, R., Haddow, B., Harter, L., Heafield, K., Homan, C., Huck, M., Amponsah-Kaakyire, K., Kasai, J., Khashabi, D., Knight, K., Kocmi, T., Koehn, P., Lourie, N., Monz, C., Morishita, M., Nagata, M., Nagesh, A., Nakazawa, T., Negri, M., Pal, S., Tapo, A. A., Turchi, M., Vydrin, V., and Zampieri, M · 2021
Cited alongside, same era.
Meliusnet: An improved network architecture for binary neural networks
Bethge, J., Bartz, C., Yang, H., Chen, Y., and Meinel, C · 2021
Cited alongside, same era.
Characterizing signal propagation to close the performance gap in unnormalized resnets
Brock, A., De, S., and Smith, S. L · 2021
Cited alongside, same era.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C · 2021
Cited alongside, same era.
Later among the works it cites.
High quality rather than high model probability: Minimum Bayes risk decoding with neural metrics
Freitag, M., Grangier, D., Tan, Q., and Liang, B · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Aqt: Accurate quantized training), 2022
Lew, L., Feinberg, V., Agrawal, S., Lee, J., Malmaud, J., Wang, L., Dormiani, P., and Pope, R · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al · 2022
Later among the works it cites.
Bit: Robustly binarized multi-distilled transformer
Liu, Z., Oguz, B., Pappu, A., Xiao, L., Yih, S., Li, M., Krishnamoorthi, R., and Mehdad, Y · 2022
Later among the works it cites.
Bibert: Accurate fully binarized bert
Qin, H., Ding, Y., Zhang, M., Yan, Q., Liu, A., Dang, Q., Liu, Z., and Liu, X · 2022
Later among the works it cites.
Pokebnn: A binary pursuit of lightweight accuracy
Zhang, Y., Zhang, Z., and Lew, L · 2022
Later among the works it cites.