Language Models are Few-Shot Learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
MLPerf: An Industry Standard Benchmark Suite for Machine Learning Performance
Mattson, P., Reddi, V. J., Cheng, C., Coleman, C., Diamos, G., Kanter, D., Micikevicius, P., Patterson, D., Schmuelling, G., Tang, H., Wei, G., and Wu, C · 2020
Later among the works it cites.
v0.7 Results
MLCommons · 2020
Later among the works it cites.
Reference numbers for BERT un-padding results
NVIDIA · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Later among the works it cites.
Visual transformers: Token-based image representation and processing for computer vision, 2020
Wu, B., Xu, C., Dai, X., Wan, A., Zhang, P., Yan, Z., Tomizuka, M., Gonzalez, J., Keutzer, K., and Vajda, P · 2020
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors · 2020
Later among the works it cites.
Datasets
Wolf, T., Lhoest, Q., von Platen, P., Jernite, Y., Drame, M., Plu, J., Chaumond, J., Delangue, C., Ma, C., Thakur, A., Patil, S., Davison, J., Scao, T. L., Sanh, V., Xu, C., Patry, N., McMillan-Major, A., Brandeis, S., Gugger, S., Lagunas, F., Debut, L., Funtowicz, M., Moi, A., Rush, S., Schmidd, P., Cistac, P., Muštar, V., Boudier, J., and Tordjmann, A · 2020
Later among the works it cites.
Mathematica, Version 12.2
Wolfram Research Inc · 2020
Later among the works it cites.
Effective Transformer
ByteDance Inc · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Closest in time.
Faster Transformer
NVIDIA · 2021
Closest in time.
XLA: Optimizing Compiler for Machine Learning
XLA, T · 2021
Closest in time.
Performance catalogue for BERT on Pytorch
NVIDIA · 2021
Closest in time.
Supplemental Material for “Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance’, 2022
Anonymous · 2022
Closest in time.