GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Original
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Later among the works it cites.
Evaluating entity disambiguation and the role of popularity in retrieval-based NLP
Chen, A., Gudipati, P., Longpre, S., Ling, X., and Singh, S · 2021
Later among the works it cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M · 2021
Later among the works it cites.
Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation
Elangovan, A., He, J., and Verspoor, K · 2021
Later among the works it cites.
Competency problems: On finding and removing artifacts in language data
Gardner, M., Merrill, W., Dodge, J., Peters, M., Ross, A., Singh, S., and Smith, N. A · 2021
Later among the works it cites.
Datasheets for datasets, 2021
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., au2, H. D. I., and Crawford, K · 2021
Later among the works it cites.
Question and answer test-train overlap in open-domain question answering datasets
Lewis, P., Stenetorp, P., and Riedel, S · 2021
Later among the works it cites.
What’s in the box? an analysis of undesirable content in the Common Crawl corpus
Luccioni, A. and Viviano, J · 2021
Later among the works it cites.
Are nlp models really able to solve simple math word problems?, 2021
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Sinha, K., Jia, R., Hupkes, D., Pineau, J., Williams, A., and Kiela, D · 2021
Later among the works it cites.
We need to talk about random splits
Søgaard, A., Ebert, S., Bastings, J., and Filippova, K · 2021
Later among the works it cites.
Concealed data poisoning attacks on NLP models
Wallace, E., Zhao, T., Feng, S., and Singh, S · 2021
Later among the works it cites.
Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX
Wang, B · 2021
Later among the works it cites.
Frequency effects on syntactic rule learning in transformers
Wei, J., Garrette, D., Linzen, T., and Pavlick, E · 2021
Later among the works it cites.
Counterfactual memorization in neural language models
Original
Zhang, C., Ippolito, D., Lee, K., Jagielski, M., Tramèr, F., and Carlini, N · 2021
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Original
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Closest in time.
Hypothesis only baselines in natural language inference
Poliak, A., Naradowsky, J., Haldar, A., Rudinger, R., and Van Durme, B · 2023
Closest in time.