Errudite: Scalable, reproducible, and testable error analysis
Wu, T., Ribeiro, M. T., Heer, J., and Weld, D. (2019) · 2019
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias” in nlp
Blodgett, S. L., Barocas, S., Daum’e, H., and Wallach, H. M. (2020) · 2020
Cited alongside, same era.
Closing the ai accountability gap: Defining an end-to-end framework for internal algorithmic auditing
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., and Barnes, P. (2020) · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S. (2020) · 2020
Cited alongside, same era.
Problematic machine behavior: A systematic literature review of algorithm audits
Bandy, J. (2021) · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Original
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021) · 2021
Cited alongside, same era.
Discovering and validating ai errors with crowdsourced failure reports
Cabrera, A. A., Druck, A. J., Hong, J. I., and Perer, A. (2021) · 2021
Cited alongside, same era.
Whose ground truth? accounting for individual and collective identities underlying dataset annotation
Original
Denton, E., Díaz, M., Kivlichan, I., Prabhakaran, V., and Rosen, R. (2021) · 2021
Cited alongside, same era.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Gordon, M. L., Zhou, K., Patel, K., Hashimoto, T., and Bernstein, M. S. (2021) · 2021
Cited alongside, same era.
On the efficacy of adversarial data collection for question answering: Results from a large-scale randomized study
Kaushik, D., Kiela, D., Lipton, Z. C., and Yih, W.-t. (2021) · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A. (2021a) · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A. (2021b) · 2021
Cited alongside, same era.