Tjong Kim Sang, E.F., De Meulder, F.: Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In: In Proc. NAACL-HLT Conf. pp. 142–147 (2003)
2003
Earlier work this paper cites.
Segura, S., Fraser, G., Sanchez, A.B., Ruiz-Cortés, A.: A survey on metamorphic testing. IEEE Transactions on software engineering 42
2016
Earlier work this paper cites.
Su, Y., Sun, H., Sadler, B., Srivatsa, M., Gür, I., Yan, Z., Yan, X.: On generating characteristic-rich question sets for qa evaluation. In: In Proc. EMNLP Conf. pp. 562–572 (2016)
2016
Earlier work this paper cites.
Yih, W.t., Richardson, M., Meek, C., Chang, M.W., Suh, J.: The value of semantic parse labeling for knowledge base question answering. In: In Proc. ACL Conf. pp. 201–206 (2016)
2016
Earlier work this paper cites.
Ngomo, N.: 9th challenge on question answering over linked data (qald-9). language 7
2018
Earlier work this paper cites.
Talmor, A., Berant, J.: The web as a knowledge-base for answering complex questions. In: In Proc. ACL Conf. pp. 641–651 (2018)
2018
Earlier work this paper cites.
Belinkov, Y., Glass, J.: Analysis methods in neural language processing: A survey. Transactions of the Association for Computational Linguistics 7
2019
Earlier work this paper cites.
Kenton, J.D.M.W.C., Toutanova, L.K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: In Proc. NAACL-HLT Conf. pp. 4171–4186 (2019)
2019
Earlier work this paper cites.
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., Miller, A.: Language models as knowledge bases? In: In Proc. IJCAI Conf. pp. 2463–2473 (2019)
2019
Earlier work this paper cites.
Rychalska, B., Basaj, D., Gosiewska, A., Biecek, P.: Models in the wild: On corruption robustness of neural nlp systems. In: In Proc. ICONIP Conf. pp. 235–247 (2019)
2019
Earlier work this paper cites.
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.R.: Superglue: a stickier benchmark for general-purpose language understanding systems. In: In Proc. NeurIPS Conf. pp. 3266–3280 (2019)
2019
Earlier work this paper cites.
Wu, T., Ribeiro, M.T., Heer, J., Weld, D.S.: Errudite: Scalable, reproducible, and testable error analysis. In: In Proc. ACL Conf. pp. 747–763 (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. Advances in neural information processing systems 33
2020
Earlier work this paper cites.
Jiang, Z., Xu, F.F., Araki, J., Neubig, G.: How can we know what language models know? Transactions of the Association for Computational Linguistics 8
2020
Earlier work this paper cites.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21
2020
Earlier work this paper cites.
Ribeiro, M.T., Wu, T., Guestrin, C., Singh, S.: Beyond accuracy: Behavioral testing of nlp models with checklist. In: In Proc. ACL Conf. pp. 4902–4912 (2020)
2020
Earlier work this paper cites.
Gu, Y., Kase, S., Vanni, M., Sadler, B., Liang, P., Yan, X., Su, Y.: Beyond iid: three levels of generalization for question answering on knowledge bases. In: In Proc. WWW Conf. pp. 3477–3488 (2021)
2021
Earlier work this paper cites.
He, H., Choi, J.D.: The stem cell hypothesis: Dilemma behind multi-task learning with transformer encoders. In: In Proc. EMNLP Conf. pp. 5555–5577 (2021)
2021
Earlier work this paper cites.